The program consists of talks on the accepted papers and two keynotes.
We plan to have a poster session to spark more discussion and networking.
August 31st (all times are in Boston Time)
09:00 - 09:15
Opening Remarks
09:15 - 10:15
Session 1: Morning Keynote - Chair: Cinzia Cappiello
Keynote Title: The Four Paradigms of Data Quality Assessment
Keynote Speaker: Divesh Srivastava (AT&T Chief Data Office)
Keynote Abstract: Scientific discovery uses a variety of methods which have been characterized through four paradigms: empirical (i.e., using descriptions of natural phenomena), theoretical (i.e., based on mathematical formulas), computational (i.e., using computationally intensive methods), and data-driven (i.e., based on analysis of large, complex data sets). This talk will draw parallels between methods used for scientific discovery and those used for discovering data errors, and present challenges and opportunities within the four paradigms of data quality assessment.
Keynote Speaker Bio: Divesh Srivastava is the head of Database Research at AT&T. He is an AT&T Fellow, a Fellow of the ACM, on the Board of Directors of the Computing Research Association (CRA), co-chair of the ACM Publications Board, and former President of the VLDB Endowment. He is a recipient of the 2021 IEEE TCDE Impact Award, the 2021 ACM SIGMOD Contributions Award, and the 2023 Distinguished Alumnus Award from the Indian Institute of Technology, Bombay. He has conducted research on a wide variety of topics in data management for three decades and has presented keynotes at several international conferences. He received his Ph.D. from the University of Wisconsin, Madison, and his Bachelor of Technology from the Indian Institute of Technology, Bombay, India.
10:15 - 10:45
Break
10:45 - 12:15
Session 2: Morning Research Session - Chair: Cinzia Cappiello
Can LLMs Serve as a Data Error Detection Engine? Trade-offs in Accuracy, Cost, and Hallucination Across Datasets
Maximilian Plazotta and Meike Klettke
Assessing Data Quality in Relational Database Migration: The Data Quality Score (DQS)
Abhinav Srivastava and Tanya Chaudhary
Measuring Credibility in Social Media Platforms through Data Quality and Provenance
Sebastian García and Adriana Marotta
Automatic Consistency Assessment through Partial Functional Dependency Mining
Marcian Seeger, Thorsten Papenbrock, Felix Naumann, and Lisa Ehrlinger
12:15 - 13:45
Lunch
13:45 - 14:45
Session 3: Afternoon Keynote - Chair: Steven Whang
Keynote Title: Ad-hoc Data Integration in Data Lakes
Keynote Speaker: Ziawasch Abedjan (TU Berlin)
Keynote Abstract: As data lakes have become a prominent foundation for enterprise and scientific data management, organizations increasingly face the challenge of locating relevant datasets and building ad-hoc integration pipelines across heterogeneous, poorly documented, and rapidly evolving data collections. In this setting, ad-hoc data discovery and integration becomes a critical capability for turning raw, distributed data assets into usable knowledge. In this talk, I will give insights into our research on holistic discovery systems, task-dependent discovery, and lake cleaning.
Keynote Speaker Bio: Ziawasch Abedjan is Chair of the Data Integration and Data Preparation group at TU Berlin and research group lead at Berlin Institute for Foundations of Learning and Data (BIFOLD). His research is focused on developing generalizable and scalable techniques for various data integration problems, such as data cleaning, data science pipeline generation and understanding, and data discovery. Previously, he chaired the Database and Information systems group at the Leibniz University Hannover. He held positions as Assistant Professor at TU Berlin, Postdoctoral Associate at MIT, Research Associate at QCRI, Senior Researcher at the German Center for Artificial Intelligence (DFKI), and Visiting Academic at Amazon Search. He received his PhD from the Hasso Plattner Institute in Potsdam and was awarded the University of Potsdam's best Dissertation Prize in 2014. He has co-authored more than 80 peer-reviewed papers in leading database and software engineering venues and received recognition with several academic awards. His research is supported by the German Research Council (DFG) and the German Federal Ministry of Science and Education.
14:45 - 15:45
Poster session (continues during the coffee break)
15:15 - 15:45
Break
15:45 - 17:15
Session 4: Afternoon Research Talks - Chair: Steven Whang
Scalable Goal-Oriented Source Selection
Ambarish Singh and Romila Pradhan
Improving Data Preparation for CSV Files with LLMs
Alexander van Renen, Moritz Rengert, Macallyster Edmondson, and Andreas Kipf
Towards Inference-Aware Privacy Guidance for Data Preparation
Vishal Chakraborty and Felix Naumann
KGpipe: Generation of Pipelines for Data Integration into Knowledge Graphs
Marvin Hofer and Erhard Rahm
17:15 -17:30
Closing Remarks
↑ top