Data Profiling

preview-18
  • Data Profiling Book Detail

  • Author : Ziawasch Abedjan
  • Release Date : 2022-06-01
  • Publisher : Springer Nature
  • Genre : Computers
  • Pages : 136
  • ISBN 13 : 3031018656
  • File Size : 71,71 MB

Data Profiling by Ziawasch Abedjan PDF Summary

Book Description: Data profiling refers to the activity of collecting data about data, {i.e.}, metadata. Most IT professionals and researchers who work with data have engaged in data profiling, at least informally, to understand and explore an unfamiliar dataset or to determine whether a new dataset is appropriate for a particular task at hand. Data profiling results are also important in a variety of other situations, including query optimization, data integration, and data cleaning. Simple metadata are statistics, such as the number of rows and columns, schema and datatype information, the number of distinct values, statistical value distributions, and the number of null or empty values in each column. More complex types of metadata are statements about multiple columns and their correlation, such as candidate keys, functional dependencies, and other types of dependencies. This book provides a classification of the various types of profilable metadata, discusses popular data profiling tasks, and surveys state-of-the-art profiling algorithms. While most of the book focuses on tasks and algorithms for relational data profiling, we also briefly discuss systems and techniques for profiling non-relational data such as graphs and text. We conclude with a discussion of data profiling challenges and directions for future work in this area.

Disclaimer: www.yourbookbest.com does not own Data Profiling books pdf, neither created or scanned. We just provide the link that is already available on the internet, public domain and in Google Drive. If any way it violates the law or has any issues, then kindly mail us via contact us page to request the removal of the link.

Data Profiling

Data Profiling

File Size : 47,47 MB
Total View : 6938 Views
DOWNLOAD

Data profiling refers to the activity of collecting data about data, {i.e.}, metadata. Most IT professionals and researchers who work with data have engaged in

An Introduction to Duplicate Detection

An Introduction to Duplicate Detection

File Size : 34,34 MB
Total View : 7838 Views
DOWNLOAD

With the ever increasing volume of data, data quality problems abound. Multiple, yet different representations of the same real-world objects in data, duplicate

The Four Generations of Entity Resolution

The Four Generations of Entity Resolution

File Size : 49,49 MB
Total View : 7155 Views
DOWNLOAD

Entity Resolution (ER) lies at the core of data integration and cleaning and, thus, a bulk of the research examines ways for improving its effectiveness and tim