Understanding the Challenges and Providing Logging Support to Monitor Data Processing in Big Data Application

preview-18
  • Understanding the Challenges and Providing Logging Support to Monitor Data Processing in Big Data Application Book Detail

  • Author : Zehao Wang
  • Release Date : 2022
  • Publisher :
  • Genre :
  • Pages : 0
  • ISBN 13 :
  • File Size : 82,82 MB

Understanding the Challenges and Providing Logging Support to Monitor Data Processing in Big Data Application by Zehao Wang PDF Summary

Book Description: To analyze large-scale data efficiently, developers have created various big data processing frameworks (e.g., Apache Spark). These big data processing frameworks provide abstractions to developers so that they can focus on implementing the logic for data analysis. In traditional software systems, developers leverage logging to monitor applications and record intermediate states to assist workload understanding and issue diagnosis. However, due to the abstraction and the peculiarity of big data frameworks, there is currently no effective monitoring approach for big data applications. In this thesis, we first manually study 1,000 randomly sampled Spark-related questions on Stack Overflow to study their root causes and the type of information, if recorded, that can assist developers with motioning and diagnosis. Then, we design an approach, DPLOG, which assists developers with monitoring Spark applications. DPLOG leverages statistical sampling to minimize performance overhead and provides intermediate information and hint/warning messages for each data processing step of a chained method pipeline. We evaluate DPLOG on six benchmarking programs and find that DPLOG has a relatively small overhead (i.e., less than 10% increase in response time when processing 5GB data) compared to without using DPLOG, and reduce the overhead by over 500% compared to the baseline. Our user study with 20 developers shows that DPLOG can reduce the needed time to debug big data applications by 63% and the participants give DPLOG 4.85/5 for its usefulness on average. Moreover, the idea of DPLOG may be applied to other big data processing frameworks, and our study sheds light on future research opportunities in assisting developers with monitoring big data applications.

Disclaimer: www.yourbookbest.com does not own Understanding the Challenges and Providing Logging Support to Monitor Data Processing in Big Data Application books pdf, neither created or scanned. We just provide the link that is already available on the internet, public domain and in Google Drive. If any way it violates the law or has any issues, then kindly mail us via contact us page to request the removal of the link.

Statistics in a Nutshell

Statistics in a Nutshell

File Size : 54,54 MB
Total View : 257 Views
DOWNLOAD

A clear and concise introduction and reference for anyone new to the subject of statistics.

Big Data

Big Data

File Size : 22,22 MB
Total View : 8885 Views
DOWNLOAD

This Springer Brief provides a comprehensive overview of the background and recent developments of big data. The value chain of big data is divided into four ph