Name: Mining Imperfect Data: With Examples in R and Python by Ronald K Pearson 9781611976267
SKU: 9781611976267
Price: 90.92 EUR
Availability: InStock

Description
Reviews

Description

It has been estimated that as much as 80% of the total effort in a typical data analysis project is taken up with data preparation, including reconciling and merging data from different sources, identifying and interpreting various data anomalies, and selecting and implementing appropriate treatment strategies for the anomalies that are found. This book focuses on the identification and treatment of data anomalies, including examples that highlight different types of anomalies, their potential consequences if left undetected and untreated, and options for dealing with them.

As both data sources and free, open-source data analysis software environments proliferate, more people and organizations are motivated to extract useful insights and information from data of many different kinds (e.g., numerical, categorical, and text). The book emphasizes the range of open-source tools available for identifying and treating data anomalies, mostly in R but also with several examples in Python.

Mining Imperfect Data: With Examples in R and Python, Second Edition

presents a unified coverage of 10 different types of data anomalies (outliers, missing data, inliers, metadata errors, misalignment errors, thin levels in categorical variables, noninformative variables, duplicated records, coarsening of numerical data, and target leakage);
includes an in-depth treatment of time-series outliers and simple nonlinear digital filtering strategies for dealing with them; and
provides a detailed introduction to several useful mathematical characteristics of important data characterizations that do not appear to be widely known among practitioners, such as functional equations and key inequalities.

About the Author
is a senior data scientist at GeoVera Holdings, a U.S. based property insurance company. He has held positions in both academia and industry and has been actively involved in both research and applications in several data-related fields, including industrial process control and monitoring, signal processing, bioinformatics, drug safety data analysis, property-casualty insurance, and software development. The author of over 100 conference and journal papers as well as six books, he is a member of SIAM and a Senior Life Member of IEEE, holds two patents, and is an author of two R packages.

Book Information
ISBN 9781611976267
Author Ronald K. Pearson
Format Paperback
Page Count 481
Imprint Society for Industrial & Applied Mathematics,U.S.
Publisher Society for Industrial & Applied Mathematics,U.S.
Weight(grams) 1035g

Reviews

No reviews yet Write a Review

Related Titles

Exploratory Data Analysis Using R by Ronald K. Pearson 9780367571566

RRP: €60.68

Booksplease Price: €55.79

Exploratory Data Analysis Using R provides a classroom-tested introduction to exploratory data analysis (EDA) and introduces the range of "interesting" – good, bad, and ugly – features that can be...

Add to Cart

Nonlinear Digital Filtering with Python: An Introduction by Ronald K. Pearson

RRP: €148.75

Booksplease Price: €132.86

Nonlinear Digital Filtering with Python: An Introduction discusses important structural filter classes including the median filter and a number of its extensions (e.g., weighted and recursive median...

Add to Cart

R and Data Mining: Examples and Case Studies by Yanchang Zhao 9780123969637

RRP: €74.96

Booksplease Price: €68.52

R and Data Mining introduces researchers, post-graduate students, and analysts to data mining using R, a free software environment for statistical computing and graphics. The book provides practical...

Add to Cart

Recently Viewed

Mining Imperfect Data: With Examples in R and Python by Ronald K Pearson 9781611976267

Frequently Bought Together:

Description

Goodreads reviews

Reviews