Linear Digressions

Data Contamination

Linear Digressions

Supervised machine learning assumes that the features and labels used for building a classifier are isolated from each other--basically, that you can't cheat by peeking. Turns out this can be easier said than done. In this episode, we'll talk about the many (and diverse!) cases where label information contaminates features, ruining data science competitions along the way. Relevant links: https://www.researchgate.net/profile/Claudia_Perlich/publication/221653692_Leakage_in_data_mining_Formulation_detection_and_avoidance/links/54418bb80cf2a6a049a5a0ca.pdf

Next Episodes

Linear Digressions

Model Interpretation (and Trust Issues) @ Linear Digressions

📆 2016-04-25 02:45 / 00:16:57


Linear Digressions

Updates! Political Science Fraud and AlphaGo @ Linear Digressions

📆 2016-04-18 04:48 / 00:31:43


Linear Digressions

Ecological Inference and Simpson's Paradox @ Linear Digressions

📆 2016-04-11 04:43 / 00:18:32


Linear Digressions

Discriminatory Algorithms @ Linear Digressions

📆 2016-04-04 04:30 / 00:15:21


Linear Digressions

Recommendation Engines and Privacy @ Linear Digressions

📆 2016-03-28 04:46 / 00:31:33