BEGIN:VCALENDAR
PRODID:-//hacksw/handcal//NONSGML v1.0//EN
METHOD:PUBLISH
BEGIN:VEVENT
DESCRIPTION:Click for Latest Location Information: http://smartdata2015.dataversity.net/sessionPop.cfm?confid=91&proposalid=7767\nFraud detection is a classic adversarial analytics challenge: As soon as an automated system successfully learns to stop one scheme, fraudsters move on to attack another way. Each scheme requires looking for different signals (i.e. features) to catch, is relatively rare (one in millions for finance or ecommerce, for example), and it may take months to investigate a single case (in healthcare or tax, for example) – making quality training data scarce.\nThis talk will cover, via live demo & code walk-through, the key lessons we’ve learned while building such real-world software systems over the past few years. We’ll be looking for fraud signals in public email datasets, using IPython and popular open-source libraries (scikit-learn, statsmodel, nltk, etc.) for data science and Apache Spark as the compute engine for scalable parallel processing.\nWe will iteratively build a machine-learned hybrid model – combining features from different data sources & algorithmic approaches, to catch diverse aspects of suspect behavior.\nWe will discuss:\nNatural language processing: Finding keywords in relevant context within unstructured text\nStatistical NLP: sentiment analysis, via supervised machine learning\nTime series analysis: understanding daily/weekly cycles and changes in habitual behavior\nGraph analysis: finding actions outside the usual or expected network of people\nHeuristic rules: finding suspect actions based on past schemes or external datasets\nTopic modeling: highlighting use of keywords outside an expected context\nAnomaly detection: Fully unsupervised ranking of unusual behavior\nThis talk assumes basic understanding of these data science tools, so that we can focus on their applicability for this use case, and on how they complement each other.\nApache Spark is used to run these models at scale – in batch mode for model training and with Spark Streaming for production use. We’ll discuss the data model, computation & feedback workflows, as well as some tools & libraries built on top of the open-source components to enable faster experimentation, optimization & productization of the models.
DTSTART:20150819T110000
SUMMARY:Hunting Criminals with Hybrid Analytics, Semi-supervised Learning & Agent Feedback
DTEND:20150819T112959
LOCATION: See Description
END:VEVENT
END:VCALENDAR 