February 13, 2018
SURFsara
CET timezone
Apache Spark is one of the most popular computing frameworks for large-scale data processing. It also includes a machine learning library (MLlib) with distributed versions of many machine learning algorithms. In this workshop we give an introduction to Apache Spark and explain how to use it for distributed machine learning. For the hands-on we will be using PySpark, Sparks Python API, from a Jupyter notebook environment.

Conference information

Date/Time

Starts

Ends

All times are in CET

Location

SURFsara
VK1/VK2
Science Park 140, 1098 XG Amsterdam

Extra information

Please bring your own laptop (with an ssh client installed) for the hands-on sessions! Requirements: * Experience with the Python programming language * Basic knowledge of supervised machine learning methods Practical information: To travel with public transport in The Netherlands, it is very convenient if you have an anonymous OV-chipkaart. Please find the detailed information at https://www.ov-chipkaart.nl/purchase-an-ov-chipkaart/anonymous-ov-chipkaart.htm. Information on how to reach SURFsara can be found at https://www.surf.nl/en/about-surf/contact/directions-to-surfsara/index.html.