Machine Learning (ML) - Notes¶

Src : towardsdatascience - Hunter Heidenreich
Agenda¶
- Agenda
- Intro
- ML vs Traditional coding
- Training the Model (CLassifier)
- ML algorithms
- ML Frameworks/tools
- ML Problem solving in 7 steps
- References
Intro¶
ML is a field of AI algorithms that learn from EXAMPLES and EXPERIENCES instead of traditional HARDCODE and RULES
Ex : Apple and oranges detection
ML vs Traditional coding¶
- Tradictional coding : too much rules
-
ML : Model / Classifier and train it to generate the RULES instead of writing them
-
classifier (as function) : takes a data as input (FEATURES) and signs LABEL to it as output(LABELS)
- Fruit(apple or orange?) => Classifier => apple (if apple chose)
- email (spam/mail_ok) => Classifier => spam (if mail_nok)
Training the Model (CLassifier)¶
To Train the Classifier we use :
- Supervised Learning(SL) : it learns from examples/experiences
- Unsuspervised Learning(USL) : it learns from events ?
- Reinforcement Learning (RL): Conceptually similar to human learning processes
- ex: a robot learning to walk
- Strategy games : Go, Chess etc
The more the training data exists => the better the classifier Will be
ML train
|--------| |-----------| |-----------|
|Collect | |Train | |Make |
|Training| => |Classifier | => |Predictions|
| Data | | | | |
|--------| |-----------| |-----------|
ML algorithms¶
Supervised:- Regression : Predicting a continuous-valued attribute associated with an object
- Multiple Linear Regression(MLR)
- Polynomial Regression (PR)
-
Classification : Identifying which category an object belongs to
- k-Nearest Neighbor
- Decision Trees(ID3, C4.5, C5.0)
- logistic regression
- Naïve Bayes
- Linear Discriminant Analysis
- Neural Networks
- Support Vector Machines (SVM)
- Random Forest(RF)
-
Unsupervised:- Clustering : Automatic grouping of similar objects into sets.
- k-Means
- Mean-shift
- Hierarchical Clustering (HC)
- Density-based Clustering (DBSCAN)
- Gaussian Mixture Models(GMM)
- Clustering : Automatic grouping of similar objects into sets.
Reinforcement:- Q-Learning
- Deep Q-Network (DQN)
- A3C
ML Frameworks/tools¶
- TensorFlow
- PyTorch
- Scikit-learn
- Spark ML
- Torch
- Huggingface
- Keras
ML Problem solving in 7 steps¶
- GATHERING / COLLECTING DATA : The more data we collect the more accurate will the model.
-
We collect datas to train the model of the system we want to deploy
-
DATA PREPARATION :
- Features ? : the input of the system
- Labels ? : the output of the system
- visualization of datas
- balances, relationship between datas
- split : training/evaluation(performance of the model)
-
Choosing the MODEL : There are already alot of model created by DataScientists :
- For : Music, image, number, text, text based data, linear model (y=ax+b)
-
TRAINING (the model : Y = mx+b) :
- Y : output
- m : SLope (many m possible, as many features)
- X : input
- b : Y-intercept Training process
Each iteration it's called, a training steps.
/!\ #Residus = Biais ( θ ^ ) ≡ E [ θ ^ ] − θ #Définition — Si θ ^ est l'estimateur de θ
- EVALUATION : after the model is good time to evaluate
|----------| |-----------| |------------|
| | | Model | | |
|EVALUATION| => | (W,b) | => | Prediction |
| Data | | | | |
|----------| |-----------| |------------|
|-----------| ||
/\ | | \/
|| <= |Test | <=
| (W,b) |
|-----------|
This metric allows the model to see the data that has not yet seen. This is to test how the model might act in the real world
-
PARAMETER TUNING
- To improve the training
- Repeat the training data several time to increase the accuracy
- Learning rate : limit of the train / how far we shift the line between two input datas
- initial conditions : for complexes models (value = 0 ...)
- Hyperparameters.
/!\ : it's important to choose the good parameters to be changed
- PREDICTIONS : ML uses datas to answer questions
Input : features Output : Labels
References¶
Scikit learn : https://scikit-learn.org/stable/#
TensorFlow : https://www.tensorflow.org/resources/learn-ml
Spark ML : https://spark.apache.org/docs/latest/ml-guide.html
PyTorch : https://pytorch.org/tutorials/beginner/deep_learning_60min_blitz.html https://docs.microsoft.com/en-us/learn/paths/pytorch-fundamentals/
Google course/ Josh gordon : https://www.youtube.com/watch?v=cKxRvEZd3Mw&list=PLOU2XLYxmsIIuiBfYad6rFYQU_jL2ryal
Google Cloud Plateform / Yufeng Guo: https://www.youtube.com/watch?v=nKW8Ndu7Mjw
IBM Cloud : https://www.ibm.com/cloud/blog/supervised-vs-unsupervised-learning