Flashcard Deck · 23 cards · Public

Machine Learning - Supervised Learning & Ensemble Methods

Master supervised learning and powerful ensemble methods with this comprehensive flashcard deck! Dive into algorithms like Linear Regression, SVMs, and Decision Trees, then unlock advanced techniques such as Random Forests, Gradient Boosting, and XGBoost to build highly accurate predictive models.

Cards in this deck

(23 cards)

Preview terms and definitions before starting your study session.

#1
Term
What is Supervised Learning?
Definition
A machine learning task where an algorithm learns from a dataset of labeled examples, meaning each input feature vector has a corresponding correct output value (label). The goal is to learn a mapping function from inputs to outputs to make predictions on unseen data.
#2
Term
What is 'labeled data' in supervised learning?
Definition
Data where each input example (features) is paired with its correct output value or target variable (label). For instance, in an image classification task, an image (input) would be labeled with the object it contains (output, e.g., 'cat' or 'dog').
#3
Term
Differentiate between Regression and Classification tasks in supervised learning.
Definition
1. Regression: Predicts a continuous output value (e.g., house price, temperature).
2. Classification: Predicts a categorical output value (e.g., spam/not spam, disease/no disease, digit recognition).
#4
Term
Explain the Bias-Variance Tradeoff.
Definition
A fundamental concept describing the conflict in simultaneously minimizing two sources of error that prevent supervised learning algorithms from generalizing beyond their training data:
  • Bias: Error from erroneous assumptions in the learning algorithm (underfitting).
  • Variance: Error from sensitivity to small fluctuations in the training set (overfitting).
The tradeoff implies that reducing one typically increases the other.
#5
Term
What is overfitting in machine learning?
Definition
Occurs when a model learns the training data too well, including its noise and specific patterns, leading to excellent performance on the training set but poor generalization to new, unseen data. High variance is a characteristic of overfitting.
#6
Term
What is underfitting in machine learning?
Definition
Occurs when a model is too simple to capture the underlying patterns in the training data, resulting in poor performance on both the training set and new data. High bias is a characteristic of underfitting.
#7
Term
What is K-Fold Cross-Validation and why is it used?
Definition
A technique to evaluate a model's performance and assess its generalization ability. The data is split into equally sized folds. The model is trained times, each time using folds for training and the remaining fold for validation. The results are averaged. It helps mitigate overfitting to the validation set and provides a more robust estimate of performance.
#8
Term
Explain Linear Regression and its objective function.
Definition
A supervised learning algorithm used for regression tasks. It models the relationship between a dependent variable and one or more independent variables by fitting a linear equation to the observed data.
  • Equation:
  • Objective (MSE): Minimize the Mean Squared Error (MSE), which is .
Showing 8 of 23 cards in this deck.
Study All Now