Training data and bias flashcards

Understanding training data and bias in AI and machine learning.

Dragon92·17 flashcards·7 questions
high schoolcomputer_scienceai_ml
0
Known
1 / 17
0
Learning
Front

What is training data?

Tap to flip
Back

Data used to teach a machine learning model how to predict outcomes.

Tap to flip
Got it
Still learning

Quiz(7 questions)

Question 1 of 7

1. What is a primary concern of biased training data?

Terms in this Study Set(17)

What is training data?

Data used to teach a machine learning model how to predict outcomes.

True or false: Bias in data is always negative.

False, because bias can sometimes lead to useful simplifications.

Difference between supervised and unsupervised learning.

Supervised learning uses labeled data; unsupervised learning does not.

What is model bias?

A model's tendency to consistently deviate from true values due to training data.

Fill in the blank: __________ can cause bias in AI models.

Imbalanced training data

What is overfitting?

When a model learns noise in training data and performs poorly on new data.

True or false: All datasets are perfectly unbiased.

False, because most datasets contain some form of bias.

Question: How can bias affect AI outcomes?

Bias can lead to unfair or inaccurate predictions, affecting real-world applications.

What is underfitting?

When a model is too simple to capture the underlying patterns in data.

Comparison: Bias vs. Variance.

Bias is error due to oversimplification; variance is error due to complexity.

What is data augmentation?

A technique to increase training data by creating modified versions of existing data.

True or false: Larger datasets always eliminate bias.

False, because larger datasets can still reflect existing biases.

Fill in the blank: __________ is used to correct bias in AI.

Fairness algorithms

Question: What is the role of validation data?

To check how well a model generalizes to new, unseen data.

What does feature selection involve?

Choosing the most relevant attributes from data that contribute to predictions.

True or false: Training data must always be diverse.

True, because diversity helps reduce bias and improve model performance.

What is a data pipeline?

A set of processes that collects, cleans, and prepares data for analysis.

Questions in this Study Set(7)

1. What is a primary concern of biased training data?

A.Unfair outcomes
B.Better accuracy
C.Increased complexity
D.More features

2. Which is NOT a type of bias?

A.Label bias
B.Sample bias
C.Data variance
D.Measurement bias

3. What can data augmentation help with?

A.Reducing overfitting
B.Increasing bias
C.Creating noise
D.Decreasing data

4. What is the effect of overfitting on model performance?

A.Improves accuracy
B.No effect
C.Lowers accuracy on new data
D.Enhances features

5. What does a fairness algorithm aim to achieve?

A.Increase model speed
B.Eliminate all biases
C.Correct unfair predictions
D.Maximize accuracy

6. Which type of learning uses labeled data?

A.Unsupervised learning
B.Reinforcement learning
C.Supervised learning
D.Semi-supervised learning

7. What is model variance?

A.Error from bias
B.Inconsistent predictions on new data
C.Correct predictions
D.Stable performance

Related Study Sets

Create Your Own Study Set

Upload a PDF, paste your notes, or describe a topic – AI generates flashcards, quizzes and more in seconds.