When to Exclude Unhelpful Training Data
Training examples should be excluded when they add no useful signal for the target task and merely increase training cost. A mismatch with the evaluation distribution is not sufficient by itself: auxiliary data from another distribution may still provide useful signal and can sometimes be given less weight instead of being removed. Data unrelated to the task, such as office emails in a damaged-highway-sign classifier, wastes computation and can make the model spend capacity on irrelevant patterns.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
When Development and Test Sets Reflect Different Populations
How Model Capacity Changes the Risk of Mixing Data Sources
One Predictor Can Work Across Multiple Data Sources
Choose evaluation data to match the real-world target
Mismatched Auxiliary Data Source
Building Dev and Test Sets Before Real Users Exist
Refreshing Evaluation Sets After a Product Launch
Using Public Web Images When No Better Future-Like Data Exists
Judging How Much to Invest in Dev and Test Sets
What should determine dev and test set selection?
True or False: You can assume the training set and test set always come from the same distribution.
Development and test sets should reflect the conditions you expect after deployment, not only the _____ available in your training pool.
Why can a simple random test split be a poor choice when the data you expect in the future is different from the data you have now?
You can usually assume the data used for training and the data used for testing come from the same distribution.
Design Dev and Test Sets for the Future
Match each concept about development and test sets to its description.
Order the steps for choosing development and test sets when future data differs from training data.
What should dev and test examples be designed to resemble?
A validation and test set must exactly match the training distribution in every project.
How should a test set be chosen when deployment data will differ?
Match each data scenario to the best dev/test set choice.
Order the reasoning steps for deciding whether a dev/test split is appropriate.
Why a Random 30 Percent Split Can Be Misleading When Future Data Will Differ
Dev and Test Splits for a Field-Photo Classifier
How to Choose Dev and Test Data When Future Data Will Differ
When to Exclude Unhelpful Training Data
Learn After
Why Skip Data That Does Not Match the Evaluation Set?
True or False: If the training, validation, and test sets are drawn from the same distribution, collecting more training examples will always make the model perform better.
According to machine learning strategy, if the dev error curve has _____, adding more training data is unlikely to help you reach your target.
Why might you exclude data that does not add useful information for training?
True or False: More training data always improves validation accuracy.
If the dev error curve has _____, adding more training data is unlikely to move you toward the target performance.
Match each machine-learning concept with the description that fits it best when deciding whether collecting more data is worthwhile.
Order the steps for using a learning curve to judge whether collecting more training data is worthwhile.
In a leaf-disease classifier, why should a large archive of handwritten invoices be left out of training?
True or False: Examining the learning curve can help you avoid spending months collecting more data that later turns out not to improve validation performance.
When compute is limited, examples that add no _____ should be left out of training.
Match each data scenario to the best action.
Order the steps for deciding whether to add a new data source to training.
Why Irrelevant Training Data Should Be Excluded
Whether to Add Contract Scans to a Plant Photo Classifier
Irrelevant Training Data and Model Capacity