Automated Data Augmentation Policy Learning for Cross-Dataset Transfer

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning models often require manually designed data augmentation policies, which can be inefficient and resource-intensive, and lack transferability across different training datasets.

Innovation Solution

A training system that automatically learns a data augmentation policy by searching a space of possible policies to identify an effective policy for training machine learning models, enhancing prediction accuracy and reducing resource consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manually designed data augmentation policies are used, then implementation is straightforward, but model performance is suboptimal and requires more training data

Engineering Contradiction:
Improvemodel performanceVSAvoidpolicy design complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system enables automatic learning of data augmentation policies through self-service mechanisms. The policy learning module autonomously searches the policy space and identifies optimal augmentation policies without requiring manual intervention, thereby improving model performance while eliminating the complexity of manual policy design

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical policy design with an automated learning system. The policy learning module uses machine learning algorithms to automatically discover optimal data augmentation policies, substituting the manual mechanical process with an intelligent automated system that improves both performance and efficiency

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If manually designed data augmentation policies are used, then resource consumption is high, but automatic policy learning requires more computational resources during training

Engineering Contradiction:
Improvetraining efficiencyVSAvoidcomputational resource consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary action by automatically learning optimal data augmentation policies before the actual model training process. The policy learning module pre-computes the best augmentation strategies, which are then reused during training to improve efficiency and reduce computational resource consumption during the main training phase

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies parameter changes by optimizing data augmentation policy parameters automatically. The policy learning module adjusts augmentation parameters such as transformation types, probabilities, and intensities to find the optimal configuration that maximizes training efficiency while minimizing computational resource consumption

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If manually designed data augmentation policies are used, then they can be applied to any dataset, but they lack transferability across different training datasets

Engineering Contradiction:
Improvepolicy transferabilityVSAvoidpolicy application ease
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The system achieves universality by learning data augmentation policies that are transferable across multiple datasets. The policy learning module discovers generalizable augmentation strategies that can be applied to different datasets and tasks, making the learned policies multi-functional and adaptable to various machine learning scenarios

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements feedback mechanisms where the policy learning module continuously evaluates the performance of learned policies across different datasets. This feedback loop enables the system to refine and adapt policies to achieve better transferability while maintaining ease of application across diverse datasets

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP4462306B1Learning data augmentation policies
Publication Date: 2025.07.02 GOOGLE LLC
  • EP4462306B1 patent drawingFigure 1
  • EP4462306B1 patent drawingFigure 2
  • EP4462306B1 patent drawingFigure 3

AI summary

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for learning a data augmentation policy for training a machine learning model. In one aspect, a method includes: receiving training data for training a machine learning model to perform a particular machine learning task; determining multiple data augmentation policies, comprising, at each of multiple time steps: generating a current data augmentation policy based on quality measures of data augmentation policies generated at previous time steps; training a machine learning model on the training data using the current data augmentation policy; and determining a quality measure of the current data augmentation policy using the machine learning model after it has been trained using the current data augmentation policy; and selecting a final data augmentation policy based on the quality measures of the determined data augmentation policies.