Synthetic AI Training Datasets from Automated Screenshot Capture

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge of acquiring sufficient data for training and validating AI-based software models that predict human-computer interactions is hindered by privacy concerns and the lack of available training data, limiting their accuracy and effectiveness.

Innovation Solution

An automated screenshot capture engine mimics real-world human-computer interactions to generate large volumes of annotated training data, including metadata, which is used to train or validate machine learning models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If real-world human-computer interaction data is used for training, then model accuracy is improved, but user privacy and copyright issues arise

Engineering Contradiction:
Improvemodel accuracyVSAvoiduser privacy concerns
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic copies of human-computer interaction data through automated screenshot capture and annotation. Instead of using real user data, the system generates artificial screenshots that mimic real interaction patterns, thereby preserving model training quality while eliminating privacy and copyright concerns associated with actual user data.

Inventive Principle:
Principle #26Copying

2Quantity of substance

If sufficient training data is obtained through manual collection, then model effectiveness is improved, but time and cost increase significantly

Engineering Contradiction:
Improvetraining data volumeVSAvoiddata collection time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The system performs self-service data generation by automatically capturing screenshots, extracting metadata, and annotating images without human intervention. The automated annotation process uses AI models to generate captions and descriptions, eliminating the need for manual data collection and annotation while producing large volumes of training data efficiently.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary actions by pre-configuring automated screenshot capture workflows and annotation pipelines before actual model training begins. This preparation enables rapid generation of training data without requiring time-consuming manual collection during the model development phase.

Inventive Principle:
Principle #10Preliminary action

3Ease of manufacture

If automated screenshot capture is used to generate training data, then data collection cost is reduced, but data annotation complexity increases

Engineering Contradiction:
Improvedata collection costVSAvoidannotation system complexity
Core Design Contradiction:
Ease of manufactureVSDevice complexity

Solution Approach 1:

The patent replaces manual mechanical annotation processes with automated AI-based annotation systems. Instead of human annotators manually labeling screenshots, the system uses machine learning models to automatically generate captions, extract metadata, and annotate images, thereby reducing costs while managing complexity through automation rather than manual processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20250218206A1Ai-generated datasets for ai model training and validation
Publication Date: 2025.07.03 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20250218206A1 patent drawing
  • US20250218206A1 patent drawing
  • US20250218206A1 patent drawing

AI summary

Disclosed are techniques for synthesizing large amounts of human-computer interaction data that is representative of real-world user data. An automated screenshot capture engine may cause an automated agent to use an application or a website in a manner designed to mimic real-world human-computer interaction. Screenshots are captured to record how a user might interact with the application. Metadata, such as window location and size, may be obtained for each screenshot. Screenshots and corresponding metadata may be automatically annotated with a large language model to indicate the context of the application and/or computer system when the screenshot was captured. Data created in this way may be used to validate AI-based software application features or to train (or retrain) a machine learning model that predicts human-computer interactions. Automated synthesis of training data significantly increases the scale of data that can be obtained for training while also reducing computing and financial costs.