Virtual Replicas for Context-Aware Machine Learning Data Labeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine learning technologies face challenges in obtaining and processing large datasets without excessive human involvement, and they often rely on fragmented data sources, which limits decision-making to contextual information, hindering holistic inference and optimization of real-world systems.

Innovation Solution

A system and method that provide machine learning algorithms with multi-source, real-time, and context-aware real-world data by using a server computer system connected to sensory mechanisms that capture data from various sources, including contextual information, and integrate this data into a persistent virtual world system for training and inference.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If large datasets are used for machine learning training, then the accuracy and reliability of machine learning models is improved, but the time and human involvement required for data labeling increases significantly

Engineering Contradiction:
Improveaccuracy of machine learning modelsVSAvoidtime required for data labeling
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system enables machines to label data automatically by processing sensory data from virtual replicas and generating labeled datasets without human intervention. The machine learning system trains on data from virtual replicas and automatically labels similar real-world data, making the system self-sufficient in data preparation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent creates virtual replicas of real-world entities that capture and store sensory data. These virtual copies serve as training data that can be processed and labeled automatically, replacing the need for humans to manually label extensive datasets from real-world sources.

Inventive Principle:
Principle #26Copying

2Speed

If data is collected from limited sources, then the processing speed and reaction time is improved, but the comprehensiveness and contextual awareness of decision-making deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidcontextual awareness
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The system collects sensory data from multiple diverse sources including cameras, microphones, temperature sensors, and other devices across different locations. This multi-source data collection enables the system to maintain fast processing while achieving comprehensive contextual awareness through aggregation of information from various perspectives and environments.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent transitions from two-dimensional data collection to three-dimensional spatial data gathering by deploying sensors at multiple locations and heights. This spatial dimensionality enables the system to capture comprehensive environmental context while maintaining processing efficiency through structured data organization in the virtual replica environment.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Ease of manufacture

If virtual models include only shape data, then the simplicity and ease of creation is improved, but the capability for management and optimization of real-world entities deteriorates

Engineering Contradiction:
Improveease of creating virtual modelsVSAvoidcapability for optimization
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The system pre-configures virtual replicas with comprehensive data structures before real-world optimization is needed. Virtual models are prepared in advance with sensory data from multiple sources and contextual information, enabling rapid deployment for optimization tasks without requiring complex real-time processing during actual optimization operations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The virtual replica serves as an intermediary between simple shape data and complex optimization requirements. It transforms basic geometric information into enriched datasets by automatically collecting and integrating sensory data from multiple sources, bridging the gap between model simplicity and optimization capability.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Measurement precision

If human labeling is used for training data, then the accuracy of labeled data is improved, but the cost and time consumption increases significantly

Engineering Contradiction:
Improveaccuracy of data labelingVSAvoidresource consumption
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The machine learning system performs self-labeling by processing sensory data from virtual replicas and automatically generating labeled datasets. The system uses its own trained models to label data without external human resources, eliminating the need for paid annotators while maintaining consistent labeling quality through algorithmic processing.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of human labeling with automated computational processes. Instead of human eyes and brains processing and labeling data, the system uses computer vision algorithms, sensor processing, and machine learning models to automatically generate accurate labels for training data.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20250068981A1Virtual intelligence and optimization through multi-source, real-time, and context-aware real-world data
Publication Date: 2025.02.27 THE CALANY HLDG S.ÀR L
  • US20250068981A1 patent drawing
  • US20250068981A1 patent drawing
  • US20250068981A1 patent drawing

AI summary

A system for managing and optimizing real world entities with machine learning algorithms are described including a server computer system configured to store and process input data, the server computer system comprising a memory and a processor; and wherein the memory of the server computer system stores a persistent virtual world system comprising virtual replicas of the real world entities, and wherein the server computer system is configured to generate explicit data sets representing functioning and behavior of the real world entities; train the machine learning algorithms with the explicit data sets to generate trained machine learning data sets; and apply the trained machine learning data sets in an artificial intelligence application to manage operation of the real world entities and generate real behavior data of the real world entities during the operation of the real world entities. Corresponding methods are also described.