Pre-training Data Search for Machine Learning Model Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional natural language processing models face challenges in improving performance due to the difficulty in preparing large amounts of specific field data for pre-training, which is essential for enhancing model accuracy and efficiency.

Innovation Solution

An information processing apparatus that receives pre-training data and search conditions to automatically search for similar data, performing pre-training using the retrieved data to generate a trained model, thereby compensating for the lack of specific domain data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If pre-training is performed using general-purpose data, then model performance is improved, but the model cannot achieve high accuracy in specific fields

Engineering Contradiction:
Improvemodel performanceVSAvoidspecific field accuracy
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary action by automatically searching for and acquiring similar pre-training data before the actual training process. The search unit retrieves data similar to the provided sample data, and the acquisition unit downloads this pre-training data from external sources, preparing everything in advance so that when training begins, the model already has access to both general-purpose and specific-field data for effective pre-training

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary mechanism (the search unit and acquisition unit) that bridges the gap between general-purpose data and specific-field requirements. This intermediary automatically finds and acquires similar pre-training data from external sources, mediating between the user's sample data and the necessary training data without requiring manual data collection

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If large amounts of specific field data are prepared for pre-training, then model accuracy in specific fields is improved, but data collection becomes difficult and time-consuming

Engineering Contradiction:
Improvespecific field accuracyVSAvoiddata collection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system implements self-service by enabling the search unit and acquisition unit to automatically search for, acquire, and prepare pre-training data without requiring manual intervention. The system serves itself by automatically finding similar data from external sources and downloading it, eliminating the time-consuming manual data collection process while still obtaining large amounts of specific-field data

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary action by automatically searching for and acquiring similar pre-training data before the training process begins. The search unit retrieves data similar to the sample data provided, and the acquisition unit downloads this data in advance, so that when training starts, all necessary data is already prepared and available

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If manual data collection is performed, then data quality can be controlled, but the process is complex and requires significant user effort

Engineering Contradiction:
Improvedata qualityVSAvoiddata collection process
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system implements self-service by enabling the search unit and acquisition unit to automatically search for, acquire, and prepare pre-training data without requiring manual intervention. The system serves itself by automatically finding similar data from external sources and downloading it, eliminating the time-consuming manual data collection process while still obtaining large amounts of specific-field data

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system introduces an intermediary mechanism (the search unit and acquisition unit) that bridges the gap between general-purpose data and specific-field requirements. This intermediary automatically finds and acquires similar pre-training data from external sources, mediating between the user's sample data and the necessary training data without requiring manual data collection

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11928142B2Information processing apparatus and information processing method
Publication Date: 2024.03.12 SONY GROUP CORP
  • US11928142B2 patent drawing
  • US11928142B2 patent drawing
  • US11928142B2 patent drawing

AI summary

An information processing apparatus according to the present disclosure includes a reception unit that receives pre-training data that is data used for pre-training in machine learning, and a search condition for similar pre-training data that is data similar to the pre-training data, a search unit that searches for similar pre-training data in accordance with the search condition, and a generation unit that performs pre-training based on the retrieved similar pre-training data, and generates a trained model by using a result obtained through the pre-training.