Learning Data Quality Scoring for NLP Conversion Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In natural language processing, machine learning techniques used for discrete series-to-discrete series conversion often incorporate incorrect data in the learning data, which adversely affects the accuracy of the conversion process.

Innovation Solution

A learning quality estimation device and method that utilize a forward direction learned model of a discrete series converter to calculate a quality score for pairs of learning data likely to include errors, thereby identifying and removing incorrect data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If machine learning techniques are used for discrete series-to-discrete series conversion, then processing efficiency and automation are improved, but incorrect data in the learning data adversely affects conversion accuracy

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidconversion accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies preliminary action by estimating the quality of learning data before the machine learning process. A quality estimation unit calculates quality scores for candidate learning data using a learned model, and a selection unit selects high-quality data before it is used for training. This preliminary quality assessment prevents incorrect data from adversely affecting conversion accuracy while maintaining the efficiency benefits of automated machine learning processing.

Inventive Principle:
Principle #10Preliminary action

2Quantity of substance

If learning data is mechanically collected to improve data availability, then quantity of learning data increases, but wrong data is often mixed in the learning data

Engineering Contradiction:
Improvequantity of learning dataVSAvoiddata correctness
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent implements feedback by using a learned model to evaluate the quality of candidate learning data. The quality estimation unit calculates quality scores based on the output of the learned model when processing candidate learning data. This feedback mechanism allows the system to identify and select high-quality data from mechanically collected sources, ensuring data correctness while maintaining large quantities of learning data for training.

Inventive Principle:
Principle #23Feedback

3Reliability

If quality score calculation is performed for all candidate learning data to identify errors, then data quality improves, but processing time and computational resources increase

Engineering Contradiction:
Improvedata qualityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies parameter changes by adjusting the quality score threshold parameter to balance data quality and processing time. The selection unit compares quality scores against a threshold to determine which candidate learning data to select. By optimizing this threshold parameter, the system can achieve high data quality while minimizing the number of quality score calculations required, thus reducing processing time and computational resource consumption.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12271410B2Learning quality estimation device, method, and program
Publication Date: 2025.04.08 NIPPON TELEGRAPH & TELEPHONE CORP
  • US12271410B2 patent drawing
  • US12271410B2 patent drawing
  • US12271410B2 patent drawing

AI summary

This disclosure relates to a device, a method, and a program capable of removing erroneous data from learning data used for machine learning used in natural language processing, for example. The method includes storing a forward direction learned model of a discrete series converter. The model is trained based on a plurality of pairs of discrete series of texts. Each pair comprises a first discrete series indicates an input of discrete series. A second discrete series indicates an output of discrete series. The first discrete series and the second discrete series are correctly associated. The method further includes converting the first discrete series to the second discrete series, and generating a quality score using the forward direction learned model, using a second learning pair of discrete series texts including an error in relationship.