Learning Data Quality Scoring for NLP Conversion Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In natural language processing, machine learning techniques used for discrete series-to-discrete series conversion often incorporate incorrect data in the learning data, which adversely affects the accuracy of the conversion process.
Innovation Solution
A learning quality estimation device and method that utilize a forward direction learned model of a discrete series converter to calculate a quality score for pairs of learning data likely to include errors, thereby identifying and removing incorrect data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If machine learning techniques are used for discrete series-to-discrete series conversion, then processing efficiency and automation are improved, but incorrect data in the learning data adversely affects conversion accuracy
Solution Approach 1:
The patent applies preliminary action by estimating the quality of learning data before the machine learning process. A quality estimation unit calculates quality scores for candidate learning data using a learned model, and a selection unit selects high-quality data before it is used for training. This preliminary quality assessment prevents incorrect data from adversely affecting conversion accuracy while maintaining the efficiency benefits of automated machine learning processing.
2Quantity of substance
If learning data is mechanically collected to improve data availability, then quantity of learning data increases, but wrong data is often mixed in the learning data
Solution Approach 1:
The patent implements feedback by using a learned model to evaluate the quality of candidate learning data. The quality estimation unit calculates quality scores based on the output of the learned model when processing candidate learning data. This feedback mechanism allows the system to identify and select high-quality data from mechanically collected sources, ensuring data correctness while maintaining large quantities of learning data for training.
3Reliability
If quality score calculation is performed for all candidate learning data to identify errors, then data quality improves, but processing time and computational resources increase
Solution Approach 1:
The patent applies parameter changes by adjusting the quality score threshold parameter to balance data quality and processing time. The selection unit compares quality scores against a threshold to determine which candidate learning data to select. By optimizing this threshold parameter, the system can achieve high data quality while minimizing the number of quality score calculations required, thus reducing processing time and computational resource consumption.
Data Source
AI summary
This disclosure relates to a device, a method, and a program capable of removing erroneous data from learning data used for machine learning used in natural language processing, for example. The method includes storing a forward direction learned model of a discrete series converter. The model is trained based on a plurality of pairs of discrete series of texts. Each pair comprises a first discrete series indicates an input of discrete series. A second discrete series indicates an output of discrete series. The first discrete series and the second discrete series are correctly associated. The method further includes converting the first discrete series to the second discrete series, and generating a quality score using the forward direction learned model, using a second learning pair of discrete series texts including an error in relationship.


