Text Quality Assessment Model Training via Automated Labeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for training text quality assessment models rely heavily on manual annotation samples, which are time-consuming and inefficient, especially when dealing with large volumes of text data.
Innovation Solution
A method is proposed that involves determining negative and positive text samples based on indicators such as views, likes, and dislikes, and then using these labeled samples to train a text quality assessment model.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual annotation samples are used to train the classification model, then the model can be trained to predict text quality, but the process is time-consuming and inefficient
Solution Approach 1:
The system automatically generates training data by having the classification model evaluate texts against quality indicators, eliminating the need for manual annotation. The model serves itself by generating the labeled training data it needs through automated quality assessment processes.
Solution Approach 2:
The patent changes the parameter of data labeling from manual human annotation to automated model-based labeling. By using the classification model to automatically assign quality labels to texts based on multiple indicators, the system transforms the time-consuming manual process into an efficient automated process while maintaining assessment accuracy.
2Reliability
If manual annotation is used for training data, then text quality can be assessed, but the efficiency is low when dealing with large volumes of text data
Solution Approach 1:
The classification model automatically generates training data by evaluating texts against quality indicators, enabling the system to handle large volumes of text data efficiently without sacrificing reliability. The automated process maintains consistent assessment standards across all texts.
Solution Approach 2:
The patent replaces the mechanical process of manual human annotation with an automated computational system. The classification model uses algorithmic processes to evaluate texts and generate training labels, substituting human labor with automated mechanical computation that can process large volumes of data efficiently while maintaining reliability.
3Productivity
If automatic sample generation is used, then training efficiency is improved, but the balance of sample data proportions for each category may be affected
Solution Approach 1:
The system incorporates feedback mechanisms where the classification model's assessments are continuously refined based on the generated training data. The automated process monitors and adjusts the distribution of training samples across different quality categories to maintain balanced proportions, using feedback from assessment results to improve future sampling.
Solution Approach 2:
The patent adjusts parameters of the automated sampling process to control the distribution of training samples. By modifying the quality indicator thresholds and selection criteria, the system maintains balanced category proportions in the training data while preserving high generation efficiency through automated processes.
Data Source
AI summary
A method of training a text quality assessment model, a method of determining text quality, an electronic device, and a storage medium are provided. The method of training the text quality assessment model includes: determining a first text satisfying a condition of being a negative sample and a second text satisfying a condition of being a positive sample from a plurality of texts based on indicators for the texts; for any text of the first text and the second text, adding a label to the text based on the condition satisfied by the text, wherein the label indicates a category of the text, and the category includes a low-quality category for the negative sample and a non-low-quality category for the positive sample; and constituting a training set by the first text having a label and the second text having a label, to train the text quality assessment model.


