Text Quality Assessment Model Training via Automated Labeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for training text quality assessment models rely heavily on manual annotation samples, which are time-consuming and inefficient, especially when dealing with large volumes of text data.

Innovation Solution

A method is proposed that involves determining negative and positive text samples based on indicators such as views, likes, and dislikes, and then using these labeled samples to train a text quality assessment model.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation samples are used to train the classification model, then the model can be trained to predict text quality, but the process is time-consuming and inefficient

Engineering Contradiction:
Improvetext quality assessment accuracyVSAvoidtraining data preparation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system automatically generates training data by having the classification model evaluate texts against quality indicators, eliminating the need for manual annotation. The model serves itself by generating the labeled training data it needs through automated quality assessment processes.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the parameter of data labeling from manual human annotation to automated model-based labeling. By using the classification model to automatically assign quality labels to texts based on multiple indicators, the system transforms the time-consuming manual process into an efficient automated process while maintaining assessment accuracy.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If manual annotation is used for training data, then text quality can be assessed, but the efficiency is low when dealing with large volumes of text data

Engineering Contradiction:
Improvetext quality assessment reliabilityVSAvoidtext processing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The classification model automatically generates training data by evaluating texts against quality indicators, enabling the system to handle large volumes of text data efficiently without sacrificing reliability. The automated process maintains consistent assessment standards across all texts.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of manual human annotation with an automated computational system. The classification model uses algorithmic processes to evaluate texts and generate training labels, substituting human labor with automated mechanical computation that can process large volumes of data efficiently while maintaining reliability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If automatic sample generation is used, then training efficiency is improved, but the balance of sample data proportions for each category may be affected

Engineering Contradiction:
Improvetraining data generation efficiencyVSAvoidsample data distribution balance
Core Design Contradiction:
ProductivityVSStability of the object's composition

Solution Approach 1:

The system incorporates feedback mechanisms where the classification model's assessments are continuously refined based on the generated training data. The automated process monitors and adjusts the distribution of training samples across different quality categories to maintain balanced proportions, using feedback from assessment results to improve future sampling.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent adjusts parameters of the automated sampling process to control the distribution of training samples. By modifying the quality indicator thresholds and selection criteria, the system maintains balanced category proportions in the training data while preserving high generation efficiency through automated processes.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12283124B2Method of training text quality assessment model and method of determining text quality
Publication Date: 2025.04.22 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US12283124B2 patent drawing
  • US12283124B2 patent drawing
  • US12283124B2 patent drawing

AI summary

A method of training a text quality assessment model, a method of determining text quality, an electronic device, and a storage medium are provided. The method of training the text quality assessment model includes: determining a first text satisfying a condition of being a negative sample and a second text satisfying a condition of being a positive sample from a plurality of texts based on indicators for the texts; for any text of the first text and the second text, adding a label to the text based on the condition satisfied by the text, wherein the label indicates a category of the text, and the category includes a low-quality category for the negative sample and a non-low-quality category for the positive sample; and constituting a training set by the first text having a label and the second text having a label, to train the text quality assessment model.