Intelligent labeling and model training reasoning system for long text multi-dimensional evaluation

By designing an intelligent annotation and model training inference system, using feature extraction and intelligent estimation modules, the subjectivity in the long text annotation process is reduced, the objectivity and efficiency of the annotation results are improved, and the problem of unstable annotation results in the existing technology is solved.

CN120124616APending Publication Date: 2025-06-10AISPEECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510115741.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the prior art, the multi-dimensional annotation process of long texts heavily relies on the subjective judgment of the labeler, resulting in unstable quality and poor consistency of the annotation results, which affects the generalization ability and prediction accuracy of subsequent model training.

Method used

An intelligent annotation and model training inference system for multi-dimensional evaluation of long texts is designed, including feature extraction module, intelligent scoring module, manual annotation module, model prediction module and annotation processing module. Through automated feature extraction and score estimation, the subjectivity is reduced and the objectivity of the annotation results is improved.

Benefits of technology

It significantly reduces the problem of labeling subjectivity and inconsistency, improves the labeling efficiency and objectivity of the results, provides support for the rapid reflow of long text, data cleaning and incremental training, and improves the generalization ability and prediction accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124616A_ABST
    Figure CN120124616A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of text labeling, in particular to an intelligent labeling and model training inference system oriented to long text multi-dimensional evaluation, which realizes comprehensive and objective evaluation of long text quality by designing four general labeling dimensions (theme, structure, grammar and readable) and corresponding feature sets thereof. The system comprises a feature extraction module, an intelligent score estimation module, a manual annotation module, a model prediction module and an annotation processing module, feature extraction and score estimation can be efficiently completed, and the accuracy and stability of an annotation result are improved through manual correction. Compared with an existing labeling method depending on subjective judgment, the method has the advantages that the problems of labeling subjectivity and inconsistency are remarkably reduced, labeling efficiency and result objectivity are improved, support is provided for rapid backflow, data cleaning and incremental training of long texts, and the method is suitable for large-scale popularization and application. Therefore, the technical problem that quality is unstable due to the fact that the labeling process depends on subjective judgment is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of text annotation, and particularly to an intelligent annotation and model training and inference system for multi-dimensional evaluation of long texts. Background Art

[0002] With the rapid development of natural language processing technology, automatic essay scoring has gradually become an important research direction in the field of education. English writing occupies an important position in various exams. Through writing, students can not only exercise their language expression, theme argumentation, and grammar application abilities, but also promote the improvement of their comprehensive language abilities. To overcome the subjectivity, lag, and high labor cost problems of manual scoring, automatic essay scoring technology based on natural language processing has been widely studied. In recent years, automatic essay scoring has gradually developed from a single overall quality assessment to a multi-dimensional scoring direction, providing a more comprehensive perspective for writing evaluation and feedback. In addition, in addition to essay evaluation, the processing and analysis of long texts is also an important and challenging field. Similar to essay texts, long texts also face similar problems and challenges during the evaluation process, especially in how to comprehensively and accurately evaluate their quality and various dimensions.

[0003] In the prior art, a multi-dimensional annotation system provides score annotations for long texts in multiple dimensions (such as theme and structure) by designing different scoring criteria and annotation manuals. Specifically, the system first requires annotators to perform multi-dimensional scoring on long texts based on subjective judgment. Long texts of different themes need to be annotated for different dimensions. To ensure the reliability of the annotation results, each long text needs to be annotated by at least two annotators. When the scores of the two annotators differ significantly, a third-party annotator performs the final annotation. In addition, the multi-dimensional annotation system supports exporting the annotation data for the training and score prediction tasks of neural network models.

[0004] However, the main problem of the prior art is that the annotation process highly depends on the subjective judgment of annotators. Due to the different experiences and understandings of annotators, even for the annotation of the same text, the results may vary significantly. This subjectivity directly affects the quality and consistency of the annotation results, making the subsequent model training possibly based on unstable annotation data, thereby reducing the generalization ability and prediction accuracy of the model. Summary of the Invention

[0005] This application provides an intelligent annotation and model training and inference system for multi-dimensional evaluation of long texts, which can significantly reduce the problems of annotation subjectivity and inconsistency, improve the annotation efficiency and objectivity of the results, and provide support for the rapid feedback, data cleaning, and incremental training of long texts, thus solving the technical problem of unstable quality caused by the dependence on subjective judgment in the annotation process. This application provides the following technical solutions:

[0006] In a first aspect, the present application provides an intelligent annotation and model training and inference method for multi-dimensional evaluation of long texts. The system includes a feature extraction module, an intelligent scoring module, a manual annotation module, a model prediction module, and an annotation processing module;

[0007] The feature extraction module is used to parse the original long text content, perform feature adaptation according to business requirements and long text types, and extract features corresponding to corresponding dimensions;

[0008] The intelligent scoring module is used to estimate scores based on the feature set extracted by the feature extraction module;

[0009] The manual annotation module is used to correct the feature set and the estimated score based on the feature set extracted by the feature extraction module and the score calculated by the intelligent scoring module, and finally submit the annotation result;

[0010] The single-dimension annotation module consists of a feature extraction module, an intelligent scoring module, and a manual annotation module, and is used to be responsible for the score annotation of a single dimension;

[0011] The model prediction module is used to generate automatically by the feature extraction module / manually correct the feature set by the manual annotation module and the long text content, and the model performs score prediction to provide additional reference for manual annotation;

[0012] The annotation processing module is used to automatically calculate the total score corresponding to the long text based on the theme, structure, grammar, and readable score annotation result, and store the multi-dimensional scores, multi-dimensional features, and long text total score as the annotation result of the current long text.

[0013] In a specific feasible implementation, the feature extraction module includes three steps: text parsing, business adaptation, and feature extraction;

[0014] The text parsing uses the latest fastopic tool to parse the theme distribution of the long text, and uses the Chatgpt tool to extract the set of topic sentences in the long text;

[0015] The business adaptation selects a feature set according to the business requirements corresponding to the long text, and sets the feature set as a set of topic words, a set of topic sentences, the number of uses of topic words, the number of uses of topic sentences, the long text theme distribution vector, and the topic similarity;

[0016] The feature extraction is based on the theme distribution and the set of topic words, and extracts features such as the set of topic words in the best theme distribution, the occurrence of topic words in the composition long text, and the overall theme distribution; based on the set of topic sentences, extracts the text of each topic sentence, the number of topic sentences, and the topic word situation in each topic sentence; combines the features into a feature set of a theme dimension and transmits it to the manual annotation module and the intelligent scoring module.

[0017] In a specific feasible implementation, the intelligent score estimation module processes the feature results of the feature extraction module or the corrected features of the manual annotation module, including two steps: feature parsing and score calculation;

[0018] The feature parsing operation filters out the corresponding features for estimating scores;

[0019] The score calculation automatically calculates the score based on the selected topic features, as shown in the following formula:

[0020]

[0021] where alpha is a penalty factor used to penalize composition inputs with too short text lengths; the TS function estimates the topic score of the long text based on the set of topic words, the set of topic sentences, and the sentences in the long composition text; m is the number of sentences in the long composition text;

[0022]

[0023] where t is a preset threshold.

[0024] In a specific feasible implementation, the manual annotation module performs annotation based on the additional information provided by the feature extraction model and the intelligent score estimation module, including:

[0025] Manually correct the relevant features obtained by the feature extraction module;

[0026] Whenever the content of the feature set in the system is manually corrected, the intelligent score estimation module updates the calculated estimated score in real time, and the annotator submits the corrected estimated score as the quality score of the long text in the topic dimension; if the score given by the intelligent score estimation module is correct, the annotator does not need to correct it;

[0027] When the feature set and the annotation score are corrected, the annotator submits them, integrates the two annotation results, and transmits them to the system background.

[0028] In a specific feasible implementation, the annotation processing module performs post-processing based on the multi-dimensional annotation results obtained by the single-dimensional annotation module. The feature set corresponding to the current long text is post-processed and integrated, and the annotation scores of four different dimensions are automatically calculated as the total score of the long text, as shown in the following formula:

[0029]

[0030] Obtain the annotation result format as a json file.

[0031] In a specific feasible implementation, the intelligent annotation system for multi-dimensional evaluation of long texts further includes:

[0032] Perform business adaptation, feature extraction, and score estimation based on the input long text to obtain the long text, the corresponding feature set of the long text, and the estimated score calculated based on the feature set;

[0033] If there are no problems with the automated feature set obtained by the feature extraction module and the estimated score calculated by the intelligent score estimation module, the annotator directly submits the automated result as the annotation result; otherwise, the annotator needs to correct the problems in the feature set to obtain the updated estimated score;

[0034] When the feature set has been corrected to the correct state, the annotator can directly adopt the estimated score as the annotation score; can also manually modify a more accurate score as the annotation score; can also use the model prediction module provided by the model training and inference system to perform model inference;

[0035] After the annotator submits the annotation results in four dimensions, the overall total score of the long text is automatically calculated; the long text, multi-dimensional scores, multi-dimensional features, and overall total score are all stored in the database as the complete annotation result of the current long text.

[0036] In a second aspect, the present application provides a model training and inference system for multi-dimensional evaluation of long texts, which is applied to the intelligent annotation system for multi-dimensional evaluation of long texts as described in the first aspect, and adopts the following technical solutions:

[0037] A model training and inference system for multi-dimensional evaluation of long texts includes:

[0038] Obtain long text multi-dimensional annotation data through the intelligent annotation system for multi-dimensional evaluation of long texts, use the data for model training and testing, and adopt five-fold cross-validation to test the model performance;

[0039] Train the model based on the manually specified model configuration and parameter settings, so that the model learns the multi-dimensional scoring task of long texts;

[0040] Export the model obtained by the training module to onnx and perform int8 quantization, and use the model for five-fold cross-validation to obtain the preliminary model performance results;

[0041] The exported model can automatically predict the score of the current long text based on the feature set automatically generated by the feature extraction module / manually corrected by the manual annotation module and the long text content.

[0042] In a specific feasible implementation, training the model based on the manually specified model configuration and parameter settings such that the model learns the multi-dimensional scoring task of long texts includes:

[0043] Support two training methods: Support the long text multi-dimensional evaluation method based on the pre-trained model BigBird and the long text multi-dimensional evaluation method based on the large language model llama.

[0044] In a specific feasible implementation, the long text multi-dimensional evaluation method based on the pre-trained model BigBird includes:

[0045] Taking the long text content and feature set as inputs, and taking the multi-dimensional scores and total scores of the long text as labels;

[0046] The model simultaneously learns the score prediction of the composition long text in four different dimensions and the score prediction of the overall long text;

[0047] Avoid overfitting problems through the R-drop mechanism.

[0048] In a specific feasible implementation, the long text multi-dimensional evaluation method based on the large language model llama includes:

[0049] Processing the input data and combining it with the fine-tuning template as the input of the model, and combining the multi-dimensional scores and total scores with the label model as the training labels;

[0050] Instructing the fine-tuning of the llama model to learn the long text multi-dimensional scoring task;

[0051] Using the lora mechanism to reduce the resource overhead required for model training, where the fine-tuning template and label template are defined by themselves in the parameter import step.

[0052] In summary, the beneficial effects of this application at least include:

[0053] 1) Designed a more general multi-dimensional scoring standard, which can effectively improve the utilization efficiency of long text multi-dimensional annotation results and the annotation efficiency of manual annotation;

[0054] 2) The designed intelligent annotation system helps to better and faster process the reflux long text data, helps to achieve the rapid reflux, data cleaning, and incremental training of long text data, and thus improves the relevant business efficiency;

[0055] 3) Additionally designed a model training and inference system for long text multi-dimensional evaluation, which serves the annotation system and can provide more convenient experimental verification and annotation references for the annotated data.

[0056] By designing four general annotation dimensions (theme, structure, grammar, readability) and their corresponding feature sets, a comprehensive and objective evaluation of the quality of long texts is achieved. The system includes modules for feature extraction, intelligent scoring, manual annotation, model prediction, and annotation processing, which can efficiently complete feature extraction and score estimation, and improve the accuracy and stability of annotation results through manual correction. Compared with existing annotation methods that rely on subjective judgment, this application significantly reduces the problems of annotation subjectivity and inconsistency, improves annotation efficiency and the objectivity of results, provides support for the rapid feedback, data cleaning, and incremental training of long texts, thus solving the technical problem of unstable quality caused by the dependence on subjective judgment in the annotation process.

[0057] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and implement it according to the content of the specification, the following describes it in detail with preferred embodiments of this application and accompanying drawings. Brief Description of the Drawings

[0058] Figure 1 It is a schematic flowchart of an intelligent annotation system for multi-dimensional evaluation of long texts in an embodiment of this application.

[0059] Figure 2 It is a schematic flowchart of a model training and inference system for multi-dimensional evaluation of long texts in an embodiment of this application. Detailed Description of the Embodiments

[0060] The following combines the drawings and embodiments to further describe the detailed implementation of this application in detail. The following embodiments are used to illustrate this application, but are not used to limit the scope of this application.

[0061] To solve the technical problems in the background art, this application first provides an intelligent annotation system for multi-dimensional evaluation of long texts. It should be noted that, therefore, this application extends the composition object to long text objects and designs a more general multi-dimensional standard for long texts: (1) theme; (2) structure; (3) grammar; (4) readability.

[0062] Among them, the range of the theme score is designed to be from 1 to 10 points. The higher the score, the better the text quality of the composition in the theme dimension. For long text compositions, the theme feature set designed by this application includes: theme distribution, theme word set, the number of times theme words are used in the composition, theme sentence set, the number of times theme sentences are used in the composition, and the semantic similarity between the composition text and the writing prompt text.

[0063] The range of the structure score is designed to be from 1 to 10 points. The higher the score, the better the text quality of the composition in terms of structure dimension. For long text compositions, the structure feature set designed in this application includes: the segmentation results of discourse basic units, rhetorical structure trees, rhetorical relations of rhetorical structure trees, nuclearity types of rhetorical structure trees, dependency structure trees, and syntactic structure trees.

[0064] The range of the grammar score is designed to be from 1 to 10 points. The higher the score, the better the text quality of the composition in terms of grammar dimension. For long text compositions, the grammar feature set designed in this application includes: the grammar error detection results, which contain 9 different types of grammar errors (misspelling, typographical, whitespace, grammar, style, locale - violation, duplication, uncategorized, inconsistency).

[0065] The range of the readability score is designed to be from 1 to 10 points. The higher the score, the better the text quality of the composition in terms of readability dimension. For long text compositions, the readability feature set designed in this application includes the following readability metrics: Coleman–Liau Index, Flesch Reading Ease, Gunning Fog Index, Simple Measure of Gobbledygook, and Linsear Write Formula. These readability metrics are used in the Scholastic Assessment Test (SAT) commissioned by the College Board of the United States and have thus been proven to be able to reflect the readability level of the text.

[0066] Combining previous experimental studies and the above - mentioned scoring criteria, this application further designs the corresponding feature sets, and the content of these feature sets can effectively reflect the quality level of long texts in the corresponding dimensions. Since manual annotation is subjective and unstable, this may increase the cost and time overhead of data annotation. For this reason, this application further designs an intelligent estimation method based on the feature sets. This method gives a predicted score according to the content of the feature sets, which can effectively alleviate the problem of unstable annotation caused by subjectivity and also greatly improve the efficiency of manual annotation.

[0067] In addition to the above four general multi - dimensional scores, this application automatically calculates the total score of the composition, which is the sum of the four multi - dimensional scores. Therefore, the annotation results after annotation by this application include the following content: multi - dimensional annotation scores, multi - dimensional feature sets, and the total score of the composition.

[0068] The object of this application extends the intelligent annotation system to long texts, not limited to composition long texts. This is because different types of long texts have commonalities in the four annotation dimensions designed in this application, which helps the intelligent annotation system designed in this application to be applied to a wider range of scenarios. Through the business adaptation process, this application helps to design different feature sets for different types of long texts, and also supports managers to define the feature sets supported by the system by themselves. For the convenience of description, the following will continue to use composition long texts as typical examples for introduction.

[0069] Referring to Figure 1 , which is a schematic flowchart of an intelligent annotation system for multi-dimensional evaluation of long texts provided by an embodiment of this application. The system includes five modules: feature extraction, intelligent scoring, manual annotation, model prediction, and annotation processing module. Among them, feature extraction, intelligent scoring, and manual annotation together constitute a single-dimensional annotation module. The single-dimensional annotation module is compatible with the differences of four scoring criteria and provides similar annotation services for annotators.

[0070] The feature extraction module is used to parse the content of the original long text, perform feature adaptation according to business requirements and long text types, and extract features corresponding to the corresponding dimension;

[0071] The intelligent scoring module is used to estimate scores based on the feature set extracted by the feature extraction module;

[0072] The manual annotation module is used to correct the feature set and the estimated score based on the feature set extracted by the feature extraction module and the score calculated by the intelligent scoring module, and finally submit the annotation result;

[0073] The single-dimensional annotation module, composed of the feature extraction module, the intelligent scoring module, and the manual annotation module, is responsible for the score annotation of a single dimension; there are certain differences between the four general annotation dimensions (theme, structure, grammar, readability) designed in this system. Therefore, the single-dimensional annotation module can not only support the annotation behaviors of each dimension, but also be compatible with the annotation differences between each dimension.

[0074] The model prediction module is used to generate automatically by the feature extraction module / manually correct the feature set and the long text content by the manual annotation module, and the model performs score prediction to provide additional reference for manual annotation.

[0075] The annotation processing module: is used to automatically calculate the total score corresponding to the long text based on the score annotation results of the four dimensions, and store the multi-dimensional scores, multi-dimensional features, and long text total scores as the annotation results of the current long text.

[0076] As an implementation mode of this application, the feature extraction module includes three steps: text parsing, business adaptation, and feature extraction. Given that different tools are required for different dimensions and different features are needed for different businesses, the following will take the topic dimension annotation of long composition texts as an example for introduction.

[0077] Text parsing: The latest fastopic tool is used to parse the topic distribution of long texts, and the Chatgpt tool is used to extract the set of topic sentences in long texts. For example, if the long composition text is “Dance connects us to society and culture in many universal and personal ways. It deepens our understanding of the world and ourselves. Synthesising personal knowledge and experiences with dance movements reinforces us to perceive the feelings and ideas evoked in a dance form…” and its writing topic is “dance”. Then, sentences such as “Dance connects us to society and culture in many universal and personal ways.” will be parsed as the topic sentences of the long composition text, and the best topic in the topic distribution of the long composition text will be composed of topic words such as “dance” and “scoiety”.

[0078] Business adaptation: Select the feature set according to the business requirements corresponding to the long text. The current input long text is an English composition, so the feature set is set to the set of topic words, the set of topic sentences, the number of uses of topic words, the number of uses of topic sentences, the topic distribution vector of the long text, topic similarity, etc. In addition, business adaptation also supports managers to customize the content of the feature set, but these features need to be within the scope supported by the system.

[0079] Feature extraction: Based on the topic distribution and the set of topic words, extract features such as the set of topic words in the best topic distribution, the occurrence of these topic words in the long composition text, and the overall topic distribution; based on the set of topic sentences, extract features such as the text of each topic sentence, the number of topic sentences, and the topic word situation in each topic sentence. Finally, combine these features into a feature set of the topic dimension and pass it to the manual annotation module and the intelligent scoring module for their use.

[0080] As an implementation manner of the present application, the intelligent score estimation module processes the feature results of the feature extraction module or the corrected features of the manual annotation module, including two steps: feature parsing and score calculation.

[0081] Feature parsing: The feature extraction module / manual annotation module will generate many features related to the theme dimension, and the automatic score estimation method does not require all features. Therefore, the feature parsing operation will screen out the corresponding features for estimating scores, such as the number of topic words feature, the occurrence situation of topic words in the long composition text feature, the number of topic sentences feature, etc.

[0082] Score calculation: Based on the screened theme features, automatic score calculation is performed, as shown in the following formula.

[0083]

[0084] Among them, alpha is a penalty factor used to penalize the composition input with too short text length; the TS function estimates the theme score of the long text based on the topic word set, the topic sentence set, and the sentences in the long composition text; m is the number of sentences in the long composition text.

[0085]

[0086] Among them, t is a preset threshold.

[0087] As an implementation manner of the present application, the manual annotation module performs annotation based on the additional information provided by the feature extraction model and the intelligent score estimation module, including:

[0088] Feature modification annotation: Given that the automatically extracted relevant features may be incorrect, manual correction is performed on the relevant features obtained by the feature extraction module. For example, if chatgpt fails to recognize the topic sentence "Synthesising personal knowledge and experiences with dance movements reinforces us to perceive the feelings and ideas evoked in a dance form", the annotator needs to correct this; if the automatically extracted relevant features are not incorrect, the annotator does not need to correct them, which will improve the efficiency of manual annotation.

[0089] Final score annotation: Whenever the content of the feature set in the system is manually corrected, the intelligent scoring module will update the calculated score in real time. The annotator can finally submit the score obtained from the correction as the quality score of the long text in the theme dimension, or manually input the score to submit, so as to obtain the annotation score corresponding to the long text. If the score given by the intelligent scoring module is correct, the annotator does not need to correct it, thus improving the efficiency of manual annotation.

[0090] Submit two annotation results: When the feature set and the annotation score are corrected, the annotator submits them. This operation requires integrating the two annotation results and transmitting them to the system background. At this time, the feature set in the annotation result can effectively alleviate the subjectivity and instability problems of the annotation score, thereby improving the objectivity and stability of the annotation result. In addition, the feature set will serve as a basis for annotating the annotation score, thus improving the traceability and rationality of the annotation result.

[0091] As an implementation manner of the present application, the model prediction module performs automated prediction on the feature set automatically extracted / manually corrected in the manual annotation module and the composition long text. The prediction result can provide additional reference for manual annotation. It should be noted that the annotator can also choose not to use this module to provide additional reference, thereby improving the annotation efficiency.

[0092] As an implementation manner of the present application, the annotation processing module performs post-processing based on the multi-dimensional annotation results obtained by the single-dimensional annotation module. First, the feature set corresponding to the current long text is post-processed and integrated. Then, the annotation scores of four different dimensions are automatically calculated as the total score of the long text, as shown in the following formula:

[0093]

[0094] Finally, the obtained annotation result format is a json file, and the format is as follows: {"LongText":"Dance connects us to society and culture in many universal and…","LongTextS":30,"TopicF":{"TopicWord":["dance","culture","insterst",…],"TopicS":8,…}}.

[0095] In implementation, based on the effective cooperation among the above-mentioned numerous modules, the main annotation services provided by the entire intelligent annotation system for annotators and administrators include the following steps:

[0096] S1: The intelligent annotation system performs business adaptation, feature extraction, and score estimation on the input long text, and then obtains the long text, the corresponding feature set of the long text, and the estimated score calculated based on the feature set.

[0097] S2: If there are no problems with the automated feature set obtained by the feature extraction module and the estimated score calculated by the intelligent score estimation module, the annotator can directly submit the automated result as the annotation result; otherwise, the annotator needs to correct the problems in the feature set to obtain the updated estimated score.

[0098] S3: When the feature set has been corrected to the correct state, the annotator can directly adopt the estimated score as the annotation score; can also manually modify a more accurate score as the annotation score; can also use the model prediction module provided by the model training and inference system to perform model inference to obtain more reference information.

[0099] S4: After the annotator submits the annotation results in four dimensions, the intelligent annotation system will automatically calculate the overall total score of the long text; finally, the long text, multi-dimensional scores, multi-dimensional features, and overall total score will all be stored in the database as the complete annotation result of the current long text for subsequent business / research use.

[0100] In summary, the intelligent annotation system is the core system of this application and can provide more comprehensive and convenient annotation services for annotators. First, this application designs four more general annotation dimensions, which enables the annotation of different long texts without replacing specific other annotators. Second, the feature set designed in this application can more objectively reflect the quality level of the long text in each dimension, effectively alleviating the subjectivity problem existing in manual annotation. Third, the automatic extraction of features and the automatic calculation of estimated scores designed in this application can quickly initialize the long text. The annotator only needs to correct the feature set and the estimated score to complete the annotation task, which effectively improves the efficiency of manual annotation and also enhances the objectivity, stability, and rationality of the annotation results. These designs are not available in other similar works. Therefore, the intelligent annotation system of this application is faster and more intelligent than them. For this reason, the intelligent annotation system also helps to achieve the rapid return, data cleaning, and incremental training of long text data, thereby improving the relevant business efficiency.

[0101] The intelligent annotation system designed in this application has a feature extraction module and an intelligent score estimation module, which are not available in other systems. Therefore, the intelligent annotation system of this application supports fully automated rough annotation, that is, directly taking the automated feature set and estimated score as the annotation result. This fully automated rough annotation can help achieve the rapid cleaning of the long text return data, which also helps to improve the efficiency of relevant businesses.

[0102] Reference Figure 2 Figure 2 , this application also provides an intelligent annotation and model training and inference system for multi-dimensional evaluation of long texts. This system is applied to an intelligent annotation system for multi-dimensional evaluation of long texts and includes the following steps:

[0103] S1: Obtain long text multi-dimensional annotation data through an intelligent annotation system for multi-dimensional evaluation of long texts, use the data for model training and testing, and adopt five-fold cross-validation to test the model performance.

[0104] S2: Train the model based on manually specified model configurations and parameter settings so that the model can fully learn the multi-dimensional scoring task of long texts; it should be noted that the current training module supports the following two training methods: support for the multi-dimensional evaluation method of long texts based on the pre-trained model BigBird and the multi-dimensional evaluation method of long texts based on the large language model llama.

[0105] S3: Export the model obtained by the training module to onnx and perform int8 quantization, and use the model for five-fold cross-validation to obtain preliminary model performance results.

[0106] S4: The exported model will be able to further serve the intelligent annotation system, that is, the exported model can automatically predict the score of the current long text based on the feature set automatically generated by the feature extraction module / manually corrected by the manual annotation module and the long text content, so as to provide additional reference opinions for manual annotation.

[0107] As an implementation manner of this application, the data format in step S1 mainly includes the following fields: long text content LongText, topic feature TopicF, structural feature StructureF, grammatical feature GrammarF, readability feature ReadabilityF, topic score TopicS, structural score StructureS, grammatical score GrammarS, readability score ReadabilityS, total long text score LongTextS.

[0108] The specific data format is as follows. Given that the content of the feature set and the long text content are too long, the overly long content will be omitted here:

[0109]

[0110] After obtaining the long text data of the composition in the above format, these data are used for subsequent model training. In the field of long text evaluation, the method based on the pre-trained model is a classic approach, while the method based on the large language model is an emerging approach, and relevant experiments on ChatGPT have demonstrated the feasibility of applying the large language model to the composition scoring task. Therefore, the training module supports two different methods, namely a multi-dimensional evaluation method for long text based on the pre-trained model BigBird and a multi-dimensional evaluation method for long text based on the large language model Llama.

[0111] The multi-dimensional evaluation method for long text based on the pre-trained model BigBird uses the BigBird pre-trained model as the base model, and its goal is to learn the regression task of the multi-dimensional scores of the composition. The specific training process is as follows: First, the long text content and the feature set (it is also possible not to use the feature set or use a partial feature set) are used as inputs, and the multi-dimensional scores and the total score of the long text are used as labels. Then, the model needs to simultaneously learn the score prediction of the long text of the composition in four different dimensions and the score prediction of the overall long text. In addition, the system also uses the R-drop mechanism to avoid overfitting problems.

[0112] The multi-dimensional evaluation method for long text based on the large language model Llama uses the open-source large language model Llama as the base model and uses the LoRA algorithm for instruction fine-tuning. The specific training process is as follows: First, the input data is processed and combined with the fine-tuning template as the input of the model, and the multi-dimensional scores and the total score are also combined with the label model as training labels. Then, the Llama model is fine-tuned by instructions to learn the multi-dimensional scoring task of long text. In addition, the LoRA mechanism is used to reduce the resource overhead required for model training. Among them, the fine-tuning template and the label template can be defined by oneself in the parameter import step.

[0113] In the training module, the user can specify the multi-dimensional scoring method for long text, the feature processing method, the model configuration parameters, the training-related parameters, the evaluation metrics, etc., that is, the parameter import operation. After obtaining the above parameters and configurations, the system will automatically start model training. After the model training is completed, the system automatically exports the model to ONNX and performs int8 quantization, and obtains the preliminary performance results through five-fold cross-validation.

[0114] Finally, the model will be updated to the cache of the system and used for subsequent manual annotation. Specifically, when the annotator needs additional annotation references, the feature set automatically extracted by the feature extraction module / manually corrected by the manual annotation module and the long text content are transmitted to the model prediction module. The model will construct these information into the input required by the model, so as to use the model to infer the quality scores of the current long text in the corresponding dimensions. This score is finally fed back to the annotator to help them obtain more accurate annotation results.

[0115] In summary, the model training and inference system is an auxiliary system for the intelligent annotation system, and it has two main functions: one is to quickly train and evaluate the annotation results to obtain a preliminary performance report, and the other is to update the inference model of the model prediction module for the intelligent annotation system. Specifically, the system will update the long text inference model in this application, which can provide additional annotation references for annotators during manual annotation. This measure can also help improve the objectivity and stability of the annotation results. These designs are not available in other similar works, which can not only provide greater convenience for managers, but also improve the efficiency of manual annotation and contribute to the objectivity of the annotation results.

[0116] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0117] The above embodiments only represent several implementation manners of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of the patent of this application should be subject to the appended claims.

Claims

1. An intelligent annotation system for multi-dimensional evaluation of long texts, characterized in that: The system includes a feature extraction module, an intelligent scoring module, a manual labeling module, a model prediction module and a labeling processing module; The feature extraction module is used to parse the original long text content, perform feature adaptation according to business requirements and long text types, and extract features of corresponding dimensions; The intelligent scoring module is used to estimate the score based on the feature set extracted by the feature extraction module; The manual labeling module is used to modify the feature set and the estimated score based on the feature set extracted by the feature extraction module and the score calculated by the intelligent scoring module, and finally submit the labeling result; The single dimension labeling module is composed of a feature extraction module, an intelligent score estimation module, and a manual labeling module, and is responsible for the score labeling of a single dimension; The model prediction module is used to automatically generate feature sets and long text content based on the feature extraction module / manually corrected by the manual annotation module. The model performs score prediction to provide additional reference for manual annotation. The annotation processing module is used to automatically calculate the total score corresponding to the long text based on the annotation results of the subject, structure, grammar and readability score, and store the multi-dimensional score, multi-dimensional features and the long text total score as the annotation result of the current long text.

2. The intelligent annotation system for multi-dimensional evaluation of long text according to claim 1 is characterized in that: The feature extraction module includes three steps: text parsing, business adaptation and feature extraction; The text analysis uses the latest fastopic tool to analyze the topic distribution of the long text, and uses the Chatgpt tool to extract the topic sentence set in the long text; The business adapter selects a feature set according to the business requirements corresponding to the long text, and sets the feature set to a subject word set, a subject sentence set, the number of subject words used, the number of subject sentences used, a long text subject distribution vector, and a subject similarity; The feature extraction is based on the topic distribution and the topic word set, and extracts the topic word set in the best topic distribution, the appearance of the topic word in the long text of the composition, the overall topic distribution and other features; Based on the topic sentence set, the text of each topic sentence, the number of topic sentences, and the topic words in each topic sentence are extracted; the features are combined into a feature set of a topic dimension and passed to the manual annotation module and the intelligent scoring module.

3. The intelligent annotation system for multi-dimensional evaluation of long text according to claim 1 is characterized in that: The intelligent scoring module processes the feature results of the feature extraction module or the corrected features of the manual labeling module, including two steps: feature analysis and score calculation; The feature parsing operation selects corresponding features for estimating the score; The score calculation is based on the automatically calculated score of the screened subject features, as follows: Among them, alpha is a penalty factor used to penalize essay inputs that are too short; the TS function estimates the topic score of a long text based on the topic word set, topic sentence set, and sentences in the long text of the essay; m is the number of sentences in the long text of the essay; Among them, t is a preset threshold.

4. The intelligent annotation system for multi-dimensional evaluation of long text according to claim 1 is characterized in that: The manual labeling module performs labeling based on the feature extraction model and the additional information provided by the intelligent scoring module, including: Manually correct the relevant features obtained by the feature extraction module; Whenever the feature set content in the system is manually revised, the intelligent scoring module updates the calculated score in real time, and the annotator submits the revised score as the quality score of the long text in the topic dimension; if the score given by the intelligent scoring module is correct, the annotator does not need to revise it; When the feature set and annotation scores are corrected, the annotator submits them, and the two annotation results are integrated and passed to the system backend.

5. The intelligent annotation system for multi-dimensional evaluation of long text according to claim 1 is characterized in that: The annotation processing module performs post-processing based on the multi-dimensional annotation results obtained by the single-dimensional annotation module. The feature set corresponding to the current long text is post-processed and integrated. The annotation scores of the four different dimensions will be automatically calculated as the total score of the long text, as shown in the following formula: The annotation result format is a json file.

6. The intelligent annotation and model training reasoning method for multi-dimensional evaluation of long text according to claim 1 is characterized in that: The intelligent annotation system for multi-dimensional evaluation of long texts also includes: Perform business adaptation, feature extraction, and score estimation based on the input long text to obtain the long text, the feature set corresponding to the long text, and the estimated score calculated based on the feature set; If there is no problem between the automated feature set obtained by the feature extraction module and the estimated score calculated by the intelligent scoring module, the annotator directly submits the automated result as the annotation result; otherwise, the annotator needs to correct the problem in the feature set and obtain an updated estimated score; When the feature set has been corrected to the correct state, the annotator can directly adopt the estimated score as the annotation score; or manually modify a more accurate score as the annotation score; or use the model prediction module provided by the model training and inference system to perform model inference; After the annotator submits the annotation results of the four dimensions, the overall score of the long text is automatically calculated; the long text, multi-dimensional scores, multi-dimensional features, and overall score are all stored in the database as the complete annotation results of the current long text.

7. A model training and reasoning system for multi-dimensional evaluation of long texts, applied to the intelligent annotation system for multi-dimensional evaluation of long texts, characterized in that: include: Obtaining long text multi-dimensional annotation data through the intelligent annotation system for long text multi-dimensional evaluation, using the data for model training and testing, and using five-fold cross validation to test model performance; The model is trained based on manually specified model configuration and parameter settings, so that the model can learn the multi-dimensional scoring task of long texts; The model obtained from the training module is exported by onnx and quantized by int8, and the model is used for five-fold cross validation to obtain preliminary model performance results; The exported model can automatically predict the score of the current long text based on the feature set automatically generated by the feature extraction module / manually corrected by the manual annotation module and the long text content.

8. The model training and reasoning system for long text multi-dimensional evaluation according to claim 7 is characterized in that: The training model based on the manually specified model configuration and parameter setting enables the model to learn the multi-dimensional scoring task of long texts, including: Two training methods are supported: the long text multi-dimensional evaluation method based on the pre-trained model BigBird and the long text multi-dimensional evaluation method based on the large language model llama.

9. The intelligent annotation system for multi-dimensional evaluation of long text according to claim 8, characterized in that: The long text multi-dimensional evaluation method based on the pre-trained model BigBird includes: Take the long text content and feature set as input, and the multi-dimensional score and total score of the long text as labels; The model simultaneously learns the score prediction of the composition long text in four different dimensions and the score prediction of the long text as a whole; The overfitting problem is avoided through the R-drop mechanism.

10. The intelligent annotation system for multi-dimensional evaluation of long text according to claim 8, characterized in that: The long text multi-dimensional evaluation method based on the large language model llama includes: The input data is processed and combined with the fine-tuning template as the input of the model. The multi-dimensional scores and total scores are combined with the label model as training labels. Instructions to fine-tune the llama model to learn long text multi-dimensional scoring tasks; The LoRa mechanism is used to reduce the resource overhead required for model training, where the fine-tuning template and label template are defined by themselves in the parameter import step.