Data annotation method and device, scoring system, electronic device and computer medium

Through large-scale language models and semi-supervised learning methods, active learning is used to annotate datasets and pre-trained models to generate pseudo-label sample sets, which solves the problem of high complexity in generating self-explanatory data and achieves efficient and accurate self-explanatory data expansion and scoring model training.

CN119990242BActive Publication Date: 2025-09-12BEIJING SANSAN SMART EDUCATION TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510071317.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-09-12
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

Existing technologies have difficulty in efficiently generating high-quality and diverse self-explanatory data, especially in mathematics disciplines, where the time-intensive nature and writing requirements of self-explanation lead to high complexity in data collection and system development.

Method used

Using a large language model and semi-supervised learning method, we obtain an active learning labeled dataset, train an initial scoring model, combine it with the pre-trained large language model to generate an unlabeled dataset, and use the unlabeled dataset and the first scoring model to generate a pseudo-label sample set to expand the training data and improve the scoring accuracy.

Benefits of technology

Effectively utilizing limited annotation resources to generate high-quality and diverse pseudo-label sample sets improves the generalization ability and scoring accuracy of the scoring model and simplifies the process of generating self-explanatory data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990242B_ABST
    Figure CN119990242B_ABST
Patent Text Reader

Abstract

This disclosure provides a data annotation method and device, relating to the field of educational technology, specifically natural language processing, large models, deep learning, and other technical fields. The specific implementation scheme comprises: obtaining an active learning annotated dataset; training an initial scoring model based on the active learning annotated dataset to obtain a trained first scoring model; obtaining an unlabeled dataset based on the active learning annotated dataset and a pre-trained large language model; and obtaining a pseudo-labeled sample set based on the unlabeled dataset and the first scoring model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the field of educational technology, specifically to technical fields such as natural language processing, large models, and deep learning, and in particular to a data annotation method and device, an electronic device, and a computer-readable storage medium. Background Art

[0002] The rise of digital learning platforms has opened up a vast field for researchers, enabling them to deeply explore and understand learning behaviors through vast amounts of data from system interactions. Among numerous learning domains, self-explanation has garnered particular attention as an effective active learning strategy. It has shown remarkable effectiveness in improving comprehension in subjects like mathematics. Self-explanation can be viewed as a learning mechanism whereby learners deepen their understanding and absorb new insights by elaborating on their ideas, clarifying concepts, expanding their problem-solving approaches, and engaging in the problem-solving process.

[0003] Due to the time-intensive nature of self-explanations, large-scale data collection presents feasibility challenges. Furthermore, producing high-quality self-explanations requires both a deep understanding of the subject matter and strong writing skills. These challenges make it particularly difficult to amass a large and diverse repository of self-explanation examples. This difficulty further complicates the development of systems that utilize large numbers of self-explanation examples to aid learning. Summary of the Invention

[0004] The present disclosure provides a data annotation method and apparatus, an electronic device, and a computer-readable storage medium.

[0005] According to a first aspect, a data labeling method is provided, which includes: obtaining an active learning labeled dataset; training an initial scoring model based on the active learning labeled dataset to obtain a trained first scoring model; obtaining an unlabeled dataset based on the active learning labeled dataset and a pre-trained large language model; and obtaining a pseudo-label sample set based on the unlabeled dataset and the first scoring model.

[0006] According to a second aspect, a data labeling device is provided, which includes: an acquisition unit, configured to acquire an active learning labeled dataset; a model acquisition unit, configured to train an initial scoring model based on the active learning labeled dataset to obtain a trained first scoring model; a data acquisition unit, configured to obtain an unlabeled dataset based on the active learning labeled dataset and a pre-trained large language model; and a labeling unit, configured to obtain a pseudo-label sample set based on the unlabeled dataset and the first scoring model.

[0007] According to a third aspect, a scoring system is provided, which includes: a data generation unit for generating an active learning annotated dataset; an automatic scoring unit connected to the data generation unit, for implementing the method described in any implementation manner of the first aspect based on the active learning annotated dataset.

[0008] According to a fourth aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any implementation manner of the first aspect.

[0009] According to a fifth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause a computer to execute the method as described in any implementation of the first aspect.

[0010] The data labeling method provided by the embodiment of the present disclosure first obtains an active learning labeled dataset; secondly, based on the active learning labeled dataset, an initial scoring model is trained to obtain a trained first scoring model; then, based on the active learning labeled dataset and the pre-trained large language model, an unlabeled dataset is obtained; finally, based on the unlabeled dataset and the first scoring model, a pseudo-labeled sample set is obtained. Thus, through the large language model and semi-supervised technology, limited labeling resources can be effectively utilized to generate a pseudo-labeled sample set to expand the training data, thereby improving the accuracy of the generation of the pseudo-labeled sample set. Semi-supervised learning can efficiently utilize limited labeled data and rich unlabeled data, thereby improving the generalization ability and scoring accuracy of the first scoring model; and the large language model ensures the quality, diversity and consistency of the generated pseudo-labeled sample set, so that the pseudo-labeled sample set provides a solid foundation for retraining the first scoring model.

[0011] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0013] Figure 1 is a flow chart of an embodiment of a data annotation method according to the present disclosure;

[0014] Figure 2 This is a structural diagram obtained from the pseudo-label sample set in this disclosure;

[0015] Figure 3is a schematic diagram of the update of the scoring model in multiple training stages in the present disclosure;

[0016] Figure 4 It is a schematic diagram of the training of the scoring model at each stage presented in the form of data in this disclosure;

[0017] Figure 5 This is a structural diagram obtained by actively learning and labeling data sets in this disclosure;

[0018] Figure 6 is a structural diagram of an embodiment of a data tagging device according to the present disclosure;

[0019] Figure 7 is a schematic diagram of a structure of a scoring system according to the present disclosure;

[0020] Figure 8 3 is a block diagram of an electronic device used to implement the data tagging method of an embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] Unless expressly stated otherwise, throughout the specification and claims, the term "comprise" or variations such as "include" or "comprising", etc., will be understood to include the stated elements or components but not to exclude other elements or other components.

[0022] The technical solutions of the present disclosure are described below through specific examples. It should be understood that one or more steps mentioned in the present disclosure do not exclude the existence of other methods and steps before and after the combination step, or other methods and steps may be inserted between these explicitly mentioned steps. It should also be understood that these examples are only used to illustrate the present disclosure and are not used to limit the scope of the present disclosure. Unless otherwise specified, the numbering of each method step is only for the purpose of identifying each method step, and does not limit the order of arrangement of each method or limit the scope of implementation of the present disclosure. Changes or adjustments in their relative relationships can also be regarded as the scope of implementation of the present disclosure without substantial changes in the technical content.

[0023] The sources of the raw materials and instruments used in the examples are not particularly limited and can be purchased from the market or prepared according to conventional methods known to those skilled in the art.

[0024] Self-explanation, as an active learning strategy, involves learners proactively explaining, clarifying, and reasoning about information they've learned. This strategy involves verbally or in writing detailing their thought processes, concepts they understand, problem-solving methods, and the steps involved in solving a problem. Self-explanation helps learners better organize and consolidate knowledge, identify gaps in understanding, and promote deeper learning and cognitive development.

[0025] With the widespread adoption of technologies like digital learning platforms, self-explanation as a learning method has once again become a focal point and has found practical application. Modern learning innovation emphasizes the role of self-explanation by designing more intuitive interfaces, building assessment models based on self-explanation behavior, and developing strategies that can tap into deep self-explanation. For example, tools developed using traditional technologies highlight the central role of self-explanation in systematic problem solving. Ongoing research continues to expand the application of self-explanation in education, such as using template-based self-explanation methods. These templates provide learners with pre-defined frameworks, acting as built-in guides to enhance their explanation and thinking processes.

[0026] Furthermore, the practice of self-explanation has expanded beyond traditional boundaries. These methods involve not only conceptual understanding but also the promotion of a variety of educational tools, such as building feedback systems, carefully designing answers to practice quizzes, and generating valuable datasets for automated assessments. Automated assessment plays a crucial role in this system. By deeply analyzing and interpreting self-explanations, educators and automated systems can gain more accurate insights into learners' thinking patterns. This valuable information allows them to tailor more appropriate educational strategies to meet the learning needs of different individuals. Such insights are crucial for tasks such as classifying learners' responses, making it easy to identify common mistakes made by students or topics that they consistently struggle to grasp.

[0027] Self-explanation is an important learning strategy that promotes students' understanding and internalization of knowledge, and is particularly effective in logically intensive subjects such as mathematics. Related research shows that self-explanation helps:

[0028] 1) Deep understanding: By explaining the solutions to problems and the underlying principles, students can better grasp the knowledge.

[0029] 2) Problem-solving ability: Self-explanation promotes the explicitness of thinking processes and improves students' ability to analyze and solve problems.

[0030] Assessing the quality of students' self-explanations is an important step in achieving personalized teaching and feedback. Existing technologies mainly focus on manual scoring and partially automated scoring:

[0031] 1) Manual Grading: Education experts or teachers evaluate students’ self-explanations based on pre-set grading criteria. However, manual grading is time-consuming, subject to subjective influence, and inefficient.

[0032] 2) Automatic Scoring Attempts: Some studies have attempted to use NLP (Natural Language Processing) technology and machine learning algorithms to automatically score self-explanations. For example, methods such as text similarity, grammatical analysis, and semantic understanding can be used to assess the logic, completeness, and quality of expression of self-explanations. However, NLP is less capable of generating self-explanations or text related to self-explanations.

[0033] 3) Language Model-Based Generation: In recent years, the use of large language models (LLMs) (such as GPT-3 and GPT-4) to generate high-quality self-explanatory text has become an important means of data augmentation. LLMs can generate well-structured, accurate, and diverse self-explanatory sentences based on a given math problem or knowledge point.

[0034] Although some studies have used LLM to generate education-related data, high-quality generation of mathematical self-explanations is not yet common, especially generation strategies that focus on diversity and logical consistency.

[0035] To address the shortcomings of traditional techniques, this paper proposes a data annotation method, a semi-supervised approach based on a large language model (LLM). This method aims to exploit the model's potential in generating active learning text (e.g., sentences and articles), which will form the basis of a predictive active learning scoring model. By adopting a semi-supervised strategy and leveraging advanced language models, the accuracy and effectiveness of the model's automatic scoring are improved. Figure 1 A process 100 according to an embodiment of a data annotation method of the present disclosure is shown. The data annotation method includes the following steps:

[0036] Step 101: Obtain an active learning annotation dataset.

[0037] In this embodiment, the active learning annotation dataset is the annotation data of the active learning strategy, wherein the active learning strategy is a strategy for learners to actively learn knowledge during the learning process. The active learning annotation dataset includes at least one active learning annotation data, the active learning annotation data includes already annotated self-explanation data, the self-explanation data includes a self-explanation text and a quality score of the self-explanation text, wherein the self-explanation text is the data of the self-explanation strategy, wherein the self-explanation strategy is a strategy of providing a detailed description of one's own thinking process, the concepts understood, the method of solving a problem, and the steps of solving the problem. The self-explanation strategy is an important learning strategy that promotes learning understanding and knowledge internalization; the quality score of the self-explanation text can be evaluated based on multiple core criteria, including three core criteria: logical coherence, clarity of expression, and content relevance. The evaluation of multiple core criteria is achieved by combining natural language processing technology and machine learning models. These criteria do not have completely independent evaluation conditions or algorithms, but are achieved by comprehensively analyzing the semantics, structure, and content of the self-explanation text.

[0038] In traditional technologies, the generation of high-quality active learning data for active learning strategies in different subjects is not common, as it requires learners to think deeply and describe the problem-solving steps and thinking process in detail, which takes a lot of time and effort.

[0039] In this embodiment, the above step 101 includes: obtaining learners' feedback data on different courses during the learning process; inputting the feedback data into a large language model to obtain an active learning annotation dataset output by the large language model.

[0040] Step 102: Based on the active learning labeled dataset, an initial scoring model is trained to obtain a trained first scoring model.

[0041] In this embodiment, the initial scoring model is trained based on an active learning annotated dataset and is used to score an unlabeled dataset, thereby generating a model for a pseudo-labeled sample set. In this process, the main function of the initial scoring model is to generate preliminary labels for the unlabeled data. These labels will be used to further train or improve the initial scoring model, and the labels can be obtained through the scores of different types of samples output by the initial scoring model, that is, the input of the initial scoring model is the number of active learning annotations, and the output is the scores of different types of samples. When the active learning dataset includes self-explanatory data, the initial scoring model can generate a corresponding quality score based on the input self-explanatory text. It should be noted that when the initial scoring model is not trained with an active learning annotated dataset, the quality score it outputs may be inaccurate.

[0042] In this embodiment, the active learning annotation data set includes annotated self-explanatory data. Figure 2 As shown, an initial scoring model (not shown) is trained by actively learning a labeled dataset. After the initial scoring model is adjusted or optimized, if the initial scoring model meets the training completion condition, a first scoring model corresponding to the initial scoring model M1 is obtained. The first scoring model M1 can be used to perform quality scoring on the self-explanatory text in the unlabeled dataset, thereby generating a pseudo-label sample set B.

[0043] Step 103: Obtain an unlabeled dataset based on the active learning labeled dataset and the pre-trained large language model.

[0044] In this embodiment, the pre-trained large language model refers to a deep learning model trained using a large amount of text data, so that the model can generate natural language text or understand the meaning of language text. These models can provide in-depth knowledge and language about various topics by training on huge data sets.

[0045] In this embodiment, the unlabeled dataset is an unlabeled dataset related to the active learning strategy and used to train multiple-stage scoring models (such as the initial scoring model and the first scoring model). Specifically, the unlabeled dataset includes at least one unlabeled data point, each of which is related to the active learning strategy. For example, each unlabeled data point is self-explanatory text. The unlabeled dataset can be data obtained through data augmentation technology. The active learning annotated dataset is pre-annotated data for the active learning strategy. The unlabeled dataset can be directly generated by using a pre-trained large-scale language model to identify and imitate the content related to the active learning strategy in the annotated dataset.

[0046] In this embodiment, since active learning annotation data is difficult to obtain, the amount of active learning annotation data in the active learning annotation data set is relatively small. Figure 2 As shown, the active learning annotation dataset D is directly trained by pre-training a large language model L * 训 By performing data augmentation, we can obtain an unlabeled dataset W with a large amount of data.

[0047] Optionally, in order to increase the data volume of the generated unlabeled dataset, the domain knowledge text of the corresponding field can be combined to generate an unlabeled dataset, and the domain knowledge text and the active learning annotated dataset can be input into the pre-trained large language model to obtain the unlabeled dataset output by the pre-trained large language model.

[0048] Step 104: Obtain a pseudo-label sample set based on the unlabeled dataset and the first scoring model.

[0049] In this embodiment, the first scoring model is a trained initial scoring model. Since the first scoring model is trained by accurately actively learning annotated datasets, compared with the initial scoring model, the first scoring model can generate more accurate labels, thereby achieving labeling of unlabeled datasets.

[0050] In this embodiment, the above-mentioned step 104 includes: inputting each unlabeled data in the unlabeled data set into the first scoring model to obtain a quality score of each unlabeled data output by the first scoring model; based on the quality score of each unlabeled data, obtaining a label for each unlabeled data, and using each unlabeled data in the unlabeled data set and each unlabeled data label as a pseudo-label sample set.

[0051] In this embodiment, the pseudo-label sample set is used to train the scoring model at different training stages. Compared with the active learning dataset, the label annotation accuracy of each pseudo-label sample in the pseudo-label sample set is lower. The pseudo-label sample set includes at least one pseudo-label sample, each of which has corresponding text and labels.

[0052] like Figure 2 As shown, the first scoring model M1 performs quality scoring on the unlabeled dataset W, obtains the label of each unlabeled data in the unlabeled dataset based on the score, adds corresponding labels to each unlabeled data, and obtains a pseudo-label sample set B.

[0053] In this embodiment, the input and output of the initial scoring model and the first scoring model can be the same. Since the first scoring model is a model obtained by retraining based on the initial scoring model, the accuracy of the first scoring model in scoring the quality of the text is higher than that of the initial scoring model.

[0054] In this embodiment, the active learning annotation dataset can be a manually annotated dataset, and the pseudo-label sample set is an automatically generated dataset. The combination of pseudo-label data and manually annotated data expands the training set, improves the performance of the automatic scoring system, and ensures data quality.

[0055] The data labeling method provided by the embodiment of the present disclosure first obtains an active learning labeled dataset; secondly, based on the active learning labeled dataset, an initial scoring model is trained to obtain a trained first scoring model; then, based on the active learning labeled dataset and the pre-trained large language model, an unlabeled dataset is obtained; finally, based on the unlabeled dataset and the first scoring model, a pseudo-labeled sample set is obtained. Thus, through the large language model and semi-supervised technology, limited labeling resources can be effectively utilized to generate a pseudo-labeled sample set to expand the training data, thereby improving the accuracy of the generation of the pseudo-labeled sample set. Semi-supervised learning can efficiently utilize limited labeled data and rich unlabeled data, thereby improving the generalization ability and scoring accuracy of the first scoring model; and the large language model ensures the quality, diversity and consistency of the generated pseudo-labeled sample set, so that the pseudo-labeled sample set provides a solid foundation for retraining the first scoring model.

[0056] In some embodiments of the present disclosure, the above method also includes: training a first scoring model based on a pseudo-label sample set and an active learning annotation dataset; obtaining a trained second scoring model in response to detecting that the performance of the first scoring model has converged or the number of training times of the first scoring model has reached a predetermined number of iterations; and obtaining a new pseudo-label sample set based on the unlabeled dataset and the second scoring model.

[0057] In this embodiment, in each iterative training of the first scoring model, data can be selected from the pseudo-label sample set and the active learning annotation data set, the selected data can be input into the first scoring model, the loss value of the first scoring model is calculated, and whether the performance of the first scoring model has converged is detected based on the loss value. The convergence of the performance of the first scoring model can be reflected in various aspects, for example, in the iterative training for a specified number of consecutive times (for example, 100 times), the increase in the loss value of the first scoring model exceeds the preset amplitude.

[0058] In this embodiment, obtaining a new pseudo-label sample set based on the unlabeled data set and the second scoring model includes: using the second scoring model to perform quality scoring on the unlabeled data set, and based on the labels of each unlabeled data in the unlabeled data set, adding labels to each unlabeled data to obtain a new pseudo-label sample set.

[0059] In this embodiment, the number of training times for the first scoring model refers to the number of iterative training times for the first scoring model, which is also the number of times the loss value of the first scoring model is calculated. During each iterative training of the first scoring model, a sample is input to the first scoring model to obtain a score output by the first scoring model; a loss value for the scoring model is calculated based on the score; the loss value is used to detect the convergence of the performance of the first scoring model, or the number of training times is used to detect whether the training of the first scoring model has reached a predetermined number of iterations and has reached the training stage for obtaining a completed second scoring model. The predetermined number of iterations for the first scoring model can be set based on development requirements, for example, the predetermined number of iterations is 100,000.

[0060] In this embodiment, Figure 2 As shown, the first scoring model M1 performs quality scoring on the unlabeled dataset W, obtains the label of each unlabeled data in the unlabeled dataset based on the score, adds corresponding labels to each unlabeled data, and obtains a pseudo-label sample set B.

[0061] In this embodiment, Figure 4 As shown, the first scoring model M1 is based on the active learning annotation dataset D * 训 The model obtained by training is as follows: Figure 4 In the example, the first scoring model M1 is used in the active learning annotation dataset D * 训 , obtained by multiple iterative training X, and the second scoring model M2 is based on the first scoring model M1 and further uses the pseudo-label sample set D * 训 +D * 样 The model obtained by iterative training X. Since the pseudo-label sample set is generated by predicting the unlabeled data set through the first scoring model, and a certain amount of noise and uncertainty will be introduced when generating pseudo labels, the training data of the second scoring model is expanded and enhanced compared to the training data of the first scoring model. It should be noted that the first scoring model M1 predicts Y on a type of data to obtain the unlabeled sample D 样 , by analyzing the unlabeled samples D 样 Mark B and get the marked sample D * 样 , for unlabeled samples D 样 The second scoring model M2 predicts Y for another type of data to obtain unlabeled samples D 样 , by analyzing the unlabeled samples D 样 Mark B and get the marked sample D * 样 .

[0062] In practical applications, the second scoring model generally outperforms the first because its training data is richer, including more manually annotated data and generated pseudo-labeled data. Therefore, the second scoring model has improved performance compared to the first scoring model.

[0063] In this embodiment, the inputs of the initial scoring model, the first scoring model, and the second scoring model in each training stage can be the same, and the outputs of the initial scoring model, the first scoring model, and the second scoring model in each training stage can also be the same. Since the second scoring model in each training stage is a model obtained by retraining on the basis of the first scoring model, the accuracy of the quality scoring of the text by the second scoring model in each training stage is higher than the accuracy of the quality scoring by the first scoring model and the initial scoring model.

[0064] Optionally, the initial scoring model is mainly used for scoring tasks, while the first scoring model and the second scoring model in each training stage can be mainly used to generate pseudo-label sample sets. To this end, the input and output of the initial scoring model can also be different from those of the first scoring model and the second scoring model, that is, the initial scoring model is used to output the score of the sample data, while the first scoring model and the second scoring model are used to generate pseudo-label data.

[0065] The data labeling method provided in this embodiment trains a first scoring model based on a pseudo-label sample set and an active learning labeled data set; in response to detecting that the performance of the first scoring model has converged or the number of training times of the first scoring model has reached a predetermined number of iterations, a trained second scoring model is obtained; based on the unlabeled data set and the second scoring model, a new pseudo-label sample set is obtained, and the pseudo-label sample set and the active learning labeled data set are used to improve the generalization of the second scoring model for the first scoring model, thereby improving the accuracy of the new pseudo-label sample set.

[0066] Optionally, the above method further includes: in response to detecting that the performance of the first scoring model has not converged and the number of training times of the first scoring model has not reached a predetermined number of iterations, continuing to train the first scoring model.

[0067] In some optional implementations of the present disclosure, the training of the first scoring model based on the pseudo-label sample set and the active learning annotation dataset includes: selecting pseudo-label samples from the pseudo-label sample set; inputting the selected pseudo-label samples into the first scoring model to obtain a scoring result of the first scoring model; and based on the scoring result of the first scoring model, detecting whether the performance of the first scoring model has converged or whether the number of training times of the first scoring model has reached a predetermined number of iterations.

[0068] In this optional implementation, there are many ways to select pseudo-label samples from the pseudo-label sample set, such as random selection, sequential selection, etc. The scoring result of the first scoring model is the result of the first scoring model performing a quality score on the selected pseudo-label samples.

[0069] In this optional implementation, when training the first scoring model, whether the performance of the first scoring model has converged can be detected in a variety of ways, such as detecting whether the performance of the first scoring model has converged by the performance of the loss value of the first scoring model in multiple iterative trainings.

[0070] The method for training the first scoring model provided by this optional implementation manner includes selecting pseudo-label samples from a pseudo-label sample set; inputting the selected pseudo-label samples into the first scoring model to obtain a scoring result of the first scoring model; and based on the scoring result of the first scoring model, detecting whether the performance of the first scoring model has converged or whether the number of training times of the first scoring model has reached a predetermined number of iterations. Thus, the first scoring model is trained by the convergence of the performance of the first scoring model and the number of training times of the first scoring model, thereby improving the reliability of the training of the first scoring model.

[0071] In some optional implementations of the present disclosure, the above-mentioned detecting whether the performance of the first scoring model has converged or whether the number of training times of the first scoring model has reached a predetermined number of iterations based on the scoring result of the first scoring model includes: calculating the loss value of the first scoring model based on the scoring result of the first scoring model, and recording the number of training times of the first scoring model; detecting whether the loss value of the first scoring model is less than a first loss value threshold; in response to detecting that the loss value of the first scoring model is less than the first loss value threshold, determining that the performance of the first scoring model has converged; in response to detecting that the loss value of the first scoring model is greater than the first loss value threshold, detecting whether the number of training times of the first scoring model has reached a predetermined number of iterations.

[0072] In this optional implementation, the loss value of the first scoring model may be calculated using a loss function set for the first scoring model, such as a cross entropy function.

[0073] In this optional implementation, the number of training times can be recorded by a counter. For example, the number of training times of the first scoring model can be recorded each time the loss value of the first scoring model is calculated. Alternatively, the number of training times can be recorded by a counter after the scoring result of the first scoring model is obtained.

[0074] The method for detecting the first scoring model provided by this optional implementation calculates the loss value of the first scoring model based on the scoring result of the first scoring model and records the number of training times of the first scoring model; detects whether the loss value of the first scoring model is less than a first loss value threshold; in response to detecting that the loss value of the first scoring model is less than the first loss value threshold, determines that the performance of the first scoring model has converged; in response to detecting that the loss value of the first scoring model is greater than the first loss value threshold, detects whether the number of training times of the first scoring model reaches a predetermined number of iterations, thereby improving the reliability of the training of the first scoring model.

[0075] Optionally, the above-mentioned detecting whether the performance of the first scoring model has converged or whether the training times of the first scoring model has reached a predetermined number of iterations based on the scoring result of the first scoring model includes: calculating the loss value of the first scoring model based on the scoring result of the first scoring model, and recording the training times of the first scoring model; detecting whether the training times of the first scoring model has reached a predetermined number of iterations, and in response to detecting that the training times of the first scoring model has not reached the predetermined number of iterations, detecting whether the loss value of the first scoring model is less than a first loss value threshold; and in response to detecting that the loss value of the first scoring model is less than the first loss value threshold, determining that the performance of the first scoring model has converged.

[0076] In some optional implementations of the present disclosure, the above method also includes: training the second scoring model of the next training stage based on the new pseudo-label sample set and the second scoring model; obtaining an updated pseudo-label sample set based on the second scoring model of the next stage and the new pseudo-label sample set; continuing to train the second scoring model of other stages based on the updated pseudo-label sample set, and updating the pseudo-label sample set of each training stage after the end of each training stage.

[0077] In this embodiment, each training stage corresponds to a second scoring model, such as Figure 4 In the training phase, the second scoring model corresponding to each stage includes: M3 128 、M3 265 、M3 k 、M3 4096 , where k is a natural number greater than zero, and the second scoring model in the current training stage is the pseudo-label sample set given by the first scoring model or the second scoring model in the previous training stage (such as Figure 4 The pseudo-label sample set D of each stage in * 128 、D * 265 ,…,D * k 、D * 4096) is obtained through basic training. A second scoring model is obtained after each training stage, and each training stage has a fixed set of pseudo-labeled samples.

[0078] like Figure 3 As shown, in the current training phase, the pseudo-label sample set B is used t and active learning annotation dataset D * 训 Train to obtain the second scoring model M t+1 , using the second scoring model M t+1 Instead of the first scoring model, the unlabeled dataset W is scored, and the label of each unlabeled data in the unlabeled dataset is obtained based on the score. The corresponding label is added to each unlabeled data to obtain a new pseudo-label sample set.

[0079] The data labeling method provided in this embodiment trains the second scoring model for the next training phase based on a new pseudo-labeled sample set and a second scoring model. An updated pseudo-labeled sample set is obtained based on the second scoring model for the next phase and the new pseudo-labeled sample set. The second scoring model for the next phase is then trained based on the updated pseudo-labeled sample set, and the pseudo-labeled sample set for each training phase is updated after each training phase. Thus, the second scoring model for the next training phase is trained based on the updated pseudo-labeled sample set obtained from the second scoring model for the previous training phase, thereby maximizing the optimization of the pseudo-labeled sample set for each training phase.

[0080] In some optional implementations of the present disclosure, the above method also includes: determining the different amounts of pseudo-label data generated in each training stage based on the updated pseudo-label sample set; obtaining the second scoring model for each training stage; determining the prediction accuracy of the second scoring model for each training stage; and taking the amount of pseudo-label data of the second scoring model in the training stage with the highest prediction accuracy as the optimal amount.

[0081] In this embodiment, the goal of determining the optimal number is to find a balance between generating pseudo self-explanatory data that maximizes the predictive performance of the second scoring model while avoiding the introduction of noise due to excessive data volume or the underutilization of the generated data due to insufficient data volume. Determining an optimal number through experimentation can provide clear guidance for practical applications, enabling optimal decisions during data generation and model training.

[0082] In this embodiment, the number of pseudo-labeled data in the pseudo-labeled sample set at each training stage can be directly counted. The prediction accuracy of the second scoring model at each training stage can be calculated by performing the prediction accuracy calculation on the second scoring model using the test sample set after the second scoring model training at each stage is completed. Specifically, the prediction results of the second scoring model at each stage are compared with the true labels of the test dataset. Different evaluation metrics, such as accuracy, precision, and recall, can be used to calculate the prediction accuracy of the model. Common evaluation metrics include mean square error (MSE), root mean square error (RMSE), mean absolute error (MAE), etc. The smaller these metrics are, the higher the prediction accuracy of the model.

[0083] The method for obtaining the optimal number provided in this embodiment determines the different amounts of pseudo-labeled data generated in each training stage based on an updated pseudo-labeled sample set; obtains a second scoring model for each training stage; determines the prediction accuracy of the second scoring model for each training stage; and uses the amount of pseudo-labeled data of the second scoring model in the training stage with the highest prediction accuracy as the optimal number. This can provide clear guidance for practical applications and help make optimal decisions during data generation and model training.

[0084] Optionally, the data annotation method further includes: after determining the optimal number, determining a second scoring model with the optimal number, and using the second scoring model with the optimal number to evaluate the test data set D 测 Make the final prediction C and get the final pseudo-label sample set D * 测, like Figure 4 shown.

[0085] In some optional implementations of the present disclosure, the method further includes: detecting whether to continue increasing the amount of pseudo-labeled data based on the performance of the second scoring model on the validation set at each training stage. For example, when the performance improvement of the second scoring model on the validation set slows, increasing the amount of generated data is stopped; or when performance begins to decline, the proportion of generated data is reduced. This dynamic adjustment approach avoids manually searching for the "optimal amount" and may result in better model performance.

[0086] In addition to optimal data quantity, data quality is also a key factor influencing model performance. Optionally, the above approach can also include ensuring high-quality data through more refined generation strategies, stricter quality control mechanisms, or the inclusion of expert review. As long as quality is guaranteed, there's no need to generate excessive amounts of data to improve model performance.

[0087] Data augmentation is another direction worth exploring. Optionally, the above methods can also include increasing data diversity by performing various transformations on the generated data (such as data mixing and perturbation), rather than simply increasing the amount of data. This diverse data may be more helpful for the generalization ability of the model.

[0088] Active learning is a learning method that can effectively utilize limited labeled data. Optionally, the above method also includes: controlling the second scoring model to select those generated data that it considers to be the most difficult to predict or the most uncertain for training, so that the generated data can be more effectively utilized.

[0089] In some optional implementations of the present disclosure, the above-mentioned acquisition of the active learning annotation dataset includes: acquiring the active learning strategy text and the title of the active learning strategy text; inputting the active learning strategy text into the first feature extraction layer to obtain text features; inputting the title into the second feature extraction layer to obtain title features; obtaining splicing features based on the text features and the title features; inputting the splicing features into the regression test model to obtain the quality score of the active learning strategy text, and using the active learning strategy text and the quality score as the active learning annotation data in the active learning annotation dataset.

[0090] In this optional implementation, the first feature extraction layer and the second feature extraction layer can be implemented using a BERT (Bidirectional Encoder Representations from Transformers) model, wherein the first feature extraction layer and the second feature extraction layer are both implemented using the embedding layer and encoding layer in the BERT model. The BERT model is based on an advanced transformer structure and has demonstrated superior performance that surpasses previous models in a variety of natural language processing tasks.

[0091] In this optional implementation, the regression test model is a regression layer that receives the concatenated features from the BERT encoder and further processes them to predict the quality score of the self-explanatory text. Specifically, this regression layer can be a fully connected layer (FC layer), which maps the output features of BERT to specific score values ​​through linear transformation and activation function. Therefore, the regression test model is a component independent of BERT and is used for the final regression prediction task. Due to the excellent performance of the regression test model and its good adaptability to Chinese, the regression test model can directly receive the pre-processed self-explanatory text and the corresponding test title as input, and then output the quality score of each self-explanatory text.

[0092] In this optional implementation, the active learning strategy text may include a self-explanatory text, and the title is the title of the active learning strategy text. The test title comes from the math test questions in the digital courseware learning platform. During the data collection phase, not only the self-explanatory texts of the scholars were collected, but also the corresponding test questions were recorded. These test questions are titles, which provide contextual information about the specific problems targeted by the self-explanation. Therefore, during the model training and prediction process, the self-explanatory text and its corresponding title are input into the regression test model together to make full use of the question information to improve the accuracy of the score.

[0093] In this optional implementation, Figure 5 The first feature extraction layer in can be the embedding layer and encoder layer of the BERT model. The BERT model first embeds the input text (i.e., self-explanatory text) and then extracts contextual features through a multi-layer Transformer encoder to obtain text features. Figure 5 The second feature extraction layer in can be the embedding layer and encoder layer of the BERT model. The BERT model first embeds the input text (i.e., title) and then extracts contextual features through a multi-layer Transformer encoder to obtain title features.

[0094] exist Figure 5 In this example, text features are the feature representation of the self-explanatory text after being embedded and encoded by the BERT model, and title features are the feature representation of the title after being embedded and encoded by the BERT model. Concatenated features are the concatenation of text features and title features. These concatenated features combine the feature information of the self-explanatory text and the title. The concatenated features are used in the subsequent regression test model to predict the quality score of the self-explanatory text.

[0095] This optional implementation provides a method for obtaining an active learning annotation dataset, obtaining active learning strategy text and the title of the active learning strategy text; inputting the active learning strategy text into the first feature extraction layer to obtain text features; inputting the title into the second feature extraction layer to obtain title features; obtaining splicing features based on the text features and the title features; inputting the splicing features into the regression test model to obtain the quality score of the active learning strategy text, and using the active learning strategy text and the quality score as the active learning annotation data in the active learning annotation dataset. This design can make full use of the contextual information of the self-explanatory text and the title, thereby improving the accuracy of the quality score.

[0096] In some optional implementations of the present disclosure, the above-mentioned training of the initial scoring model based on the active learning labeled dataset to obtain the trained first scoring model includes: selecting active learning labeled data from the active learning labeled dataset; inputting the selected active learning labeled data into the initial scoring model to obtain a scoring result of the initial scoring model; based on the scoring result, detecting whether the initial scoring model meets the training completion condition; in response to detecting that the initial scoring model meets the training completion condition, obtaining the trained first scoring model.

[0097] In this optional implementation, there are various ways to select active learning annotation data from the active learning annotation dataset, such as random selection, sequential selection, etc. The scoring result of the initial scoring model is the result of the quality scoring of the selected pseudo-label samples by the initial scoring model.

[0098] In this optional implementation, when training the first scoring model, whether the initial scoring model meets the training completion conditions can be detected in a variety of ways, such as detecting whether the first scoring model meets the training completion conditions by the performance of the loss value of the initial scoring model in multiple iterative trainings.

[0099] The method for obtaining a trained first scoring model provided by this optional implementation method includes selecting active learning annotated data from an active learning annotated data set; inputting the selected active learning annotated data into an initial scoring model to obtain a scoring result of the initial scoring model; based on the scoring result, detecting whether the initial scoring model meets the training completion condition; in response to detecting that the initial scoring model meets the training completion condition, obtaining a trained first scoring model, thereby improving the reliability of obtaining the first scoring model.

[0100] In some optional implementations of the present disclosure, the above-mentioned training completion condition includes: the prediction accuracy of the initial scoring model on the validation set is greater than or equal to a predetermined accuracy threshold.

[0101] In this optional implementation, the validation set is a set of samples selected for the initial scoring model. The prediction results of the initial scoring model are compared with the true labels of the validation set. Different evaluation metrics such as accuracy, precision, and recall can be used to calculate the prediction accuracy of the model.

[0102] In this optional implementation, the predetermined accuracy threshold may be set based on development requirements, for example, the predetermined accuracy threshold is 89%.

[0103] This optional implementation provides a training completion condition, which can effectively verify whether the first scoring model has been trained, thereby improving the reliability of the first scoring model.

[0104] In some optional implementations of the present disclosure, obtaining an unlabeled dataset based on an active learning labeled dataset and a pre-trained large language model includes: selecting a random dataset based on the active learning labeled dataset; extracting active learning keywords based on the random dataset; and inputting the active learning keywords and guiding prompt words into the pre-trained large language model to obtain an unlabeled dataset output by the pre-trained large language model.

[0105] In this optional implementation, randomly selecting data means first randomly selecting 30% of the data from the manually labeled training data set to fully utilize the rich diversity of students' self-explanations.

[0106] Keyword extraction refers to extracting ten keywords from each self-explanation, which capture its core meaning and guide the pre-trained large language model to generate context-related data.

[0107] Data generation for a pre-trained large language model: Using the extracted keywords, we provide prompts to the pre-trained large language model. Specifically, we use each group of 10 keywords as initial input to guide the pre-trained large language model to generate contextually consistent pseudo self-explanations. The prompt could be "elaborate based on the provided keywords," ensuring that the generated content remains relevant to the context of the original self-explanation.

[0108] The method for obtaining an unlabeled dataset provided by this optional implementation method selects a random dataset based on an active learning labeled dataset; extracts active learning keywords based on the random dataset; and inputs the active learning keywords and guiding prompt words into a pre-trained large language model to obtain an unlabeled dataset output by the pre-trained large language model, thereby improving the quality of the unlabeled dataset.

[0109] In some optional implementations of the present disclosure, the above-mentioned obtaining of a pseudo-label sample set based on an unlabeled data set and a first scoring model includes: inputting each unlabeled data in the unlabeled data set into the first scoring model to obtain a predicted quality score of each unlabeled data output by the first scoring model; obtaining at least one to-be-labeled data based on the predicted quality score of each unlabeled data; and adding respective predicted quality scores to each to-be-labeled data in the at least one to-be-labeled data to obtain a pseudo-label sample set.

[0110] In this optional implementation, the predicted quality score is the quality score of each unlabeled data. The above-mentioned obtaining at least one data to be labeled based on the predicted quality score of each unlabeled data includes: in response to the predicted quality score being in a score segment among different score segments, querying the label corresponding to the score segment, and adding the label to the unlabeled data.

[0111] The method for obtaining a pseudo-label sample set provided by this optional implementation inputs each unlabeled data in the unlabeled data set into a first scoring model to obtain a predicted quality score of each unlabeled data output by the first scoring model; based on the predicted quality score of each unlabeled data, at least one to-be-labeled data is obtained; and each to-be-labeled data in the at least one to-be-labeled data is added with its own predicted quality score to obtain a pseudo-label sample set, thereby improving the reliability of obtaining the pseudo-label sample set.

[0112] In some optional implementations of the present disclosure, the above-mentioned obtaining at least one data to be labeled based on the predicted quality scores of each unlabeled data includes: comparing the predicted quality scores of each unlabeled data with a predetermined threshold to obtain a comparison result; and taking the unlabeled data with a predicted quality score greater than or equal to the predetermined threshold in the comparison result as the data to be labeled.

[0113] In this optional implementation, the data to be labeled is data for generating labeled samples. The data to be labeled may be data that meets the sample requirements but is not labeled accordingly. The predetermined threshold may be set based on development requirements.

[0114] This optional implementation provides a method for obtaining data to be labeled, comparing the predicted quality scores of each unlabeled data with a predetermined threshold to obtain a comparison result; the unlabeled data with a predicted quality score greater than or equal to the predetermined threshold in the comparison result is used as the data to be labeled, thereby improving the reliability of obtaining the data to be labeled.

[0115] Further references Figure 6 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a data annotation device. Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0116] like Figure 6 As shown, the data labeling device 600 provided in this embodiment includes: an acquisition unit 601, a model acquisition unit 602, a data acquisition unit 603, and a labeling unit 604. The acquisition unit 601 can be configured to acquire an active learning labeled dataset. The model acquisition unit 602 can be configured to train an initial scoring model based on the active learning labeled dataset to obtain a trained first scoring model. The data acquisition unit 603 can be configured to obtain an unlabeled dataset based on the active learning labeled dataset and a pre-trained large language model. The labeling unit 604 can be configured to obtain a pseudo-label sample set based on the unlabeled dataset and the first scoring model.

[0117] In this embodiment, the data annotation device 600 includes the acquisition unit 601, the model acquisition unit 602, the data acquisition unit 603, and the annotation unit 604. The specific processing and technical effects thereof can be referred to in the respective Figure 1 The relevant descriptions of step 101, step 102, step 103, and step 104 in the corresponding embodiment are not repeated here.

[0118] In some embodiments of the present disclosure, the above-mentioned data labeling device 600 also includes: obtaining a new pseudo-label sample set unit (not shown in the figure), and the new pseudo-label sample set unit is configured to: train a first scoring model based on the pseudo-label sample set and the active learning labeling data set; in response to detecting that the performance of the first scoring model converges or the training number of the first scoring model reaches a predetermined number of iterations, obtain a trained second scoring model; obtain a new pseudo-label sample set based on the unlabeled data set and the second scoring model.

[0119] In some embodiments of the present disclosure, the above-mentioned device further includes a first model training unit (not shown in the figure), and the first model training unit is configured to: train a first scoring model based on a pseudo-label sample set and an active learning annotation data set; obtain a trained second scoring model in response to detecting that the performance of the first scoring model converges or the number of training times of the first scoring model reaches a predetermined number of iterations; obtain a new pseudo-label sample set based on the unlabeled data set and the second scoring model.

[0120] In some embodiments of the present disclosure, the above-mentioned first model training unit is configured to: select pseudo-label samples from the pseudo-label sample set; input the selected pseudo-label samples into the first scoring model to obtain the scoring results of the first scoring model; based on the scoring results of the first scoring model, detect whether the performance of the first scoring model converges or whether the number of training times of the first scoring model reaches a predetermined number of iterations.

[0121] In some embodiments of the present disclosure, the above-mentioned first model training unit is configured to: calculate the loss value of the first scoring model based on the scoring result of the first scoring model, and record the number of training times of the first scoring model; detect whether the loss value of the first scoring model is less than a first loss value threshold; in response to detecting that the loss value of the first scoring model is less than the first loss value threshold, determine that the performance of the first scoring model has converged; in response to detecting that the loss value of the first scoring model is greater than the first loss value threshold, detect whether the number of training times of the first scoring model reaches a predetermined number of iterations.

[0122] In some embodiments of the present disclosure, the above-mentioned device 600 also includes: a second model training unit (not shown in the figure), and the above-mentioned second model training unit is configured to: train the second scoring model of the next training stage based on the new pseudo-label sample set and the second scoring model; obtain an updated pseudo-label sample set based on the second scoring model of the next stage and the new pseudo-label sample set; continue to train the second scoring model of other stages based on the updated pseudo-label sample set, and update the pseudo-label sample set of each training stage after the end of each training stage.

[0123] In some embodiments of the present disclosure, the above-mentioned device 600 also includes a quantity acquisition unit (not shown in the figure), which is configured to: determine the different quantities of pseudo-label data generated in each training stage based on the updated pseudo-label sample set; obtain the second scoring model of each training stage; determine the prediction accuracy of the second scoring model of each training stage; and take the amount of pseudo-label data of the second scoring model of the training stage with the highest prediction accuracy as the optimal quantity.

[0124] In some embodiments of the present disclosure, the above-mentioned acquisition unit 601 is configured to: obtain the active learning annotation data unit is configured to: obtain the active learning strategy text and the title of the active learning strategy text; input the active learning strategy text into the first feature extraction layer to obtain text features; input the title into the second feature extraction layer to obtain title features; obtain splicing features based on text features and title features; input the splicing features into the regression test model to obtain the quality score of the active learning strategy text, and use the active learning strategy text and the quality score as the active learning annotation data in the active learning annotation dataset.

[0125] In some embodiments of the present disclosure, the above-mentioned model obtaining unit 602 is configured to: select active learning annotation data from the active learning annotation dataset; input the selected active learning annotation data into the initial scoring model to obtain a scoring result of the initial scoring model; based on the scoring result, detect whether the initial scoring model meets the training completion condition; in response to detecting that the initial scoring model meets the training completion condition, obtain a first scoring model that has completed training.

[0126] In some embodiments of the present disclosure, the above-mentioned training completion condition includes: the prediction accuracy of the initial scoring model on the validation set is greater than or equal to a predetermined accuracy threshold.

[0127] In some embodiments of the present disclosure, the data acquisition unit 603 is configured to: select a random data set based on an active learning annotated data set; extract active learning keywords based on the random data set; input the active learning keywords and guiding prompt words into a pre-trained large language model to obtain an unlabeled data set output by the pre-trained large language model.

[0128] In some embodiments of the present disclosure, the above-mentioned labeling unit 604 is configured to: input each unlabeled data in the unlabeled data set into the first scoring model to obtain the predicted quality score of each unlabeled data output by the first scoring model; obtain at least one to-be-labeled data based on the predicted quality score of each unlabeled data; add the respective predicted quality scores to each to-be-labeled data in the at least one to-be-labeled data to obtain a pseudo-label sample set.

[0129] In some embodiments of the present disclosure, the above-mentioned labeling unit 604 is configured to: compare the prediction quality score of each unlabeled data with a predetermined threshold to obtain a comparison result; and take the unlabeled data with a prediction quality score greater than or equal to the predetermined threshold in the comparison result as the data to be labeled.

[0130] The data labeling device provided by the embodiment of the present disclosure is as follows: first, the acquisition unit 601 acquires an active learning labeled data set; second, the model acquisition unit 602 trains an initial scoring model based on the active learning labeled data set to obtain a trained first scoring model; then, the data acquisition unit 603 obtains an unlabeled data set based on the active learning labeled data set and the pre-trained large language model; finally, the labeling unit 604 obtains a pseudo-label sample set based on the unlabeled data set and the first scoring model. Thus, through the large language model and semi-supervised technology, limited labeling resources can be effectively utilized to generate a pseudo-label sample set to expand the training data, thereby improving the accuracy of the generation of the pseudo-label sample set. Semi-supervised learning can efficiently utilize limited labeled data and rich unlabeled data, thereby improving the generalization ability and scoring accuracy of the first scoring model; and the large language model ensures the quality, diversity and consistency of the generated pseudo-label sample set, so that the pseudo-label sample set provides a solid foundation for retraining the first scoring model.

[0131] Further references Figure 7 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a scoring system. Figure 1 The method embodiment shown corresponds to the embodiment shown.

[0132] like Figure 7 As shown, the scoring system 700 provided in this embodiment includes: a data generation unit 701 and an automatic scoring unit 702. The data generation unit 701 is used to generate an active learning annotated dataset; the automatic scoring unit 702 is connected to the data generation unit 701 and is used to implement the data annotation method of the above embodiment based on the active learning annotated dataset.

[0133] In this embodiment, the active learning annotation dataset generated by the data generation unit 701 can be an annotation dataset obtained manually by annotating initial data. Optionally, the active learning annotation dataset can also be a small amount of annotation data generated by a machine. After the machine generates the active annotation dataset, the active learning annotation dataset can be quality evaluated. The quality evaluation is mainly based on three core standards: logical coherence, clarity of expression, and content relevance. Specifically, "logical coherence" is used to measure the orderliness of the explanation; "clarity of expression" is used to evaluate the ease of understanding of the explanation; and "content relevance" ensures that all relevant knowledge points and procedural details are included in the explanation. In order to further ensure the consistency of quality evaluation, when the active learning annotation dataset is obtained through self-explanatory text, the consistency scoring criteria and definitions can be used to perform an overall evaluation of the following steps Step_1 to Step_5. This overall evaluation can be applicable to tasks with multiple solutions or strategies.

[0134] Step 1: First, read through the entire self-explanation text to gain a preliminary impression of the learner's overall thinking and understanding. The focus of this stage is to determine whether the learner's overall thinking is clear, whether they can accurately understand the question, and whether they can explain it in an organized manner.

[0135] Step 2, based on the overall reading, focuses on the logic of the self-explanatory text. Specifically, this assessment will assess whether the learner can develop their explanation in a reasonable and logical order, whether the various sections flow smoothly together, and whether there are any logical errors or loopholes. For example, whether the learner can correctly draw conclusions or whether there are any contradictions in the explanation.

[0136] Step 3: Evaluators assess whether learners use clear, understandable language in their self-explanations and whether they accurately express their ideas. This phase focuses on whether learners use appropriate technical terminology, explain complex issues in clear and concise language, and avoid ambiguous expressions.

[0137] Step 4: The evaluator will check whether the learner's self-explanation covers all the key knowledge points and problem-solving steps involved in the question and whether it can fully answer the question. For example, whether the learner has omitted important problem-solving steps or deviated from the core requirements of the question.

[0138] Step 5: Based on the evaluation of the above three dimensions, the evaluator will comprehensively consider the learner's overall performance and give a comprehensive score. The scoring will refer to the pre-set scoring criteria and definitions to ensure the objectivity and consistency of the scoring.

[0139] This disclosure adopts a more holistic evaluation method for active learning annotated data, and comprehensively understands the learner's understanding of a specific topic through the overall evaluation of each test.

[0140] In this embodiment, in the scoring system 700, the specific processing of the automatic scoring unit 702 and the technical effects thereof can be referred to in Figure 1 The relevant descriptions of step 101, step 102, step 103, and step 104 in the corresponding embodiment are not repeated here.

[0141] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0142] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their modes are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0143] like Figure 8 As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0144] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0145] The computing unit 801 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the data labeling method. For example, in some embodiments, the data labeling method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the data labeling method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the data labeling method by any other appropriate means (e.g., by means of firmware).

[0146] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0147] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. Such program code may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data annotation device, such that when the program code is executed by the processor or controller, the modes / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on the machine, partially on the machine, as a standalone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0148] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0149] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0150] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0151] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0152] The foregoing descriptions of specific exemplary embodiments of the present disclosure are for purposes of illustration and description. These descriptions are not intended to limit the present disclosure to the precise forms disclosed, and it is apparent that many variations and modifications are possible in light of the foregoing teachings. The exemplary embodiments have been selected and described for the purpose of explaining the specific principles of the present disclosure and their practical application, thereby enabling those skilled in the art to realize and utilize a variety of exemplary embodiments of the present disclosure and various options and modifications. The scope of the present disclosure is intended to be defined by the claims and their equivalents.

Claims

1. A data annotation method, comprising: Obtain active learning annotated dataset; Based on the active learning annotated dataset, an initial scoring model is trained to obtain a trained first scoring model; An unlabeled dataset is obtained based on the active learning annotated dataset and the pre-trained large language model; the active learning annotated dataset is pre-annotated data of the active learning strategy, and the data includes annotated self-explanatory data, and the self-explanatory data includes self-explanatory text and a quality score of the self-explanatory text. The obtaining of the unlabeled dataset based on the active learning annotated dataset and the pre-trained large language model includes: using the pre-trained large language model to identify and imitate active learning strategy-related content in the active learning annotated dataset to generate the unlabeled dataset, wherein the unlabeled dataset includes at least one unlabeled data, and each unlabeled data is related to the active learning strategy; Based on the unlabeled dataset and the first scoring model, a pseudo-label sample set is obtained.

2. The method according to claim 1, further comprising: Training the first scoring model based on the pseudo-label sample set and the active learning annotation dataset; In response to detecting that the performance of the first scoring model has converged or the number of training times of the first scoring model has reached a predetermined number of iterations, obtaining a trained second scoring model; Based on the unlabeled dataset and the second scoring model, a new pseudo-label sample set is obtained.

3. The method according to claim 2, wherein training the first scoring model based on the pseudo-label sample set and the active learning annotated dataset comprises: Selecting a pseudo-label sample from the pseudo-label sample set; Inputting the selected pseudo-label samples into the first scoring model to obtain a scoring result of the first scoring model; Based on the scoring result of the first scoring model, it is detected whether the performance of the first scoring model has converged or whether the number of training times of the first scoring model has reached a predetermined number of iterations.

4. The method according to claim 3, wherein: The detecting, based on the scoring result of the first scoring model, whether the performance of the first scoring model has converged or whether the number of training times of the first scoring model has reached a predetermined number of iterations includes: Based on the scoring result of the first scoring model, calculating the loss value of the first scoring model and recording the number of training times of the first scoring model; Detecting whether the loss value of the first scoring model is less than a first loss value threshold; In response to detecting that the loss value of the first scoring model is less than a first loss value threshold, determining that the performance of the first scoring model has converged; In response to detecting that the loss value of the first scoring model is greater than the first loss value threshold, detecting whether the number of training times of the first scoring model reaches a predetermined number of iterations.

5. The method according to claim 2, further comprising: Training a second scoring model for the next training phase based on the new pseudo-label sample set and the second scoring model; Obtaining an updated pseudo-label sample set based on the second scoring model of the next stage and the new pseudo-label sample set; Based on the updated pseudo-label sample set, the second scoring model of other stages is continuously trained, and after each training stage is completed, the pseudo-label sample set of each training stage is updated.

6. The method according to claim 5, further comprising: Based on the updated pseudo-label sample set, determine the different amount of pseudo-label data generated in each training phase; Obtain the second scoring model for each training stage; determining the predictive accuracy of the second scoring model at each training stage; The amount of pseudo-labeled data of the second scoring model in the training stage with the highest prediction accuracy is taken as the optimal amount.

7. The method according to any one of claims 1 to 6, wherein: The obtaining of the active learning annotation dataset includes: Obtaining an active learning strategy text and a title of the active learning strategy text; Inputting the active learning strategy text into the first feature extraction layer to obtain text features; Inputting the title into the second feature extraction layer to obtain title features; Obtaining a splicing feature based on the text feature and the title feature; The splicing features are input into a regression test model to obtain a quality score of the active learning strategy text, and the active learning strategy text and the quality score are used as active learning annotation data in an active learning annotation dataset.

8. The method according to any one of claims 1 to 6, wherein: The training of the initial scoring model based on the active learning labeled dataset to obtain a trained first scoring model includes: Selecting active learning annotation data from the active learning annotation dataset; Inputting the selected active learning annotated data into the initial scoring model to obtain a scoring result of the initial scoring model; Based on the scoring result, detecting whether the initial scoring model meets the training completion condition; In response to detecting that the initial scoring model meets the training completion condition, a first scoring model that has completed training is obtained.

9. The method according to claim 8, wherein The training completion conditions include: The prediction accuracy of the initial scoring model on the validation set is greater than or equal to a predetermined accuracy threshold.

10. The method according to any one of claims 1 to 6, wherein: The step of obtaining an unlabeled dataset based on the active learning labeled dataset and the pre-trained large language model includes: Based on the active learning labeled dataset, a random dataset is selected; Extracting active learning keywords based on the random data set; The active learning keywords and guiding prompt words are input into a pre-trained large language model to obtain an unlabeled data set output by the pre-trained large language model.

11. The method according to any one of claims 1 to 6, wherein: The obtaining of a pseudo-label sample set based on the unlabeled dataset and the first scoring model comprises: Inputting each unlabeled data in the unlabeled data set into the first scoring model to obtain a prediction quality score of each unlabeled data output by the first scoring model; Based on the predicted quality scores of each unlabeled data, at least one unlabeled data is obtained; A respective prediction quality score is added to each to-be-labeled data in the at least one to-be-labeled data to obtain a pseudo-label sample set.

12. The method according to claim 11, wherein The obtaining of at least one piece of data to be labeled based on the predicted quality scores of each piece of unlabeled data includes: Compare the prediction quality scores of each unlabeled data with a predetermined threshold to obtain a comparison result; The unlabeled data with a prediction quality score greater than or equal to a predetermined threshold in the comparison result is used as the data to be labeled.

13. A data annotation device, comprising: An acquisition unit, configured to acquire an active learning labeled dataset; A model obtaining unit is configured to train an initial scoring model based on the active learning annotated dataset to obtain a trained first scoring model; A data acquisition unit is configured to obtain an unlabeled dataset based on the active learning annotated dataset and the pre-trained large language model; the active learning annotated dataset is pre-labeled data of the active learning strategy, which includes annotated self-explanatory data, and the self-explanatory data includes self-explanatory text and a quality score of the self-explanatory text. The data acquisition unit is further configured to: use the pre-trained large language model to identify and imitate active learning strategy-related content in the active learning annotated dataset to generate an unlabeled dataset, wherein the unlabeled dataset includes at least one unlabeled data, each unlabeled data being related to the active learning strategy; The labeling unit is configured to obtain a pseudo-label sample set based on the unlabeled data set and the first scoring model.

14. A scoring system, comprising: A data generation unit, used to generate an active learning labeled data set; An automatic scoring unit, connected to the data generating unit, is used to implement the data labeling method according to any one of claims 1 to 12 based on the active learning labeling dataset.

15. An electronic device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.

16. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Question and answer data construction method and device based on large language model

    CN117591661A