Data labeling method and device, scoring system, electronic equipment and computer medium
By training the initial scoring model and generating a pseudo-label sample set using pre-trained large language models, the time and quality challenges of self-interpreting data annotation are solved, and efficient self-interpreting text scoring and model generalization are achieved.
Patent Information
- Application Number
- CN202510071317.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-16
AI Technical Summary
The existing technology is difficult to effectively solve the time-intensive and quality requirements of self-interpretation data annotation, resulting in the difficulty of accumulating a huge and diverse self-interpretation sample library, which in turn increases the development complexity of auxiliary learning systems.
By obtaining the actively learning labeled dataset, training the initial scoring model, combining pre-trained large language models to generate unlabeled datasets, and using semi-supervised learning to generate pseudo-label sample sets, expanding the training data and improving scoring accuracy.
Effectively utilize limited annotation resources to generate high-quality pseudo-label sample sets, improve the scoring accuracy of self-interpretation text and the generalization ability of the model, and simplify the data annotation process.
Smart Images

Figure CN119990242A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the field of educational technology, and specifically relates to technical fields such as natural language processing, large models, and deep learning, and in particular to a data annotation method and device, an electronic device, and a computer-readable storage medium. Background Art
[0002] The rise of digital learning platforms has opened up a vast space for researchers to deeply explore and understand learning behaviors through massive amounts of system interaction data. Among many learning fields, self-explanation has received special attention as an effective active learning strategy. In particular, self-explanation has shown significant effects in improving comprehension in subjects such as mathematics. Self-explanation can be seen as a learning mechanism where learners deepen their understanding of knowledge and absorb new insights by explaining their ideas, clarifying concepts, expanding problem-solving methods, and deepening the problem-solving process.
[0003] Due to the time-intensive nature of self-explanations, there are feasibility challenges in large-scale data collection. In addition, writing high-quality self-explanations requires both proficiency in subject-specific content and good writing skills. Faced with these challenges, it is particularly difficult to accumulate a large and diverse sample library of self-explanations. This difficulty further exacerbates the complexity of developing systems to assist learning with large numbers of self-explanation examples. Summary of the invention
[0004] The present disclosure provides a data labeling method and device, an electronic device, and a computer-readable storage medium.
[0005] According to a first aspect, a data labeling method is provided, the method comprising: obtaining an active learning labeled dataset; training an initial scoring model based on the active learning labeled dataset to obtain a trained first scoring model; obtaining an unlabeled dataset based on the active learning labeled dataset and a pre-trained large language model; obtaining a pseudo-label sample set based on the unlabeled dataset and the first scoring model.
[0006] According to a second aspect, a data labeling device is provided, which includes: an acquisition unit, configured to acquire an active learning labeled data set; a model acquisition unit, configured to train an initial scoring model based on the active learning labeled data set to obtain a trained first scoring model; a data acquisition unit, configured to obtain an unlabeled data set based on the active learning labeled data set and a pre-trained large language model; and a labeling unit, configured to obtain a pseudo-label sample set based on the unlabeled data set and the first scoring model.
[0007] According to a third aspect, a scoring system is provided, the system comprising: a data generation unit, for generating an active learning annotated data set; an automatic scoring unit, connected to the data generation unit, for implementing a method as described in any implementation manner of the first aspect according to the active learning annotated data set.
[0008] According to a fourth aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any implementation manner of the first aspect.
[0009] According to a fifth aspect, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method described in any implementation of the first aspect.
[0010] The data labeling method provided by the embodiment of the present disclosure first obtains an active learning labeling data set; secondly, based on the active learning labeling data set, an initial scoring model is trained to obtain a trained first scoring model; then, based on the active learning labeling data set and the pre-trained large language model, an unlabeled data set is obtained; finally, based on the unlabeled data set and the first scoring model, a pseudo-label sample set is obtained. Thus, through a large language model and semi-supervised technology, limited labeling resources can be effectively utilized to generate a pseudo-label sample set to expand the training data, thereby improving the accuracy of the generation of the pseudo-label sample set. Semi-supervised learning can efficiently utilize limited labeled data and abundant unlabeled data, thereby improving the generalization ability and scoring accuracy of the first scoring model; and the large language model ensures the quality, diversity and consistency of the generated pseudo-label sample set, so that the pseudo-label sample set provides a solid foundation for retraining the first scoring model.
[0011] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0013] Figure 1 is a flow chart of an embodiment of a data annotation method according to the present disclosure;
[0014] Figure 2 is a structural schematic diagram obtained from the pseudo-label sample set in the present disclosure;
[0015] Figure 3is a schematic diagram of updating the scoring model of multiple training stages in the present disclosure;
[0016] Figure 4 It is a training diagram of the scoring model at each stage presented in the form of data in the present disclosure;
[0017] Figure 5 It is a structural schematic diagram obtained by actively learning and labeling data sets in the present disclosure;
[0018] Figure 6 is a structural schematic diagram of an embodiment of a data annotation device according to the present disclosure;
[0019] Figure 7 is a schematic diagram of a structure of a scoring system according to the present disclosure;
[0020] Figure 8 It is a block diagram of an electronic device used to implement the data labeling method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0021] Unless explicitly stated otherwise, throughout the specification and claims, the term “comprise” or variations such as “include” or “comprising”, etc., will be understood to include the stated elements or components but not to exclude other elements or components.
[0022] The technical solution of the present disclosure is described below by means of specific embodiments. It should be understood that one or more steps mentioned in the present disclosure do not exclude the existence of other methods and steps before and after the combination step, or other methods and steps may be inserted between these explicitly mentioned steps. It should also be understood that these examples are only used to illustrate the present disclosure and are not used to limit the scope of the present disclosure. Unless otherwise specified, the numbering of each method step is only for the purpose of identifying each method step, and does not limit the order of arrangement of each method or limit the scope of implementation of the present disclosure. Changes or adjustments in their relative relationships can also be regarded as the scope of implementation of the present disclosure without substantial changes in the technical content.
[0023] The sources of the raw materials and instruments used in the examples are not particularly limited and can be purchased from the market or prepared according to conventional methods known to those skilled in the art.
[0024] Self-explanation, as an active learning strategy, refers to learners actively explaining, clarifying, and reasoning about the information they have learned during the learning process. This strategy involves learners providing detailed verbal or written explanations of their thinking processes, concepts they understand, problem-solving methods, and problem-solving steps. Through self-explanation, learners are able to better organize and consolidate knowledge, identify gaps in understanding, and promote deeper learning and cognitive development.
[0025] With the widespread popularity of technologies such as digital courseware learning platforms, self-explanation as a learning method has once again become the focus and has been put into practical use. The field of modern learning innovation attaches great importance to the role of self-explanation, by designing more intuitive interfaces, building evaluation models based on self-explanation behavior, and developing strategies that can tap into deep self-explanation. For example, tools developed by traditional technologies highlight the centrality of self-explanation in systematic problem solving. Ongoing research continues to expand the scope of application of self-explanation in education, such as the use of template-based self-explanation methods. These templates provide learners with a pre-set framework, acting as a built-in guide to enhance their explanation and thinking process.
[0026] Furthermore, the practice of self-explanation has expanded beyond traditional boundaries. These methods involve not only conceptual understanding, but also promoting the application of a variety of educational tools, such as building feedback systems, carefully designing answers to practice quizzes, and generating valuable data sets for automated assessments. In this system, the role of automated assessment is crucial. By deeply analyzing and interpreting self-explanations, educators and automated systems can gain more accurate insights into learners' thinking patterns. This valuable information allows them to customize more appropriate educational strategies to meet the learning needs of different individuals. Such insights are crucial for tasks such as classifying learners' responses, making it easy to identify the mistakes that students often make or the topics that they have been struggling to master.
[0027] Self-explanation is an important learning strategy that promotes students’ understanding and knowledge internalization, especially in subjects with strong logic such as mathematics. Related research shows that self-explanation helps:
[0028] 1) Deep understanding: By explaining the solutions to the problems and the principles behind them, students can better grasp the knowledge.
[0029] 2) Problem-solving ability: Self-explanation promotes the explicitness of thinking processes and improves students' ability to analyze and solve problems.
[0030] Assessing the quality of students' self-explanation is an important part of personalized teaching and feedback. Existing technologies mainly focus on manual scoring and partially automated scoring:
[0031] 1) Manual scoring: Education experts or teachers evaluate students’ self-explanations based on preset scoring criteria. However, manual scoring is time-consuming, subject to the subjective influence of the scorer, and has low scoring efficiency.
[0032] 2) Automatic scoring attempts: Some studies have attempted to use NLP (Natural Language Processing) technology and machine learning algorithms to automatically score self-explanations. For example, methods such as text similarity, grammatical analysis, and semantic understanding can be used to evaluate the logic, completeness, and expression quality of self-explanations, but NLP has poor ability to generate self-explanations or text related to self-explanations.
[0033] 3) Generation based on language models: In recent years, the use of LLM (Large Language Model) (such as GPT-3, GPT-4, etc.) to generate high-quality self-explanatory text has become an important means of data enhancement. LLM can generate self-explanatory sentences with reasonable structure, accurate content and diversity based on given math problems or knowledge points.
[0034] Although there are studies using LLM to generate education-related data, high-quality generation of mathematical self-explanations is not yet common, especially generation strategies that focus on diversity and logical consistency.
[0035] In view of the defects in traditional technologies, this paper proposes a data annotation method, a semi-supervised method based on a large language model (LLM). The purpose of this method is to explore the potential of the model in generating active learning texts (such as sentences and articles), which will form the basis of the scoring model for predicting active learning, and by adopting a semi-supervised strategy and utilizing advanced language models, the accuracy and effectiveness of the model's automatic scoring are improved. Figure 1 A process 100 according to an embodiment of a data annotation method of the present disclosure is shown. The data annotation method comprises the following steps:
[0036] Step 101: Obtain an active learning annotation dataset.
[0037] In this embodiment, the active learning annotation data set is the annotation data of the active learning strategy, wherein the active learning strategy is a strategy for learners to actively learn knowledge during the learning process, the active learning annotation data set includes at least one active learning annotation data, the active learning annotation data includes the annotated self-explanation data, the self-explanation data includes the self-explanation text and the quality score of the self-explanation text, wherein the self-explanation text is the data of the self-explanation strategy, wherein the self-explanation strategy is a strategy for detailed description of one's own thinking process, the concepts understood, the method of solving the problem and the steps of solving the problem, and the self-explanation strategy is an important learning strategy for promoting learning understanding and knowledge internalization; the quality score of the self-explanation text can be evaluated based on multiple core standards, three core standards: logical coherence, clarity of expression and content relevance. The evaluation of multiple core standards is achieved by combining natural language processing technology and machine learning models. These standards do not have completely independent evaluation conditions or algorithms, but are achieved by comprehensive analysis of the semantics, structure and content of the self-explanation text.
[0038] In traditional technologies, the active learning data of active learning strategies require learners to think deeply and describe the problem-solving steps and thinking process in detail, which takes a lot of time and effort. Therefore, the high-quality generation of active learning data for active learning strategies in different subjects is not yet common.
[0039] In this embodiment, the above step 101 includes: obtaining learners' feedback data on different courses during the learning process; inputting the feedback data into a large language model to obtain an active learning annotation data set output by the large language model.
[0040] Step 102: Based on the active learning annotated data set, an initial scoring model is trained to obtain a trained first scoring model.
[0041] In this embodiment, the initial scoring model is trained based on the active learning annotated data set and is used to score the unlabeled data set to generate a model of a pseudo-label sample set. In this process, the main function of the initial scoring model is to generate preliminary labels for the unlabeled data, which will be used to further train or improve the initial scoring model, and the labels can be obtained by the scores of different types of samples output by the initial scoring model, that is, the input of the initial scoring model is the number of active learning annotations, and the output is the scores of different types of samples. When the active learning data set includes self-explanatory data, the initial scoring model can generate a corresponding quality score based on the input self-explanatory text. It should be noted that when the initial scoring model is not trained with an active learning annotated data set, the quality score it outputs may be inaccurate.
[0042] In this embodiment, the active learning annotation data set includes annotated self-explanatory data. Figure 2 As shown, the initial scoring model (not shown in the figure) is trained by actively learning the labeled data set. After the initial scoring model is adjusted or trained and optimized, if the initial scoring model meets the training completion condition, a first scoring model corresponding to the initial scoring model M1 is obtained. The first scoring model M1 can be used to perform quality scoring on the self-explanatory text in the unlabeled data set, thereby generating a pseudo-label sample set B.
[0043] Step 103: obtaining an unlabeled dataset based on the active learning labeled dataset and the pre-trained large language model.
[0044] In this embodiment, the pre-trained large language model refers to a deep learning model trained using a large amount of text data, so that the model can generate natural language text or understand the meaning of language text. These models can provide in-depth knowledge and language about various topics by training on huge data sets.
[0045] In this embodiment, the unlabeled data set is an unlabeled data set related to the active learning strategy and used to train the scoring models of multiple stages (such as the initial scoring model and the first scoring model). Specifically, the unlabeled data set includes at least one unlabeled data, each of which is related to the active learning strategy. For example, each unlabeled data is a self-explanatory text, and the unlabeled data set can be data obtained through data enhancement technology. The active learning labeled data set is the data of the active learning strategy that has been labeled in advance. The active learning labeled data set is used to identify and imitate the content related to the active learning strategy by using a pre-trained large language model, and an unlabeled unlabeled data set can be directly generated.
[0046] In this embodiment, since active learning annotation data is difficult to obtain, the amount of active learning annotation data in the above active learning annotation data set is relatively small. Figure 2 As shown, by pre-training a large language model L, we can directly use the active learning annotation dataset D * 训 By performing data enhancement, we can obtain an unlabeled dataset W with a large amount of data.
[0047] Optionally, in order to increase the data volume of the generated unlabeled data set, the domain knowledge text of the corresponding field can be combined to generate an unlabeled data set, and the domain knowledge text and the active learning annotated data set can be input into a pre-trained large language model to obtain an unlabeled data set output by the pre-trained large language model.
[0048] Step 104: obtaining a pseudo-label sample set based on the unlabeled data set and the first scoring model.
[0049] In this embodiment, the first scoring model is a trained initial scoring model. Since the first scoring model is trained by accurately actively learning and labeling the data set, compared with the initial scoring model, the first scoring model can generate more accurate labels, thereby achieving labeling of unlabeled data sets.
[0050] In this embodiment, the above step 104 includes: inputting each unlabeled data in the unlabeled data set into the first scoring model to obtain a quality score of each unlabeled data output by the first scoring model; based on the quality score of each unlabeled data, obtaining a label for each unlabeled data, and using each unlabeled data in the unlabeled data set and each unlabeled data label as a pseudo-label sample set.
[0051] In this embodiment, the pseudo-label sample set is used to train the scoring model at different training stages. Compared with the active learning data set, the labeling accuracy of each pseudo-label sample in the pseudo-label sample set is lower. The pseudo-label sample set includes at least one pseudo-label sample, and each pseudo-label sample has corresponding text and label.
[0052] like Figure 2 As shown, the first scoring model M1 performs quality scoring on the unlabeled data set W, obtains labels for each unlabeled data in the unlabeled data set based on the scoring, adds corresponding labels to each unlabeled data, and obtains a pseudo-label sample set B.
[0053] In this embodiment, the input and output of the initial scoring model and the first scoring model may be the same. Since the first scoring model is a model obtained by retraining based on the initial scoring model, the accuracy of the quality scoring of the text by the first scoring model is higher than the accuracy of the quality scoring by the initial scoring model.
[0054] In this embodiment, the active learning annotation data set can be a manually annotated data set, and the pseudo-label sample set is an automatically generated data set. The combination of pseudo-label data and manually annotated data expands the training set, improves the performance of the automatic scoring system, and ensures data quality.
[0055] The data labeling method provided by the embodiment of the present disclosure first obtains an active learning labeling data set; secondly, based on the active learning labeling data set, an initial scoring model is trained to obtain a trained first scoring model; then, based on the active learning labeling data set and the pre-trained large language model, an unlabeled data set is obtained; finally, based on the unlabeled data set and the first scoring model, a pseudo-label sample set is obtained. Thus, through a large language model and semi-supervised technology, limited labeling resources can be effectively utilized to generate a pseudo-label sample set to expand the training data, thereby improving the accuracy of the generation of the pseudo-label sample set. Semi-supervised learning can efficiently utilize limited labeled data and abundant unlabeled data, thereby improving the generalization ability and scoring accuracy of the first scoring model; and the large language model ensures the quality, diversity and consistency of the generated pseudo-label sample set, so that the pseudo-label sample set provides a solid foundation for retraining the first scoring model.
[0056] In some embodiments of the present disclosure, the above method also includes: training a first scoring model based on a pseudo-label sample set and an active learning annotation data set; obtaining a trained second scoring model in response to detecting that the performance of the first scoring model has converged or the number of training times of the first scoring model has reached a predetermined number of iterations; and obtaining a new pseudo-label sample set based on the unlabeled data set and the second scoring model.
[0057] In this embodiment, in each iterative training of the first scoring model, data can be selected from the pseudo-label sample set and the active learning annotation data set, the selected data can be input into the first scoring model, the loss value of the first scoring model is calculated, and whether the performance of the first scoring model has converged is detected based on the loss value, wherein the convergence of the performance of the first scoring model can be reflected from various aspects, for example, in the iterative training for a specified number of consecutive times (for example, 100 times), the increase in the loss value of the first scoring model exceeds the preset amplitude.
[0058] In this embodiment, based on the unlabeled data set and the second scoring model, obtaining a new pseudo-label sample set includes: using the second scoring model to perform quality scoring on the unlabeled data set, and based on obtaining the label of each unlabeled data in the unlabeled data set, adding a label to each unlabeled data to obtain a new pseudo-label sample set.
[0059] In this embodiment, the number of training times of the first scoring model refers to the number of iterative training times of the first scoring model, which is also the number of times the loss value of the first scoring model is calculated. During each iterative training of the first scoring model, a sample is input into the first scoring model to obtain a score output by the first scoring model; the loss value of the scoring model is calculated based on the score; the performance convergence of the first scoring model is detected by the loss value, or the training times are detected by the training times to determine whether the training times of the first scoring model have reached a predetermined number of iterations and whether it has reached the training stage of obtaining a second scoring model that has been trained. The predetermined number of iterations of the first scoring model can be set based on development requirements, for example, the predetermined number of iterations is 100,000 times.
[0060] In this embodiment, Figure 2 As shown, the first scoring model M1 performs quality scoring on the unlabeled data set W, obtains labels for each unlabeled data in the unlabeled data set based on the scoring, adds corresponding labels to each unlabeled data, and obtains a pseudo-label sample set B.
[0061] In this embodiment, Figure 4 As shown, the first scoring model M1 is based on the active learning annotation dataset D * 训 The model obtained by training is as follows: Figure 4 In the example, the first scoring model M1 is used in the active learning annotation dataset D * 训 , obtained by multiple iterative training X, and the second scoring model M2 is based on the first scoring model M1 and further uses the pseudo-label sample set D * 训 +D * 样 The model obtained by iterative training X. Since the pseudo-label sample set is generated by predicting the unlabeled data set through the first scoring model, and a certain amount of noise and uncertainty will be introduced when generating pseudo-labels, the training data of the second scoring model is expanded and enhanced compared with the training data of the first scoring model. It should be noted that the first scoring model M1 predicts Y on a type of data to obtain the unlabeled sample D 样 , by using the unlabeled samples D 样 Mark B and get the marked sample D * 样 , for unlabeled samples D 样 The second scoring model M2 predicts another type of data Y to obtain unlabeled samples D 样 , by using the unlabeled samples D 样 Mark B and get the marked sample D * 样 .
[0062] In practical applications, the performance of the second scoring model is generally better than that of the first scoring model because its training data is richer, including more manually annotated data and generated pseudo-labeled data. Therefore, the performance of the second scoring model is improved compared to the first scoring model.
[0063] In this embodiment, the inputs of the initial scoring model, the first scoring model, and the second scoring model in each training stage may be the same, and the outputs of the initial scoring model, the first scoring model, and the second scoring model in each training stage may also be the same. Since the second scoring model in each training stage is a model obtained by retraining on the basis of the first scoring model, the accuracy of the quality scoring of the text by the second scoring model in each training stage is higher than the accuracy of the quality scoring by the first scoring model and the initial scoring model.
[0064] Optionally, the initial scoring model is mainly used for scoring tasks, while the first scoring model and the second scoring model in each training stage can be mainly used to generate pseudo-label sample sets. To this end, the input and output of the initial scoring model can also be different from the first scoring model and the second scoring model, that is, the initial scoring model is used to output the score of the sample data, while the first scoring model and the second scoring model are used to generate pseudo-label data.
[0065] The data labeling method provided in this embodiment trains a first scoring model based on a pseudo-label sample set and an active learning labeled data set; in response to detecting that the performance of the first scoring model has converged or the number of training times of the first scoring model has reached a predetermined number of iterations, a trained second scoring model is obtained; based on the unlabeled data set and the second scoring model, a new pseudo-label sample set is obtained, and the pseudo-label sample set and the active learning labeled data set are used to improve the generalization of the second scoring model obtained by the first scoring model, thereby improving the accuracy of the new pseudo-label sample set.
[0066] Optionally, the method further includes: in response to detecting that the performance of the first scoring model has not converged and the number of training times of the first scoring model has not reached a predetermined number of iterations, continuing to train the first scoring model.
[0067] In some optional implementations of the present disclosure, the training of the first scoring model based on the pseudo-label sample set and the active learning annotation data set includes: selecting pseudo-label samples from the pseudo-label sample set; inputting the selected pseudo-label samples into the first scoring model to obtain a scoring result of the first scoring model; based on the scoring result of the first scoring model, detecting whether the performance of the first scoring model has converged or whether the number of training times of the first scoring model has reached a predetermined number of iterations.
[0068] In this optional implementation, there are many ways to select pseudo-label samples from the pseudo-label sample set, such as random selection, sequential selection, etc. The scoring result of the first scoring model is the result of the first scoring model performing quality scoring on the selected pseudo-label samples.
[0069] In this optional implementation, when training the first scoring model, whether the performance of the first scoring model has converged can be detected in a variety of ways, such as detecting whether the performance of the first scoring model has converged by the performance of the loss value of the first scoring model in multiple iterative trainings.
[0070] The method for training the first scoring model provided by this optional implementation manner selects pseudo-label samples from a pseudo-label sample set; inputs the selected pseudo-label samples into the first scoring model to obtain a scoring result of the first scoring model; based on the scoring result of the first scoring model, detects whether the performance of the first scoring model converges or whether the number of training times of the first scoring model reaches a predetermined number of iterations, thereby training the first scoring model through the convergence of the performance of the first scoring model and the number of training times of the first scoring model, thereby improving the reliability of the training of the first scoring model.
[0071] In some optional implementations of the present disclosure, the above-mentioned detecting whether the performance of the first scoring model converges or whether the training number of the first scoring model reaches a predetermined number of iterations based on the scoring result of the first scoring model includes: calculating the loss value of the first scoring model based on the scoring result of the first scoring model, and recording the training number of the first scoring model; detecting whether the loss value of the first scoring model is less than a first loss value threshold; in response to detecting that the loss value of the first scoring model is less than the first loss value threshold, determining that the performance of the first scoring model converges; in response to detecting that the loss value of the first scoring model is greater than the first loss value threshold, detecting whether the training number of the first scoring model reaches a predetermined number of iterations.
[0072] In this optional implementation, the loss value of the first scoring model may be calculated using a loss function set for the first scoring model, such as a cross entropy function.
[0073] In this optional implementation, the number of training times can be recorded by a counter. For example, the number of training times of the first scoring model is recorded each time the loss value of the first scoring model is calculated. Alternatively, the number of training times can be recorded by a counter after the scoring result of the first scoring model is obtained.
[0074] The method for detecting the first scoring model provided by this optional implementation calculates the loss value of the first scoring model based on the scoring result of the first scoring model, and records the number of training times of the first scoring model; detects whether the loss value of the first scoring model is less than a first loss value threshold; in response to detecting that the loss value of the first scoring model is less than the first loss value threshold, determines that the performance of the first scoring model has converged; in response to detecting that the loss value of the first scoring model is greater than the first loss value threshold, detects whether the number of training times of the first scoring model reaches a predetermined number of iterations, thereby improving the reliability of the training of the first scoring model.
[0075] Optionally, the above-mentioned detecting whether the performance of the first scoring model converges or whether the training number of the first scoring model reaches a predetermined number of iterations based on the scoring result of the first scoring model includes: calculating the loss value of the first scoring model based on the scoring result of the first scoring model, and recording the training number of the first scoring model; detecting whether the training number of the first scoring model reaches a predetermined number of iterations, and in response to detecting that the training number of the first scoring model does not reach the predetermined number of iterations, detecting whether the loss value of the first scoring model is less than a first loss value threshold; in response to detecting that the loss value of the first scoring model is less than the first loss value threshold, determining that the performance of the first scoring model converges.
[0076] In some optional implementations of the present disclosure, the above method also includes: training a second scoring model for the next training stage based on the new pseudo-label sample set and the second scoring model; obtaining an updated pseudo-label sample set based on the second scoring model for the next stage and the new pseudo-label sample set; continuing to train the second scoring model for other stages based on the updated pseudo-label sample set, and updating the pseudo-label sample sets for each training stage after each training stage.
[0077] In this embodiment, each training stage corresponds to a second scoring model, such as Figure 4 In the above example, the second scoring models corresponding to each training stage include: M3 128 、M3 265 、M3 k 、M3 4096 , where k is a natural number greater than zero, and the second scoring model in the current training stage is the pseudo-label sample set given by the first scoring model or the second scoring model in the previous training stage (such as Figure 4 The pseudo-label sample set D of each stage in * 128 , D * 265 ,…,D * k , D * 4096) is obtained through basic training. A second scoring model is obtained after each training stage, and each training stage has a fixed set of pseudo-label samples.
[0078] like Figure 3 As shown, in the current training phase, the pseudo-label sample set B is used t And the active learning annotation dataset D * 训 Train to get the second scoring model M t+1 , using the second scoring model M t+1 Instead of the first scoring model, the unlabeled data set W is scored, and the label of each unlabeled data in the unlabeled data set is obtained based on the score, and a corresponding label is added to each unlabeled data to obtain a new pseudo-label sample set.
[0079] The data labeling method provided in this embodiment trains the second scoring model of the next training stage based on the new pseudo-label sample set and the second scoring model; obtains an updated pseudo-label sample set based on the second scoring model of the next stage and the new pseudo-label sample set; continues to train the second scoring model of other stages based on the updated pseudo-label sample set, and updates the pseudo-label sample set of each training stage after each training stage. Thus, the second scoring model of the next training stage is trained based on the updated pseudo-label sample set obtained from the second scoring model of the previous training stage, and the pseudo-label sample set of each training stage can be optimized to the maximum extent.
[0080] In some optional implementations of the present disclosure, the above method also includes: determining the different amounts of pseudo-label data generated in each training stage based on the updated pseudo-label sample set; obtaining the second scoring model for each training stage; determining the prediction accuracy of the second scoring model for each training stage; and taking the amount of pseudo-label data of the second scoring model in the training stage with the highest prediction accuracy as the optimal amount.
[0081] In this embodiment, the purpose of determining the "optimal number" is to find a balance point so that the generated pseudo self-explanatory data can maximize the prediction performance of the second scoring model while avoiding the introduction of noise due to excessive data volume or the inability to fully utilize the potential of the generated data due to too small data volume. Determining an "optimal number" through experiments can provide a clear guide for practical applications to make optimal decisions during data generation and model training.
[0082] In this embodiment, the number of pseudo-label data in the pseudo-label sample set of each training stage can be directly obtained by statistics, and the prediction accuracy of the second scoring model of each training stage can be calculated by the test sample set after the training of the second scoring model of each stage is completed. Specifically, the prediction results of the second scoring model of each stage are compared with the true labels of the test data set, and different evaluation indicators such as accuracy, precision, and recall can be used to calculate the prediction accuracy of the model. Commonly used evaluation indicators include mean square error (MSE), root mean square error (RMSE), mean absolute error (MAE), etc. The smaller these indicators are, the higher the prediction accuracy of the model is.
[0083] The method for obtaining the optimal number provided in this embodiment determines the different numbers of pseudo-label data generated in each training stage based on the updated pseudo-label sample set; obtains the second scoring model in each training stage; determines the prediction accuracy of the second scoring model in each training stage; and takes the amount of pseudo-label data of the second scoring model in the training stage with the highest prediction accuracy as the optimal number, which can provide a clear guide for practical applications and make the best decision in the data generation and model training process.
[0084] Optionally, the data labeling method further comprises: after determining the optimal number, determining a second scoring model with the optimal number, and using the second scoring model with the optimal number to evaluate the test data set D 测 Make the final prediction C and get the final pseudo-label sample set D * 测, like Figure 4 shown.
[0085] In some optional implementations of the present disclosure, the above method further includes: detecting whether to continue to increase the amount of pseudo-label data according to the performance of the second scoring model on the validation set at each training stage. For example, when the performance improvement of the second scoring model on the validation set slows down, stop increasing the amount of generated data; or when the performance begins to decline, reduce the proportion of generated data. This dynamic adjustment method can avoid manually finding the "optimal number" and may obtain better model performance.
[0086] In addition to the optimal quantity, the quality of the data is also a key factor affecting model performance. Optionally, the above method also includes: ensuring the generated data has high quality through more sophisticated generation strategies, stricter quality control mechanisms, or introducing expert review. As long as the quality is guaranteed, there is no need to generate excessive amounts of data to improve model performance.
[0087] Data enhancement is another direction worth exploring. Optionally, the above method also includes: increasing the diversity of data by performing different transformations on the generated data (such as data mixing, data perturbation, etc.) instead of simply increasing the amount of data. This diverse data may be more helpful for the generalization ability of the model.
[0088] Active learning is a learning method that can effectively utilize limited labeled data. Optionally, the above method also includes: controlling the second scoring model to select those generated data that it considers to be the most difficult to predict or the most uncertain for training, so that the generated data can be utilized more effectively.
[0089] In some optional implementations of the present disclosure, the above-mentioned acquisition of the active learning annotation dataset includes: acquiring the active learning strategy text and the title of the active learning strategy text; inputting the active learning strategy text into the first feature extraction layer to obtain text features; inputting the title into the second feature extraction layer to obtain title features; obtaining splicing features based on the text features and the title features; inputting the splicing features into the regression test model to obtain the quality score of the active learning strategy text, and using the active learning strategy text and the quality score as the active learning annotation data in the active learning annotation dataset.
[0090] In this optional implementation, the first feature extraction layer and the second feature extraction layer can be implemented using a BERT (Bidirectional Encoder Representations from Transformers) model, wherein the first feature extraction layer and the second feature extraction layer are both implemented using an embedding layer and an encoding layer in the BERT model. The BERT model is based on an advanced transformer structure and has demonstrated superior performance that exceeds previous models in a variety of natural language processing tasks.
[0091] In this optional implementation, the regression test model is a regression layer that receives the concatenated features from the BERT encoder and further processes them to predict the quality score of the self-explanatory text. Specifically, this regression layer can be a fully connected layer (FC layer), which maps the output features of BERT to specific score values through linear transformation and activation function. Therefore, the regression test model is a component independent of BERT and is used for the final regression prediction task. Due to the excellent performance of the regression test model and its good adaptability to Chinese, the regression test model can directly receive the preprocessed self-explanatory text and the corresponding test title as input, and then output the quality score of each self-explanatory text.
[0092] In this optional implementation, the active learning strategy text may include a self-explanatory text, and the title is the title of the active learning strategy text. The test title comes from the math test questions in the digital courseware learning platform. During the data collection phase, not only the self-explanatory texts of scholars were collected, but also the corresponding test questions were recorded. These test questions are titles, which provide contextual information about the specific problems that the self-explanation targets. Therefore, during the model training and prediction process, the self-explanatory text and its corresponding title are input into the regression test model together to make full use of the question information to improve the accuracy of the score.
[0093] In this optional implementation, Figure 5 The first feature extraction layer in can be the embedding layer and encoder layer of the BERT model. The BERT model first embeds the input text (i.e., self-explanatory text) and then extracts context features through a multi-layer Transformer encoder to obtain text features. Figure 5 The second feature extraction layer in can be the embedding layer and encoder layer of the BERT model. The BERT model first embeds the input text (i.e., title), and then extracts context features through a multi-layer Transformer encoder to obtain title features.
[0094] exist Figure 5 In the figure, the text feature is the feature representation of the self-explanatory text after being embedded and encoded by the BERT model, and the title feature is the feature representation of the title after being embedded and encoded by the BERT model. The concatenated feature represents the feature after the text feature and the title feature are concatenated. The concatenated feature can combine the feature information of the self-explanatory text and the title. The concatenated feature is used in the subsequent regression test model to predict the quality score of the self-explanatory text.
[0095] This optional implementation provides a method for obtaining an active learning annotation dataset, obtaining an active learning strategy text and a title of the active learning strategy text; inputting the active learning strategy text into a first feature extraction layer to obtain text features; inputting the title into a second feature extraction layer to obtain title features; obtaining concatenation features based on text features and title features; inputting the concatenation features into a regression test model to obtain a quality score of the active learning strategy text, and using the active learning strategy text and the quality score as active learning annotation data in an active learning annotation dataset. This design can make full use of the contextual information of the self-explanatory text and the title, thereby improving the accuracy of the quality score.
[0096] In some optional implementations of the present disclosure, the above-mentioned training of the initial scoring model based on the active learning annotation data set to obtain the trained first scoring model includes: selecting active learning annotation data from the active learning annotation data set; inputting the selected active learning annotation data into the initial scoring model to obtain a scoring result of the initial scoring model; based on the scoring result, detecting whether the initial scoring model meets the training completion condition; in response to detecting that the initial scoring model meets the training completion condition, obtaining the trained first scoring model.
[0097] In this optional implementation, there are many ways to select active learning annotation data from the active learning annotation data set, such as random selection, sequential selection, etc. The scoring result of the initial scoring model is the result of the initial scoring model scoring the quality of the selected pseudo-label samples.
[0098] In this optional implementation, when training the first scoring model, it is possible to detect whether the initial scoring model meets the training completion conditions in a variety of ways, such as detecting whether the first scoring model meets the training completion conditions by the performance of the loss value of the initial scoring model in multiple iterations of training.
[0099] The method for obtaining a trained first scoring model provided by this optional implementation manner comprises selecting active learning annotation data from an active learning annotation data set; inputting the selected active learning annotation data into an initial scoring model to obtain a scoring result of the initial scoring model; based on the scoring result, detecting whether the initial scoring model satisfies a training completion condition; in response to detecting that the initial scoring model satisfies the training completion condition, obtaining a trained first scoring model, thereby improving the reliability of obtaining the first scoring model.
[0100] In some optional implementations of the present disclosure, the above-mentioned training completion condition includes: the prediction accuracy of the initial scoring model on the validation set is greater than or equal to a predetermined accuracy threshold.
[0101] In this optional implementation, the validation set is a sample set selected for the initial scoring model. The prediction results of the initial scoring model are compared with the true labels of the validation set. Different evaluation indicators, such as accuracy, precision, and recall, can be used to calculate the prediction accuracy of the model.
[0102] In this optional implementation, the predetermined accuracy threshold may be set based on development requirements, for example, the predetermined accuracy threshold is 89%.
[0103] This optional implementation provides a training completion condition, which can effectively verify whether the first scoring model has been trained, thereby improving the reliability of the first scoring model.
[0104] In some optional implementations of the present disclosure, the above-mentioned obtaining of the unlabeled dataset based on the active learning labeled dataset and the pre-trained large language model includes: selecting a random dataset based on the active learning labeled dataset; extracting active learning keywords based on the random dataset; inputting the active learning keywords and guiding prompt words into the pre-trained large language model to obtain the unlabeled dataset output by the pre-trained large language model.
[0105] In this optional implementation, randomly selecting data means first randomly selecting 30% of the data from the manually labeled training data set to fully utilize the rich diversity of students' self-explanations.
[0106] Keyword extraction refers to extracting ten keywords from each self-explanation, which capture its core meaning and guide the pre-trained large language model to generate context-related data.
[0107] Pre-trained large language model generates data: Using the extracted keywords, prompts are given to the pre-trained large language model. Specifically, each group of 10 keywords is used as the initial input to guide the pre-trained large language model to generate context-consistent pseudo self-explanation data, where the prompt word can be "elaborate based on the provided keywords" to ensure that the generated content remains relevant to the context of the original self-explanation.
[0108] The method for obtaining an unlabeled data set provided by this optional implementation method selects a random data set based on an active learning labeled data set; extracts active learning keywords based on the random data set; inputs the active learning keywords and guiding prompt words into a pre-trained large language model to obtain an unlabeled data set output by the pre-trained large language model, thereby improving the quality of the unlabeled data set.
[0109] In some optional implementations of the present disclosure, the above-mentioned obtaining of a pseudo-label sample set based on an unlabeled data set and a first scoring model includes: inputting each unlabeled data in the unlabeled data set into the first scoring model to obtain a predicted quality score of each unlabeled data output by the first scoring model; obtaining at least one to-be-labeled data based on the predicted quality score of each unlabeled data; and adding respective predicted quality scores to each to-be-labeled data in the at least one to-be-labeled data to obtain a pseudo-label sample set.
[0110] In this optional implementation, the predicted quality score is the quality score of each unlabeled data. The above-mentioned obtaining at least one data to be labeled based on the predicted quality score of each unlabeled data includes: in response to the predicted quality score being located in a score segment among different score segments, querying the label corresponding to the score segment, and adding the label to the unlabeled data.
[0111] The method for obtaining a pseudo-label sample set provided by this optional implementation inputs each unlabeled data in the unlabeled data set into a first scoring model to obtain a predicted quality score of each unlabeled data output by the first scoring model; based on the predicted quality score of each unlabeled data, at least one data to be labeled is obtained; and each to-be-labeled data in the at least one to-be-labeled data is added with its own predicted quality score to obtain a pseudo-label sample set, thereby improving the reliability of obtaining the pseudo-label sample set.
[0112] In some optional implementations of the present disclosure, the above-mentioned obtaining at least one data to be labeled based on the predicted quality scores of each unlabeled data includes: comparing the predicted quality scores of each unlabeled data with a predetermined threshold to obtain a comparison result; and taking the unlabeled data whose predicted quality scores in the comparison result are greater than or equal to the predetermined threshold as the data to be labeled.
[0113] In this optional implementation, the data to be labeled is data for which labeled samples are to be generated. The data to be labeled may be data that meets sample requirements but is not labeled accordingly. The predetermined threshold may be set based on development requirements.
[0114] This optional implementation provides a method for obtaining data to be labeled, comparing the predicted quality scores of each unlabeled data with a predetermined threshold to obtain a comparison result; the unlabeled data with a predicted quality score greater than or equal to the predetermined threshold in the comparison result is used as the data to be labeled, thereby improving the reliability of obtaining the data to be labeled.
[0115] Further references Figure 6 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a data annotation device. Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0116] like Figure 6 As shown, the data labeling device 600 provided in this embodiment includes: an acquisition unit 601, a model acquisition unit 602, a data acquisition unit 603, and a labeling unit 604. Among them, the above-mentioned acquisition unit 601 can be configured to be configured to acquire an active learning labeling data set. The above-mentioned model acquisition unit 602 can be configured to train an initial scoring model based on the active learning labeling data set to obtain a trained first scoring model. The above-mentioned data acquisition unit 603 can be configured to obtain an unlabeled data set based on the active learning labeling data set and a pre-trained large language model. The above-mentioned labeling unit 604 can be configured to obtain a pseudo-label sample set based on the unlabeled data set and the first scoring model.
[0117] In this embodiment, in the data labeling device 600, the specific processing of the acquisition unit 601, the model acquisition unit 602, the data acquisition unit 603, and the labeling unit 604 and the technical effects thereof can be referred to respectively. Figure 1 The relevant descriptions of step 101, step 102, step 103, and step 104 in the corresponding embodiment are not repeated here.
[0118] In some embodiments of the present disclosure, the above-mentioned data labeling device 600 also includes: obtaining a new pseudo-label sample set unit (not shown in the figure), and the new pseudo-label sample set unit is configured to: train a first scoring model based on the pseudo-label sample set and the active learning labeling data set; in response to detecting that the performance of the first scoring model converges or the training number of the first scoring model reaches a predetermined number of iterations, obtain a trained second scoring model; based on the unlabeled data set and the second scoring model, obtain a new pseudo-label sample set.
[0119] In some embodiments of the present disclosure, the above-mentioned device also includes a first model training unit (not shown in the figure), and the first model training unit is configured to: train a first scoring model based on a pseudo-label sample set and an active learning annotation data set; obtain a trained second scoring model in response to detecting that the performance of the first scoring model converges or the number of training times of the first scoring model reaches a predetermined number of iterations; obtain a new pseudo-label sample set based on the unlabeled data set and the second scoring model.
[0120] In some embodiments of the present disclosure, the first model training unit is configured to: select pseudo-label samples from a pseudo-label sample set; input the selected pseudo-label samples into the first scoring model to obtain a scoring result of the first scoring model; and based on the scoring result of the first scoring model, detect whether the performance of the first scoring model converges or whether the number of training times of the first scoring model reaches a predetermined number of iterations.
[0121] In some embodiments of the present disclosure, the first model training unit is configured to: calculate the loss value of the first scoring model based on the scoring result of the first scoring model, and record the number of training times of the first scoring model; detect whether the loss value of the first scoring model is less than a first loss value threshold; in response to detecting that the loss value of the first scoring model is less than the first loss value threshold, determine that the performance of the first scoring model has converged; in response to detecting that the loss value of the first scoring model is greater than the first loss value threshold, detect whether the number of training times of the first scoring model reaches a predetermined number of iterations.
[0122] In some embodiments of the present disclosure, the above-mentioned device 600 also includes: a second model training unit (not shown in the figure), and the above-mentioned second model training unit is configured to: train the second scoring model of the next training stage based on the new pseudo-label sample set and the second scoring model; obtain an updated pseudo-label sample set based on the second scoring model of the next stage and the new pseudo-label sample set; continue to train the second scoring model of other stages based on the updated pseudo-label sample set, and update the pseudo-label sample set of each training stage after the end of each training stage.
[0123] In some embodiments of the present disclosure, the above-mentioned device 600 also includes a quantity acquisition unit (not shown in the figure), and the quantity acquisition unit is configured to: determine the different quantities of pseudo-label data generated in each training stage based on the updated pseudo-label sample set; obtain the second scoring model of each training stage; determine the prediction accuracy of the second scoring model of each training stage; and take the amount of pseudo-label data of the second scoring model of the training stage with the highest prediction accuracy as the optimal quantity.
[0124] In some embodiments of the present disclosure, the above-mentioned acquisition unit 601 is configured to: obtain the active learning annotation data unit is configured to: obtain the active learning strategy text and the title of the active learning strategy text; input the active learning strategy text into the first feature extraction layer to obtain text features; input the title into the second feature extraction layer to obtain title features; obtain splicing features based on text features and title features; input the splicing features into the regression test model to obtain the quality score of the active learning strategy text, and use the active learning strategy text and the quality score as the active learning annotation data in the active learning annotation dataset.
[0125] In some embodiments of the present disclosure, the above-mentioned model obtaining unit 602 is configured to: select active learning annotation data from the active learning annotation data set; input the selected active learning annotation data into the initial scoring model to obtain the scoring result of the initial scoring model; based on the scoring result, detect whether the initial scoring model meets the training completion condition; in response to detecting that the initial scoring model meets the training completion condition, obtain a first scoring model that has completed training.
[0126] In some embodiments of the present disclosure, the above-mentioned training completion condition includes: the prediction accuracy of the initial scoring model on the validation set is greater than or equal to a predetermined accuracy threshold.
[0127] In some embodiments of the present disclosure, the data obtaining unit 603 is configured to: select a random data set based on an active learning annotated data set; extract active learning keywords based on the random data set; input the active learning keywords and guiding prompt words into a pre-trained large language model to obtain an unlabeled data set output by the pre-trained large language model.
[0128] In some embodiments of the present disclosure, the above-mentioned labeling unit 604 is configured to: input each unlabeled data in the unlabeled data set into the first scoring model to obtain the prediction quality score of each unlabeled data output by the first scoring model; obtain at least one data to be labeled based on the prediction quality score of each unlabeled data; add the respective prediction quality score to each data to be labeled in the at least one data to be labeled to obtain a pseudo-label sample set.
[0129] In some embodiments of the present disclosure, the above-mentioned labeling unit 604 is configured to: compare the prediction quality score of each unlabeled data with a predetermined threshold to obtain a comparison result; and take the unlabeled data in the comparison result whose prediction quality score is greater than or equal to the predetermined threshold as the data to be labeled.
[0130] The data annotation device provided by the embodiment of the present disclosure, first, the acquisition unit 601 acquires the active learning annotation data set; secondly, the model acquisition unit 602 trains the initial scoring model based on the active learning annotation data set to obtain the trained first scoring model; then, the data acquisition unit 603 obtains the unlabeled data set based on the active learning annotation data set and the pre-trained large language model; finally, the annotation unit 604 obtains the pseudo-label sample set based on the unlabeled data set and the first scoring model. Thus, through the large language model and semi-supervised technology, limited annotation resources can be effectively utilized to generate pseudo-label sample sets to expand training data, thereby improving the accuracy of the generation of pseudo-label sample sets. Semi-supervised learning can efficiently utilize limited labeled data and rich unlabeled data, improve the generalization ability and scoring accuracy of the first scoring model; and the large language model ensures the quality, diversity and consistency of the generated pseudo-label sample set, so that the pseudo-label sample set provides a solid foundation for retraining the first scoring model.
[0131] Further references Figure 7 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a scoring system. Figure 1 The method embodiments shown correspond.
[0132] like Figure 7 As shown, the scoring system 700 provided in this embodiment includes: a data generation unit 701, and an automatic scoring unit 702. The data generation unit 701 is used to generate an active learning annotation data set; the automatic scoring unit 702 is connected to the data generation unit 701, and is used to implement the data annotation method of the above embodiment according to the active learning annotation data set.
[0133] In this embodiment, the active learning annotation data set generated by the data generation unit 701 can be an annotation data set obtained manually by annotating initial data. Optionally, the active learning annotation data set can also be a small amount of annotation data generated by a machine. After the machine generates the active annotation data set, the active learning annotation data set can be quality evaluated. The quality evaluation is mainly based on three core standards: logical coherence, clarity of expression, and content relevance. Specifically, "logical coherence" is used to measure the orderliness of the explanation; "clarity of expression" is used to evaluate the comprehensibility of the explanation; and "content relevance" ensures that all relevant knowledge points and procedural details are included in the explanation. In order to further ensure the consistency of quality evaluation, when the active learning annotation data set is obtained through self-explanatory text, a consistency scoring standard and definition can be used to perform an overall evaluation of the following steps Step_1-Step Step_5. This overall evaluation can be applicable to tasks with multiple solutions or strategies.
[0134] Steps Step 1 First, read the entire self-explanation text to get a preliminary impression of the learner's overall thinking and understanding. The focus of this stage is to grasp whether the learner's overall thinking is clear, whether he can accurately understand the meaning of the question, and whether he can explain it in an organized manner.
[0135] Step 2: Based on the overall reading, focus on the logic of the self-explanatory text. Specifically, it will assess whether the learner can explain in a reasonable logical order, whether the parts are connected smoothly, and whether there are logical errors or loopholes. For example, whether the learner can correctly draw the conclusion, or whether there are contradictions in the explanation process.
[0136] Step 3 The assessor will assess whether the learner uses clear and understandable language in self-explanation and whether he or she can accurately express his or her ideas. This stage will focus on whether the learner uses appropriate professional terms, whether he or she can explain complex issues in concise and clear language, and whether he or she avoids ambiguous expressions.
[0137] Step 4 The evaluator will check whether the learner's self-explanation covers all the key knowledge points and problem-solving steps involved in the question and whether it can answer the question comprehensively. For example, whether the learner has omitted important problem-solving steps or deviated from the core requirements of the question.
[0138] Step_5 Based on the evaluation of the above three dimensions, the evaluator will comprehensively consider the learner's overall performance and give a comprehensive score. When scoring, the pre-set scoring criteria and definitions will be referred to to ensure the objectivity and consistency of the scoring.
[0139] The present disclosure adopts a more holistic evaluation method for active learning annotated data, and comprehensively understands the learner's understanding of a specific topic through an overall evaluation of each test.
[0140] In this embodiment, in the scoring system 700, the specific processing of the automatic scoring unit 702 and the technical effects thereof can be referred to in Figure 1 The relevant descriptions of step 101, step 102, step 103, and step 104 in the corresponding embodiment are not repeated here.
[0141] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0142] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their modes are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0143] like Figure 8 As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0144] A number of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0145] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as data annotation methods. For example, in some embodiments, the data annotation method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the data annotation method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the data annotation method in any other appropriate manner (e.g., by means of firmware).
[0146] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0147] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data annotation device, so that the program code, when executed by the processor or controller, implements the modes / operations specified in the flow chart and / or block diagram. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0148] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0149] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0150] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0151] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0152] The foregoing description of specific exemplary embodiments of the present disclosure is for the purpose of illustration and demonstration. These descriptions are not intended to limit the present disclosure to the precise form disclosed, and it is clear that many changes and variations can be made based on the above teachings. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the present disclosure and its practical application, so that those skilled in the art can realize and utilize various different exemplary embodiments of the present disclosure and various different selections and changes. The scope of the present disclosure is intended to be defined by the claims and their equivalents.
Claims
1. A data annotation method, the method comprising: Get active learning annotated dataset; Based on the active learning annotated data set, an initial scoring model is trained to obtain a trained first scoring model; Based on the active learning labeled dataset and the pre-trained large language model, an unlabeled dataset is obtained; Based on the unlabeled data set and the first scoring model, a pseudo-label sample set is obtained.
2. The method according to claim 1, further comprising: Training the first scoring model based on the pseudo-label sample set and the active learning annotation dataset; In response to detecting that the performance of the first scoring model has converged or the number of training times of the first scoring model has reached a predetermined number of iterations, obtaining a trained second scoring model; Based on the unlabeled data set and the second scoring model, a new pseudo-label sample set is obtained.
3. According to the method of claim 2, the training of the first scoring model based on the pseudo-label sample set and the active learning annotation dataset comprises: Selecting a pseudo-label sample from the pseudo-label sample set; Inputting the selected pseudo-label samples into the first scoring model to obtain a scoring result of the first scoring model; Based on the scoring result of the first scoring model, it is detected whether the performance of the first scoring model converges or whether the number of training times of the first scoring model reaches a predetermined number of iterations.
4. The method according to claim 3, wherein: The detecting, based on the scoring result of the first scoring model, whether the performance of the first scoring model converges or whether the number of training times of the first scoring model reaches a predetermined number of iterations comprises: Based on the scoring result of the first scoring model, calculating the loss value of the first scoring model, and recording the number of training times of the first scoring model; Detecting whether the loss value of the first scoring model is less than a first loss value threshold; In response to detecting that the loss value of the first scoring model is less than a first loss value threshold, determining that the performance of the first scoring model has converged; In response to detecting that the loss value of the first scoring model is greater than the first loss value threshold, detecting whether the number of training times of the first scoring model reaches a predetermined number of iterations.
5. The method according to claim 2, further comprising: Based on the new pseudo-label sample set and the second scoring model, training a second scoring model for the next training phase; Based on the second scoring model of the next stage and the new pseudo-label sample set, an updated pseudo-label sample set is obtained; Based on the updated pseudo-label sample set, the second scoring model of other stages is continuously trained, and after each training stage is finished, the pseudo-label sample set of each training stage is updated.
6. The method according to claim 5, further comprising: Based on the updated pseudo-label sample set, determine the different amount of pseudo-label data generated in each training phase; Obtaining a second scoring model at each training stage; determining the prediction accuracy of the second scoring model at each training stage; The amount of pseudo-labeled data of the second scoring model in the training phase with the highest prediction accuracy was taken as the optimal amount.
7. The method according to any one of claims 1 to 6, wherein: The obtaining of the active learning annotation data set comprises: Obtaining an active learning strategy text and a title of the active learning strategy text; Inputting the active learning strategy text into a first feature extraction layer to obtain text features; Inputting the title into a second feature extraction layer to obtain title features; Based on the text feature and the title feature, obtaining a splicing feature; The splicing features are input into a regression test model to obtain a quality score of the active learning strategy text, and the active learning strategy text and the quality score are used as active learning annotation data in an active learning annotation data set.
8. The method according to any one of claims 1 to 6, wherein: The training of the initial scoring model based on the active learning annotated data set to obtain the trained first scoring model comprises: Selecting active learning annotated data from the active learning annotated data set; Inputting the selected active learning annotated data into the initial scoring model to obtain a scoring result of the initial scoring model; Based on the scoring result, detecting whether the initial scoring model meets the training completion condition; In response to detecting that the initial scoring model meets the training completion condition, a first scoring model that has completed training is obtained.
9. The method according to claim 8, wherein: The training completion conditions include: The prediction accuracy of the initial scoring model on the validation set is greater than or equal to a predetermined accuracy threshold.
10. The method according to any one of claims 1 to 6, wherein: The unlabeled data set obtained based on the active learning labeled data set and the pre-trained large language model includes: Based on the active learning labeled data set, selecting a random data set; Extracting active learning keywords based on the random data set; The active learning keywords and guiding prompt words are input into a pre-trained large language model to obtain an unlabeled data set output by the pre-trained large language model.
11. The method according to any one of claims 1 to 6, wherein: The obtaining of a pseudo-label sample set based on the unlabeled data set and the first scoring model comprises: Inputting each unlabeled data in the unlabeled data set into the first scoring model to obtain a prediction quality score of each unlabeled data output by the first scoring model; Based on the predicted quality scores of each unlabeled data, at least one to-be-labeled data is obtained; A respective prediction quality score is added to each of the at least one to-be-labeled data to obtain a pseudo-label sample set.
12. The method according to claim 11, wherein: The obtaining of at least one to-be-labeled data based on the predicted quality scores of each unlabeled data includes: Compare the prediction quality scores of each unlabeled data with a predetermined threshold to obtain a comparison result; The unlabeled data with a prediction quality score greater than or equal to a predetermined threshold in the comparison result is used as the data to be labeled.
13. A data annotation device, comprising: An acquisition unit, configured to acquire an active learning annotated dataset; A model obtaining unit is configured to train an initial scoring model based on the active learning annotated data set to obtain a trained first scoring model; A data acquisition unit is configured to acquire an unlabeled data set based on the active learning labeled data set and the pre-trained large language model; The labeling unit is configured to obtain a pseudo-label sample set based on the unlabeled data set and the first scoring model.
14. A scoring system, comprising: A data generation unit, used to generate an active learning labeled data set; An automatic scoring unit, connected to the data generating unit, is used to implement the data labeling method according to any one of claims 1 to 12 based on the active learning labeling data set.
15. An electronic device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 12.
16. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Question type vertical field literature retrieval method and system based on semi-supervised learning
CN116775883A
Question and answer data construction method and device based on large language model
CN117591661A
Follow-up visit data acquisition method and system based on large language model and knowledge distillation
CN118352097A
Learning method based on fusion ensemble learning, electronic equipment and storage medium
CN118821975A
AI training paradigm based on Personalized Heuristic QA 3D Self-study Method trains AI for Personalized education and General Rational AI System: Hybrid AGRINN (Artificial General Rational Intelligent Neural Network)
US20240395162A1
Cited By
Patent value evaluation method and system based on deep learning
CN120163507A