Active learning model training method and device for constructing dataset
By acquiring a set of keywords in the target domain for literature retrieval and clustering, using a large language model to filter positive and negative samples, constructing a training dataset and optimizing the model, the problems of long time consumption and high cost of traditional methods are solved, and efficient neural network model training is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INST OF SCI & TECHN INFORMATION OF CHINA
- Filing Date
- 2025-10-10
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies, when constructing training datasets for neural network models, are time-consuming and costly using traditional retrieval methods. They also struggle to cover important data patterns and potential correlations in emerging fields and are difficult to construct when sufficient historical data is lacking.
Literature retrieval is performed by obtaining a keyword set in the target field, the literature set is clustered, positive and negative samples are filtered using a large language model, a training dataset is constructed, and the model is optimized.
It achieves efficient literature screening covering the target domain, saves computational costs and time, and improves the model's ability to identify the target domain.
Smart Images

Figure CN121434770B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more specifically, to an active learning model training method, apparatus, electronic device, and storage medium for constructing datasets. Background Technology
[0002] Neural network models are a hot topic in today's society. Neural networks have broad and attractive prospects in fields such as system identification, pattern recognition, and intelligent control. For example, in the field identification of documents, neural network models can not only quickly identify the field to which the documents belong, but also save manpower costs. Therefore, how to train a better neural network model is a topic worth exploring.
[0003] To train a neural network model with strong recognition capabilities, it is essential to train it using a high-quality training dataset. Current technologies often employ traditional retrieval methods to obtain training datasets, requiring domain experts to accurately define the dataset's scope. However, for projects requiring continuous iteration or large-scale data construction, retrieval-based construction is time-consuming and costly. Furthermore, in emerging fields lacking sufficient historical data or mature knowledge systems, retrieval-based construction is highly susceptible to overlooking important data patterns, marginal cases, or potential connections not currently covered by expert knowledge. Summary of the Invention
[0004] The purpose of this application is to at least solve one of the aforementioned technical defects. The technical solution provided by the embodiments of this application is as follows:
[0005] In a first aspect, embodiments of this application provide a model training method, including:
[0006] Obtain a set of keywords related to the target field, and retrieve a set of documents from a pre-defined document database based on the keyword set;
[0007] Cluster the document set to obtain multiple document subsets;
[0008] For each document subset, at least one document is sampled from the document subset level based on a preset ratio and used as the first document.
[0009] For each first document, the probability that the first document belongs to the target domain is predicted by multiple first-largest language models. If the first probability predicted by each first-largest language model is greater than the first threshold, the first document is regarded as the first positive sample. If the first probability predicted by each first-largest language model is less than the second threshold, the first document is regarded as the first negative sample, and the first threshold is greater than the second threshold.
[0010] The target model is obtained by training the initial model based on each first positive sample and first negative sample.
[0011] Secondly, embodiments of this application provide a model training apparatus, including:
[0012] The document retrieval module is used to obtain a set of keywords related to the target field, and retrieve a set of documents from a preset document database based on the keyword set;
[0013] The document clustering module is used to cluster a document set to obtain multiple document subsets;
[0014] The document extraction module is used to sample at least one document from the document subset level based on a preset ratio for each document subset, and use it as the first document.
[0015] The sample selection module is used to predict the first probability that each first document belongs to the target domain through multiple first-largest language models. If the first probability predicted by each first-largest language model is greater than the first threshold, the first document is regarded as the first positive sample. If the first probability predicted by each first-largest language model is less than the second threshold, the first document is regarded as the first negative sample. The first threshold is greater than the second threshold.
[0016] The model training module is used to train the initial model based on each first positive sample and the first negative sample to obtain the target model.
[0017] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory;
[0018] The processor executes a computer program to implement the method provided in the first aspect embodiment or any alternative embodiment of the first aspect.
[0019] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method provided in the first aspect embodiment or any optional embodiment of the first aspect.
[0020] The beneficial effects of the technical solutions provided in this application are:
[0021] First, in this embodiment of the application, by obtaining a keyword set for the target field and retrieving documents using the keyword set, the retrieved documents can cover the target field as much as possible, ensuring the richness of the document information.
[0022] Secondly, in this embodiment, the documents are clustered and a portion of the documents in each cluster are extracted and submitted to multiple large language models for voting to predict whether they belong to the target domain. Based on the probability output by the large language models, the documents belonging to the target domain and those not belonging to the target domain are accurately selected. At the same time, it is not necessary to predict all the massive amount of documents retrieved, which can save computational and time costs.
[0023] Finally, literature belonging to the target domain and literature not belonging to the target domain are used as positive and negative samples for model training to train the initial model and obtain the target model.
[0024] The solution provided in this application, through literature retrieval, literature classification, literature extraction, and literature prediction, can accurately select suitable literature to construct a dataset for model training. Then, the model is trained using this dataset, so that the trained model has a high ability to identify the target domain of the literature. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.
[0026] Figure 1 A schematic flowchart illustrating a model training method provided in an embodiment of this application;
[0027] Figure 2 This is a flowchart illustrating a few-shot model training method in one example of an embodiment of this application.
[0028] Figure 3 This is a flowchart illustrating a model optimization method in one example of an embodiment of this application.
[0029] Figure 4 This is a flowchart illustrating the discussion process of the second largest language model in one example of an embodiment of this application.
[0030] Figure 5 A structural block diagram of a model training device provided in an embodiment of this application;
[0031] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0032] The embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions of the embodiments of this application.
[0033] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the terms “comprising” and “including” as used in embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, operation, element, and / or component, but do not exclude implementation as other features, information, data, step, operation, element, component, and / or combinations thereof supported by the art. It should be understood that when we say that an element is “connected” or “coupled” to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. Furthermore, “connected” or “coupled” as used herein can include wireless connection or wireless coupling. The term “and / or” as used herein indicates at least one of the items defined by the term; for example, “A and / or B” can be implemented as “A,” or as “B,” or as “A and B.”
[0034] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0035] The technical solutions of this application and their effects are described below through several exemplary embodiments. It should be noted that the following embodiments can be referenced, borrowed from, or combined with each other. Identical terms, similar features, and similar implementation steps in different embodiments will not be repeated.
[0036] Figure 1 This application provides a flowchart illustrating a model training method, the execution subject of which can be a terminal (e.g., a computer, mobile phone, etc.). Figure 1 As shown, the method may include:
[0037] Step S101: Obtain a set of keywords related to the target field, and retrieve a set of documents from a preset document database based on the keyword set.
[0038] In the embodiments of this application, the target field can be a certain technical field, such as "super special steel", "low-altitude economy", "semiconductor", etc., and the embodiments of this application are not limited here.
[0039] The keyword set includes multiple keywords, which can be divided into primary keywords and extended keywords. Primary keywords should cover as many subfields as possible within the target field, while extended keywords provide a summary of each subfield within the target field. The preset literature database can include literature related to the technologies of various fields. The literature set can be a collection of multiple documents retrieved from the preset literature database. The documents can be descriptive texts about the technologies of the target field, such as papers, journals, patents, etc., which are not limited in this embodiment.
[0040] Specifically, to ensure that the retrieved literature in the target field comprehensively covers the target field, this embodiment of the application can use multiple keywords to retrieve literature in the target field from a preset literature database. When designing keywords, they are divided into primary keywords and extended keywords. The primary keywords need to cover all subfields of the target field as much as possible. For example, for the field of "super special steel", the primary keyword could be "steel". The extended keywords can cover only a specific subfield within this field. For example, for the field of "super special steel", the extended keywords could be "high adaptability", "high toughness", etc. When performing the search, all keywords included in the keyword set can be entered together (of which the primary keyword must be entered), or they can be output in batches. This embodiment of the application does not limit this. After the keyword set is entered, the search system will retrieve literature containing the keyword set from the preset literature database. Optionally, to reduce search time, it can search only the title and abstract of the literature for relevant keywords.
[0041] Step S102: Cluster the document set to obtain multiple document subsets.
[0042] Specifically, after retrieving documents from the pre-defined literature database, since the subfields associated with each document are different, it is necessary to first perform a clustering operation on the retrieved documents. In this embodiment, the BertTopic algorithm (a clustering algorithm) is used for clustering. By reasonably configuring the model parameters of the clustering model, an appropriate number of initial categories is determined, ultimately obtaining each document subset. The purpose of document clustering is to select the clustered topic cluster dataset for subsequent annotation in an unsupervised manner, effectively reducing annotation costs.
[0043] Step S103: For each document subset, at least one document is sampled from the document subset level based on a preset ratio and used as the first document.
[0044] Specifically, after clustering each document to obtain multiple document subsets, since the content pointed to by the documents included in each document subset is similar, in order to reduce the computational cost, a portion of the documents can be extracted from each document subset for subsequent training dataset selection. When extracting from each document subset, a certain preset ratio can be used to randomly select from all the documents included in that document subset.
[0045] Step S104: For each first document, predict the first probability that the first document belongs to the target domain using multiple first-largest language models. If the first probability predicted by each first-largest language model is greater than the first threshold, then the first document is taken as the first positive sample. If the first probability predicted by each first-largest language model is less than the second threshold, then the first document is taken as the first negative sample, and the first threshold is greater than the second threshold.
[0046] In the embodiments of this application, the first probability can characterize the probability that a document belongs to the target domain. Generally speaking, the higher the first probability, the greater the probability that the document belongs to the target domain. The first threshold and the second threshold can be thresholds defined based on practical experience. When the first probability is higher than the threshold, it can be determined that the document does indeed belong to the target domain. When the first probability is lower than the second threshold, it can be determined that the document does not belong to the target domain. When the first probability is between the first threshold and the second threshold, it means that it is impossible to accurately determine whether the document belongs to the target domain using the first large language model. The first large language model can be a model with a certain reasoning ability, such as qwen-32b, moonshot-v1-8k, and deepseek-v3, etc. This embodiment of the application does not limit the model.
[0047] Specifically, after extracting each first document, each first document can be input into multiple first language models. The input method can be to submit the entire first document as a file to the first language model, or to input only the title and abstract of the first document to the first language model. This embodiment of the application does not limit this method. Then, each first language model will analyze the first document and predict whether the first document belongs to the target domain based on the analysis. The output of the first language model can include the first probability that the first document belongs to the target domain and the prediction logic. After the first document is predicted by multiple top-ranking language models, the first probabilities predicted by all top-ranking language models can be combined to comprehensively determine whether the first document belongs to the target domain. Specifically, when the first probabilities predicted by all top-ranking language models are greater than a first threshold (e.g., the first probabilities predicted by all top-ranking language models are greater than 65%), it can be considered with a high degree of confidence that the first document does indeed belong to the target domain. When the first probabilities predicted by all top-ranking language models are less than a second threshold (e.g., the first probabilities predicted by all top-ranking language models are less than 5%), it can be considered with a high degree of confidence that the first document does not belong to the target domain. These first documents can be used as training datasets (i.e., the first positive sample and the first negative sample) in the subsequent training process, respectively. As for the remaining first documents, it is considered that it is not possible to accurately determine whether they belong to the target domain, and they can be discarded.
[0048] Step S105: Train the initial model based on each first positive sample and first negative sample to obtain the target model.
[0049] In the embodiments of this application, the initial model can be an untrained but characteristic embedding model, such as bce (BCEmbedding), bge (BAAI General Embedding), minilm (Miniature Language Model), mxbai-embed, etc. For the above embedding models, they can be evaluated separately first, and the evaluation results are shown in Table 1:
[0050]
[0051] Table 1
[0052] Accuracy, representing the percentage of correctly predicted samples out of the total number of samples, is one of the most intuitive performance metrics for classification models. Its calculation formula is:
[0053]
[0054] Where TP (True Positive) represents the number of true positives, TN (True Negative) represents the number of true negatives, FP (False Positive) represents the number of false positives, and FN (False Negative) represents the number of false negatives. Accuracy values range from [0,1], with values closer to 1 indicating better model classification performance.
[0055] Precision, or accuracy, measures the proportion of samples that the model predicts to be positive, but which are actually positive. Its formula is:
[0056]
[0057] Precision reflects the reliability of the model's prediction of positive classes. A high precision means that most of the samples predicted as positive by the model are indeed positive. Its value range is also [0,1].
[0058] Recall, or recall rate, measures the proportion of samples that are actually positive that the model correctly predicts as positive. Its formula is:
[0059]
[0060] F1 is the harmonic mean of precision and recall, used to comprehensively evaluate the performance of a model. Its calculation formula is:
[0061]
[0062] The F1 score strikes a balance between precision and recall, making it particularly suitable for situations with an imbalanced ratio of positive to negative samples. Its value ranges from [0,1], with values closer to 1 indicating better overall model performance.
[0063] By examining the above evaluation indicators and Table 1, it can be found that the bge model performs well in all aspects. Therefore, in this embodiment, the bge model can be selected as the type of initial model.
[0064] Specifically, after selecting the first positive samples and the first negative samples, these samples can be used together as the training dataset for training the initial model. The label of the first positive sample is "belongs to the target domain" (which can be represented by "1" in the specific implementation), and the label of the first negative sample is "does not belong to the target domain" (which can be represented by "0" in the specific implementation). Each sample can be input into the initial model in turn. After the initial model obtains the corresponding output, it compares the output with the label corresponding to the sample, and fine-tunes the model parameters of the initial model based on the difference in the comparison, thus obtaining the target model. It should be noted that the target model obtained by the embodiment of this application is associated with a specific target domain. That is to say, for different target domains, multiple different target models need to be trained to predict whether it belongs to that target domain.
[0065] The solution provided in this application embodiment firstly involves obtaining a keyword set for the target field and retrieving documents using the keyword set, which enables the retrieved documents to cover the target field as much as possible, thus ensuring the richness of the document information.
[0066] Secondly, in this embodiment, the documents are clustered and a portion of the documents in each cluster are extracted and fed into a large language model to predict whether they belong to the target domain. Based on the probability output by the large language model, the documents belonging to the target domain and those not belonging to the target domain are accurately selected. At the same time, it is not necessary to predict all the massive amount of documents retrieved, which can save computational and time costs.
[0067] Finally, literature belonging to the target domain and literature not belonging to the target domain are used as positive and negative samples for model training to train the model and obtain the target model.
[0068] The solution provided in this application, through literature retrieval, literature classification, literature extraction, and literature prediction, can accurately select suitable literature to construct a dataset for model training. Then, the model is trained using this dataset, so that the trained model has a high ability to identify the target domain of the literature.
[0069] Based on the above embodiments, as an optional embodiment, the method further includes:
[0070] The target model is optimized through multiple optimization steps until preset conditions are met, resulting in an optimized target model. Each optimization step specifically includes:
[0071] For each document in the literature set, the document is input into the target model of this optimization to obtain the second probability that the document belongs to the target domain. If the second probability is within the preset interval, the document is determined as the second document.
[0072] For each second document, the third probability of the second document belonging to the target domain is predicted by multiple first-largest language models respectively; if the third probability predicted by each first-largest language model is greater than the first threshold, the second document is used as the sample to be analyzed; if the third probability predicted by each first-largest language model is less than the second threshold, the second document is used as the second negative sample.
[0073] For each sample to be analyzed, the sample is predicted using two second-largest language models to obtain the first prediction result; the analytical power of the second-largest language model is stronger than that of the first-largest language model; the first prediction result indicates whether the sample to be analyzed belongs to the target domain.
[0074] The first prediction result represents the sample to be analyzed belonging to the target domain as the second positive sample. The target model is trained using each second positive sample and each second negative sample to obtain the target model for the next optimization.
[0075] In the embodiments of this application, the preset interval can be set based on practical experience. When the second probability of a certain document falls within the preset interval, it indicates that the target model cannot make an accurate judgment on this document. The second language model has higher model parameter settings than the first language model, thus making its analytical capabilities stronger than the first language model. The sample to be analyzed is a document that needs further confirmation as to whether it belongs to the target domain. In the embodiments of this application, to improve the reliability of the first positive sample data, a debate mechanism of the second language model is further introduced to filter the subset obtained by the first language model "voting" (i.e., judging whether the document belongs to the target domain based on the second probability output by each first language model) to improve the quality of the training dataset. The first prediction result can be a judgment conclusion on whether the sample to be analyzed belongs to the target domain (such as belonging to the target domain or not belonging to the target domain) jointly given by multiple second language models after prediction analysis of the sample to be analyzed. It can also include the probability of "belonging to the target domain". This embodiment of the application does not limit this.
[0076] Specifically, although a high-quality training dataset is used during the training of the target model, the identification ability of the target domain is still relatively one-sided because the literature included in the training dataset has been screened. Therefore, the solution of this application embodiment also provides a step to further optimize the trained target model. Specifically, all first documents retrieved from a preset literature database can be used as inputs to the target model in sequence. The target model judges whether all first documents belong to the target domain. The output of the target model can be the probability that the first document belongs to the target domain (i.e., the second probability). The output probability is between [0,1] (where 0 means it is completely impossible to belong to the target domain, and 1 means it definitely belongs to the target domain). If the output second probability is around 0.5, it can be considered that the target model cannot accurately judge whether the first document belongs to the target domain. Therefore, it is necessary to optimize the target model's judgment ability for this part of the first documents. In this application embodiment, a preset interval for the second probability can be set (such as [0.4,0.6]). If the second probability output by the target model for a certain first document falls within the preset interval, then the first document can be determined as the second document used to optimize the judgment ability of the target model.
[0077] After obtaining each second document, the probability (i.e., the third probability) of each second document belonging to the target domain can be predicted by each of the first largest language models during the training process of the target model. As mentioned above, for second documents whose third probabilities output by each of the first largest language models are all less than the second threshold, they can be used as negative samples (i.e., second negative samples) in the optimization process. For second documents whose third probabilities output by each of the first largest language models are all greater than the first threshold, in order to ensure the quality of the dataset in the optimization process, these second documents can be identified as samples to be analyzed first. Then, these samples to be analyzed can be further analyzed by at least two second largest language models with stronger analytical capabilities. After the second largest language models perform "joint" analysis at the same time, they will give the first prediction result for each second document. Then, the samples to be analyzed that belong to the target domain represented by the first prediction result can be used as positive samples (i.e., second positive samples) in the optimization process, thereby ensuring the quality of the dataset in the optimization process.
[0078] After obtaining each second positive sample and each second negative sample, these samples can be used together as the dataset in the target model optimization process. Then, the parameters of the target model are further adjusted in the same way as the training process described above until the preset conditions are met, and the optimized target model is obtained.
[0079] Based on the above embodiments, as an optional embodiment, prediction is performed on the sample to be analyzed using two second-largest language models to obtain a first prediction result, specifically including:
[0080] The samples to be analyzed are input into two second-largest language models respectively, and the corresponding second prediction results and first prediction logic are obtained respectively;
[0081] If the two second prediction results are the same, then the second prediction result is taken as the first prediction result;
[0082] If the two first prediction results are different, then the second prediction result of any of the second largest language models for the sample to be analyzed, the first prediction logic of any of the second largest language models for the sample to be analyzed, and the sample to be analyzed are respectively input into another second largest language model, so that the other second largest language model combines the second prediction result and the first prediction logic to predict the sample to be analyzed, and obtains the third prediction result.
[0083] If two third prediction results are the same, the third prediction result is taken as the first prediction result.
[0084] In the embodiments of this application, the second and third prediction results can be the same as the first prediction result described above, that is, the judgment conclusion of each second language model on whether the sample to be analyzed belongs to the target domain. Similarly, the second or third prediction result can also include the probability of "belonging to the target domain". This embodiment of the application does not limit this. The specific thought process of the second language model performing predictive analysis on the second document to obtain the second prediction result is as follows:
[0085] Specifically, the embodiments of this application will be described below using two second major language models as an example. For each second document, the embodiments of this application will perform prediction analysis on each second document using two second major language models respectively. Each second document will obtain two second prediction results and two first prediction logics that obtain the second prediction results through the two second major language models respectively. Then, the second prediction results output by the two second major language models will be compared. If the second prediction results output by the two second major language models are consistent (i.e., both are considered to belong to the target domain or not to the target domain), then the consistent second prediction results can be used as the first prediction results. If the second prediction results output by the two second-largest language models are inconsistent, they can engage in a "discussion." This discussion can involve each model sharing its second prediction results and first prediction logic with the other, who then combines these two results with the first logic to perform another prediction analysis on the second document. Optionally, while exchanging the second prediction results and first logic, some relevant knowledge and technology from the target domain can also be input into the second-largest language model as output. During this process, the two second-largest language models may change their second prediction results due to the other's thought process, or they may stick to their own judgment. In this case, the prediction result output by the second-largest language model will be used as the third prediction result.
[0086] After a round of "discussion", if the third prediction results output by the two second-largest language models are the same, it means that the two second-largest language models have reached a "consensus" after the "discussion", and the consistent third prediction result can be used as the first prediction result.
[0087] Optionally, if the third prediction results output by the two second-largest language models are still different after a round of "discussion", a new round of "discussion" can be started again. In each round of "discussion", each second-largest language model will input all its previous prediction results and prediction logic into the other.
[0088] Based on the above embodiments, as an optional embodiment, the output of the second language model also includes a second prediction logic to obtain a third prediction result;
[0089] If the two third-party predictions differ, the method further includes:
[0090] The third language model is obtained, and its analytical capabilities are stronger than those of the second language model.
[0091] The first prediction result is obtained by combining the second prediction results, the first prediction logic, the third prediction results, and the second prediction logic of the third language model with the second prediction logic.
[0092] In the embodiments of this application, the third language model has higher model parameter settings than the second language model, thereby making its analytical capabilities stronger than those of the second language model.
[0093] Specifically, if two second-largest language models still cannot reach a consensus after "discussion", then a third-largest language model can be introduced to perform predictive analysis on the second document. During the analysis, all prediction results and prediction logic output by the two second-largest language models can be input into the third-largest language model to provide sufficient analytical basis. Since the analytical capabilities of the third-largest language model are stronger than those of the second-largest language models, the prediction results output by the third-largest language model can be directly used as the first prediction result.
[0094] It should be noted that, due to the higher parameter settings of the third language model, the computational cost is also higher. Therefore, to save computational costs, this embodiment prioritizes the second language model for predictive analysis of the second document. Preferably, the solution provided in this embodiment can set an upper limit on the number of "discussion" rounds. That is, before the upper limit is reached, the two second language models can continuously "discuss" until a consensus is reached. When the upper limit is reached, the third language model is then used for predictive analysis of the second document.
[0095] Based on the above embodiments, as an optional embodiment, the preset conditions include: the optimization index corresponding to the current optimization step is greater than the optimization index corresponding to the previous optimization step, and the number of second documents with the second probability within the preset interval in the current optimization step is less than the preset threshold.
[0096] The optimization metrics are determined in the following ways:
[0097] Determine the precision and recall of the target model; where precision is used to characterize the proportion of literature in the target domain that the target model predicts belongs to the target domain, and recall is used to characterize the proportion of literature in the target domain that the target model correctly predicts.
[0098] Determine the harmonic mean between precision and recall, and then weight the harmonic mean to obtain the optimized metric.
[0099] In this embodiment of the application, the optimization metric is the "F1" mentioned above.
[0100] Specifically, in this embodiment, upon completion of each optimization round, the accuracy, precision, recall, and optimization metrics are calculated. Simultaneously, the amount of new data (i.e., the number of documents whose second probability predicted by the target model falls within a preset range) after each optimization round is also tallied. When the optimization metrics increase and the amount of new data is relatively small, the optimization of the target model is considered to be nearing saturation, and the optimization process can be terminated. Specific examples are shown in Table 2.
[0101]
[0102] Table 2
[0103] As shown in Table 2, the conditions of increasing the optimization index and having a small amount of new data were not met simultaneously in the first to sixth optimization processes. Only in the seventh optimization process were the above conditions met simultaneously. In the eighth optimization process, the optimization index decreased again. Therefore, the target model obtained after the seventh optimization process can be considered as a model close to saturation. The target model after the seventh optimization can be used as the target model after the optimization is completed.
[0104] Based on the above embodiments, as an optional embodiment, the number of first negative samples is greater than the number of first positive samples;
[0105] The initial model is trained based on each first positive sample and each first negative sample, and this process includes a step of augmenting the model with the first positive samples:
[0106] Determine the difference between the number of the first positive samples and the number of the first negative samples, and select the documents whose number is equal to the difference from all documents except the first document in each document set as the first positive samples.
[0107] Specifically, in practice, the number of first positive samples is often much smaller than the number of first negative samples. However, during training, it is necessary to ensure that the target model has a certain ability to distinguish between "belonging to the target domain" and "not belonging to the target domain." Therefore, it is necessary to supplement the first positive samples to make their number the same as or close to the number of first negative samples, in order to ensure the balance of positive and negative samples in the training dataset. Specifically, the difference between the number of first positive samples determined by the first language model and the number of first negative samples can be obtained. Then, documents with the difference in number can be selected from the documents not extracted in the retrieved literature set as the first positive samples.
[0108] Based on the above embodiments, as an optional embodiment, documents with a difference in quantity are obtained from each document other than the first document in each document set as the first positive sample, specifically including:
[0109] For each document in the literature set except for the first document, determine the similarity between the document and each first positive sample;
[0110] The similarity scores are sorted in descending order, and the documents corresponding to the highest similarity scores with the largest number of differences are taken as the first positive sample.
[0111] In the embodiments of this application, similarity can refer to the semantic similarity of the text content in the documents. This similarity can be obtained by converting the text content contained in each document into corresponding vectors and then calculating the cosine value between the vectors.
[0112] Specifically, since the number of first positive samples is being expanded, the documents obtained from the literature set must also belong to the target domain. Since a portion of the first positive samples belonging to the target domain has already been identified using the first major language model, documents with content similar to this portion of the first positive samples can be selected. Specifically, each first positive sample and each document not extracted can be converted into a corresponding text vector. Preferably, to reduce the amount of text content that needs to be converted, only the document titles and abstracts can be converted into corresponding text vectors. After obtaining the vectors of the documents not extracted, the similarity between each document and each first positive sample is calculated sequentially. Then, the similarities are sorted from largest to smallest, and the documents with the highest difference in the sorted results are selected as the first positive samples.
[0113] The model training method provided in the embodiments of this application will be described below with reference to the accompanying drawings. Figure 2 This is a flowchart illustrating a few-shot model training method provided in an embodiment of this application, as shown below. Figure 2 As shown, this method can be divided into the following stages:
[0114] 1. In the literature acquisition stage, a set of keywords related to the target field can be obtained based on relevant knowledge in the target field. Then, the keyword set can be used to search from a preset literature database to obtain a set of literature that may be related to or belong to the target field.
[0115] 2. In the literature extraction stage, the literature in the literature set is first clustered to obtain multiple literature subsets with category labels. In order to save computational costs, a preset proportion of literature can be randomly selected from each literature subset as the first literature.
[0116] 3. In the literature prediction stage, each first document is used to predict whether it belongs to the target domain through multiple first-class language models. For each first document, if multiple first-class language models believe that the first document belongs to the target domain, then the first document is regarded as the first positive sample. If multiple first-class language models believe that the first document does not belong to the target domain, then the first document is regarded as the first negative sample. If multiple first-class language models do not make the same judgment, then the first document is not regarded as a sample.
[0117] 4. Sample Supplementation Stage: This stage mainly involves determining whether the number of first positive samples and first negative samples is balanced. Generally, the number of first positive samples will be much smaller than the number of first negative samples. Therefore, it is necessary to determine whether the number of first positive samples is sufficient. If it is insufficient, similar documents can be selected from the unextracted documents to supplement the first positive samples and obtain more first positive samples. Then, all the first positive samples and first negative samples are combined into a training dataset. If it is sufficient, the process of supplementing the first positive samples can be skipped, and the determined first positive samples and first negative samples can be directly combined into a training dataset.
[0118] 5. Model training stage: This stage involves training the initial model using the training dataset from the previous stage. Once training is complete, the target model can be obtained.
[0119] Figure 3 This is a flowchart illustrating a model optimization method provided in an embodiment of this application, as shown below. Figure 3 As shown, this method can be divided into the following stages:
[0120] 1. In the target determination stage, each document in the previously retrieved literature set is input into the target model, which then predicts whether each document belongs to the target domain and outputs a second probability of belonging to the target domain. When the second probability falls within a preset range close to 0.5, it indicates that the predictive ability of the target model for that document needs to be optimized, and thus this part of the documents is determined as the second document.
[0121] 2. In the sample analysis stage, the second documents obtained in the previous stage are input into multiple first-level language models. The first-level language models then make preliminary judgments. The first-level language models can directly identify the second negative samples in the optimization process from the second documents. For the second positive samples, based on the preliminary judgment of the first-level language models, two second-level language models with stronger analytical capabilities are introduced for further discussion and judgment. If the second-level language models also cannot make a judgment, a third-level language model with stronger analytical capabilities can be introduced for judgment. Finally, the documents whose judgment results indicate that they belong to the target domain are identified as the second positive samples.
[0122] 3. In the model optimization stage, the target model is optimized and trained using the second positive samples and second negative samples obtained in the previous stage until the preset conditions are met, thus completing the optimization of the target model.
[0123] Figure 4 A flowchart illustrating a discussion process for a second major language model, as provided in this application embodiment, is shown below. Figure 4 As shown, for the bibliographic text of the given literature, "Comparison of three-phase corrosion behavior: Performance of SiN and 304L stainless steel in 6M nitric acid solution at different temperatures---In this study, the three-phase corrosion behavior of SiN and 304L stainless steel in 6M nitric acid solution at different temperatures was compared...", this text content was first sent to two second-largest language models to start the initial debate. One second-largest language model output the result "relevant", while the other second-largest language model output the result "irrelevant". Using regular expression judgment, the conclusion of "inconsistent" is obtained. At this time, the subsequent debate needs to be started.
[0124] In subsequent debates, the bibliographic text and the outputs of the two second-largest language models from the initial debate are input to the opposing side. The two second-largest language models then combine the opposing side's output from the initial debate to continue debating their own viewpoints. If the two second-largest language models still reach inconsistent conclusions, a new round of debate is initiated, in which a third-largest evidence model is added to provide key supporting evidence from the bibliographic text. This continues until the two second-largest language models reach the same conclusion in one round of debate or the number of debates reaches three. If the two second-largest language models still do not reach the same conclusion after three rounds of debate, the entire debate record is handed over to the third-largest language model as the judge model to provide a conclusion.
[0125] Figure 5 A structural block diagram of a model training device provided in an embodiment of this application is shown below. Figure 5 As shown, the model training device 500 may include: a document retrieval module 501, a document clustering module 502, a document extraction module 504, a sample screening module 505, and a model training module 506, wherein,
[0126] The document retrieval module 501 is used to obtain a set of keywords related to the target field, and retrieve a set of documents from a preset document database based on the keyword set;
[0127] The document clustering module 502 is used to cluster the document set to obtain multiple document subsets;
[0128] The document extraction module 503 is used to sample at least one document from the document subset level based on a preset ratio for each document subset, and use it as the first document.
[0129] The sample screening module 504 is used to predict the first probability that the first document belongs to the target domain for each first document through multiple first language models. If the first probability predicted by each first language model is greater than the first threshold, the first document is regarded as the first positive sample. If the first probability predicted by each first language model is less than the second threshold, the first document is regarded as the first negative sample. The first threshold is greater than the second threshold.
[0130] The model training module 505 is used to train the initial model based on each first positive sample and the first negative sample to obtain the target model.
[0131] The solution provided in this application embodiment firstly involves obtaining a keyword set for the target field and retrieving documents using the keyword set, which enables the retrieved documents to cover the target field as much as possible, thus ensuring the richness of the document information.
[0132] Secondly, in this embodiment, the documents are clustered and a portion of the documents in each cluster are extracted and fed into a large language model to predict whether they belong to the target domain. Based on the probability output by the large language model, the documents belonging to the target domain and those not belonging to the target domain are accurately selected. At the same time, it is not necessary to predict all the massive amount of documents retrieved, which can save computational and time costs.
[0133] Finally, literature belonging to the target domain and literature not belonging to the target domain are used as positive and negative samples for model training to train the model and obtain the target model.
[0134] The solution provided in this application, through literature retrieval, literature classification, literature extraction, and literature prediction, can accurately select suitable literature to construct a dataset for model training. Then, the model is trained using this dataset, so that the trained model has a high ability to identify the target domain of the literature.
[0135] Based on the above embodiments, as an optional embodiment, the device further includes a model optimization module, specifically used for:
[0136] For each document in the literature set, the document is input into the target model of this optimization to obtain the second probability that the document belongs to the target domain. If the second probability is within the preset interval, the document is determined as the second document.
[0137] For each second document, the third probability of the second document belonging to the target domain is predicted by multiple first-largest language models respectively; if the third probability predicted by each first-largest language model is greater than the first threshold, the second document is used as the sample to be analyzed; if the third probability predicted by each first-largest language model is less than the second threshold, the second document is used as the second negative sample.
[0138] For each sample to be analyzed, the sample is predicted using two second-largest language models to obtain the first prediction result; the analytical power of the second-largest language model is stronger than that of the first-largest language model; the first prediction result indicates whether the sample to be analyzed belongs to the target domain.
[0139] The first prediction result represents the sample to be analyzed belonging to the target domain as the second positive sample. The target model is trained using each second positive sample and each second negative sample to obtain the target model for the next optimization.
[0140] Based on the above embodiments, as an optional embodiment, the model optimization module is further used for:
[0141] The samples to be analyzed are input into two second-largest language models respectively, and the corresponding second prediction results and first prediction logic are obtained respectively;
[0142] If the two second prediction results are the same, then the second prediction result is taken as the first prediction result;
[0143] If the two first prediction results are different, then the second prediction result of any of the second largest language models for the sample to be analyzed, the first prediction logic of any of the second largest language models for the sample to be analyzed, and the sample to be analyzed are respectively input into another second largest language model, so that the other second largest language model combines the second prediction result and the first prediction logic to predict the sample to be analyzed, and obtains the third prediction result.
[0144] If two third prediction results are the same, the third prediction result is taken as the first prediction result.
[0145] Based on the above embodiments, as an optional embodiment, the output of the second language model also includes a second prediction logic to obtain a third prediction result;
[0146] If the two third-party predictions are different, the model optimization module can also be used for:
[0147] The third language model is obtained, and its analytical capabilities are stronger than those of the second language model.
[0148] The first prediction result is obtained by combining the second prediction results, the first prediction logic, the third prediction results, and the second prediction logic of the third language model with the second prediction logic.
[0149] Based on the above embodiments, as an optional embodiment, the preset conditions include: the optimization index corresponding to the current optimization step is greater than the optimization index corresponding to the previous optimization step, and the number of second documents with the second probability within the preset interval in the current optimization step is less than the preset threshold.
[0150] The device also includes an optimization index determination module, specifically used for:
[0151] Determine the precision and recall of the target model; where precision is used to characterize the proportion of literature in the target domain that the target model predicts belongs to the target domain, and recall is used to characterize the proportion of literature in the target domain that the target model correctly predicts.
[0152] Determine the harmonic mean between precision and recall, and then weight the harmonic mean to obtain the optimized metric.
[0153] Based on the above embodiments, as an optional embodiment, the number of first negative samples is greater than the number of first positive samples;
[0154] The device also includes a sample augmentation module, specifically used for:
[0155] Determine the difference between the number of the first positive samples and the number of the first negative samples, and select the documents whose number is equal to the difference from all documents except the first document in each document set as the first positive samples.
[0156] Based on the above embodiments, as an optional embodiment, the sample expansion module is further used for:
[0157] For each document in the literature set except for the first document, determine the similarity between the document and each first positive sample;
[0158] The similarity scores are sorted in descending order, and the documents corresponding to the highest similarity scores with the largest number of differences are taken as the first positive sample.
[0159] The following is for reference. Figure 6 It illustrates an electronic device suitable for implementing embodiments of this application (e.g., performing...). Figure 1 The diagram shows the structure of the terminal device or server 600 of the method shown. The electronic devices in the embodiments of this application may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (e.g., vehicle navigation terminals), wearable devices, etc., as well as fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0160] The electronic device includes a memory and a processor. The memory stores a program for executing the methods described in the various method embodiments above. The processor is configured to execute the program stored in the memory. The processor may be referred to as processing device 601 as described below. The memory may include at least one of read-only memory (ROM) 602, random access memory (RAM) 603, and storage device 608 as described below, as follows:
[0161] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0162] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0163] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 609, or installed from storage device 608, or installed from ROM 602. When the computer program is executed by processing device 601, it performs the functions defined in the methods of embodiments of this application.
[0164] It should be noted that the computer-readable storage medium described above in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0165] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0166] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0167] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0168] Obtain a keyword set related to the target domain, and retrieve a document set from a pre-defined document database based on the keyword set; cluster the document set to obtain multiple document subsets; for each document subset, sample at least one document from the document subset hierarchy based on a pre-defined ratio, as the first document; for each first document, predict the first probability that the first document belongs to the target domain using multiple first-largest language models; if the first probability predicted by each first-largest language model is greater than a first threshold, then the first document is used as the first positive sample; if the first probability predicted by each first-largest language model is less than a second threshold, then the first document is used as the first negative sample, where the first threshold is greater than the second threshold; train the initial model based on each first positive sample and the first negative sample to obtain the target model.
[0169] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0170] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0171] The modules or units described in the embodiments of this application can be implemented in software or hardware. The names of modules or units do not necessarily limit the specific unit; for example, a first constraint acquisition module can also be described as a "module for acquiring the first constraint".
[0172] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0173] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0174] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0175] The above description is only a partial embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for training an active learning model that constructs a dataset, characterized in that, include: Obtain a set of keywords related to the target field, and retrieve a set of documents from a preset document database based on the keyword set; Cluster the document set to obtain multiple document subsets; For each document subset, at least one document is sampled from the document subset hierarchically based on a preset ratio and used as the first document; For each first document, the first probability that the first document belongs to the target domain is predicted by multiple first language models. If the first probability predicted by each first language model is greater than a first threshold, the first document is regarded as a first positive sample. If the first probability predicted by each first language model is less than a second threshold, the first document is regarded as a first negative sample. The first threshold is greater than the second threshold. The target model is obtained by training the initial model based on each first positive sample and first negative sample; The method further includes: The target model is subjected to multiple optimization steps until a preset condition is met, resulting in an optimized target model; wherein each optimization step includes: For each document in the document set, the document is input into the target model for this optimization to obtain a second probability that the document belongs to the target domain. If the second probability is within a preset interval, the document is identified as the second document. For each second document, the third probability of the second document belonging to the target domain is predicted by multiple first language models respectively; if the third probability predicted by each first language model is greater than the first threshold, the second document is used as a sample to be analyzed; if the third probability predicted by each first language model is less than the second threshold, the second document is used as a second negative sample. For each sample to be analyzed, the sample is predicted using two second-largest language models to obtain a first prediction result; the analytical capabilities of the second-largest language models are stronger than those of the first-largest language models; the first prediction result indicates whether the sample to be analyzed belongs to the target domain. The first prediction result represents the sample to be analyzed belonging to the target domain as the second positive sample. The target model is trained using each second positive sample and each second negative sample to obtain the target model for the next optimization.
2. The method according to claim 1, characterized in that, The process of predicting the sample to be analyzed using two second-largest language models to obtain a first prediction result includes: The sample to be analyzed is input into two second-largest language models respectively, and the corresponding second prediction results and first prediction logic are obtained respectively; If the two second prediction results are the same, then the second prediction result shall be used as the first prediction result; If the two first prediction results are different, then the second prediction result of any second language model for the sample to be analyzed, the first prediction logic of any second language model for the sample to be analyzed, and the sample to be analyzed are respectively input into another second language model, so that the other second language model combines the second prediction result and the first prediction logic to predict the sample to be analyzed, and obtain a third prediction result; If the two third prediction results are the same, then the third prediction result is taken as the first prediction result.
3. The method according to claim 2, characterized in that, The output of the second large language model also includes a second prediction logic to obtain the third prediction result; If the two third prediction results are different, the method further includes: Obtain a third major language model, whose analytical capabilities are stronger than those of the second major language model; The first prediction result is obtained by combining the third language model with each second prediction result, each first prediction logic, each third prediction result, and each second prediction logic to predict the sample to be analyzed.
4. The method according to claim 1, characterized in that, The preset conditions include: the optimization index corresponding to the current optimization step is greater than the optimization index corresponding to the previous optimization step, and the number of second documents with the second probability within the preset interval in the current optimization step is less than the preset threshold. The optimization index is determined in the following way: Determine the accuracy and recall of the target model; wherein, the accuracy is used to characterize the proportion of documents that the target model predicts belong to the target domain, which actually belong to the target domain; the recall is used to characterize the proportion of documents that the target model correctly predicts, which actually belong to the target domain. Determine the harmonic mean between the precision and the recall, and weight the harmonic mean to obtain the optimized metric.
5. The method according to claim 1, characterized in that, The number of the first negative samples is greater than the number of the first positive samples; The initial model is trained based on each first positive sample and each first negative sample, and this process includes a step of augmenting the first positive samples: Determine the difference between the number of the first positive samples and the number of the first negative samples, and obtain a number of documents equal to the difference from each document in each document set other than the first document in each document set as the first positive samples.
6. The method according to claim 5, characterized in that, The step of obtaining documents equal to the difference from each document in each of the document sets, excluding the first document in each of the document sets, as the first positive sample includes: For each document in the document set other than the first document, determine the similarity between the document and each first positive sample; The similarity scores are sorted in descending order, and the documents corresponding to the highest similarity scores with the largest number of differences are taken as the first positive samples.
7. A training device for an active learning model that constructs a dataset, characterized in that, include: The document retrieval module is used to obtain a set of keywords related to the target field, and retrieve a set of documents from a preset document database based on the set of keywords; The document clustering module is used to cluster the document set to obtain multiple document subsets; The document extraction module is used to sample at least one document from the document subset hierarchically based on a preset ratio for each document subset, and use it as the first document. The sample screening module is used to predict the first probability that the first document belongs to the target domain for each first document through multiple first major language models. If the first probability predicted by each first major language model is greater than a first threshold, the first document is regarded as a first positive sample. If the first probability predicted by each first major language model is less than a second threshold, the first document is regarded as a first negative sample. The first threshold is greater than the second threshold. The model training module is used to train the initial model based on each first positive sample and first negative sample to obtain the target model. The device also includes a model optimization module, specifically used for: The target model is subjected to multiple optimization steps until a preset condition is met, resulting in an optimized target model; wherein each optimization step includes: For each document in the document set, the document is input into the target model for this optimization to obtain a second probability that the document belongs to the target domain. If the second probability is within a preset interval, the document is identified as the second document. For each second document, the third probability of the second document belonging to the target domain is predicted by multiple first language models respectively; if the third probability predicted by each first language model is greater than the first threshold, the second document is used as a sample to be analyzed; if the third probability predicted by each first language model is less than the second threshold, the second document is used as a second negative sample. For each sample to be analyzed, the sample is predicted using two second-largest language models to obtain a first prediction result; the analytical capabilities of the second-largest language models are stronger than those of the first-largest language models; the first prediction result indicates whether the sample to be analyzed belongs to the target domain. The first prediction result represents the sample to be analyzed belonging to the target domain as the second positive sample. The target model is trained using each second positive sample and each second negative sample to obtain the target model for the next optimization.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the method of any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1-6.