Sample data screening method and device, equipment and medium

By answering multiple large language models and testing models to be trained, combined with statistical indicators to filter data, the problem of uneven data quality in large-scale question banks was solved, efficient and intelligent data screening was achieved, and the training effect of the question-solving model was improved.

CN120687834APending Publication Date: 2025-09-23BEIJING YUANLI WEILAI SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510824506.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In existing technologies, the data quality in large-scale question banks varies greatly, making it difficult to accurately reflect the actual training value of the questions for the problem-solving model. It consumes huge computing resources, and a large amount of simple or repeated data has a diminishing marginal effect on improving model capabilities. Existing quality discrimination models are unable to fully identify data quality issues.

Method used

Use multiple large language models to answer sample test questions, clean the data based on the answer results, combine the test answer information of the model to be trained, determine the target data suitable for training, measure the difficulty of the questions through statistical indicators, and realize intelligent data screening.

Benefits of technology

It significantly improves the efficiency and coverage of data cleaning, ensures data quality, reduces computing resource consumption, and selects target data that can improve model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687834A_ABST
    Figure CN120687834A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a sample data screening method and device, equipment and a medium, and the method comprises the steps: obtaining first data; the first data are sample test questions including question information and reference answer information; answering the question information by using the plurality of large language models to obtain a first answering result; cleaning the first data according to the first solution result to obtain second data meeting a preset condition; inputting question information in the second data into the to-be-trained model to obtain test answer information; and based on the reference answer information and the test answer information, determining target data suitable for training the to-be-trained model. And verifying the quality of the sample test questions based on the first solution result. The problem that various types of quality problems cannot be identified by using a quality discrimination model is solved; and the difficulty of the question information on the to-be-trained model is measured by using the answer performance of the to-be-trained model on the question information, so that the accuracy of the target data obtained by screening is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and more particularly to a method, apparatus, device, and medium for screening sample data. Background Art

[0002] Currently, problem-solving models are typically trained using data from large-scale question banks. However, this data is typically obtained from a variety of sources, resulting in inconsistent data quality and a large amount of low-quality, unusable data. Furthermore, the scale of current large-scale data is enormous (e.g., billions of questions). Using all of this data to train problem-solving models consumes significant computing resources. Furthermore, large amounts of simple or repetitive data have diminishing returns on improving the problem-solving model's capabilities. Therefore, when training problem-solving models with stronger reasoning capabilities, it is necessary to filter out accurate data from large-scale databases that is more difficult or more challenging for the problem-solving model.

[0003] In existing technologies, the difficulty of a question is usually determined based on the surface features of the data (questions) (such as question length) or manually annotated difficulty information (information determined by experts based on the difficulty of the knowledge points contained in the question). The training set used to train the problem-solving model is then screened based on the difficulty information of the question. However, the "difficulty" of a question is not an absolutely unchanging attribute; it is closely related to the ability level of the solver (the problem-solving model). For example, for problem-solving models of different sizes (for example, with 7B, 32B, 72B parameters, etc.) and different capabilities, the same question may pose completely different challenges to them. Therefore, static difficulty labels set based on the surface features of the question or macro-knowledge points cannot accurately reflect the actual training value or "learnability" of the question for the problem-solving model.

[0004] Based on this, how to provide a method for screening sample data to improve the optimization effect of sample data on problem-solving model training has become a technical problem that needs to be solved urgently. Summary of the Invention

[0005] In view of this, embodiments of this specification provide a method for screening sample data. One or more embodiments of this specification also relate to a method and apparatus for screening sample data, a computing device, and a computer-readable storage medium to address technical deficiencies in the prior art.

[0006] According to a first aspect of an embodiment of this specification, a method for screening sample data is provided, comprising: Acquire first data; the first data is a sample test question extracted from a sample test question database; the sample test question includes question information and reference answer information; answering the sample test questions using multiple large language models to obtain first answer results output by each of the large language models; Cleaning the first data according to the first answer results output by each of the large language models to obtain second data that meets a preset condition; the preset condition is that the question information is clear and complete, and the reference answer information is correct; Inputting the question information in the second data into the model to be trained to obtain the test answer information output by the model to be trained; Determining statistical indicators of the to-be-trained model for the question information based on the reference answer information and the test answer information; Based on the statistical indicators, target data suitable for training the model to be trained is determined.

[0007] According to a second aspect of an embodiment of this specification, a device for screening sample data is provided, comprising: An acquisition module is configured to acquire first data; the first data is a sample test question extracted from a sample test question database; the sample test question includes question information and reference answer information; an answering module, configured to answer the question information in the sample test question using a plurality of large language models, and obtain a first answer result output by each of the large language models; a data cleaning module, configured to clean the first data based on the first answer results output by each of the large language models to obtain second data that meets a preset condition; the preset condition being that the question information is clear and complete, and the reference answer information is correct; A testing module, configured to input the question information in the second data into the model to be trained, and obtain test answer information output by the model to be trained; A statistical module, configured to determine statistical indicators of the to-be-trained model for the question information based on the reference answer information and the test answer information; The screening module determines target data suitable for training the model to be trained based on the statistical indicators.

[0008] According to a third aspect of an embodiment of this specification, a computing device is provided, including: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned sample data screening method are implemented.

[0009] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the above-mentioned sample data screening method are implemented.

[0010] At least one embodiment provided in this specification can achieve the following beneficial effects: by obtaining first data; wherein the first data is a sample test question extracted from a sample test question database and includes question information and reference answer information; for the question information in the sample test question, multiple large language models are used to answer it, and a first answer result output by each large language model is obtained; based on the first answer result, the first data is cleaned to obtain second data that meets preset conditions; the question information in the second data is input into the model to be trained, and the test answer information is output by the model to be trained is obtained; based on the reference answer information and the test answer information, the statistical indicators of the model to be trained for the question information are determined; based on the statistical indicators, the target data suitable for training the model to be trained is determined. In the embodiment of this specification, multiple industry-leading large language models are used to answer sample test questions, and the quality of the sample test questions is verified based on the first answer results output by each large language model, which effectively solves the problem that a special quality discrimination model can only identify specific types of quality problems, but cannot identify other types of quality problems, and significantly improves the efficiency and coverage of data cleaning.

[0011] Furthermore, the difficulty of the question for the model being trained can be measured directly using the model's own performance (statistical indicators) on the question. This allows for dynamic question difficulty information that is highly correlated with the model's capabilities, enabling intelligent target data screening based on the model's capabilities, effectively selecting target data that will improve the model's performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 This is a flow chart of a method for screening sample data provided by one embodiment of this specification; Figure 2 This is a schematic structural diagram of a sample data screening device provided by one embodiment of this specification; Figure 3 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION

[0013] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0014] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0015] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0016] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0017] First, the terms involved in one or more embodiments of this specification are explained.

[0018] Large Language Models: Large Language Models (LLMs), also known as large language models or large models, are a type of natural language processing technology based on deep learning. They can better understand natural language and generate high-quality text based on a given context. Examples include GPT-4, Claude, LLaMA, Qwen, PaLM, Gemini Ultra, and Galactica. Large language models typically contain hundreds of billions (or more) of references, trained on large amounts of textual data. Through training, large language models can learn the grammatical structure of a language and the relationships between words, enabling them to perform a wide range of tasks.

[0019] In this specification, a method for screening sample data is provided. This specification also relates to a sample data screening device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.

[0020] Currently, problem-solving models are typically trained using data from large-scale question banks. However, this data is typically obtained from a variety of sources. Due to the diverse data sources of sample questions, this data can be subject to quality issues. Question information can be missing, with unclear descriptions and even logical contradictions. Furthermore, the reference answers can mismatch the question requirements, contain errors, or have insufficient reasoning. These issues severely impact the usability of the sample questions, potentially leading to the problem-solving model learning incorrect knowledge and reducing model training effectiveness. Furthermore, current large-scale data sets are enormous (e.g., in the billions). Using all the data in these databases to train problem-solving models consumes significant computing resources. Furthermore, large amounts of simple or repetitive data have diminishing returns on improving the problem-solving model's capabilities. Therefore, when training problem-solving models with stronger reasoning capabilities, it is necessary to filter out accurate, challenging, or challenging data from these databases.

[0021] In the prior art, different data quality issues are typically detected using corresponding detection methods or models. For example, to address issues such as incomplete, malformed, or incomprehensible question information (especially after conversion from images to text using methods such as OCR), models for identifying incomplete questions are developed and applied. For example, advanced natural language processing models such as BERT can be used to analyze the grammatical structure, completeness, and similarity of question text to standard questions to determine the validity of the question information. For another example, to verify the correctness of reference answer information, when the reference answer information is expressed in a mathematical typesetting language such as LaTeX, answer verification models or validation rules can be used. This approach of using corresponding models to detect different data quality issues focuses on specific types of quality defects in sample question databases. However, large-scale databases still contain a large number of data quality issues of varying types but with relatively low frequency. For these data that do have relatively low frequency but quality issues, it is difficult to collect sufficient and representative training data, making it impossible to train models to identify these quality issues. Consequently, it is impossible to comprehensively and thoroughly clean the data in large-scale databases, and thus it is impossible to ensure the quality of the data used to train problem-solving models.

[0022] See also Figure 1 , Figure 1A flowchart of a method for screening sample data according to an embodiment of the present specification is shown. From a program perspective, the execution entity of the process can be a server or a program installed on a training platform or training device for training a problem-solving model.

[0023] like Figure 1 As shown, the process may specifically include the following steps: Step 102: Obtain first data; the first data is a sample test question extracted from a sample test question database; the sample test question includes question information and reference answer information.

[0024] In the embodiments of this specification, the sample test question database may be a massive test question bank containing a large number of sample test questions. The sample test questions included in the sample test question database may be automatically or semi-automatically obtained from different databases or websites; or they may be manually written or generated based on a large language model.

[0025] In practical applications, the first data can be obtained by randomly extracting sample test questions from a sample test question database, or by extracting sample test questions from a sample test question database based on the purpose of the first data. For example, if the first data is subsequently used to train a math problem-solving model, one or several math test questions can be obtained from the sample test question database as the first data. For another example, if the first data is subsequently used to train a Chinese problem-solving model, one or several Chinese test questions can be obtained from the sample test question database as the first data. For another example, if the first data is subsequently used to train a comprehensive problem-solving model, one or more test questions of any subject can be randomly obtained from the sample test question database.

[0026] In practical applications, the first data may include multiple sample test questions or only one sample test question. Each sample test question in the obtained first data includes question information and reference answer information.

[0027] Step 104: answer the question information in the sample test question using multiple large language models to obtain a first answer result output by each of the large language models.

[0028] In the embodiments of this specification, multiple large language models can be used to answer the same sample test question, obtaining the answers given by each large language model. For example, if there are seven large language models, the seven large language models can be used to answer the same question, resulting in seven answers output by the seven large language models.

[0029] In the embodiments of this specification, the large language model used to answer the question information in the sample test questions is a relatively advanced large language model, such as models GPT-4, Claude, LLaMA, Qwen, PaLM, Gemini Ultra, Galactica, etc.

[0030] In actual applications, the first answer output by each large language model is usually the same, but can also be different. In the embodiments of this specification, the first answer can be the predicted answer information obtained by the large language model based on the question information, or it can be explanatory information used to explain that the question information cannot be answered. For example, if the large language model can answer the question information, the large language model can output the predicted answer information obtained by answering the question information; if the large language model cannot answer the question information, it can output explanatory information explaining that the question information cannot be answered.

[0031] It is understandable that if the first data includes multiple sample test questions, each large language model can be used to answer the question information in each sample test question, thereby obtaining the first answer result output by each large language model for the question information in each sample test question. For example, assuming that there are seven large language models and the first data includes two sample test questions, the 7 large language models can be used to answer the question information in the first sample test question first, and the 7 answer results for the question information in the first sample test question output by the 7 large language models can be obtained; then, the 7 large language models can be used to answer the question information in the second sample test question, and the 7 answer results for the question information in the second sample test question output by the 7 large language models can be obtained. Of course, the question information of the above two sample test questions can also be input into the 7 large language models at the same time, thereby obtaining the answer results output by the 7 large language models for the above two sample test questions.

[0032] Step 106: Based on the first answer results output by each of the large language models, the first data is cleaned to obtain second data that meets preset conditions; the preset conditions are that the question information is clear and complete, and the reference answer information is correct.

[0033] In an embodiment of the present specification, the first data is cleaned according to the first answer result output by each of the large language models, which may specifically include determining whether each large language model can answer the question information based on the first answer result. If the above-mentioned more advanced large language models are unable to answer the question information, it means that the question information has a defect that cannot be answered. Therefore, it is possible to determine whether the question information has a defect based on the first answer result output by each large language model, so as to clean the first data. If at least one of the large language models can answer the question information, it is possible to further determine whether the reference answer information is correct. If the reference answer information is correct, it is possible to determine that the first data is the second data that meets the preset conditions. In actual applications, it is also possible not to further determine whether the reference answer information is correct. In this case, if at least one of the large language models can answer the question information, it is possible to determine that the first data meets the preset conditions.

[0034] In the embodiment of this specification, the second data may be the first data without obvious defects. Specifically, the second data may be the first data with clear question information (clear logic and no ambiguity), complete (no missing necessary conditions and questions), and correct reference answer information.

[0035] In the embodiments of this specification, multiple industry-leading large language models are used to independently answer sample test questions. The quality of the sample test questions is verified based on the first answer output by each large language model. This effectively addresses the problem that traditional quality discrimination models can only identify specific types of quality issues but not other types of quality issues. This can significantly improve the efficiency and coverage of data cleaning. In addition, it can effectively reduce the technical costs of developing, training, deploying, and continuously maintaining multiple independent quality discrimination models.

[0036] Step 108: Input the question information in the second data into the model to be trained to obtain the test answer information output by the model to be trained.

[0037] In the embodiments of this specification, the model to be trained may be a problem-solving model for answering test questions; specifically, the model to be trained may be a problem-solving model for answering test questions of a specified subject; for example, a mathematics problem-solving model, a Chinese problem-solving model, a chemistry problem-solving model. It is understandable that the model to be trained may also be a problem-solving model for other subjects, which will not be described in detail here. As an embodiment, the model to be trained may also be a problem-solving model that can answer test questions of all subjects; for example, the model to be trained may be a problem-solving model that can answer test questions of multiple subjects such as mathematics, Chinese, chemistry, and physics. The question information in the second data is input into the model to be trained, and the test answer information for the question information output by the model to be trained can be obtained.

[0038] Step 110: Based on the reference answer information and the test answer information, determine the statistical indicators of the model to be trained for the question information.

[0039] In the embodiment of this specification, the statistical indicator may be the accuracy rate of the test answer information output by the to-be-trained model for the question information.

[0040] Step 112: Based on the statistical indicators, determine target data suitable for training the model to be trained.

[0041] In the embodiment of this specification, the target data may be data that can further improve the current problem-solving ability of the model to be trained. In the embodiment of this specification, if the accuracy of the test answer information output by the model to be trained for the question information is relatively high, it may indicate that the model to be trained has a strong ability to solve the question information in the second data. Therefore, if the second data is used to train the model to be trained, the improvement in the ability of the model to be trained will be relatively small. That is, the second data has a relatively small optimization effect on the model to be trained. On the contrary, if the accuracy of the test answer information output by the model to be trained for the question information is relatively low, it may indicate that the model to be trained has a relatively weak ability to solve the question information in the second data. Using the second data to train the model to be trained can effectively improve the ability of the model to be trained to process the second data. That is, the second data has a relatively large optimization effect on the model to be trained.

[0042] In the embodiments of this specification, the difficulty of the question information for the model being trained is directly measured by the model's own performance (statistical indicators) on the question information. This allows for dynamic question difficulty information that is highly correlated with the model's capabilities, enabling intelligent target data screening based on the model's capabilities, effectively selecting target data that will improve the model's performance.

[0043] It should be understood that the order of some steps in the methods described in one or more embodiments of this specification can be interchanged according to actual needs, or some steps can be omitted or deleted.

[0044] Figure 1 The method, by obtaining first data; wherein the first data is a sample test question extracted from a sample test question database and includes question information and reference answer information; for the question information in the sample test question, multiple large language models are used to answer, and the first answer result output by each large language model is obtained; according to the first answer result, the first data is cleaned to obtain second data that meets the preset conditions; wherein the preset conditions are that the question information is clear and complete, and the reference answer information is correct; the question information in the second data is input into the model to be trained, and the test answer information output by the model to be trained is obtained; based on the reference answer information and the test answer information, the statistical indicators of the model to be trained for the question information are determined; based on the statistical indicators, the target data suitable for training the model to be trained is determined. In the embodiment of this specification, multiple industry-leading large language models are used to answer sample test questions, and the quality of the sample test questions is verified based on the first answer result output by each large language model, which effectively solves the problem that the use of a special quality discrimination model can only identify specific types of quality problems, but cannot identify other types of quality problems, and significantly improves the efficiency and coverage of data cleaning.

[0045] Furthermore, the difficulty of the question for the model being trained can be measured directly using the model's own performance (statistical indicators) on the question. This allows for dynamic question difficulty information that is highly correlated with the model's capabilities, enabling intelligent target data screening based on the model's capabilities, effectively selecting target data that will improve the model's performance.

[0046] based on Figure 1 The present specification also provides some specific implementation plans of the method, which are described below.

[0047] In actual applications, since the sample test questions in the sample test question database have different data sources, the identifiers of the same type of information content in each sample test question in the sample test question database may be different, and the data formats in each sample test question may also be different.

[0048] Based on this, in order to further improve the accuracy of the first answer results output by each large language model, information normalization and identification processing and format standardization processing can also be performed on the first data.

[0049] Based on this, optionally, after obtaining the first data, the following steps may also be included: The same type of information content in the first data is marked with the same preset identifier to obtain updated first data.

[0050] Based on a preset data format, the data format of the updated first data is converted into the preset data format.

[0051] The same type of information content may refer to information with the same attributes. Specifically, in the embodiments of this specification, the question information in each sample test question may be the same type of information content; and the reference answer information in each sample test question may be another type of information content. Since the data sources of each sample test question in the sample test question database are different, the identifiers of the question information in each sample test question may be different. For example, the identifier of the question information of some sample test questions may be "Question", the identifier of the question information of some sample test questions may be "problem", and the identifier of the question information of some sample test questions may be "question". Similarly, the identifiers of the reference answer information in each sample test question may also be different; for example, the identifier of the reference answer information of some sample test questions may be "Answer", the identifier of the reference answer information of some sample test questions may be "Solution", and the identifier of the reference answer information of some sample test questions may be "Answer".

[0052] In the embodiments of this specification, information content of the same type in the first data is marked using the same preset identifier. This may be the case where the question information in the first data is marked using the same identifier, and the reference answer information in the first data is marked using the same identifier. To facilitate distinguishing question information from reference answer information, the identifier used to mark the question information is different from the identifier used to mark the reference answer information. Specifically, the question information in the first data may be marked using the first preset identifier, and the reference answer information in the first data may be marked using the second preset identifier.

[0053] The forms of the first identifier and the second identifier are not limited, and the first identifier and the second identifier can be any characters or identifiers composed of any characters. Specifically, the first identifier can be set based on the "question information". For example, characters such as "question", "problem", "Question" that are directly related to the "question information" can be used as the first identifier. In addition, characters such as "B", "Q" that are not directly related to the "question information" can also be selected as the first identifier. Similarly, the second identifier can be set based on the "answer information". For example, characters such as "answer", "Answer", "Solution" that are directly related to the "answer information" can be used as the second identifier. In addition, characters such as "A", "C", "S" that are not directly related to the "answer information" can also be selected as the second identifier.

[0054] In practice, due to the different data sources of the sample questions in this question bank, the data formats of the question information or reference answer information in each sample question may vary. For example, the representation of mathematical symbols in math questions may vary. For example, some sample questions use LaTeX format (such as "\sqrt{x}" to represent square roots), while others use plain text format (such as "square root x"). For another example, some sample questions use superscript formatting such as "x²", while others use superscript formatting such as "x^2".

[0055] In the embodiments of this specification, in the embodiments of this specification, the first data can be uniformly processed using format standardization rules or semantic mapping mechanisms to unify the data format of the first data. The preset data format can be a data format of any specified type. The core requirement is that all data in the first data set must follow a unified format specification. For example: assuming that the preset data format of the sample test questions in the mathematics subject is LaTeX format, all formulas in the first data must use a unified symbol representation rule (such as using \sqrt{x} instead of "square root x").

[0056] In the embodiments of this specification, unifying the data format in the first data can ensure the consistency and parsability of the information in the first data, thereby avoiding semantic ambiguity caused by format confusion.

[0057] To facilitate understanding, the embodiments of this specification provide a specific method for answering question information in sample test questions using a large language model.

[0058] Optionally, answering the question information in the sample test question using multiple large language models to obtain a first answer result output by each of the large language models may specifically include: Based on the prompt word template, a prompt word for the sample test question is generated; the prompt word is used to instruct the large language model to generate a first answer result for the sample test question according to the task requirements.

[0059] The prompt words are input into each of the large language models to obtain a first answer result of the sample test question output by each of the large models.

[0060] In the embodiments of this specification, a prompt template may be a structured text used to generate prompts. Prompts may be generated based on the prompt template. Prompts may be information that instructs the large language model to generate relevant content based on task information. The task information may be information related to a task that the large language model needs to perform, and the task information may include task requirements.

[0061] In actual application, the prompt word template can be obtained first, and then the task information is inserted into the prompt word template, thereby obtaining the prompt word template containing the task information. In addition, the task information can also be the content already in the prompt word template.

[0062] The task information is the core component of the prompt, clearly and explicitly indicating the specific task the user wants the large language model to complete. It's important to note that task information should be concise and clear, avoiding lengthy or complex sentences to reduce the complexity of the large language model. Furthermore, task information should be specific and unambiguous, avoiding vague or ambiguous sentences.

[0063] In the embodiments of this specification, a prompt word template can be first obtained, and then the question information from the sample test questions can be inserted into the prompt word template to obtain the prompt word. The prompt word can guide the large language model to better understand the task to be performed, thereby facilitating the large language model to generate and output content that better meets the task requirements.

[0064] In practical applications, the prompt template may also include model role information used to represent the model role. Model role information is a key element in the prompt. In the embodiments of this specification, model role information may refer to the occupation or identity that the user wants the large language model to assume when performing a task. This model role information enables the large language model to better simulate the behaviors and language habits of a specific occupation or identity, thereby enhancing the professionalism and credibility of the content.

[0065] In practical applications, the model's role can be specified through explicit statements. For example, a statement like "You are a math teacher, responsible for solving math problems" not only provides role information but also provides the corresponding context for the large language model.

[0066] For ease of understanding, this specification provides an example of a prompt word template, which is as follows: You are a third-grade math teacher; Please answer the following questions based on the information, calculate the answer, and write down the answer process step by step.

[0067] topic: On the weekend, Xiao Ming and his mother went to a stationery store to buy school supplies. The prices in the store were as follows: The pencils come in boxes of 12 and are sold for 8 yuan each; Notebooks are 5 yuan each, buy 3 and get 1 free; Each backpack is 68 yuan, and there is a 10% discount (90% of the original price).

[0068] Xiao Ming bought 2 boxes of pencils. How many pencils are there in total? If he wants to share them among his 6 classmates, how many pencils can each person get? As shown above, "You are a third-grade math teacher" can represent the role information of the large language model. Here, setting the role information of the large language model to "third-grade math teacher" can constrain the large language model to consider the knowledge points used during the solution process, avoiding using knowledge above the third grade level to solve the problem. "Please solve the following problem information, calculate the answer, and explain the solution step by step" can be the task information used to instruct the large language model to complete the task.

[0069] In the embodiments of this specification, multiple industry-leading large language models are used to independently answer sample test questions, and the quality of the sample test questions is verified based on the first answer results output by each large language model. This solves the problem that traditional quality discrimination models can only identify specific types of quality problems but cannot identify other types of quality problems, and significantly improves the efficiency and scope of data cleaning.

[0070] For ease of understanding, the embodiments of this specification further provide specific descriptions of a solution for obtaining the second data that meets the preset conditions.

[0071] Optionally, the cleaning of the first data according to the first answer results output by each of the large language models to obtain second data meeting a preset condition may specifically include: Based on the first answer results output by each of the large language models, it is determined whether the sample test question meets the preset conditions.

[0072] If so, the sample test question is determined to be second data that meets the preset conditions.

[0073] In the embodiments of this specification, the preset conditions serve as the core basis for determining the quality of the first data. Whether the first data has obvious defects can be determined based on the preset conditions. Specifically, the preset conditions may include clarity of the question information (clear logic, no ambiguity or contradiction), completeness (no missing necessary conditions or questions), and correct reference answer information.

[0074] For ease of understanding, the embodiment of this specification specifically describes determining whether the sample test question meets the preset conditions based on the first answer results output by each of the large language models.

[0075] Optionally, judging whether the sample test question meets the preset condition based on the first answer result output by each of the large language models may specifically include: Semantic recognition is performed on the first answer results of the sample test questions output by each of the large language models to determine a second answer result in which the question information representing the sample test questions is unqualified.

[0076] If the number of the second answer results is greater than the preset number, it is determined that the sample test question does not meet the preset condition.

[0077] In the embodiments of this specification, the first answer result can be predicted answer information obtained by solving the question information, or it can be explanatory information used to explain why the question information cannot be answered. In actual applications, if the question information is not flawed, the large language model can usually solve the question information and obtain predicted answer information. If the question information is flawed, it may cause the large language model to be unable to solve the question. In this case, the large language model can output explanatory information to explain why the question information cannot be solved.

[0078] In practical applications, task information can include instructions for the large language model to output explanations when it cannot answer a question. This allows the large language model to output explanations when it cannot answer a question. For example, the task information mentioned above, "Please answer the following question, calculate the answer, and describe the answer step by step," could also include "If the question cannot be answered, please explain the reason."

[0079] As one embodiment, for any large language model, if the question information cannot be answered, the large language model may output text indicating that the question cannot be answered, such as "Question cannot be answered." As another embodiment, if the question information cannot be answered, the large language model may also output an explanation of the inability to answer the question information, determined based on the actual reasons obtained from analyzing the question information. For example, if the question information has multiple interpretations, the large language model may output an explanation message such as "The question information is ambiguous"; if the question information is incomplete, the large language model may output an explanation message such as "The question information is incomplete"; if the question information is inconsistent, the large language model may output an explanation message such as "The question information is logically unclear."

[0080] Thus, semantic recognition can be performed on the first answer result output by the large language model to determine whether the first answer result includes text information indicating that the question information cannot be answered. It is understandable that ambiguous question information, incomplete question information, or unclear question information logic all constitute situations where the question information cannot be answered. If the number of text information indicating that the question information cannot be answered, that is, the number of second answer results, is greater than a preset number, it can indicate that multiple large language models believe that the question information is flawed and cannot be answered. Therefore, it can be determined that the sample test question does not meet the preset conditions and cannot be used to train the model to be trained. If the number of text information indicating that the question information cannot be answered, that is, the number of second answer results, is less than the preset number, it can be preliminarily determined that the sample test question meets the preset conditions. In actual applications, the preset number can be set according to actual needs. Specifically, the preset number can be half the number of large language models. For example, assuming that seven large language models are used to answer the question information, the preset number can be 3. If four or more of the seven first answer results obtained indicate that the question information cannot be answered, it can be determined that the sample test question does not meet the preset conditions.

[0081] In the embodiment of this specification, whether the sample test question meets the preset conditions can also be judged based on the consistency between the first answer result output by each large language model and the reference answer information.

[0082] Optionally, the reference answer information includes a first answer process and a first answer; the first answer result includes a second answer process and a second answer; and judging whether the sample test question meets the preset condition based on the first answer result output by each of the large language models may specifically include: Based on a majority voting mechanism, third answers are selected from the second answers output by each of the large language models; the number of the third answers is greater than the number of other answers in the second answers.

[0083] It is determined whether the first answer is consistent with the third answer to obtain a first determination result.

[0084] If the first judgment result indicates that the first answer is inconsistent with the third answer, then it is determined whether the first answer is consistent with other answers in the second answer to obtain a second judgment result.

[0085] If the second judgment result indicates that the first answer is inconsistent with other answers in the second answer, it is determined that the sample test question does not meet the preset condition.

[0086] In the embodiment of this specification, the reference answer information may include a first answer process and a first answer; the first answer result may include a second answer process and a second answer.

[0087] The third answer is the answer with the largest number or the largest proportion among the second answers. For example, suppose seven large language models are used to answer the question information. Among the seven answers obtained, there are 3 answers of answer 4, 2 answers of answer 5, and 2 answers of answer 6. Answer 4 has the largest number and the largest proportion, so answer 4 is the third answer.

[0088] As an implementation method, whether a sample test question meets a preset condition can be determined based on a hierarchical judgment mode. Specifically, the first answer included in the reference answer information can be compared with the third answer to determine whether the first answer and the third answer are consistent. If they are consistent, it can be preliminarily determined that the sample test question meets the preset condition. If they are inconsistent, the first answer can be compared with the other answers in the second answer except the third answer to determine whether the first answer is consistent with the other answers in the second answer except the third answer. If they are consistent, it can be preliminarily determined that the sample test question meets the preset condition. If the first answer and the other answers in the second answer except the third answer are also inconsistent, it can be determined that the first answer in the reference answer information is incorrect. Therefore, it can be determined that the reference answer information does not meet the preset condition, and further determined that the sample test question is not the second data that meets the preset condition.

[0089] In the embodiments of the present specification, in the graded judgment mode, the third answer is screened out from the second answer based on the voting mechanism, and the first answer is compared with the third answer. Compared with comparing the first answer with the second answer in sequence, the comparison process of the first answer and the second answer can be reduced, thereby saving computing resources; in addition, since most of the answers in the sample test question database are correct, the correctness of the first answer can usually be verified by comparing the first answer included in the reference answer information with the third answer. Therefore, the process of comparing the first answer included in the reference answer information with the third answer, and then comparing the first answer with other answers in the second answer except the third answer when the first answer and the third answer are inconsistent, can effectively improve the efficiency of determining whether the sample test question meets the preset conditions compared with the process of comparing the first answer with each second answer in sequence.

[0090] As another implementation, a strict judgment mode can be used to determine whether a sample test question meets the preset conditions. Specifically, when the first answer and the third answer are determined to be inconsistent, the first answer can be directly determined to not meet the preset conditions without comparing the first answer with the second answer except the third answer. This method can efficiently screen out high-quality sample test questions and ensure the accuracy of the training data used to train the model to be trained.

[0091] In practical applications, the judgment mode to be selected can be determined based on actual needs. For example, if the quality of sample test questions used to train the model to be trained needs to be strictly controlled, the strict judgment mode can be used to determine whether the sample test questions meet the preset conditions. For another example, if the quality requirements for sample test questions are relatively loose, and the focus is on obtaining a large amount of first data, the hierarchical judgment mode can be used to determine whether the sample test questions meet the preset conditions.

[0092] As an implementation method, it is also possible to determine whether the sample test question meets the preset conditions by comparing the first answer with each second answer in sequence. For example, if the first answer is consistent with any of the second answers, it can be preliminarily determined that the sample test question meets the preset conditions; if the first answer is inconsistent with any of the second answers, it can be determined that the sample test question does not meet the preset conditions.

[0093] As an implementation method, the process of determining whether a sample test question meets preset conditions based on the consistency between the first answer result output by each large language model and the reference answer information, and the process of determining whether a sample test question meets preset conditions based on the semantics of the first answer result output by the large language model in the previous text, can be processes executed in parallel. If one of the processes indicates that the sample test question does not meet the preset conditions, the execution of the other process can be terminated to save computing resources.

[0094] As another embodiment, the process of determining whether the sample test question meets the preset conditions based on the semantics of the first answer result output by the large language model and the process of determining whether the sample test question meets the preset conditions based on the consistency between the first answer results output by each large language model and the reference answer information can also be performed sequentially. If the number of second answer results representing unqualified question information of the sample test question is less than a preset number, the first answer can be compared with the fourth answer information in the first answer result representing qualified question information of the sample test question.

[0095] In practice, there are cases where a large language model outputs a correct second answer (the same as the first answer included in the reference answer information), but this second answer is not derived through reasonable reasoning logic. For example, when using a large language model to answer multiple-choice questions in a sample test database, the large language model may select an option as the answer even if the question information itself is flawed. Therefore, to ensure the quality of the sample test questions used to train the model, it is also possible to further determine the logical correctness of the first answer process output by the large language model and the second answer process in the reference answer information.

[0096] Optionally, judging whether the sample test question meets the preset condition based on the first answer result output by each of the large language models may specifically include: Based on the natural language processing method, it is determined whether the reasoning logic in the first solution process and the reasoning process in the second solution process is coherent, and a third judgment result is obtained.

[0097] It is determined whether the information referenced in the first solution process and the second solution process is correct to obtain a fourth determination result.

[0098] If the third judgment result indicates that the reasoning logic in the first solution process or the second solution process is incoherent, or the fourth judgment result indicates that the information cited in the first solution process or the second solution process is incorrect, it is determined that the sample test question does not meet the preset conditions.

[0099] In the embodiments of this specification, a natural language processing method can be used to determine whether the reasoning logic in the first and second answering processes is coherent. Specifically, a logical judgment prompt word can be constructed, and based on the logical judgment prompt word, a first preset large language model can be used to determine whether the reasoning logic in the first and second answering processes is coherent. The first preset large language model can be the model used to answer the question information in the sample test questions mentioned above, or it can be a model other than the model used to answer the question information in the sample test questions mentioned above, and no specific limitation is made here.

[0100] In the embodiments of this specification, an information judgment prompt word can also be constructed, and based on the information judgment prompt word, a second pre-trained large language model is used to judge whether the information cited in the first solution process and the second solution process is correct. The second pre-trained large language model and the first pre-trained large language model can be the same model or different models, and no specific limitation is made here.

[0101] As an implementation manner, if the third judgment result indicates that the reasoning logics in the first solution process and the second solution process are coherent, and the fourth judgment result indicates that the information cited in the first solution process and the second solution process is correct, then it can be preliminarily determined that the sample test question meets the preset conditions.

[0102] The information cited in the first solution process and the second solution process can include data, theorems, formulas, etc.

[0103] For the convenience of understanding, the embodiments of this specification provide specific examples of logical judgment prompt words as follows: ## Step Discrimination - Reasoning Logic Coherence Your task is to determine whether there are any skipped steps in the solution process of a question. Please carefully read the following question and the corresponding solution process, and evaluate them according to the given criteria.

[0104] First, please carefully read the following question: <Question> {{PROBLEM}} < / Question> Now, please carefully read the following solution process: <Solution Process> {{SOLUTION}} < / Solution Process> The criteria for judging skipped steps are as follows: 1. Whether key reasoning steps are omitted in the solution process.

[0105] 2. Whether there is a logical gap from one step to the next without a reasonable transition.

[0106] 3. Whether necessary calculations or derivation processes are omitted.

[0107] Please evaluate according to the following steps: 1. Carefully read the entire solution process.

[0108] [[ID=**********]]<00002**********> 2. Compare the solution process with the above criteria one by one.

[0109] 3. Consider the logical coherence between each step.

[0110] 4. Form a preliminary judgment.

[0111] 5. Check again to ensure that no important details have been overlooked.

[0112] Analyze the solution process in detail within the <Thought> tag and consider whether it violates any skipping criteria. Then give your final judgment in the <Judgment> tag, using "There are skips" or "There are no skips". Finally, explain your judgment reason in detail in the <Explanation> tag.

[0113] <Thought> [Analyze in detail here whether there are skips in the solution process] < / Thought> <Judgment> [Give the judgment of "There are skips" or "There are no skips" here] < / Judgment> <Explanation> [Provide a detailed explanation here to state the reason for the judgment] < / Explanation> Please ensure that your judgment is objective and fair and is based on the given criteria. If the content of the solution process is ambiguous, please explain your consideration process in the explanation.

[0114] As shown above, "Your task is to determine whether there are skips in the solution process of a question. Please carefully read the following question and the corresponding solution process and evaluate according to the given criteria." can represent the task information for instructing the large language model to complete. "Please follow the following steps for evaluation: 1. Carefully read the entire solution process. 2. Compare the solution process with the above criteria one by one. 3. Consider the logical coherence between each step. 4. Form a preliminary judgment. 5. Check again to ensure that no important details have been overlooked. Analyze the solution process in detail within the <Thought> tag and consider whether it violates any skipping criteria. Then give your final judgment in the <Judgment> tag, using "There are skips" or "There are no skips". Finally, explain your judgment reason in detail in the <Explanation> tag. Please ensure that your judgment is objective and fair and is based on the given criteria. If the content of the solution process is ambiguous, please explain your consideration process in the explanation." can represent the specific task requirement information.

[0115] In the embodiments of this specification, it is also possible to further determine whether the output content of the large language model is suitable for the model role information in the prompt. Continuing to use the example in the previous text for illustration, assuming that the role information in the prompt is "You are a third - grade math teacher responsible for answering math questions", it can be determined whether the knowledge points applied in the second solution process of the first solution result output by the large language model exceed the third - grade scope. As an implementation method, a third - preset large language model can be used to determine whether the output content of each large language model is suitable for the model role information in the prompt.

[0116] As an implementation method, the process of judging whether the sample test question meets the preset conditions based on the answer process and reference information, the process of judging whether the sample test question meets the preset conditions based on the consistency between the first answer result output by each large language model and the reference answer information, and the process of judging whether the sample test question meets the preset conditions based on the semantics of the first answer result output by the large language model can be executed in parallel. If one of the processes indicates that the sample test question does not meet the preset conditions, the execution of the other two processes can be terminated to save computing resources. If all three processes indicate that the sample test question meets the preset conditions, it can be determined that the sample test question meets the preset conditions. In actual applications, it is also possible to determine whether the sample test question meets the preset conditions based on one or two of the judgment processes as needed.

[0117] As another embodiment, the process of judging whether the sample test question meets the preset conditions based on the semantics of the first answer result output by the large language model, the process of judging whether the sample test question meets the preset conditions based on the consistency between the first answer result output by each large language model and the reference answer information, and the process of judging whether the sample test question meets the preset conditions based on the answer process and reference information can also be performed in sequence. If the number of second answer results representing unqualified question information of the sample test question is less than the preset number, the first answer can be compared with the fourth answer in the first answer result representing qualified question information of the sample test question; if the first answer is consistent with the fourth answer, it can be further judged whether the reasoning logic in the first answer process and the second answer process corresponding to the fourth answer is consistent, and whether the information referenced in the first answer process and the second answer process corresponding to the fourth answer is correct.

[0118] In practical applications, individual sample questions can be independently evaluated to determine whether they meet the requirements for effectively improving the capabilities of the model to be trained.

[0119] Optionally, if the second data includes a sample test question, inputting the question information in the second data into the model to be trained to obtain the test answer information output by the model to be trained may specifically include: The question information in the sample test question is inputted into the model to be trained repeatedly n times to obtain n groups of test answer information outputted by the model to be trained for the question information.

[0120] Determining the statistical indicators of the to-be-trained model for the question information based on the reference answer information and the test answer information may specifically include: Based on the n groups of test answer information and the reference answer information, the accuracy of the model to be trained for the question information is determined.

[0121] In an embodiment of the present specification, the title information of a sample test question to be judged can be repeatedly inputted into the model to be trained n times, thereby obtaining n groups of test answer information outputted by the model to be trained for the title information. Based on the reference answer information, the number of correct test answer information in the n groups of test answer information can be further determined, and the accuracy of the model to be trained for the title information can be determined based on the number of correct test answer information and the number of repetitions n. Specifically, assuming that the title information of a sample test question is repeatedly inputted into the model to be trained 100 times, and 75 of the 100 groups of test answer information obtained are consistent with the reference answer information, then the accuracy of the model to be trained for the title information is 75%.

[0122] In addition, a systematic analysis can be conducted on batches of sample test questions to consider whether the sample test question set has the value to effectively improve the capabilities of the model to be trained.

[0123] Optionally, if the second data includes a plurality of sample test questions, inputting the question information in the second data into the model to be trained to obtain the test answer information output by the model to be trained may specifically include: The question information in the plurality of sample test questions is input into the model to be trained, and the test answer information output by the model to be trained for each question information in the plurality of sample test questions is obtained.

[0124] Determining the statistical indicators of the to-be-trained model for the question information based on the reference answer information and the test answer information may specifically include: Based on the test answer information output by the to-be-trained model for each question information in the plurality of sample test questions, the accuracy rate of the to-be-trained model for the plurality of question information is determined.

[0125] In an embodiment of the present specification, question information from multiple sample test questions is input into a model to be trained, and test answer information output by the model to be trained for each question information in each sample test question can be obtained. Based on the reference answer information for each sample test question, the number of correct test answer information in the test answer information output by the model to be trained can be determined. Based on the number of correct test answer information and the number of sample test questions input to the model to be trained, the accuracy of the model to be trained for the set of sample test questions can be determined.

[0126] For ease of understanding, the embodiments of this specification also provide a specific method for determining target data suitable for training the model to be trained based on the statistical indicators.

[0127] Optionally, determining target data suitable for training the model to be trained based on the statistical indicators may specifically include: If the statistical indicator is less than a preset value, the sample test question corresponding to the question information is determined as the target data.

[0128] In the embodiments of this specification, if the accuracy of the test answer information output by the model to be trained for the question information is greater than a preset value, it can be indicated that the model to be trained has a strong ability to answer the question information in the second data. Therefore, if the second data is used to train the model to be trained, the improvement in the ability of the model to be trained will be small. That is, the second data has a small optimization effect on the model to be trained. On the contrary, if the accuracy of the test answer information output by the model to be trained for the question information is less than a preset value, it can be indicated that the model to be trained has a weak ability to answer the question information in the second data. Therefore, the second data can be used to train the model to be trained to improve the model to be trained's ability to process the second data. That is, the second data has a large optimization effect on the model to be trained.

[0129] Corresponding to the above method embodiment, this specification also provides an embodiment of a device for screening sample data. Figure 2 FIG. 1 shows a schematic diagram of a sample data screening device provided by an embodiment of this specification. Figure 2 As shown, the device includes: The acquisition module 202 is configured to acquire first data; the first data is a sample test question extracted from a sample test question database; the sample test question includes question information and reference answer information.

[0130] The answering module 204 is configured to answer the question information in the sample test question using multiple large language models to obtain a first answer result output by each of the large language models.

[0131] The data cleaning module 206 is configured to clean the first data according to the first answer results output by each of the large language models to obtain second data that meets preset conditions; the preset conditions are that the question information is clear and complete, and the reference answer information is correct.

[0132] The testing module 208 is configured to input the question information in the second data into the model to be trained to obtain the test answer information output by the model to be trained.

[0133] The statistical module 210 is configured to determine the statistical indicators of the model to be trained for the question information based on the reference answer information and the test answer information.

[0134] The screening module 212 is configured to determine target data suitable for training the model to be trained based on the statistical indicators.

[0135] based on Figure 2 The present specification also provides some specific implementation plans of the method, which are described below.

[0136] Optional, Figure 2 The device shown may further include: The normalization identification module is configured to mark the same type of information content in the first data with the same preset identifier to obtain updated first data.

[0137] The format standardization module is configured to convert the data format of the updated first data into the preset data format based on the preset data format.

[0138] Optionally, the answer module 204 may be specifically configured to: Based on the prompt word template, a prompt word for the sample test question is generated; the prompt word is used to instruct the large language model to generate a first answer result for the sample test question according to the task requirements.

[0139] The prompt words are input into each of the large language models to obtain a first answer result of the sample test question output by each of the large models.

[0140] Optionally, the data cleaning module 206 may specifically include: The judgment unit is configured to judge whether the sample test question meets the preset condition based on the first answer results output by each of the large language models.

[0141] The determining unit is configured to, if yes, determine that the sample test question is second data that meets a preset condition.

[0142] Optionally, the judging unit may be specifically configured to: Semantic recognition is performed on the first answer results of the sample test questions output by each of the large language models to determine a second answer result in which the question information representing the sample test questions is unqualified.

[0143] If the number of the second answer results is greater than the preset number, it is determined that the sample test question does not meet the preset condition. Optionally, the reference answer information includes the first answer process and the first answer; the first answer result includes the second answer process and the second answer; the judgment unit can be specifically configured to: Based on a majority voting mechanism, third answers are selected from the second answers output by each of the large language models; the number of the third answers is greater than the number of other answers in the second answers.

[0144] It is determined whether the first answer is consistent with the third answer to obtain a first determination result.

[0145] If the first judgment result indicates that the first answer is inconsistent with the third answer, then it is determined whether the first answer is consistent with other answers in the second answer to obtain a second judgment result.

[0146] If the second judgment result indicates that the first answer is inconsistent with other answers in the second answer, it is determined that the sample test question does not meet the preset condition.

[0147] Optionally, the judging unit may be specifically configured to: Based on the natural language processing method, it is determined whether the reasoning logic in the first solution process and the reasoning process in the second solution process is coherent, and a third judgment result is obtained.

[0148] It is determined whether the information cited in the first solution process and the second solution process is correct to obtain a fourth determination result.

[0149] If the third judgment result indicates that the reasoning logic in the first solution process or the second solution process is incoherent, or the fourth judgment result indicates that the information cited in the first solution process or the second solution process is incorrect, it is determined that the sample test question does not meet the preset conditions.

[0150] Optionally, if the second data includes a sample test question, the test module 208 may be specifically configured to: Repeating n times the question information in the sample test question and inputting it into the to-be-trained model, and obtaining n sets of test answer information output by the to-be-trained model for the question information; The statistics module 210 may be specifically configured to: Based on the n groups of test answer information and the reference answer information, the accuracy of the model to be trained for the question information is determined.

[0151] Optionally, if the second data includes a plurality of sample test questions, the test module 208 may be specifically configured as follows: Inputting the question information of the plurality of sample test questions into the to-be-trained model, and obtaining the test answer information output by the to-be-trained model for each question information of the plurality of sample test questions; The statistics module 210 may be specifically configured to: Based on the test answer information output by the to-be-trained model for each question information in the plurality of sample test questions, the accuracy rate of the to-be-trained model for the plurality of question information is determined.

[0152] Optionally, the screening module 212 may be specifically configured to: If the statistical indicator is less than a preset value, the sample test question corresponding to the question information is determined as the target data.

[0153] The above is a schematic diagram of a sample data screening device according to this embodiment. It should be noted that the technical solution of this sample data screening device and the technical solution of the sample data screening method described above share the same concept. For details not described in detail in the technical solution of the sample data screening device, please refer to the description of the technical solution of the sample data screening method described above.

[0154] Figure 3 1 shows a block diagram of a computing device according to one embodiment of the present disclosure. Components of the computing device 300 include, but are not limited to, a memory 310 and a processor 320. The processor 320 is connected to the memory 310 via a bus 330, and a database 350 is used to store data.

[0155] The computing device 300 also includes an access device 340 that enables the computing device 300 to communicate via one or more networks 360. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 540 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.

[0156] In one embodiment of the present specification, the above components of the computing device 300 and Figure 3 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 3 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.

[0157] Computing device 300 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 300 can also be a mobile or stationary server.

[0158] The processor 320 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned method for screening sample data.

[0159] The above is a schematic diagram of a device for generating a sample data screening method according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the sample data screening method described above are based on the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the sample data screening method described above.

[0160] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned sample data screening method.

[0161] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the sample data screening method described above are based on the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the sample data screening method described above.

[0162] An embodiment of the present specification further provides a computer program product, wherein when the computer program / instructions are executed in a computer, the computer is caused to execute the steps of the above-mentioned sample data screening method.

[0163] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0164] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium may be appropriately increased or decreased based on the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0165] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0166] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0167] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A method for screening sample data, characterized in that: include: Acquire first data; the first data is a sample test question extracted from a sample test question database; The sample test questions include question information and reference answer information; answering the sample test questions using multiple large language models to obtain first answer results output by each of the large language models; Cleaning the first data according to the first answer results output by each of the large language models to obtain second data that meets a preset condition; the preset condition is that the question information is clear and complete, and the reference answer information is correct; Inputting the question information in the second data into the model to be trained to obtain the test answer information output by the model to be trained; Determining statistical indicators of the to-be-trained model for the question information based on the reference answer information and the test answer information; Based on the statistical indicators, target data suitable for training the model to be trained is determined.

2. The method according to claim 1, characterized in that After obtaining the first data, the method further includes: Marking information content of the same type in the first data with the same preset identifier to obtain updated first data; Based on a preset data format, the data format of the updated first data is converted into the preset data format.

3. The method according to claim 1, characterized in that The method of answering the sample test questions using multiple large language models to obtain first answer results output by each of the large language models specifically includes: Generating a prompt word for the sample test question based on a prompt word template; the prompt word is used to instruct the large language model to generate a first answer to the sample test question according to the task requirements; The prompt words are input into each of the large language models to obtain a first answer result of the sample test question output by each of the large models.

4. The method according to claim 1, wherein Cleaning the first data according to the first answer results output by each of the large language models to obtain second data that meets preset conditions specifically includes: Determining whether the sample test question meets the preset condition based on the first answer result output by each of the large language models; If so, the sample test question is determined to be second data that meets the preset conditions.

5. The method according to claim 4, characterized in that The determining whether the sample test question meets the preset condition based on the first answer result output by each of the large language models specifically includes: Performing semantic recognition on the first answer results of the sample test questions output by each of the large language models, and determining a second answer result in the first answer results that represents unqualified question information of the sample test questions; If the number of the second answer results is greater than the preset number, it is determined that the sample test question does not meet the preset condition.

6. The method according to claim 4, characterized in that The reference answer information includes a first answer process and a first answer; the first answer result includes a second answer process and a second answer; and judging whether the sample test question meets the preset condition based on the first answer result output by each of the large language models specifically includes: Based on a majority voting mechanism, third answers are selected from the second answers output by each of the large language models; the number of the third answers is greater than the number of other answers in the second answers; Determine whether the first answer is consistent with the third answer, and obtain a first determination result; If the first judgment result indicates that the first answer is inconsistent with the third answer, then determining whether the first answer is consistent with other answers in the second answer to obtain a second judgment result; If the second judgment result indicates that the first answer is inconsistent with other answers in the second answer, it is determined that the sample test question does not meet the preset condition.

7. The method according to claim 4, characterized in that The determining whether the sample test question meets the preset condition based on the first answer result output by each of the large language models specifically includes: determining, based on a natural language processing method, whether the reasoning logic in the first solution process and the reasoning process in the second solution process is coherent, to obtain a third judgment result; determining whether the information referenced in the first solution process and the second solution process is correct, to obtain a fourth determination result; If the third judgment result indicates that the reasoning logic in the first solution process or the second solution process is incoherent, or the fourth judgment result indicates that the information cited in the first solution process or the second solution process is incorrect, it is determined that the sample test question does not meet the preset conditions.

8. The method according to claim 1, characterized in that If the second data includes a sample test question, inputting the question information in the second data into the model to be trained to obtain the test answer information output by the model to be trained specifically includes: Repeating n times the question information in the sample test question and inputting it into the to-be-trained model, and obtaining n sets of test answer information output by the to-be-trained model for the question information; The step of determining the statistical indicators of the to-be-trained model for the question information based on the reference answer information and the test answer information specifically includes: Based on the n groups of test answer information and the reference answer information, the accuracy of the model to be trained for the question information is determined.

9. The method according to claim 1, characterized in that If the second data includes a plurality of sample test questions, inputting the question information in the second data into the to-be-trained model to obtain the test answer information output by the to-be-trained model specifically includes: Inputting the question information of the plurality of sample test questions into the to-be-trained model, and obtaining the test answer information output by the to-be-trained model for each question information of the plurality of sample test questions; The step of determining the statistical indicators of the to-be-trained model for the question information based on the reference answer information and the test answer information specifically includes: Based on the test answer information output by the to-be-trained model for each question information in the plurality of sample test questions, the accuracy rate of the to-be-trained model for the plurality of question information is determined.

10. The method according to claim 1, characterized in that The determining, based on the statistical indicators, target data suitable for training the model to be trained specifically includes: If the statistical indicator is less than a preset value, the sample test question corresponding to the question information is determined as the target data.

11. A device for screening sample data, characterized in that: include: an acquisition module, configured to acquire first data; The first data is a sample test question extracted from a sample test question database; The sample test questions include question information and reference answer information; an answering module configured to answer the question information in the sample test question using multiple large language models and obtain a first answer result output by each of the large language models; a data cleaning module configured to clean the first data based on the first answer results output by each of the large language models to obtain second data that meets a preset condition; the preset condition is that the question information is clear and complete, and the reference answer information is correct; a testing module configured to input the question information in the second data into the model to be trained, and obtain test answer information output by the model to be trained; A statistical module is configured to determine statistical indicators of the to-be-trained model for the question information based on the reference answer information and the test answer information; The screening module is configured to determine target data suitable for training the model to be trained based on the statistical indicators.

12. A computer device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor so that the at least one processor can implement the sample data screening method according to any one of claims 1 to 10.

13. A computer-readable medium, characterized in that Computer-readable instructions are stored thereon, and the computer-readable instructions can be executed by a processor to implement the sample data screening method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method and device for carrying out sample screening on large language model for questions and answers

    CN117493890A