Data quality inspection method, device and equipment based on large model, medium and product
By using a data quality inspection method based on a large model, the system uses the first prompt to judge the mastery of model knowledge and imports external knowledge data for secondary quality inspection. This solves the problem of low accuracy caused by the illusion of a large model, achieves efficient and automated quality inspection, and improves the accuracy and efficiency of data quality inspection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI IFLYHEALTH CO LTD
- Filing Date
- 2025-12-05
- Publication Date
- 2026-05-01
AI Technical Summary
In existing technologies, large models are prone to illusions during data quality inspection, resulting in low accuracy. Furthermore, traditional manual quality inspection is inefficient and costly, while automated rule matching lacks flexibility and accuracy.
By constructing a data quality inspection method based on a large model, the first prompt is used to determine whether the model has mastered the relevant knowledge of the data to be inspected. If it has not mastered the knowledge, external knowledge data is imported for secondary quality inspection to avoid model illusion and improve accuracy with the help of external knowledge data. When the model has mastered the knowledge, the results are directly output to achieve automated quality inspection.
This improves the accuracy and efficiency of data quality inspection, reduces labor costs, avoids subjective judgment, ensures the quality of data and training datasets, and thus enhances the robustness of the model.
Smart Images

Figure CN121958918A_ABST
Abstract
Description
Data quality inspection methods, devices, equipment, media, and products based on large models Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a data quality inspection method, apparatus, equipment, medium, and product based on a large model. Background Technology
[0002] With the popularization of the internet and the advancement of medical informatization, online medical Q&A platforms have gradually become an important channel for the public to obtain medical knowledge and solve health problems. However, the answers on these platforms are typically generated by medical professionals, machines, or other users. Since medical information is complex and specialized, it is difficult to ensure the accuracy and reliability of all answers. Therefore, how to effectively detect the quality of Q&A data to ensure the quality of medical Q&A data is a pressing technical need that requires immediate attention.
[0003] Traditional data quality inspection solutions mostly rely on manual inspection or rule matching technology. However, manual review is inefficient, time-consuming, and costly, especially in highly specialized fields like medicine, where it is difficult to deploy a large number of quality inspectors. Furthermore, manual inspection requires extensive data review and comparison, significantly impacting the speed and efficiency of actual quality inspection. While simple automated rule matching can improve efficiency to some extent, it lacks flexibility, scalability, and accuracy, making it difficult to capture and process complex relationships and implicit information in the data, resulting in poor data quality inspection effectiveness.
[0004] Currently, quality inspection of data is performed using large models. This involves inputting the data to be inspected into a large model and obtaining the model's output quality inspection results. However, large models may exhibit "illusions," leading to inaccurate quality inspection results and reduced overall data quality inspection accuracy. Summary of the Invention
[0005] This invention provides a data quality inspection method, apparatus, equipment, medium, and product based on a large model, to address the shortcomings of low accuracy in existing data quality inspection technologies and to achieve a highly efficient and high-quality data quality inspection method.
[0006] This invention provides a data quality inspection method based on a large model, comprising: inputting a first prompt constructed based on the data to be inspected into a data quality inspection model to obtain an output result from the data quality inspection model; the first prompt is used to instruct the data quality inspection model to determine whether it has mastered the knowledge related to the data to be inspected, and to instruct the data quality inspection model to perform quality inspection on the data to be inspected; the data quality inspection model is constructed based on a large model; if, based on the output result, it is determined that the data quality inspection model does not master the knowledge related to the data to be inspected, the data to be inspected and the knowledge data related to the data to be inspected are input into the data quality inspection model to obtain a data quality inspection result output by the data quality inspection model; if, based on the output result, it is determined that the data quality inspection model has mastered the knowledge related to the data to be inspected, the data quality inspection result is determined based on the output result.
[0007] This invention also provides a data quality inspection device based on a large model, comprising: a first input module, used to input a first prompt constructed based on the data to be inspected into a data quality inspection model, and obtain an output result output by the data quality inspection model; the first prompt is used to instruct the data quality inspection model to determine whether it has mastered the knowledge related to the data to be inspected, and to instruct the data quality inspection model to perform quality inspection on the data to be inspected; the data quality inspection model is constructed based on a large model; a second input module, used to input the data to be inspected and the knowledge data related to the data to be inspected into the data quality inspection model when it is determined based on the output result that the data quality inspection model has not mastered the knowledge related to the data to be inspected, and obtain a data quality inspection result output by the data quality inspection model; and a result determination module, used to determine the data quality inspection result based on the output result when it is determined based on the output result that the data quality inspection model has mastered the knowledge related to the data to be inspected.
[0008] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the data quality inspection method based on a large model as described above.
[0009] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the data quality inspection method based on a large model as described above.
[0010] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the data quality inspection method based on a large model as described above.
[0011] The data quality inspection method, apparatus, device, medium, and product based on a large model provided by this invention input a first prompt message constructed based on the data to be inspected into the data quality inspection model, obtaining the output result of the data quality inspection model. The first prompt message is used to instruct the data quality inspection model to determine whether it possesses knowledge related to the data to be inspected, and to instruct the data quality inspection model to perform quality inspection on the data to be inspected. Therefore, if the output result determines that the data quality inspection model does not possess knowledge related to the data to be inspected, the data to be inspected and related knowledge data are input into the data quality inspection model, obtaining the data quality inspection result output by the data quality inspection model. Furthermore, even when the data quality inspection model does not possess knowledge related to the data to be inspected, relevant output results can be output to instruct the input of the data to be inspected and related knowledge data into the data quality inspection model, i.e., importing relevant knowledge data so that the data quality inspection model can acquire knowledge related to the data to be inspected. This allows for secondary data quality checks, preventing the output of correct data quality check results due to model illusions. In other words, the data quality check model can rely not only on its own capabilities but also on external knowledge data to improve accuracy. Whether or not to use external knowledge data is determined by the data quality check model itself. If the data quality check model possesses relevant knowledge about the data to be checked based on the output results, the data quality check result is determined accordingly. Therefore, regardless of whether the data quality check model possesses relevant knowledge, no manual judgment is required for the final data quality check result, enabling automated data quality check. This reduces labor costs and improves efficiency, while avoiding subjective human judgment to enhance accuracy. Furthermore, the knowledge already possessed by the data quality check model only requires a single call to obtain the correct data quality check result, further improving efficiency. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0013] Figure 1 is one of the flowcharts of the data quality inspection method based on a large model provided by the present invention.
[0014] Figure 2 is a second schematic diagram of the data quality inspection method based on a large model provided by the present invention.
[0015] Figure 3 is the third flowchart of the data quality inspection method based on a large model provided by the present invention.
[0016] Figure 4 is the fourth flowchart of the data quality inspection method based on a large model provided by the present invention.
[0017] Figure 5 is a schematic diagram of the data quality inspection device based on a large model provided by the present invention.
[0018] Figure 6 is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0020] Current data quality inspection solutions mostly rely on large models to inspect the data to be inspected. However, large models may produce illusions, leading to reduced accuracy in data quality inspection and consequently a decline in data quality.
[0021] Given the poor accuracy of current data quality inspection solutions, this research was conducted. The initial approach was as follows: First, a multi-dimensional quality inspection model was developed to perform quality inspection on the data to be inspected. This model consisted of multiple sub-models associated with various quality inspection dimensions, with one sub-model corresponding to each dimension, ensuring coverage of multiple data quality dimensions. Second, the data to be inspected was input into each of the sub-models, and each sub-model performed quality inspection on the data, yielding results for each quality inspection dimension. This step, through the collaborative work of multiple sub-models, provided a comprehensive quality check of the data. Finally, based on the quality inspection results across each dimension, the dataset was cleaned. The cleaned training dataset was used to determine the training samples for the business model. This step, through data cleaning and optimization, ensured high-quality training data, thereby improving the performance and output quality of the business model.
[0022] Research on the above approach reveals that while it can improve the accuracy of data quality inspection to some extent, after obtaining the data quality inspection results for each quality inspection dimension, it is still necessary to filter the results through rules or manually judge the final data quality inspection results. Therefore, the data quality inspection cost remains high and the data quality inspection efficiency remains low. Furthermore, this data quality inspection scheme only performs a single quality inspection on all data to be inspected, relying excessively on the effect of the quality inspection sub-model. For some potentially controversial data, there is no process for secondary or multiple quality inspection reviews, which has a certain impact on the quality of data quality inspection. In other words, the accuracy of data quality inspection is still not high.
[0023] To address the problems with the aforementioned approach, further research was conducted. During this process, a domain-specific approach was considered, such as fine-tuning the large model using a large amount of medical question-and-answer data to improve its accuracy in evaluating medical question-and-answer data. However, even with domain-specific fine-tuning, the large model still relies solely on its own capabilities for data quality evaluation. Knowledge gaps that the model lacks can still lead to misinterpretations and erroneous data evaluation results.
[0024] To address the shortcomings of the aforementioned approach, further research was conducted. First, single-round medical question-and-answer data was acquired from the training data. Second, the questions and answers from the single-round medical question-and-answer data were classified using a template classification model to obtain the corresponding adaptation scores of the single-round medical question-and-answer data with all prompt templates in a pre-defined prompt library. The prompt template with the highest adaptation score was then selected for subsequent quality control. Third, by setting parameters, the large model's responses were designed to generate K responses based on the same input (constructed from the aforementioned prompt templates) while maintaining diversity as much as possible. Next, the consistency of the K responses is checked. If all responses are not "yes" (a "yes" response indicates the data quality check passed), the question-and-answer data is considered unqualified and deleted. If all responses are "yes," the question-and-answer data is considered completely correct, and its quality weight is set to 1. Secondly, if inconsistencies exist, and if W responses are not "yes" and R responses are "yes," parameters are set to ensure that the responses of the large model, while maintaining diversity, explain the model with pre-defined correct reasons. The system first inputs a question and answer template into a large model, obtaining R responses O11, O12, ..., O1R. Then, it inputs another question and answer template with a preset error explanation, obtaining W responses O21, O22, ..., O2W. Next, the R+W responses are concatenated into a preset prompt template and input into the large model for intelligent scheduling. The system then selects subsequent steps based on the output. Finally, based on the output, if the result is "human review," the question and answer data is submitted to a professional doctor for further review. If the professional doctor confirms it is correct, quality control is set. If the quality weight is 1, the question-and-answer data is discarded and the process ends. If the output result is "Change prompt template", then according to the adaptation scores of the question-and-answer data and the corresponding prompt templates, the prompt template with the highest score that has not been selected is selected from high to low, and the above steps are repeated. If the output result is "Discard", then the question-and-answer data is discarded. If the output result is "Correct", then the quality weight is set according to the "confidence level" of the output result. For the question-and-answer data that is not discarded, in subsequent training, the corresponding quality weight is used for sampling and subsequent training.
[0025] Research on the above scheme reveals that while it can determine data quality based on the number of errors in the large model, it is not detailed enough. It is still prone to making unreasonable settings in areas where the model lacks knowledge. Furthermore, this data may actually be supported by existing knowledge from different levels of authority, leading to a significant discrepancy between the actual quality stratification and the large model's judgment. In other words, the accuracy of data quality inspection remains low. Similarly, the large model still relies solely on its own capabilities for data quality inspection, and it is prone to misinterpretations regarding knowledge it lacks, resulting in erroneous data quality inspection results.
[0026] To address the shortcomings of the aforementioned data quality inspection schemes, a data quality inspection method based on a large model was proposed through continuous research. This method inputs a first prompt, constructed based on the data to be inspected, into the data quality inspection model, obtaining the model's output. The first prompt instructs the model to determine whether it possesses knowledge related to the data to be inspected, and also instructs it to perform quality inspection. If, based on the output, it is determined that the model lacks knowledge of the data to be inspected, the data to be inspected and related knowledge data are input into the model, yielding the model's output. Furthermore, even when the model lacks knowledge of the data to be inspected, it can output relevant results to instruct the model to input the data to be inspected and related knowledge data, i.e., importing relevant knowledge data for the data quality inspection model. The model acquires knowledge related to the data to be inspected, enabling secondary data quality checks. This avoids model illusions and outputs correct data quality check results. In other words, the data quality check model can rely not only on its own capabilities but also on external knowledge data to improve accuracy. Whether or not to use external knowledge data is determined by the model itself. Based on the output results, if the model has acquired knowledge of the data to be inspected, the data quality check result is determined accordingly. Therefore, regardless of whether the model actually possesses the knowledge, no manual judgment is needed for the final data quality check result, achieving automated data quality check. This reduces labor costs and improves efficiency, while avoiding subjective human judgment to enhance accuracy. Furthermore, the knowledge already acquired by the model only requires a single call to obtain the correct data quality check result, further improving efficiency. In summary, improving the accuracy of data quality inspection (i.e., improving the effectiveness of data quality inspection) allows for the selection of higher-quality data, thus improving overall data quality. Furthermore, if the data to be inspected is training data within the training dataset, it can improve the quality of the training data, thereby enhancing the training performance of the model trained on the selected training dataset and ultimately improving the robustness of the trained model.
[0027] The data quality inspection method based on a large model provided by the present invention will be described below through the following embodiments. The data quality inspection method based on a large model of the present invention will be described in conjunction with Figures 1-4.
[0028] Figure 1 is one of the flowcharts of the data quality inspection method based on a large model provided by the present invention. As shown in Figure 1, the data quality inspection method based on a large model includes the following steps 110, 120 and 130.
[0029] Step 110: Input the first prompt message constructed based on the data to be inspected into the data quality inspection model to obtain the output result of the data quality inspection model.
[0030] Here, the data to be inspected is the data to be inspected; this data can be set according to actual needs, such as question-and-answer data, sample data in the training dataset, or knowledge data, etc.
[0031] In one embodiment, the data to be quality inspected is question-and-answer data, which includes questions and answers. This question-and-answer data can be single-round question-and-answer data, such as single-round medical question-and-answer data. Specifically, a question-and-answer dataset is acquired, and based on this dataset, the question-and-answer data requiring quality inspection is determined, i.e., the data requiring quality inspection is selected for subsequent processing. Further, this question-and-answer data can be sample data from the training data for training a question-and-answer model.
[0032] For example, if the question is "What medications are available for treating stomach cramps?", the answer would be "Medications for treating stomach cramps generally include omeprazole, cinnarizine, montmorillonite powder, domperidone, and amoxicillin."
[0033] In another embodiment, the data to be inspected is the answer in the question-and-answer data, that is, it may not include the question.
[0034] In another embodiment, the data to be quality inspected is the training data in the training dataset. Specifically, the training dataset is obtained, and based on the training dataset, the training data that needs to be quality inspected (i.e., the data to be inspected) is determined, that is, the data that needs to be quality inspected is selected for subsequent processing.
[0035] The first prompt is used to instruct the data quality inspection model to determine whether it has knowledge related to the data to be inspected, and to instruct the data quality inspection model to perform quality inspection on the data to be inspected.
[0036] In one specific embodiment, a first prompt is constructed based on the data to be inspected and a first prompt template. More specifically, the data to be inspected is concatenated into the first prompt template to obtain the first prompt.
[0037] For example, the first prompt template is as follows: "Please determine if the following content is correct. If there are any parts you are unsure about, please mark them. The following is the content: xxx." Here, the content is the data to be inspected, "Please determine if the following content is correct" instructs the data quality inspection model to inspect the data, and "If there are any parts you are unsure about, please mark them" instructs the data quality inspection model to determine whether it has the relevant knowledge of the data to be inspected.
[0038] The data quality inspection model is built upon a large model. This large model can be a Large Language Model (LLM). The reason for using a large language model as the initial foundational model is that, compared to traditional small models or pre-trained models, it possesses richer prior knowledge and reasoning capabilities. Furthermore, compared to implicit information extraction and modeling, the large language model can explicitly output the thinking, analysis, and reasoning steps, thereby enhancing the correctness of the final reasoning result at the semantic level and improving the accuracy of data quality inspection. Moreover, the superior performance of natural language processing technologies based on these large language models in text understanding and generation capabilities provides support for achieving efficient and highly accurate data quality inspection.
[0039] In one embodiment, if the data to be inspected is medical question-and-answer data, it can be a large model with hundreds of billions to trillions of parameters and pre-trained with a large amount of medical text.
[0040] Since the first prompt is used to instruct the data quality inspection model to determine whether it has mastered the knowledge related to the data to be inspected, the output results can be used to determine whether the data quality inspection model has mastered the knowledge related to the data to be inspected.
[0041] Step 120: If, based on the output result, it is determined that the data quality inspection model does not possess knowledge related to the data to be inspected, the data to be inspected and the knowledge data related to the data to be inspected are input into the data quality inspection model to obtain the data quality inspection result output by the data quality inspection model.
[0042] Here, the knowledge data related to the data to be inspected is external knowledge data, which is obtained by searching based on the data to be inspected. Specifically, a knowledge search is performed on the data to be inspected. If relevant knowledge data is found, the data to be inspected and the relevant knowledge data are input into the data quality inspection model to obtain the data quality inspection result output by the model. If no relevant knowledge data is found, the first preset data quality inspection result is determined as the final data quality inspection result.
[0043] In one specific embodiment, a third prompt is constructed based on the data to be inspected and related knowledge data; the third prompt is then input into the data quality inspection model to obtain the data quality inspection result output by the model. Further, a third prompt is constructed based on the data to be inspected, related knowledge data, and a third prompt template. More specifically, the data to be inspected and related knowledge data are concatenated into the third prompt template to obtain the third prompt.
[0044] The third prompt is used to instruct the data quality inspection model to evaluate the conflict between the data to be inspected and the related knowledge data.
[0045] For example, the third prompt template is as follows: "Please first assess the conflict between the relevant knowledge data and the given data to be inspected, based on your knowledge. Give a score between 0 and 1 for the conflict between the two pieces of knowledge, where 0 represents the lowest (the two pieces of knowledge do not conflict at all), and vice versa. The following is the content of the given data to be inspected: xxx. Search knowledge content: xxx."
[0046] Since the third prompt is used to instruct the data quality inspection model to evaluate the conflict between the data to be inspected and the related knowledge data, the data quality inspection result can be directly determined based on the output of the data quality inspection model.
[0047] Furthermore, based on the data to be inspected, the knowledge data related to the data to be inspected, and the source of the knowledge data related to the data to be inspected, a third prompt is constructed; this third prompt is also used to instruct the data quality inspection model to evaluate the authority of the source of the knowledge data related to the data to be inspected.
[0048] For example, the third prompt template is as follows: "Please first assess the authority of the source of the relevant knowledge data based on your knowledge, and then assess the conflict between the relevant knowledge data and the given data to be inspected. Give a score between 0 and 1 for the authority of the source of the searched knowledge and the conflict between the two pieces of knowledge, where 0 represents the lowest (lowest source authority / no conflict between the two pieces of knowledge), and vice versa. The following is the content of the given data to be inspected: xxx. Source: xxx. Searched knowledge content: xxx."
[0049] It should be understood that, considering the data quality inspection model may be stateless and memoryless, it is necessary to input the data to be inspected into the data quality inspection model again.
[0050] Step 130: If, based on the output results, it is determined that the data quality inspection model has mastered the relevant knowledge of the data to be inspected, then, based on the output results, the data quality inspection result is determined.
[0051] It should be understood that if a data quality inspection model possesses knowledge related to the data to be inspected, it can correctly perform quality inspection on that data, thus directly determining the data quality inspection result based on the output. Therefore, the knowledge already possessed by the data quality inspection model only requires a single call to obtain the correct data quality inspection result, thereby improving data quality inspection efficiency.
[0052] In one specific embodiment, if the data to be inspected is determined to be incorrect based on the output result, the second preset data quality inspection result is determined as the data quality inspection result; if the data to be inspected is determined to be correct based on the output result, the third preset data quality inspection result is determined as the data quality inspection result; the quality inspection score represented by the third preset data quality inspection result is greater than the quality inspection score represented by the second preset data quality inspection result.
[0053] Furthermore, the first prompt is also used to instruct the data quality inspection model to determine whether the data to be inspected is correct, and to instruct the data quality inspection model to correct errors in the data to be inspected. Based on this, if the data to be inspected is determined to be incorrect based on the output results, the output results include the corrected data. It should be understood that this embodiment of the invention fully utilizes the capabilities of the large model itself to perform possible corrections based on the knowledge that the data to be inspected is incorrect, rather than directly losing the data to be inspected, thereby improving data utilization.
[0054] It should be understood that the above steps can be repeated to perform data quality checks on multiple datasets to be inspected. Based on the data quality check results of these datasets, the data to be used can be selected from them. For example, if the data quality check results are used to characterize the quality score of the datasets to be inspected, then based on the quality scores of the multiple datasets to be inspected, a corresponding sampling ratio is set to sample the multiple datasets. For example, if the datasets to be inspected are training data, in subsequent model training, sampling and subsequent training are performed according to the corresponding quality score (less than 1, i.e., quality weight).
[0055] The data quality inspection method based on a large model provided in this invention inputs a first prompt message constructed based on the data to be inspected into the data quality inspection model, obtaining the output result of the data quality inspection model. The first prompt message instructs the data quality inspection model to determine whether it possesses knowledge related to the data to be inspected, and also instructs the data quality inspection model to perform quality inspection on the data. Therefore, if the output result determines that the data quality inspection model does not possess knowledge related to the data to be inspected, the data to be inspected and related knowledge data are input into the data quality inspection model, obtaining the data quality inspection result output by the model. Furthermore, even when the data quality inspection model does not possess knowledge related to the data to be inspected, it can output relevant results to instruct the input of the data to be inspected and related knowledge data into the data quality inspection model, i.e., importing relevant knowledge data so that the data quality inspection model can acquire knowledge related to the data to be inspected, thereby performing quality inspection. Secondary data quality inspection avoids model illusions and outputs correct data quality inspection results. This means the data quality inspection model can rely not only on its own capabilities but also on external knowledge data to improve accuracy. Whether or not to use external knowledge data is determined automatically by the data quality inspection model. Based on the output results, if the data quality inspection model possesses relevant knowledge about the data to be inspected, the data quality inspection result is determined accordingly. Therefore, regardless of whether the data quality inspection model possesses relevant knowledge, no manual judgment is needed for the final data quality inspection result, achieving automated data quality inspection. This reduces labor costs and improves efficiency, while avoiding subjective human judgment to enhance accuracy. Furthermore, the knowledge already possessed by the data quality inspection model only requires a single call to obtain the correct data quality inspection result, further improving efficiency.
[0056] Based on any of the above embodiments, a specific embodiment of the data quality inspection method based on a large model is given below. Figure 2 is a second schematic flowchart of the data quality inspection method based on a large model provided by the present invention. As shown in Figure 2, step 110 includes steps 111, 112, 113, and 114.
[0057] Step 111: Based on the data to be inspected, construct a second prompt message.
[0058] The second prompt is used to instruct the data quality inspection model to break down the data to be inspected into knowledge points.
[0059] In one specific embodiment, a second prompt is constructed based on the data to be inspected and a second prompt template. More specifically, the data to be inspected is concatenated into the second prompt template to obtain the second prompt.
[0060] For example, if the data to be inspected is medical Q&A data, and the question in the medical Q&A data is "What are some medications for treating stomach cramps?", and the answer in the medical Q&A data is "Medications for treating stomach cramps generally include omeprazole, cinnarizine, montmorillonite powder, domperidone, and amoxicillin", then the second prompt is as follows: "The following is a pair of medical knowledge Q&A. Please extract the medical viewpoints contained in the question and answer and output them in JSON format."
[0061] The following is a question: What medications are available for treating stomach cramps? The following is an answer: Medications for treating stomach cramps generally include omeprazole, cinnarizine, montmorillonite powder, domperidone, and amoxicillin.
[0062] In one specific embodiment, the data to be inspected is question-and-answer data. Based on the answers in the data to be inspected, a second prompt is constructed. That is, the second prompt is used to instruct the data quality inspection model to perform knowledge point decomposition on the answers. Further, the second prompt is constructed based on the questions and answers in the data to be inspected. It should be understood that, while it may actually be a matter of decomposing knowledge points on the answers, the questions can assist in better decomposing the answers into knowledge points, thereby improving the accuracy of the knowledge point decomposition.
[0063] Step 112: Input the second prompt into the data quality inspection model to obtain the splitting result output by the data quality inspection model.
[0064] The breakdown result includes several knowledge points. Since the second prompt instructs the data quality inspection model to break down the data to be inspected into knowledge points, the breakdown result output by the data quality inspection model includes several knowledge points. Specifically, these knowledge points can be obtained by parsing the breakdown result.
[0065] For example, the data to be inspected includes the answer "Medications for treating stomach cramps generally include omeprazole, cinnarizine, montmorillonite powder, domperidone, and amoxicillin." Based on this, the breakdown results are as follows: ["Omeprazole can treat stomach cramps.", "Cinnarizine can treat stomach cramps.", "Montmorillonite powder can treat stomach cramps.", "Domperidone can treat stomach cramps.", "Amoxicillin can treat stomach cramps."] Based on this, the first knowledge point to be broken down is "Omeprazole can treat stomach cramps," and so on.
[0066] Step 113: Based on the splitting results, construct the first prompt message.
[0067] The first prompt includes a first sub-prompt constructed based on the plurality of split knowledge points; the first sub-prompt constructed based on any of the split knowledge points is used to instruct the data quality inspection model to determine whether it has mastered the knowledge related to the split knowledge point, and to instruct the data quality inspection model to perform quality inspection on the split knowledge point.
[0068] In one specific embodiment, for any split knowledge point, a first sub-prompt is constructed based on the split knowledge point and the first prompt template. More specifically, the split knowledge point is concatenated into the first prompt template to obtain the first sub-prompt.
[0069] For example, the first prompt template is as follows: "Please determine if the following content is correct. If there are any parts you are unsure about, please use..." Mark it. The following is the content: xxx. The content is a breakdown of knowledge points. "Please judge whether the following content is correct" instructs the data quality control model to perform quality checks on the broken-down knowledge points. "If there are parts you are unsure about, please use..." "Mark" is used to instruct the data quality inspection model to determine whether it has mastered the knowledge related to the split knowledge points.
[0070] Step 114: Input each of the first sub-prompts into the data quality inspection model to obtain the output results of the data quality inspection model.
[0071] The output results include the sub-output results corresponding to each of the first sub-prompts.
[0072] Since any first sub-prompt is used to instruct the data quality inspection model to determine whether it has mastered the knowledge related to the split knowledge points, the data quality inspection model can be determined to have mastered the knowledge related to the split knowledge points based on its corresponding sub-output results.
[0073] The data quality inspection method based on a large model provided in this invention inputs a second prompt into the data quality inspection model to obtain the splitting result output by the data quality inspection model. The second prompt is used to instruct the data quality inspection model to split the data to be inspected into knowledge points, that is, the splitting result includes several split knowledge points. Based on the splitting result, a first prompt is constructed, and the first prompt includes a first sub-prompt constructed based on several split knowledge points. The first sub-prompt constructed based on any split knowledge point is used to instruct the data quality inspection model to determine whether it has mastered the knowledge related to the split knowledge point, and to instruct the data quality inspection model to perform quality inspection on the split knowledge point. Thus, each first sub-prompt is input into the data quality inspection model to obtain the output result of the data quality inspection model. In this way, the data to be inspected is first split into knowledge points, and then each split knowledge point is inspected separately, thereby realizing fine-grained quality inspection and further improving the accuracy of data quality inspection.
[0074] Based on any of the above embodiments, a specific embodiment of the data quality inspection method based on a large model is given below. Figure 3 is a third schematic flowchart of the data quality inspection method based on a large model provided by the present invention. As shown in Figure 3, step 120 above includes steps 121, 122 and 123.
[0075] Step 121: If, based on the sub-output results, it is determined that the data quality inspection model does not possess knowledge related to at least one first target splitting knowledge point among the several splitting knowledge points, the at least one first target splitting knowledge point and the knowledge data related to the at least one first target splitting knowledge point are input into the data quality inspection model to obtain several first sub-data quality inspection results output by the data quality inspection model.
[0076] Here, the first target knowledge point is the knowledge point that the data quality inspection model has not mastered. The knowledge data related to the first target knowledge point is external knowledge data, which is obtained by searching based on the first target knowledge point. The number of quality inspection results for this first sub-data is the same as the number of the first target knowledge points.
[0077] In one specific embodiment, for any first target-splitting knowledge point, a third prompt is constructed based on the first target-splitting knowledge point and related knowledge data; the third prompt is input into a data quality inspection model to obtain the first sub-data quality inspection result output by the data quality inspection model. Further, a third prompt is constructed based on the first target-splitting knowledge point, related knowledge data, and a third prompt template. More specifically, the first target-splitting knowledge point and related knowledge data are concatenated into the third prompt template to obtain the third prompt.
[0078] The third prompt is used to instruct the data quality inspection model to evaluate the conflict between the knowledge points of the first target split and the knowledge data related to the knowledge points of the first target split.
[0079] It should be understood that, considering the data quality inspection model may be stateless and memoryless, it is necessary to input the first target split knowledge points into the data quality inspection model again.
[0080] Step 122: Based on the sub-output results corresponding to the knowledge points of several second objectives, determine several second sub-data quality inspection results respectively.
[0081] Among them, the several second target splitting knowledge points are the other splitting knowledge points other than the at least one first target splitting knowledge point.
[0082] It should be understood that if the data quality inspection model possesses knowledge related to the knowledge points of the second objective breakdown, then the model can correctly inspect these knowledge points and directly determine the quality inspection result of the second sub-data based on their corresponding sub-outputs. Therefore, the knowledge already possessed by the data quality inspection model only requires a single call to obtain the correct quality inspection result of the second sub-data, thereby improving data quality inspection efficiency.
[0083] Furthermore, the first sub-prompt is also used to instruct the data quality inspection model to determine whether the second target split knowledge point is correct, and to instruct the data quality inspection model to correct the second target split knowledge point. Based on this, if the sub-output result determines that the second target split knowledge point is incorrect, the sub-output result includes the corrected data. It should be understood that this embodiment of the invention fully utilizes the capabilities of the large model itself to correct any errors in the second target split knowledge point determination, rather than directly losing the data to be inspected, thereby improving data utilization.
[0084] Step 123: Determine the data quality inspection result based on the quality inspection results of each of the first sub-data and each of the second sub-data.
[0085] Specifically, the overall data quality inspection result is determined based on the quality inspection results of each first sub-data and each second sub-data, that is, based on the quality inspection results of each segmented knowledge point. If the quality inspection results of each sub-data are represented by quality inspection scores, the average of the quality inspection scores can be used to calculate the overall data quality inspection result.
[0086] For example, the first prompt template is as follows: "Please judge whether the following content is correct. If you find that there is an error, please mark the part you think is wrong with "_" and correct it. If there is a part you are unsure about, please use '_'." "Mark this. The following is the content: xxx."
[0087] Suppose we have a set of five sub-knowledge points: the first is "Omeprazole can treat stomach cramps," the second is "Cinnarizine can treat stomach cramps," the third is "Montmorillonite powder can treat stomach cramps," the fourth is "Domperidone can treat stomach cramps," and the fifth is "Amoxicillin can treat stomach cramps." Based on this, the sub-outputs for each sub-knowledge point are: "Omeprazole can treat stomach cramps.", "Cinnarizine can treat _stomach cramps_ (cerebral embolism).", "Montmorillonite powder can treat stomach cramps.", "Domperidone can treat stomach cramps.", and " Amoxicillin "It can treat stomach cramps." This means that the first to fourth subdivided knowledge points are all second-target subdivided knowledge points, the fifth subdivided knowledge point is the first-target subdivided knowledge point, and "Cinnarizine can treat _stomach cramps_ (cerebral embolism)" is used to indicate "Cinnarizine can treat cerebral embolism", thereby correcting the error.
[0088] Assuming the quality control result of the first sub-data of the fifth knowledge point is 0.55, when the sub-output result corresponding to the second target knowledge point does not contain '_' and '...' When ' is present, the quality inspection result of the corresponding second sub-data is 1. When the sub-output result corresponding to the knowledge point of the second target contains '_', the quality inspection result of the corresponding second sub-data is 0.4. Based on this, the data quality inspection results are as follows: (1+0.4+1+1+0.55) / 5=0.79.
[0089] It should be understood that, compared with the existing technology which only involves two operations for data processing after quality inspection: retention and cleaning, the data quality inspection results of the present invention can be subject to detailed quality assessment. That is, sampling can be performed based on the final data quality inspection score, and data is not directly cleaned if the score is less than 1, thereby improving data utilization.
[0090] The data quality inspection method based on a large model provided in this invention can perform quality inspection on each segmented knowledge point in the above manner, thereby achieving fine-grained quality inspection and further improving the accuracy of data quality inspection.
[0091] Based on any of the above embodiments, a specific embodiment of the data quality inspection method based on a large model is given below. In this method, step 121 includes: when it is determined based on each of the sub-output results that the data quality inspection model does not possess knowledge related to at least one first target split knowledge point among the plurality of split knowledge points, a knowledge search is performed on each of the at least one first target split knowledge point; for each first target split knowledge point, the following is performed: if knowledge data related to the first target split knowledge point is found, the first target split knowledge point and the knowledge data related to the first target split knowledge point are input into the data quality inspection model to obtain a first sub-data quality inspection result output by the data quality inspection model; if no knowledge data related to the first target split knowledge point is found, a first preset data quality inspection result is determined as the first sub-data quality inspection result.
[0092] In one specific embodiment, an external retrieval tool is used to perform knowledge searches on at least one first target split knowledge point.
[0093] It should be understood that this knowledge data should be the data most relevant to the knowledge points split from the first target in the search data.
[0094] Here, the first preset data quality inspection result is pre-set; if the data quality inspection result is represented by a quality inspection score, this first preset data quality inspection result can be set to 0.5, that is, a compromise value. For example, when the sub-output result contains ' 'However, if no relevant content is found, the quality inspection score is 0.5.'
[0095] The data quality inspection method based on a large model provided in this invention performs knowledge search on at least one first target split knowledge point in the manner described above. If knowledge data related to the first target split knowledge point is found, the first target split knowledge point and the related knowledge data are input into the data quality inspection model to obtain the first sub-data quality inspection result output by the data quality inspection model. If no knowledge data related to the first target split knowledge point is found, the first preset data quality inspection result is determined as the first sub-data quality inspection result. Thus, the first sub-data quality inspection result can be obtained regardless of whether related knowledge data is found, thereby ensuring that the data quality inspection process is executed normally and will not be stopped due to being stuck at a certain step, thus ensuring the stability of data quality inspection.
[0096] Based on any of the above embodiments, a specific embodiment of a data quality inspection method based on a large model is given below. In this method, the step of inputting the first target-splitting knowledge point and related knowledge data into the data quality inspection model to obtain the first sub-data quality inspection result output by the data quality inspection model includes: constructing a third prompt based on the first target-splitting knowledge point and related knowledge data; and inputting the third prompt into the data quality inspection model to obtain the first sub-data quality inspection result output by the data quality inspection model.
[0097] In one specific embodiment, a third prompt is constructed based on the first target-decomposed knowledge points, related knowledge data, and a third prompt template. More specifically, the first target-decomposed knowledge points and related knowledge data are concatenated into the third prompt template to obtain the third prompt.
[0098] The third prompt is used to instruct the data quality inspection model to evaluate the conflict between the first target-splitting knowledge point and the related knowledge data. It should be understood that the external knowledge data is correct; if the first target-splitting knowledge point has a significant conflict with the external knowledge data, it indicates that the first target-splitting knowledge point is not entirely correct.
[0099] For example, the third prompt template is as follows: "Please first assess the conflict between the relevant knowledge data and the content of the given broken knowledge points based on your knowledge. Give a score between 0 and 1 for the conflict between the two knowledge points, where 0 represents the lowest (the two knowledge points do not conflict at all), and vice versa. The following is the content of the given broken knowledge points: xxx. Search for knowledge content: xxx."
[0100] Since the third prompt is used to instruct the data quality inspection model to evaluate the conflict between the knowledge points of the first target split and the knowledge data related to the knowledge points of the first target split, the result of the first sub-data quality inspection can be directly determined based on the output of the data quality inspection model.
[0101] In one specific embodiment, a third prompt is input into the data quality inspection model to obtain the conflict assessment result output by the data quality inspection model. Based on the conflict assessment result, the first sub-data quality inspection result is determined. For example, if the conflict assessment result is represented by a conflict score, and the conflict assessment result is 0.9, then the first sub-data quality inspection result is 1-0.9=0.1.
[0102] The data quality inspection method based on a large model provided in this invention evaluates the conflict between the knowledge points of the first target split and the knowledge data related to the knowledge points of the first target split in the above manner, thereby accurately determining the quality inspection result of the first sub-data based on the conflict score of the two knowledge points, and thus improving the accuracy of data quality inspection.
[0103] Based on any of the above embodiments, a specific embodiment of a data quality inspection method based on a large model is given below. In this method, constructing a third prompt based on the first target-splitting knowledge point and related knowledge data includes: constructing a third prompt based on the first target-splitting knowledge point, related knowledge data, and the source of the related knowledge data.
[0104] In one specific embodiment, a third prompt is constructed based on the first target-decomposed knowledge point, the knowledge data related to the first target-decomposed knowledge point, the source of the knowledge data related to the first target-decomposed knowledge point, and a third prompt template. More specifically, the first target-decomposed knowledge point, the knowledge data related to the first target-decomposed knowledge point, and the source of the knowledge data related to the first target-decomposed knowledge point are concatenated into the third prompt template to obtain the third prompt.
[0105] The third prompt also instructs the data quality inspection model to evaluate the source authority of the knowledge data related to the first target split knowledge point. It should be understood that while external knowledge data is generally accurate, some data may lack sufficient authority; therefore, the source authority is also evaluated.
[0106] For example, assuming the first objective is to break down knowledge points into medical knowledge, the third prompt template would be as follows: "The following is a passage of medical knowledge. Currently, we are using underscores (two lines) in Markdown format..." The bolded portion is questionable, and a search yielded related knowledge that needs verification. First, based on your knowledge, assess the authority of the source of this related knowledge. Then, assess the conflict between this knowledge and the given medical knowledge. Output in JSON format, giving a score between 0 and 1 for both the authority of the searched knowledge source and the conflict between the two pieces of knowledge, where 0 represents the lowest (lowest authority / no conflict), and higher represents the highest. The following is the content of the given medical knowledge: xxx. The following is the searched knowledge: Source: xxx. Searched knowledge content: xxx.
[0107] For example, the third prompt is as follows: "The following is a piece of medical knowledge. Currently, we are using underscores (two lines) in Markdown format." The bolded portion is questionable, and a search yielded related information that needs verification. First, based on your knowledge, assess the authority of the source of this related information. Then, based on the content of this information and the given medical knowledge, assess the potential conflict between the two. Output in JSON format, providing a score between 0 and 1 for both the authority of the source of the searched information and the potential conflict between the two pieces of knowledge. 0 represents the lowest (lowest authority / no conflict between the two pieces of knowledge), and vice versa.
[0108] The following is the content of the given medical knowledge: Amoxicillin It can treat stomach cramps.
[0109] The following information was found: Source: Internet - Website: Miaoshou Doctor.
[0110] Searching for knowledge content: Amoxicillin cannot cure stomach ulcers. Amoxicillin is an antibiotic used for anti-infective treatment. If a stomach ulcer patient has Helicobacter pylori infection, then using amoxicillin in combination with clarithromycin and omeprazole to eradicate Helicobacter pylori helps promote ulcer healing. "Since the third prompt is used to instruct the data quality inspection model to evaluate the conflict between the first target split knowledge point and related knowledge data, and to evaluate the source authority of the knowledge data related to the first target split knowledge point, the results based on the output of the data quality inspection model can directly determine the quality inspection result of the first sub-data."
[0111] In one specific embodiment, a third prompt is input into the data quality inspection model to obtain the conflict assessment result and the source authority assessment result output by the data quality inspection model. Based on the conflict assessment result and the source authority assessment result, the first sub-data quality inspection result is determined. For example, the conflict assessment result is represented by a conflict score, and the source authority assessment result is represented by a source authority score. If the conflict assessment result is 0.9 and the source authority assessment result is 0.5, then the first sub-data quality inspection result is 1-0.5. 0.9 = 0.55.
[0112] The data quality inspection method based on a large model provided in this invention evaluates the conflict between the knowledge points of the first target split and the knowledge data related to the first target split in the above manner, and also evaluates the source authority of the knowledge data related to the first target split knowledge points, thereby further improving the accuracy of the determination of the first sub-data quality inspection result, and further improving the accuracy of data quality inspection.
[0113] Based on any of the above embodiments, a specific embodiment of the data quality inspection method based on a large model is given below. In this method, the first sub-prompt constructed based on any of the described split knowledge points is also used to instruct the data quality inspection model to determine whether the split knowledge points are correct, and to instruct the data quality inspection model to correct the split knowledge points.
[0114] For example, the first prompt template is as follows: "Please judge whether the following content is correct. If you find that there is an error, please mark the part you think is wrong with "_" and correct it. If there is a part you are unsure about, please use '_'." "Mark this. The following is the content: xxx."
[0115] Accordingly, the quality inspection result of the second sub-data corresponding to any second target split knowledge point is determined in the following manner: if the second target split knowledge point is determined to be incorrect based on the sub-output result corresponding to the second target split knowledge point, the second preset data quality inspection result is determined as the quality inspection result of the second sub-data corresponding to the second target split knowledge point; if the second target split knowledge point is determined to be correct based on the sub-output result corresponding to the second target split knowledge point, the third preset data quality inspection result is determined as the quality inspection result of the second sub-data corresponding to the second target split knowledge point.
[0116] The quality inspection score represented by the third preset data quality inspection result is greater than the quality inspection score represented by the second preset data quality inspection result.
[0117] Here, the second and third preset data quality inspection results can be set in advance.
[0118] For example, when the sub-output results corresponding to the knowledge points split into the second objective do not contain '_' and ' When ' is present, the corresponding second sub-data quality inspection result is 1, that is, the third preset data quality inspection result is 1. When the sub-output result corresponding to the second target split knowledge point contains '_', the corresponding second sub-data quality inspection result is 0.4, that is, the second preset data quality inspection result is 0.4.
[0119] Furthermore, the quality inspection score represented by the second preset data quality inspection result is less than the quality inspection score represented by the first preset data quality inspection result, thereby ensuring that the quality inspection score of the clearly identified erroneous data is less than the quality inspection score of the data whose correctness is not clearly determined.
[0120] The data quality inspection method based on a large model provided in this invention includes a first sub-prompt that instructs the data quality inspection model to determine whether the split knowledge point is correct, and to instruct the data quality inspection model to correct the split knowledge point. Based on this, when the split knowledge point is determined to be incorrect based on the sub-output result, the output result includes the corrected data, thereby making full use of the large model's own capabilities to correct the knowledge that the split knowledge point is determined to be incorrect, rather than directly losing the data to be inspected, thus improving data utilization. Furthermore, by determining the second sub-data quality inspection result corresponding to the second target split knowledge point in the above manner, it can be ensured that the quality inspection score for determining the split knowledge point to be correct is greater than the quality inspection score for determining the split knowledge point to be incorrect. Moreover, by determining the second sub-data quality inspection result through a pre-set preset data quality inspection result, the data quality inspection efficiency can be improved.
[0121] Based on any of the above embodiments, another specific embodiment of the data quality inspection method based on a large model is given below. In this method, the data quality inspection model is trained in the following way: the large model is trained based on training data to obtain the data quality inspection model.
[0122] Here, the model training method may include, but is not limited to, at least one of the following: full parameter fine-tuning, LORA fine-tuning, etc.
[0123] Accordingly, Figure 4 is the fourth flowchart of the data quality inspection method based on a large model provided by the present invention. As shown in Figure 4, the training data is constructed in the following manner.
[0124] Step 410: Input the target question into the large model and obtain the output answer from the large model.
[0125] The target test question is designed to assess knowledge related to the sample data to be inspected. This target test question is used to determine whether the large model possesses the relevant knowledge of the sample data to be inspected. For example, the target test question could be a fill-in-the-blank question: "Dopamine can act on α-receptors in blood vessels, causing blood vessels to ___, thereby achieving the effect of raising blood pressure."
[0126] Here, the sample data to be inspected refers to the data that needs to be inspected. This sample data to be inspected can be set according to actual needs, such as question-and-answer data, sample data from the training dataset, or knowledge data, etc. The number of this sample data to be inspected can be multiple, thereby continuously optimizing the large model.
[0127] In one embodiment, the sample data to be quality inspected is question-and-answer data, which includes questions and answers. This question-and-answer data can be single-round question-and-answer data, such as single-round medical question-and-answer data.
[0128] For example, if the question is "What is the mechanism by which dopamine raises blood pressure?", the answer is "Dopamine can act on α receptors in blood vessels, causing vasodilation, thereby raising blood pressure."
[0129] In another embodiment, the sample data to be quality checked is the answers in the question-and-answer data, i.e., it may not include the questions. In another embodiment, the sample data to be quality checked is the training data in the training dataset.
[0130] Here, the output answer is the answer to the target question, but this output answer is not necessarily the correct answer.
[0131] Step 420: In the case of an incorrect output answer, based on the sample data to be inspected, construct a first prompt for the sample and a first output result label corresponding to the first prompt for the sample, so that the training data includes the first prompt for the sample and the first output result label.
[0132] Here, the first prompt is used to instruct the large model to determine whether it has mastered the knowledge related to the sample data to be inspected, and to instruct the large model to perform quality inspection on the sample data to be inspected.
[0133] In one specific embodiment, a first prompt message for the sample is constructed based on the sample data to be inspected and the first prompt message template. More specifically, the sample data to be inspected is concatenated into the first prompt message template to obtain the first prompt message for the sample.
[0134] For example, the first prompt template is as follows: "Please determine if the following content is correct. If there are any parts you are unsure about, please mark them. The following is the content: xxx." Here, the content is the sample data to be inspected, "Please determine if the following content is correct" instructs the large model to inspect the sample data, and "If there are any parts you are unsure about, please mark them" instructs the large model to determine whether it has mastered the knowledge related to the sample data to be inspected.
[0135] The first output label indicates that the large model lacks knowledge related to the sample data to be inspected. When the output answer is incorrect, it signifies that the large model lacks knowledge related to the sample data to be inspected. Therefore, the first output label is also used to indicate that the large model lacks knowledge related to the sample data to be inspected. Training the large model with training data constructed based on this label is more conducive to determining the knowledge boundary of the large model, allowing it to perform knowledge quality inspection within the knowledge boundary. That is, it can accurately obtain the output result during the reasoning process. In other words, it enables the data quality inspection model to more accurately determine whether it has mastered the knowledge related to the data to be inspected, thereby improving the reliability and accuracy of data quality inspection.
[0136] Step 430: If the output answer is correct, based on the sample data to be inspected, construct the first prompt of the sample and the second output result label corresponding to the first prompt of the sample, so that the training data includes the first prompt of the sample and the second output result label.
[0137] The second output label indicates that the large model has mastered the knowledge related to the sample data to be inspected. When the output answer is correct, it means the large model has mastered the knowledge related to the sample data to be inspected. Therefore, the second output label is also used to indicate that the large model has mastered the knowledge related to the sample data to be inspected. Training the large model with training data constructed based on this label is more conducive to judging the knowledge boundary of the large model, allowing it to perform knowledge quality inspection within the knowledge boundary. That is, it can accurately obtain the output result during the reasoning process. In other words, it enables the data quality inspection model to more accurately determine whether it has mastered the knowledge related to the data to be inspected, thereby improving the reliability and accuracy of data quality inspection.
[0138] It should be understood that by determining whether the output answer of the large model for the target question is correct, if it is correct, the large model can better understand its knowledge of the target question by constructing training data; if it is incorrect, the large model can also better understand its lack of knowledge of the target question by constructing training data.
[0139] The data quality inspection method based on a large model provided in this invention determines whether the large model has mastered the knowledge related to the target test question through the above-mentioned method, and constructs corresponding training data based on the judgment result. This is more conducive to judging the knowledge boundary of the large model, allowing it to perform knowledge quality inspection within the knowledge boundary. That is, it can accurately obtain the output result during the reasoning process. In other words, it enables the data quality inspection model to more accurately judge whether it has mastered the knowledge related to the data to be inspected, thereby improving the reliability and accuracy of data quality inspection.
[0140] Based on any of the above embodiments, a specific embodiment of a data quality inspection method based on a large model is given below. In this method, the sample data to be inspected is erroneous data, and the target test question is generated in the following manner: a fourth prompt is constructed based on the sample data to be inspected and the error cause data of the sample data to be inspected; the fourth prompt is input into the large model to obtain the test question output result of the large model.
[0141] Here, error cause data is used to characterize the cause of errors in the sample data to be inspected. For example, if the sample data to be inspected includes "dopamine can act on α receptors in blood vessels, causing vasodilation, thereby raising blood pressure," then the error cause data could be "vasoconstriction is necessary to raise blood pressure." In one embodiment, this error cause data can be manually labeled; in another embodiment, this error cause data can be automatically labeled by machine (e.g., by labeling a large model).
[0142] The fourth prompt is used to instruct the large model to refer to the sample data to be inspected and the error cause data to construct test questions that examine the knowledge related to the error cause data.
[0143] In one specific embodiment, a fourth prompt is constructed based on the sample data to be inspected, the error cause data, and the fourth prompt template. More specifically, the sample data to be inspected and the error cause data are concatenated into the fourth prompt template to obtain the fourth prompt.
[0144] For example, the sample data to be inspected is question-and-answer data, and the template for the fourth prompt is as follows: "The following is a pair of knowledge questions and answers that have been confirmed as incorrect. Please construct the test questions based on the question, the answer, and the given error reason data. The following is the question: xxx. The following is the answer: xxx. The following is the error reason: xxx."
[0145] The test output includes the target test question. There can be multiple target test questions, allowing for the determination of whether the large model has mastered the knowledge related to the error causes of the data through multiple different test questions. This improves the accuracy of training data construction and ultimately enhances the reliability and accuracy of data quality inspection.
[0146] Furthermore, the test output includes multiple target test questions of different types. By using multiple test questions of different types, it is possible to determine whether the large model has mastered the knowledge related to the error cause data, thereby improving the accuracy of training data construction and ultimately improving the reliability and accuracy of data quality inspection.
[0147] Furthermore, based on the sample data to be inspected, the error cause data of the sample data to be inspected, and the test question construction requirements, a fourth prompt is constructed to better construct the target test questions, thereby better determining whether the large model has mastered the knowledge related to the error cause data, and ultimately improving the reliability and accuracy of data quality inspection.
[0148] For example, the sample data to be inspected is question-and-answer data, and the template for the fourth prompt is as follows: "The following is a pair of knowledge questions and answers that have been confirmed as incorrect. Based on the question, the answer, and the given error reason data, please construct one fill-in-the-blank question, one multiple-choice question, and one true / false question, and output them in JSON format. If you believe that you cannot construct this type of question, then this field should be an empty list. The requirements for the three types of questions are as follows: Fill-in-the-blank question: Please refer to the given error reason, remove the incorrect part of the answer, and construct the corresponding fill-in-the-blank question in combination with the question, and give the actual correct answer to the fill-in-the-blank question."
[0149] Multiple choice questions: Please refer to the given reasons for errors, questions, and answers to construct multiple choice questions based on the concepts related to the reasons for errors, and provide the actual correct answers.
[0150] True or False Questions: Please refer to the given reasons for the error, combine them with the questions, construct the corresponding true or false questions based on the concepts related to the reasons for the error, and give the actual answers to the true or false questions.
[0151] The following is a question: xxx.
[0152] The following is the answer: xxx.
[0153] The following is the reason for the error: xxx.
[0154] Furthermore, the test output also includes the correct answer to the target test question, which is used to determine whether the output answer is correct.
[0155] It should be understood that for medical knowledge Q&A, presenting medical knowledge points in appropriate question formats helps determine whether the large model has mastered the knowledge. This is more conducive to judging the knowledge boundaries of the large model, allowing it to conduct quality checks within those boundaries, thus improving the reliability of the quality checks.
[0156] The data quality inspection method based on a large model provided in this invention sets the sample data to be inspected as erroneous data. Based on the sample data to be inspected and the error cause data of the sample data to be inspected, a fourth prompt is constructed. The fourth prompt is used to instruct the large model to refer to the sample data to be inspected and the error cause data to construct test questions that examine the knowledge related to the error cause data. Thus, the fourth prompt is input into the large model to obtain the accurate test question output results from the large model, thereby improving the accuracy of the construction of training data and ultimately improving the reliability and accuracy of data quality inspection.
[0157] Based on any of the above embodiments, a specific embodiment of a data quality inspection method based on a large model is given below. In this method, the construction of a fourth prompt based on the sample data to be inspected and the error cause data of the sample data to be inspected includes: constructing a fourth prompt based on the sample data to be inspected, the error cause data of the sample data to be inspected, and several existing test questions related to the sample data to be inspected.
[0158] Here, the number of existing test questions can be preset. In one specific embodiment, based on the sample data to be inspected, a search is performed in an existing exam question bank to obtain a preset number of existing test questions that are most relevant to the sample data to be inspected. For example, if the sample data to be inspected is medical Q&A data, a search is performed in an existing medical exam question bank based on the medical Q&A data to obtain a preset number of medical test questions that are most relevant to the medical Q&A data.
[0159] The fourth prompt is used to instruct the large model to construct test questions that examine knowledge related to the error cause data, referring to the sample data to be inspected, the error cause data, and the several existing test questions.
[0160] For example, the sample data to be inspected is question-and-answer data, and the template for the fourth prompt is as follows: "The following is a pair of knowledge questions and answers that have been confirmed as incorrect. Please construct a question based on the question and answer, the given error reason data, and the given existing questions. The following is the question: xxx. The following is the answer: xxx. The following is the error reason: xxx. The following is the existing question: xxx."
[0161] The data quality inspection method based on a large model provided in this invention constructs a fourth prompt based on the sample data to be inspected, the error cause data of the sample data to be inspected, and several existing test questions related to the sample data to be inspected. The fourth prompt is used to instruct the large model to construct test questions that examine the knowledge related to the error cause data, referring to the sample data to be inspected, the error cause data, and several existing test questions. This allows for the construction of more accurate target test questions by referring to existing test questions, thereby improving the accuracy of training data construction and ultimately improving the reliability and accuracy of data quality inspection.
[0162] Based on any of the above embodiments, a specific embodiment of a data quality inspection method based on a large model is given below. In this method, the construction of a fourth prompt based on the sample data to be inspected, the error cause data of the sample data to be inspected, and several existing test questions related to the sample data to be inspected includes: constructing a fourth prompt based on the sample data to be inspected, the error cause data of the sample data to be inspected, several existing test questions related to the sample data to be inspected, and the correct answers to each of the existing test questions.
[0163] In one specific embodiment, based on the sample data to be inspected, a search is performed in the existing test question bank to obtain a preset number of existing test questions that are most relevant to the sample data to be inspected, as well as the correct answers to each existing test question.
[0164] The fourth prompt is used to instruct the large model to construct test questions that examine knowledge related to the error cause data, referring to the sample data to be inspected, the error cause data, the existing test questions, and the correct answers to each of the existing test questions.
[0165] For example, the sample data to be inspected is medical Q&A data. The question in this medical Q&A data is "What is the mechanism by which dopamine raises blood pressure?" The answer in this medical Q&A data is "Dopamine can act on α receptors in blood vessels, causing vasodilation, thereby achieving the effect of raising blood pressure." The error reason data is "Vasoconstriction is necessary to raise blood pressure." There is currently 1 question, which is "Regarding the pharmacological effects of dopamine, which of the following statements is correct? A. At low doses, it stimulates β receptors; B. At slightly higher doses, it stimulates D2 receptors; C. At high doses, it stimulates α receptors; D. It is mostly used for heart failure after myocardial infarction." The correct answer to the question is "C". Based on this, the fourth prompt is as follows: "The following is a pair of medical knowledge questions and answers that have been confirmed as incorrect. Please construct one fill-in-the-blank question, one multiple-choice question, and one true / false question based on the question, the answer, and the given reason for the error, and refer to the given relevant medical exam questions. Output them in JSON format. If you believe that the referenced medical exam questions cannot be used to construct this type of question, then this field should be an empty list. The requirements for the three types of questions are as follows: Fill-in-the-blank question: Please refer to the given reason for the error, remove the incorrect part of the answer, and construct the corresponding fill-in-the-blank question based on the question, and provide the actual correct answer to the fill-in-the-blank question."
[0166] Multiple choice questions: Please refer to the given reasons for errors, questions and answers, as well as similar medical exam questions, to construct multiple choice questions based on the concepts related to the reasons for errors, and provide the actual correct answers.
[0167] True or False Questions: Please refer to the given reasons for the error, combine them with the questions, construct the corresponding true or false questions based on the concepts related to the reasons for the error, and give the actual answers to the true or false questions.
[0168] The following is a question: What is the mechanism by which dopamine raises blood pressure?
[0169] The answer is: Dopamine can act on α receptors in blood vessels, causing vasodilation and thus raising blood pressure.
[0170] Here's why it's wrong: Blood pressure only rises when blood vessels constrict.
[0171] The following is a sample medical exam question: Regarding the pharmacological effects of dopamine, which of the following statements is correct? A. At low doses, it stimulates β receptors; B. At slightly higher doses, it stimulates D2 receptors; C. At high doses, it stimulates α receptors; D. It is mostly used for heart failure after myocardial infarction. The correct answer to this question is "C".
[0172] Accordingly, the test output results are as follows: "Multiple choice question: "How does dopamine promote the increase of blood pressure? A. Acts on α receptors in blood vessels, causing vasodilation B. Acts on D2 receptors, causing vasoconstriction C. Acts on α receptors in blood vessels, causing vasoconstriction D. Acts on β receptors in blood vessels, causing vasoconstriction Correct answer: C" Fill in the blank question: "Dopamine can act on α receptors in blood vessels, causing blood vessels to ___, thereby achieving the effect of raising blood pressure."
[0173] Correct answer: contraction. True or False: "Dopamine can act on α receptors in blood vessels, causing vasodilation, thereby raising blood pressure."
[0174] Correct answer: Incorrect.
[0175] The test output also includes the correct answer to the target test question, which is used to determine whether the output answer is correct.
[0176] The data quality inspection method based on a large model provided in this invention constructs a fourth prompt based on the sample data to be inspected, error cause data of the sample data to be inspected, several existing test questions related to the sample data to be inspected, and the correct answers of each existing test question. The fourth prompt is used to instruct the large model to construct test questions that examine knowledge related to the error cause data, referring to the sample data to be inspected, the error cause data, several existing test questions, and the correct answers of each existing test question. This allows for the construction of more accurate and comprehensive target test questions by referring to existing test questions and their correct answers, thereby improving the accuracy of training data construction and ultimately improving the reliability and accuracy of data quality inspection. Furthermore, by referring to existing test questions and their correct answers to construct more accurate correct answers for target test questions, the correctness of the output answer is accurately determined, thereby improving the accuracy of training data construction and ultimately improving the reliability and accuracy of data quality inspection.
[0177] Based on any of the above embodiments, a specific embodiment of a data quality inspection method based on a large model is given below. In this method, step 410 includes: inputting multiple target test questions of different question types into the large model respectively, and obtaining multiple output answers output by the large model.
[0178] Here, the target questions of different types may include, but are not limited to, at least one of the following: fill-in-the-blank questions, multiple-choice questions, and true / false questions, etc. Each target question corresponds to one output answer.
[0179] Accordingly, step 420 includes: if at least one of the multiple output answers is incorrect, constructing a first prompt for the sample and a first output result label corresponding to the first prompt for the sample based on the sample data to be inspected.
[0180] It should be noted that if at least one of the multiple output answers is incorrect, the large model is considered to have not mastered the knowledge related to the sample data to be inspected. Therefore, the first output result label is also used to indicate that the large model has not mastered the knowledge related to the sample data to be inspected. The training data constructed based on this is used to train the large model, so that the data quality inspection model can more accurately judge whether it has mastered the knowledge related to the data to be inspected, thereby further improving the reliability and accuracy of data quality inspection.
[0181] Accordingly, step 430 above includes: if all the output answers are correct, constructing a first prompt for the sample and a second output result label corresponding to the first prompt for the sample based on the sample data to be inspected.
[0182] It should be noted that the large model is considered to have mastered the knowledge related to the sample data to be inspected only if all output answers are correct. Therefore, the second output label is also used to indicate that the large model has mastered the knowledge related to the sample data to be inspected. The training data constructed based on this is used to train the large model, so that the data quality inspection model can more accurately judge whether it has mastered the knowledge related to the data to be inspected, thereby further improving the reliability and accuracy of data quality inspection.
[0183] The data quality inspection method based on a large model provided in this invention determines whether the large model has mastered the knowledge related to the sample data to be inspected by using multiple target questions of different question types, thereby improving the accuracy of training data construction and ultimately improving the reliability and accuracy of data quality inspection. Through the first and second output result labels set in the above manner, if at least one of the multiple output answers is incorrect, the large model is considered not to have mastered the knowledge related to the sample data to be inspected; conversely, all output answers must be correct for the large model to be considered to have mastered the knowledge related to the sample data to be inspected. This allows the data quality inspection model to more accurately determine whether it has mastered the knowledge related to the data to be inspected, thereby further improving the reliability and accuracy of data quality inspection.
[0184] Based on any of the above embodiments, a specific embodiment of a data quality inspection method based on a large model is given below. In this method, the sample data to be inspected is erroneous data, the first prompt of the sample is used to instruct the large model to correct the errors in the sample data to be inspected, and the second output result label includes the corrected data of the sample data to be inspected.
[0185] For example, the first prompt template is as follows: "Please judge whether the following content is correct. If you find that there is an error, please mark the part you think is wrong with "_" and correct it. If there is a part you are unsure about, please use '_'." "Mark this. The following is the content: xxx."
[0186] Here, the corrected data refers to the corrected data, thus making full use of the large model's own capabilities to correct the knowledge that the sample data to be inspected is incorrect, rather than directly losing the sample data to be inspected, thereby improving data utilization.
[0187] It should be understood that with a small amount of labeled data, large models can learn to identify and correct erroneous knowledge, and generalize to unseen data. This can help identify problematic parts of the data to be inspected and provide the corresponding correct knowledge.
[0188] The data quality inspection method based on a large model provided in this invention, through the above-mentioned method, makes full use of the capabilities of the large model itself to correct the knowledge that the sample data to be inspected is incorrect, rather than directly losing the sample data to be inspected, thereby improving the data utilization rate; and by setting the first prompt of the sample in this way during the model training process, the error correction effect of the data quality inspection model in the inference process can be improved, thereby further improving the data utilization rate.
[0189] Based on any of the above embodiments, a specific embodiment of a data quality inspection method based on a large model is given below. In this method, when the sample data to be inspected is erroneous data, the training data is further constructed in the following manner: based on the sample data to be inspected, a first prompt for the sample and a third output result label corresponding to the first prompt for the sample are constructed, so that the training data includes the first prompt for the sample and the third output result label.
[0190] The third output label includes the correct data corresponding to the sample data to be inspected. Therefore, when the sample data to be inspected is incorrect, this third output label is set to enable the large model to learn the correct data corresponding to the sample data to be inspected. This allows the large model to acquire as much knowledge as possible about the correct data, thereby improving the data quality inspection effect of the data quality inspection model and ultimately increasing the accuracy of data quality inspection.
[0191] For example, the sample data to be inspected is "Dopamine can act on α receptors in blood vessels, causing vasodilation, thereby raising blood pressure"; the first prompt for the sample is as follows: "Please judge whether the following content is correct. If you confirm that there is an error, please mark the part you think is wrong with "_" and correct it. If there is a part you are unsure about, please use '_'." "Tag: The following is the content: Dopamine can act on α receptors in blood vessels, causing vasodilation, thereby raising blood pressure."; The third output result is tagged as "Dopamine can act on α receptors in blood vessels, causing vasoconstriction, thereby raising blood pressure."
[0192] The data quality inspection method based on a large model provided in this invention provides the correct data corresponding to the erroneous sample data to be inspected, regardless of whether the output answer of the large model is correct, so that the large model can have as much knowledge as possible about the correct data, thereby improving the data quality inspection effect of the data quality inspection model and ultimately improving the accuracy of data quality inspection.
[0193] Based on any of the above embodiments, a specific embodiment of a data quality inspection method based on a large model is given below. In this method, the sample data to be inspected is sample question-and-answer data, which includes sample questions and sample answers. This sample question-and-answer data can be single-round question-and-answer data, such as single-round medical question-and-answer data.
[0194] For example, the sample question is "What is the mechanism by which dopamine raises blood pressure?" and the sample answer is "Dopamine can act on α receptors in blood vessels, causing vasodilation, thereby raising blood pressure."
[0195] The first prompt in the sample is constructed based on the sample answer. Therefore, only the first prompt needs to be constructed based on the sample answer; no sample question is required.
[0196] In one specific embodiment, a sample first prompt is constructed based on the sample answer and the first prompt template. More specifically, the sample answer is concatenated into the first prompt template to obtain the sample first prompt.
[0197] The first output result label is constructed based on the sample answers and the target test question.
[0198] For example, if the target question is a fill-in-the-blank question: "Dopamine can act on α receptors in blood vessels, causing blood vessels to ___, thereby achieving the effect of raising blood pressure," and the sample answer is "Dopamine can act on α receptors in blood vessels, causing blood vessels to dilate, thereby achieving the effect of raising blood pressure," then based on the target question, we can determine where the error occurred, and then better modify the sample answer to obtain the first output result label.
[0199] The second output result label is constructed based on the sample answers and the target test question.
[0200] For example, if the target question is a fill-in-the-blank question: "Dopamine can act on α receptors in blood vessels, causing blood vessels to ___, thereby achieving the effect of raising blood pressure," and the sample answer is "Dopamine can act on α receptors in blood vessels, causing blood vessels to dilate, thereby achieving the effect of raising blood pressure," then based on the target question, we can determine where the error occurred, and then better modify the sample answer to obtain the second output result label.
[0201] It should be understood that if there are multiple target questions, you can choose the best one to construct the output result label, such as choosing a fill-in-the-blank question.
[0202] For example, the sample answer is "Dopamine can act on α receptors in blood vessels, causing vasodilation, thereby raising blood pressure." If the output answer is correct, it indicates that the large model has mastered the knowledge related to the sample answer. Based on this, the first prompt for the sample is: "Please judge whether the following content is correct. If you confirm that there is an error, please mark the part you think is wrong with "_" and correct it. If there is a part you are unsure about, please use '_'." The following is the content: "Dopamine can act on α receptors in blood vessels, causing vasodilation, thereby raising blood pressure." The second output result is labeled "Dopamine can act on α receptors in blood vessels, causing vasodilation (contraction), thereby raising blood pressure." If the output answer is incorrect, it indicates that the large model has not mastered the relevant knowledge of the sample answer. Based on this, the first prompt for the sample is: "Please judge whether the following content is correct. If you confirm that it is incorrect, please mark the part you think is wrong with "_" and correct it. If there is a part you are unsure about, please use '_'." "Tag: The following is the content: Dopamine can act on α receptors in blood vessels, causing vasodilation, thereby raising blood pressure." The first output result is tagged as "Dopamine can act on α receptors in blood vessels, causing vasodilation, thereby raising blood pressure." expansion This can raise blood pressure.
[0203] The data quality inspection method based on a large model provided in this invention, through the above-mentioned method, only needs to construct the sample first prompt based on the sample answer, without the need for sample questions, thereby reducing the length of the sample first prompt and improving the quality inspection efficiency of the data quality inspection model; and the output result label is constructed based on the sample answer and the target test question, thereby accurately constructing the output result label with reference to the target test question, thereby improving the accuracy of training data construction, and ultimately improving the accuracy and reliability of data quality inspection.
[0204] The data quality inspection device based on a large model provided by the present invention is described below. The data quality inspection device based on a large model described below and the data quality inspection method based on a large model described above can be referred to and correspond to each other.
[0205] Figure 5 is a schematic diagram of the data quality inspection device based on a large model provided by the present invention. As shown in Figure 5, the data quality inspection device based on a large model includes: a first input module 510, a second input module 520, and a result determination module 530.
[0206] The first input module 510 is used to input a first prompt message constructed based on the data to be inspected into the data quality inspection model, and to obtain the output result output by the data quality inspection model; the first prompt message is used to instruct the data quality inspection model to determine whether it has mastered the knowledge related to the data to be inspected, and to instruct the data quality inspection model to perform quality inspection on the data to be inspected; the data quality inspection model is constructed based on a large model.
[0207] The second input module 520 is used to input the data to be inspected and the knowledge data related to the data to be inspected into the data quality inspection model when it is determined from the output result that the data quality inspection model does not have knowledge related to the data to be inspected, so as to obtain the data quality inspection result output by the data quality inspection model.
[0208] The result determination module 530 is used to determine the data quality inspection result based on the output result, provided that the data quality inspection model has knowledge related to the data to be inspected.
[0209] Figure 6 illustrates a schematic diagram of the physical structure of an electronic device. As shown in Figure 6, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. The processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can call logic instructions in the memory 630 to execute a data quality inspection method based on a large model. This method includes: inputting a first prompt constructed based on the data to be inspected into a data quality inspection model to obtain an output result from the data quality inspection model; the first prompt is used to instruct the data quality inspection model to determine whether it possesses knowledge related to the data to be inspected, and to instruct the data quality inspection model to perform quality inspection on the data to be inspected; the data quality inspection model is constructed based on a large model; if, based on the output result, it is determined that the data quality inspection model does not possess knowledge related to the data to be inspected, the data to be inspected and related knowledge data are input into the data quality inspection model to obtain a data quality inspection result output by the data quality inspection model; if, based on the output result, it is determined that the data quality inspection model possesses knowledge related to the data to be inspected, a data quality inspection result is determined based on the output result.
[0210] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0211] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the data quality inspection method based on a large model provided by the above methods. The method includes: inputting a first prompt constructed based on the data to be inspected into a data quality inspection model to obtain an output result output by the data quality inspection model; the first prompt is used to instruct the data quality inspection model to determine whether it has mastered the knowledge related to the data to be inspected, and to instruct the data quality inspection model to perform quality inspection on the data to be inspected; the data quality inspection model is constructed based on a large model; if it is determined based on the output result that the data quality inspection model does not master the knowledge related to the data to be inspected, inputting the data to be inspected and the knowledge data related to the data to be inspected into the data quality inspection model to obtain a data quality inspection result output by the data quality inspection model; if it is determined based on the output result that the data quality inspection model has mastered the knowledge related to the data to be inspected, determining the data quality inspection result based on the output result.
[0212] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the data quality inspection method based on a large model provided by the above methods. The method includes: inputting a first prompt constructed based on data to be inspected into a data quality inspection model to obtain an output result from the data quality inspection model; the first prompt instructing the data quality inspection model to determine whether it possesses knowledge related to the data to be inspected, and instructing the data quality inspection model to perform quality inspection on the data to be inspected; the data quality inspection model is constructed based on a large model; if, based on the output result, it is determined that the data quality inspection model does not possess knowledge related to the data to be inspected, inputting the data to be inspected and knowledge data related to the data to be inspected into the data quality inspection model to obtain a data quality inspection result output by the data quality inspection model; if, based on the output result, it is determined that the data quality inspection model possesses knowledge related to the data to be inspected, determining the data quality inspection result based on the output result.
[0213] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0214] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0215] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data quality inspection method based on a large model, characterized in that, include: The first prompt, constructed based on the data to be inspected, is input into the data quality inspection model to obtain the output result of the data quality inspection model; The first prompt is used to instruct the data quality inspection model to determine whether it has knowledge related to the data to be inspected, and to instruct the data quality inspection model to perform quality inspection on the data to be inspected. The data quality inspection model is built based on a large model. If, based on the output results, it is determined that the data quality inspection model does not possess knowledge related to the data to be inspected, the data to be inspected and related knowledge data are input into the data quality inspection model to obtain the data quality inspection results output by the data quality inspection model. If, based on the output results, it is determined that the data quality inspection model possesses relevant knowledge about the data to be inspected, then the data quality inspection result is determined based on the output results.
2. The data quality inspection method based on a large model according to claim 1, characterized in that, The step of inputting a first prompt based on the data to be inspected into a data quality inspection model to obtain the output result of the data quality inspection model includes: constructing a second prompt based on the data to be inspected; inputting the second prompt into the data quality inspection model to obtain a splitting result output by the data quality inspection model; the second prompt is used to instruct the data quality inspection model to split the data to be inspected into knowledge points, and the splitting result includes several split knowledge points; constructing a first prompt based on the splitting result; the first prompt includes first sub-prompts constructed based on the several split knowledge points respectively; the first sub-prompt constructed based on any of the split knowledge points is used to instruct the data quality inspection model to determine whether it has mastered the knowledge related to the split knowledge point, and to instruct the data quality inspection model to perform quality inspection on the split knowledge point; inputting each of the first sub-prompts into the data quality inspection model to obtain the output result output by the data quality inspection model; the output result includes the sub-output result corresponding to each of the first sub-prompts.
3. The data quality inspection method based on a large model according to claim 2, characterized in that, When it is determined, based on the output results, that the data quality inspection model does not possess knowledge related to the data to be inspected, the data to be inspected and related knowledge data are input into the data quality inspection model to obtain the data quality inspection results output by the data quality inspection model. This includes: when it is determined, based on each of the sub-output results, that the data quality inspection model does not possess knowledge related to at least one first target splitting knowledge point among the plurality of splitting knowledge points, the at least one first target splitting knowledge point and related knowledge data are input into the data quality inspection model to obtain a plurality of first sub-data quality inspection results output by the data quality inspection model; determining a plurality of second sub-data quality inspection results based on the sub-output results corresponding to a plurality of second target splitting knowledge points; the plurality of second target splitting knowledge points are other splitting knowledge points among the plurality of splitting knowledge points besides the at least one first target splitting knowledge point; and determining the data quality inspection result based on each of the first sub-data quality inspection results and each of the second sub-data quality inspection results.
4. The data quality inspection method based on a large model according to claim 3, characterized in that, When it is determined, based on the sub-output results, that the data quality inspection model does not possess knowledge related to at least one first target split knowledge point among the plurality of split knowledge points, the at least one first target split knowledge point and the knowledge data related to the at least one first target split knowledge point are input into the data quality inspection model to obtain a plurality of first sub-data quality inspection results output by the data quality inspection model. This includes: when it is determined, based on the sub-output results, that the data quality inspection model does not possess knowledge related to at least one first target split knowledge point among the plurality of split knowledge points, performing knowledge search on each of the at least one first target split knowledge point; for each first target split knowledge point, performing the following respectively: if knowledge data related to the first target split knowledge point is found, inputting the first target split knowledge point and the knowledge data related to the first target split knowledge point into the data quality inspection model to obtain the first sub-data quality inspection result output by the data quality inspection model; if no knowledge data related to the first target split knowledge point is found, determining a first preset data quality inspection result as the first sub-data quality inspection result.
5. The data quality inspection method based on a large model according to claim 4, characterized in that, The step of inputting the first target-splitting knowledge point and related knowledge data into the data quality inspection model to obtain the first sub-data quality inspection result output by the data quality inspection model includes: constructing a third prompt based on the first target-splitting knowledge point and related knowledge data; inputting the third prompt into the data quality inspection model to obtain the first sub-data quality inspection result output by the data quality inspection model; wherein, the third prompt is used to instruct the data quality inspection model to evaluate the conflict between the first target-splitting knowledge point and related knowledge data.
6. The data quality inspection method based on a large model according to claim 5, characterized in that, The step of constructing a third prompt based on the first target split knowledge point and the knowledge data related to the first target split knowledge point includes: constructing a third prompt based on the first target split knowledge point, the knowledge data related to the first target split knowledge point, and the source of the knowledge data related to the first target split knowledge point; wherein, the third prompt is also used to instruct the data quality inspection model to evaluate the source authority of the knowledge data related to the first target split knowledge point.
7. The data quality inspection method based on a large model according to claim 3, characterized in that, The first sub-prompt constructed based on any of the aforementioned split knowledge points is also used to instruct the data quality inspection model to determine whether the split knowledge point is correct, and to instruct the data quality inspection model to correct the split knowledge point; the second sub-data quality inspection result corresponding to any of the second target split knowledge points is determined in the following manner: if the second target split knowledge point is determined to be incorrect based on the sub-output result corresponding to the second target split knowledge point, the second preset data quality inspection result is determined as the second sub-data quality inspection result corresponding to the second target split knowledge point; if the second target split knowledge point is determined to be correct based on the sub-output result corresponding to the second target split knowledge point, the third preset data quality inspection result is determined as the second sub-data quality inspection result corresponding to the second target split knowledge point; wherein, the quality inspection score represented by the third preset data quality inspection result is greater than the quality inspection score represented by the second preset data quality inspection result.
8. The data quality inspection method based on a large model according to any one of claims 1 to 7, characterized in that, The data quality inspection model is trained as follows: The large model is trained using training data to obtain the data quality inspection model; wherein the training data is constructed as follows: a target question is input into the large model to obtain the output answer; the target question is a question used to examine knowledge related to the sample data to be inspected; in the case of an incorrect output answer, a first prompt and a corresponding first output result label are constructed based on the sample data to be inspected, so that the training data includes the first prompt and the first output result label; the first output result label is used to indicate that the large model has not mastered the knowledge related to the sample data to be inspected; in the case of a correct output answer, a second prompt and a corresponding second output result label are constructed based on the sample data to be inspected, so that the training data includes the first prompt and the second output result label; the second output result label is used to indicate that the large model has mastered the knowledge related to the sample data to be inspected.
9. The data quality inspection method based on a large model according to claim 8, characterized in that, The sample data to be inspected is erroneous data, and the target test question is generated in the following manner: based on the sample data to be inspected and the error cause data of the sample data to be inspected, a fourth prompt is constructed; the fourth prompt is input into the large model to obtain the test question output result of the large model; the test question output result includes the target test question; wherein, the fourth prompt is used to instruct the large model to refer to the sample data to be inspected and the error cause data to construct a test question that examines the knowledge related to the error cause data.
10. The data quality inspection method based on a large model according to claim 9, characterized in that, The step of constructing a fourth prompt based on the sample data to be inspected and the error cause data of the sample data to be inspected includes: constructing a fourth prompt based on the sample data to be inspected, the error cause data of the sample data to be inspected, and several existing test questions related to the sample data to be inspected; wherein, the fourth prompt is used to instruct the large model to refer to the sample data to be inspected, the error cause data, and the several existing test questions to construct test questions that examine the knowledge related to the error cause data.
11. The data quality inspection method based on a large model according to claim 10, characterized in that, The fourth prompt is constructed based on the sample data to be inspected, the error cause data of the sample data to be inspected, and several existing test questions related to the sample data to be inspected. This includes: constructing the fourth prompt based on the sample data to be inspected, the error cause data of the sample data to be inspected, several existing test questions related to the sample data to be inspected, and the correct answers to each of the existing test questions; wherein, the fourth prompt is used to instruct the large model to construct test questions that examine knowledge related to the error cause data, referring to the sample data to be inspected, the error cause data, the several existing test questions, and the correct answers to each of the existing test questions; the test question output also includes the correct answer to the target test question, which is used to determine whether the output answer is correct.
12. The data quality inspection method based on a large model according to claim 8, characterized in that, The step of inputting the target test question into the large model and obtaining the output answer from the large model includes: inputting multiple target test questions of different question types into the large model respectively, and obtaining multiple output answers from the large model; the step of constructing a sample first prompt and a first output result label corresponding to the sample first prompt based on the sample data to be inspected when the output answer is incorrect includes: constructing a sample first prompt and a first output result label corresponding to the sample first prompt based on the sample data to be inspected when at least one of the multiple output answers is incorrect; the step of constructing a sample first prompt and a second output result label corresponding to the sample first prompt based on the sample data to be inspected when the output answer is correct includes: constructing a sample first prompt and a second output result label corresponding to the sample first prompt based on the sample data to be inspected when all multiple output answers are correct.
13. The data quality inspection method based on a large model according to claim 8, characterized in that, The sample data to be inspected is erroneous data. The first prompt of the sample is used to instruct the large model to correct the sample data to be inspected. The second output result label includes the corrected data of the sample data to be inspected.
14. The data quality inspection method based on a large model according to claim 8, characterized in that, In the case that the sample data to be inspected is erroneous data, the training data is further constructed in the following manner: based on the sample data to be inspected, a first prompt for the sample and a third output result label corresponding to the first prompt for the sample are constructed, so that the training data includes the first prompt for the sample and the third output result label; The third output result label includes the correct data corresponding to the sample data to be inspected.
15. The data quality inspection method based on a large model according to claim 8, characterized in that, The sample data to be inspected is sample question-and-answer data, which includes sample questions and sample answers; the first prompt is constructed based on the sample answers; the first output result label is constructed based on the sample answers and the target test question; the second output result label is constructed based on the sample answers and the target test question.
16. A data quality inspection device based on a large model, characterized in that, include: The first input module is used to input the first prompt message constructed based on the data to be inspected into the data quality inspection model, and obtain the output result output by the data quality inspection model. The first prompt is used to instruct the data quality inspection model to determine whether it has knowledge related to the data to be inspected, and to instruct the data quality inspection model to perform quality inspection on the data to be inspected. The data quality inspection model is built upon a large model; The second input module is used to input the data to be inspected and the knowledge data related to the data to be inspected into the data quality inspection model when it is determined from the output result that the data quality inspection model does not have the knowledge related to the data to be inspected, so as to obtain the data quality inspection result output by the data quality inspection model. The result determination module is used to determine the data quality inspection result based on the output result, provided that the data quality inspection model has knowledge related to the data to be inspected.
17. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the data quality inspection method based on a large model as described in any one of claims 1 to 15.
18. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the data quality inspection method based on a large model as described in any one of claims 1 to 15.
19. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the data quality inspection method based on a large model as described in any one of claims 1 to 15.