Question and answer data quality inspection method and device, related equipment and computer program product

By using the target big model to identify doubtful content in the Q&A data, and searching external knowledge information in the source for verification, and determining the data quality based on the authority and conflict of the source, the problems of low efficiency and poor scalability of traditional quality inspection methods are solved, and efficient and accurate Q&A data quality inspection is achieved.

CN120218253APending Publication Date: 2025-06-27ANHUI IFLYHEALTH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510383603.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Traditional Q&A data quality inspection methods rely on manual review or rule matching technology, and have problems such as high cost, low efficiency and poor scalability, making it difficult to effectively detect the quality of Q&A data.

Method used

By obtaining the Q&A data to be quality-tested, sending it into the configured target model, outputting content that is suspected to be incorrect as questionable content, and searching external knowledge information from each source to verify the correctness of the questionable content, and determining the quality evaluation results of the question&A data based on the authority and conflict of the source.

Benefits of technology

The automated data quality inspection process is realized, which avoids the cost and efficiency of manual audits, improves the accuracy and reliability of data quality inspection, and has stronger scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218253A_ABST
    Figure CN120218253A_ABST
Patent Text Reader

Abstract

The invention discloses a question and answer data quality inspection method, a question and answer data quality inspection device, related equipment and a computer program product, realizes an automatic data quality inspection process, and avoids the defects of high cost, long consumed time and the like existing in manual auditing. Besides, by means of the capability of the target large model, content suspected to have errors is found out from the question and answer data to serve as suspected content, then external knowledge information used for verifying correctness of the suspected content can be retrieved in an information source according to the suspected content, and the correctness of the suspected content can be verified based on information source authority of external knowledge and conflict between the external knowledge information and answer data. Compared with a traditional rule matching technology, the method has the advantage that the expandability is higher. By comprehensively considering the information source authority of the external knowledge information and the conflict between the external knowledge information and the answer data, the quality evaluation result of the question and answer data can be measured more accurately, and the accuracy and reliability of data quality inspection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data quality inspection, and more specifically, to a method and apparatus for quality inspection of question-and-answer data, related devices, and computer program products. Background Art

[0002] With the popularization of the Internet, there are more and more online question-and-answer applications, such as online medical Q&A platforms, legal Q&A platforms, etc. The answers on the platform are usually generated by domain practitioners, robots, or other users. However, the professionalism of some fields is strong, and it is difficult to ensure the accuracy and reliability of all answers.

[0003] Traditional data quality inspection methods mostly rely on manual review or rule matching techniques, which have problems such as high cost, low efficiency, and poor scalability. Therefore, how to effectively detect the quality of question-and-answer data has become an urgent problem to be solved. Summary of the Invention

[0004] In view of the above problems, this application is proposed to provide a method and apparatus for quality inspection of question-and-answer data, related devices, and computer program products to solve the defects of traditional manual review or rule matching techniques. The specific solutions are as follows:

[0005] In a first aspect, a method for quality inspection of question-and-answer data is provided, including:

[0006] Obtain the question-and-answer data to be quality-inspected, where the question-and-answer data includes question data and answer data;

[0007] Send the question-and-answer data into a configured target large model to instruct the target large model to output the content suspected of being incorrect in the answer data as the suspicious content;

[0008] For the suspicious content in the answer data, retrieve external knowledge information for verifying the correctness of the suspicious content from each information source;

[0009] Determine the authority of the information source of the external knowledge information and the conflict between the external knowledge information and the answer data;

[0010] Based on the information source authority and the conflict, determine the quality evaluation result of the question-and-answer data.

[0011] In a possible design, in another implementation manner of the first aspect of the embodiments of this application, the process of retrieving external knowledge information for verifying the correctness of the suspicious content from each information source for the suspicious content in the answer data includes:

[0012] Invoke the target large model through a first prompt instruction to instruct the target large model to generate a retrieval question for the doubtful content in the answer data within the Q&A data, where the retrieval question is used to retrieve external knowledge information for verifying the correctness of the doubtful content;

[0013] Invoke a retrieval tool through the retrieval question to retrieve the external knowledge information from various information sources.

[0014] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, when the external knowledge information cannot be retrieved by invoking the retrieval tool through the retrieval question, the method further includes:

[0015] Iteratively execute the step of invoking the target large model through the first prompt instruction, and add, in each invocation, the retrieval questions that the target large model has previously generated and for which it is known that the external knowledge information cannot be retrieved, to the first prompt instruction;

[0016] Until the external knowledge information is retrieved by the retrieval question generated by the target large model.

[0017] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the process of determining the source authority of the external knowledge information and the conflict between the external knowledge information and the answer data includes:

[0018] Invoke the target large model through a second prompt instruction to instruct the target large model to evaluate the source authority of the external knowledge information, where the external knowledge information is used to verify the correctness of the doubtful content in the Q&A data, and to evaluate the conflict between the external knowledge information and the answer data;

[0019] Obtain the source authority score of the external knowledge information output by the target large model and the conflict score between the external knowledge information and the answer data.

[0020] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, there is more than one piece of doubtful content, correspondingly, the source authority includes the source authority of the external knowledge information corresponding to each piece of doubtful content, and the conflict includes the conflict between the external knowledge information corresponding to each piece of doubtful content and the answer data;

[0021] The process of determining the quality evaluation result of the Q&A data by combining the source authority and the conflict includes:

[0022] For each piece of doubtful content, calculate a quality penalty score according to the corresponding source authority and conflict.

[0023] Determine the quality score of the Q&A data by combining the quality penalty scores of the suspected content.

[0024] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the target large model is obtained by training an initial large model, and the training process of the target large model includes:

[0025] Call the initial large model through a third prompt instruction to instruct the initial large model to combine the input sample question data and the masked sample answer data to predict the predicted entity corresponding to the masked part, where the masking process is to mask the key entities in the sample answer data;

[0026] For each key entity in the sample answer data, determine whether the key entity and the corresponding predicted entity express the same meaning, and filter out the target predicted entities with different expressed meanings;

[0027] Replace each key entity in the sample answer data with the corresponding predicted entity to obtain the first data;

[0028] Determine the training input data based on the sample question data and the first data, and determine the training output data based on the target predicted entity, so as to instruct the initial large model to identify the content suspected of being incorrect in the input data as a training task to train the initial large model, and obtain the trained target large model.

[0029] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, before calling the initial large model through a third prompt instruction, it further includes:

[0030] Obtain sample Q&A data, where the sample Q&A data includes sample question data and sample answer data;

[0031] Extract the key entities in the sample answer data;

[0032] Perform a masking process on at least one key entity in the sample answer data to obtain the masked sample answer data.

[0033] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the process of extracting the key entities in the sample answer data includes:

[0034] Call the initial large model to instruct the initial large model to extract the key entities in the sample answer data according to the input sample question data and sample answer data.

[0035] In a possible design, in another implementation of the first aspect of the embodiments of the present application, for each of the key entities in the sample answer data, the process of determining whether the key entity and the corresponding predicted entity express the same meaning includes:

[0036] For any one of the key entities in the sample answer data, use the corresponding predicted entity for replacement to obtain the replaced sample answer data;

[0037] Call the initial large model to instruct the initial large model to compare whether the sample answer data and the replaced sample answer data express the same meaning;

[0038] If the comparison result output by the initial large model indicates the same meaning, it is determined that the predicted entity in the replaced sample answer data expresses the same meaning as the corresponding key entity; otherwise, it is determined that the predicted entity in the replaced sample answer data expresses a different meaning from the corresponding key entity.

[0039] In a possible design, in another implementation of the first aspect of the embodiments of the present application, the process of determining training input data based on the sample question data and the first data, and determining training output data based on the target predicted entity, and instructing the initial large model to identify the content suspected of being incorrect in the input data as a training task to train the initial large model includes:

[0040] Mark the target predicted entity in the first data to obtain the second data;

[0041] The training input data is composed of the set task description text, the sample question data, and the first data, and the training output data is composed of the second data. Use the training input data and the training output data to train the initial large model, where the set task description text is used to instruct the large model to mark and output the content suspected of being incorrect in the first data.

[0042] In a possible design, in another implementation of the first aspect of the embodiments of the present application, after obtaining the first data, it further includes:

[0043] When the number of predicted entities included in the first data is more than two, any one of the predicted entities is restored to the corresponding key entity according to a set probability to obtain the new first data.

[0044] In a possible design, in another implementation of the first aspect of the embodiments of the present application, it further includes:

[0045] If the content in doubt does not exist in the output of the target large model, it is determined that the quality of the question-and-answer data is qualified.

[0046] In a second aspect, a question-and-answer data quality inspection device is provided, including:

[0047] A data acquisition unit, configured to acquire question-and-answer data to be quality-inspected, where the question-and-answer data includes question data and answer data;

[0048] A suspicious content identification unit, configured to send the question-and-answer data to a configured target large model to instruct the target large model to output the content suspected of being incorrect in the answer data as suspicious content;

[0049] An external knowledge retrieval unit, configured to retrieve external knowledge information for verifying the correctness of the suspicious content from various information sources for the suspicious content in the answer data;

[0050] An external knowledge analysis unit, configured to determine the information source authority of the external knowledge information and the conflict between the external knowledge information and the answer data;

[0051] A quality evaluation unit, configured to determine the quality evaluation result of the question-and-answer data by combining the information source authority and the conflict.

[0052] In a third aspect, an electronic device is provided, including: a memory and a processor;

[0053] The memory is configured to store a program;

[0054] The processor is configured to execute the program to implement each step of the question-and-answer data quality inspection method described in any one of the foregoing first aspects of the present application.

[0055] In a fourth aspect, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, each step of the question-and-answer data quality inspection method described in any one of the foregoing first aspects of the present application is implemented.

[0056] In a fifth aspect, a computer program product is provided, including a computer program. When the computer program is executed by a processor, each step of the question-and-answer data quality inspection method described in any one of the foregoing first aspects of the present application is implemented.

[0057] As can be seen from the above technical solution, the present application realizes an automated data quality inspection process, avoiding the defects such as high cost and long time consumption existing in manual review. In addition, by leveraging the capabilities of the target large model, the content suspected of being incorrect is identified from the Q&A data as the content in doubt. Furthermore, external knowledge information for verifying the correctness of the content in doubt can be retrieved from the information source based on the content in doubt, and the quality evaluation result of the Q&A data is determined based on the authority of the information source of the external knowledge and the conflict between the external knowledge information and the answer data. Compared with the traditional rule matching technology, it has stronger scalability.

[0058] Furthermore, the present application also considers the problem that the large model may have hallucinations for unmastered knowledge. Instead of directly having the large model perform quality inspection and evaluation on the Q&A data, the large model is made to output the content suspected of being incorrect, obtaining the content in doubt in the answer data, which can effectively alleviate the problem of large model hallucinations. For this content in doubt, the quality evaluation is assisted by retrieving external information source knowledge information. By comprehensively considering the authority of the information source of the external knowledge information and the conflict between the external knowledge information and the answer data, the quality evaluation result of the Q&A data can be measured more accurately, improving the accuracy and reliability of data quality inspection. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0060] Figure 1 It is a schematic diagram of an implementation system architecture of the Q&A data quality inspection method provided by an embodiment of the present application;

[0061] Figure 2 It is a schematic diagram of the process flow of a Q&A data quality inspection method provided by an embodiment of the present application;

[0062] Figure 3 It is a schematic diagram of the process flow of a target large model training method provided by an embodiment of the present application;

[0063] Figure 4 It exemplifies a schematic diagram of the complete process flow of a Q&A data quality inspection method;

[0064] Figure 5 It is a schematic diagram of the structure of a Q&A data quality inspection device provided by an embodiment of the present application;

[0065] Figure 6 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0067] Traditional data quality inspection methods generally involve manual medical review or rule matching techniques, but they generally suffer from problems such as high cost, low efficiency, and poor scalability.

[0068] With the development of large model technology, the excellent performance demonstrated by model-based natural language processing technology in text understanding and generation capabilities provides new possibilities for achieving efficient and highly accurate data quality inspection. Currently, some studies have applied large models to data quality inspection work in specific fields. A simple solution is to directly send the question-and-answer data to be quality inspected into the large model and let the large model output a quality inspection score. This simple processing method completely relies on the capabilities of the large model, and the large model itself has hallucination problems for knowledge it has not mastered, so the accuracy and reliability of the directly output quality scores are not high. Therefore, there is currently a lack of a mature solution for applying large models to question-and-answer data quality inspection.

[0069] To this end, the present application provides a mature solution for applying large models to question-and-answer data quality inspection work, which can effectively solve the defects existing in traditional solutions and improve the efficiency, accuracy, and reliability of data quality inspection. The question-and-answer data quality inspection solution provided by the present application can be applied to various scenarios, such as question-and-answer data quality inspection in the medical field, question-and-answer data quality inspection in the legal field, etc. For the sake of easy understanding, only the question-and-answer data quality inspection in the medical field will be used for exemplary illustration in the subsequent embodiments of the present application.

[0070] The present application provides a question-and-answer data quality inspection method that can be applied to a system architecture as shown in Figure 1 The system may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 1 Taking the example of including one server for illustration).

[0071] Either the terminal 100 or the server 200 can be used alone to execute the question-and-answer data quality inspection method provided in the embodiments of the present application. In addition, the terminal 100 and the server 200 can also be used in cooperation to execute the question-and-answer data quality inspection method provided in the embodiments of the present application.

[0072] Next, the product form of the terminal 100 in Figure 1 will be described;

[0073] The terminal 100 in the embodiments of the present application may be a mobile phone, a tablet computer, a teaching large screen, a robot, a wearable device, a vehicle-mounted device, a conference terminal, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiments of the present application do not make any restrictions on this.

[0074] The embodiments of the present application provide a method for quality inspection of question-and-answer data. Taking the application of this method to a computer device as an example, the computer device may specifically be Figure 1 the terminal 100 therein or a system composed of the terminal 100 and the server 200. Referring to Figure 2 , the method for quality inspection of question-and-answer data specifically includes the following steps:

[0075] Step S100: Obtain the question-and-answer data to be quality-inspected, where the question-and-answer data includes question data and answer data.

[0076] Exemplarily, the question data is: "What are the treatment drugs for gastric spasm?", and the corresponding answer data is "Gastric spasm is a common symptom, usually manifested as severe contraction of the stomach muscles. The following are some recommended drugs: Omeprazole Enteric-coated Tablets, Cinnarizine, Montmorillonite Powder, Domperidone".

[0077] It should be noted that since the question-and-answer data obtained in this step is the data to be quality-inspected, the quality of the question-and-answer data may be qualified or unqualified. That is, for the question data, the answer data may be correct, may be wrong, or some parts of the answer data may be correct and some parts may be wrong. Executing the method of this embodiment will finally obtain the quality evaluation result of the question-and-answer data to be quality-inspected.

[0078] Step S110: Send the question-and-answer data into the configured target large model to instruct the target large model to output the content suspected of being incorrect in the answer data as the suspicious content.

[0079] The target large model adopted in this step may be a general large model, or a domain large model obtained by fine-tuning the general large model with the training data in the field to which the question-and-answer data to be quality-inspected belongs. In addition, in order to further improve the ability of the target large model to identify the content suspected of being incorrect in the question-and-answer data, some training strategies can also be used to fine-tune the initial large model, which will be elaborated in the subsequent embodiments. The initial large model here may be a general large model or a domain large model.

[0080] In this step, by invoking the capabilities of the target large model, the content suspected of being incorrect in the answer data is output by the target large model as the content in doubt.

[0081] In one possible implementation, the process of invoking the target large model to identify the content in doubt can be implemented as follows:

[0082] Obtain a prompt format template. The prompt format template includes a task instruction and a Q&A data slot. The task instruction is used to instruct the large model to mark and output the content suspected of being incorrect in the answer data.

[0083] Fill the Q&A data to be quality-checked into the Q&A data slot to obtain a prompt. Input the prompt into the target large model to obtain the content suspected of being incorrect in the answer data output by the target large model.

[0084] The following exemplifies a specific example of a prompt:

[0085] "The following is a piece of medical knowledge. Please use an underline in markdown format (two ) to bold the places you suspect are incorrect:

[0086] Title: What are the treatment drugs for gastric spasm?

[0087] Content: Gastric spasm is a common symptom, usually manifested as severe contraction of the stomach muscles. The following are some recommended drugs: Omeprazole Enteric-coated Tablets, Cinnarizine, Montmorillonite Powder, Domperidone."

[0088] After sending the above prompt into the target large model, the output result is as follows:

[0089] Gastric spasm is a common symptom, usually manifested as severe contraction of the stomach muscles. The following are some recommended drugs: Omeprazole Enteric-coated Tablets, Cinnarizine 、Montmorillonite Powder, Domperidone.

[0090] It can be seen that the target large model suspects that "Cinnarizine" in the answer data is incorrect and marks it in bold.

[0091] Of course, the above only exemplifies an optional form of the target large model outputting the content in doubt. In addition, the content in doubt can also be output in other forms, such as requiring the target large model to only output the content in doubt, or requiring the target large model to mark and output the content in doubt in other formats, etc.

[0092] In this step, instead of directly asking the target large model to evaluate the quality of the Q&A data, the target large model is made to output the content suspected of being incorrect in the answer data, and then external knowledge information related to the suspected content is retrieved for auxiliary verification later, which can fully utilize the capabilities of the target large model to complete the data quality evaluation work.

[0093] It should be noted that if there is no suspected content in the output of the target large model, that is, the target large model believes that there is no incorrect content in the answer data, then in this embodiment, it can be determined that the quality of the Q&A data is qualified.

[0094] Step S120: For the suspected content in the answer data, retrieve external knowledge information in each information source for verifying the correctness of the suspected content.

[0095] Specifically, each information source includes but is not limited to various information channels on the Internet, self-built knowledge bases, third-party knowledge bases, etc.

[0096] In this step, in order to further verify the correctness of the suspected content, a retrieval tool can be called to retrieve external knowledge information related to the suspected content in each information source.

[0097] In one possible implementation, the suspected content can be directly used as the retrieval term to retrieve external knowledge information related to the suspected content in each information source.

[0098] In another possible implementation, the capabilities of the target large model can also be called to generate relevant retrieval questions, and then the retrieval tool is called through the generated retrieval questions to retrieve external knowledge information in each information source.

[0099] Exemplarily, the target large model is called through a first prompt instruction to instruct the target large model to generate a retrieval question for the suspected content in the answer data of the Q&A data, and this retrieval question is used to retrieve external knowledge information for verifying the correctness of the suspected content.

[0100] Taking the answer data marked with suspected content output by the target large model in the previous example as an example, the following provides an optional example of the first prompt instruction:

[0101] "The following is a piece of medical knowledge. Currently, assume that we have doubts about the part underlined with markdown format (two ) and need to review it. Please generate a question as short as possible for retrieving relevant knowledge for verification:

[0102] Gastric spasm is a common symptom, usually manifested as severe contraction of the stomach muscles. The following are some recommended drugs: Omeprazole Enteric-coated Tablets, Cinnarizine , Montmorillonite Powder, Domperidone."

[0103] After sending the above first prompt instruction into the target large model, the retrieved question sentence output by the model is:

[0104] What are the main indications of cinnarizine?

[0105] Furthermore, the retrieved question sentence can be used to retrieve external knowledge information in each information source. For example, the retrieval results are as follows:

[0106] Source: Internet - Website: Encyclopedia

[0107] Knowledge content: The main effects of cinnarizine: cerebral insufficiency, vertebral artery ischemia, after cerebral thrombosis, etc. In addition, it can also treat tinnitus and dizziness, and can also be used for migraine prevention.

[0108] In a possible situation, when retrieving external knowledge information in each information source for the retrieved question sentence generated by the target large model, there may be a situation where no relevant external knowledge information can be retrieved. This indicates that the retrieved question sentence generated by the target large model may not be appropriate. In this case, a solution is provided in this embodiment, that is:

[0109] The step of calling the target large model through the above first prompt instruction can be iteratively executed, and in each call, the retrieved question sentence that the target large model has previously generated and for which it is known that no external knowledge information can be retrieved is added to the first prompt instruction. This is done until the retrieved question sentence generated by the target large model can retrieve external knowledge information.

[0110] The following introduces an example of the first prompt instruction used when calling the target large model for the second time when the retrieved question sentence generated by the target large model for the first time cannot retrieve external knowledge information:

[0111] "This is a task of generating a retrieved question sentence. The following is a piece of medical knowledge. Currently, assume that we have doubts about the part underlined with markdown format (two ) for bolding and need to verify it. Please generate a retrieved question sentence as short as possible for retrieving relevant knowledge for verification:

[0112] Gastric spasm is a common symptom, usually manifested as severe contraction of the gastric muscles. The following are some recommended drugs: Omeprazole Enteric - coated Tablets, Cinnarizine , Montmorillonite Powder, Domperidone.

[0113] The retrieved question sentences for which it is known that no corresponding knowledge can be retrieved include:

[0114] What is cinnarizine used for?"

[0115] Sending the above first prompt instruction into the target large model again, a new retrieval question generated by the target large model can be obtained as follows:

[0116] What are the main indications of cinnarizine?

[0117] Using this retrieval question, external knowledge information can be retrieved, as described above.

[0118] Through the method provided above, the target large model can be iteratively called to generate new retrieval questions until relevant external knowledge information can be retrieved using the newly generated retrieval questions. In this way, retrieval questions can be generated quickly and efficiently for the content in doubt, and the external knowledge information corresponding to the content in doubt can be effectively found. Compared with extracting all the key information in the answer data and performing individual matching or using the answer data for full-text search, the efficiency and accuracy are higher.

[0119] Step S130: Determine the source authority of the external knowledge information and the conflict between the external knowledge information and the answer data.

[0120] Source authority refers to the credibility, professionalism, and public credibility of the information source. The conflict between the external knowledge information and the answer data refers to the degree of inconsistency between the external knowledge information and the answer data at the semantic level. The higher the conflict, the more inconsistent the meanings expressed by the two data.

[0121] The authorities of different information sources may vary. In this step, the source authority of the retrieved external knowledge information can be determined. Further, the conflict between the external knowledge information and the answer data is determined. This is convenient for evaluating the quality of the Q&A data by combining the source authority and the conflict in subsequent steps.

[0122] The process of determining the source authority of the external knowledge information can be achieved in various different ways.

[0123] Exemplarily, experts can pre-organize and formulate the authority scores of different information sources in advance. Then, according to the source of the retrieved external knowledge information, the corresponding authority score can be found.

[0124] In another implementation, the ability of the target large model can be utilized to automatically evaluate the source authority of the external knowledge information.

[0125] The process of determining the conflict between the external knowledge information and the answer data can also be achieved in various different ways.

[0126] Exemplarily, the semantic similarity between the external knowledge information and the answer data can be calculated, and the conflict score can be determined based on the semantic similarity. Among them, the higher the semantic similarity, the lower the corresponding conflict score.

[0127] In another implementation, the ability of the target large model can be utilized to automatically evaluate the conflict between external knowledge information and answer data.

[0128] In some embodiments of the present application, an implementation method for determining the source authority of external knowledge information and the conflict between external knowledge information and answer data by means of a target large model is provided, which may specifically include:

[0129] Call the target large model through a second prompt instruction to instruct the target large model to evaluate the source authority of the external knowledge information, where the external knowledge information is used to verify the correctness of the doubtful content in the Q&A data, and to evaluate the conflict between the external knowledge information and the answer data.

[0130] Obtain the source authority score of the external knowledge information output by the target large model and the conflict score between the external knowledge information and the answer data.

[0131] The following exemplifies an optional example of the second prompt instruction:

[0132] "The following is a piece of medical knowledge. Currently, we have doubts about the underlined part (two ) in bold in markdown format and have searched for a piece of related knowledge, which needs to be verified. Please first evaluate the authority of the source of this related knowledge based on your knowledge, and then evaluate the conflict between this knowledge and the given medical knowledge according to the content of the knowledge. The output should be in JSON format, giving scores between 0 and 1 for the source authority of the searched knowledge and the conflict between the two pieces of knowledge, where 0 represents the lowest (lowest authority / complete non - conflict between the two pieces of knowledge), and vice versa represents the highest.

[0133] The following is the given medical knowledge:

[0134] Gastric spasm is a common symptom, usually manifested as severe contraction of the stomach muscles. The following are some recommended drugs: Omeprazole Enteric - coated Tablets, Cinnarizine, Montmorillonite Powder, Domperidone.

[0135] The following is the searched knowledge:

[0136] Source: Internet - Website: Encyclopedia

[0137] Knowledge content: The main efficacy of Cinnarizine: cerebral insufficiency, vertebrobasilar ischemia, after cerebral thrombosis, etc. In addition, it can also treat tinnitus, dizziness, and can also be used for migraine prevention."

[0138] After sending the above - mentioned second prompt instruction into the target large model, the output of the model is obtained:

[0139] {"Source Authority": 0.2, "Conflict": 0.9}.

[0140] In this embodiment, by leveraging the capabilities of the target large model, it is possible to accurately predict the source authority of external knowledge information and the conflict between external knowledge information and the answer data, and there is no need to additionally configure rules or train a task model, improving the processing efficiency.

[0141] Step S140: Combine the source authority and the conflict to determine the quality evaluation result of the Q&A data.

[0142] Specifically, in the above steps, for the doubtful content in the answer data, relevant external knowledge information is retrieved, and the source authority of the external knowledge information and the conflict between the external knowledge information and the answer data are obtained. On this basis, the quality evaluation result of the Q&A data can be determined by combining the source authority and the conflict.

[0143] It should be noted that the number of doubtful contents in the answer data can be 0, indicating that there is no doubtful content in the answer data. In this case, the quality of the Q&A data can be directly determined to be qualified. When there is more than one doubtful content in the answer data, for each doubtful content, the corresponding source authority and conflict can be obtained. Therefore, for each doubtful content, the quality penalty score can be calculated according to the corresponding source authority and conflict. Combining the quality penalty scores of each doubtful content in the answer data, the quality score of the Q&A data is determined.

[0144] The following exemplifies an optional calculation formula for the total quality score of the Q&A data:

[0145] Total quality score = Max(1 - ∑authority score × conflict score, 0).

[0146] Where ∑ represents the accumulation of the quality penalty scores for all doubtful contents.

[0147] It can be understood that for each doubtful content, the higher the source authority of the retrieved external knowledge information and the higher the conflict between the external knowledge information and the answer data, the higher the quality penalty score corresponding to this doubtful content, and the lower the quality score of the corresponding Q&A data.

[0148] The Q&A data quality inspection method provided by the embodiments of the present application realizes an automated data quality inspection process, avoiding the defects of high cost and long time-consuming existing in manual review. In addition, by leveraging the capabilities of the target large model, the content suspected of being incorrect is found from the Q&A data as the doubtful content. Furthermore, external knowledge information for verifying the correctness of the doubtful content can be retrieved from the source for the doubtful content. Based on the source authority of the external knowledge and the conflict between the external knowledge information and the answer data, the quality evaluation result of the Q&A data is determined, which has stronger scalability compared to traditional rule matching techniques.

[0149] Furthermore, this application also considers the problem that large models may have hallucinations about unmastered knowledge. Instead of directly having the large model perform quality inspection and evaluation on the Q&A data, it makes the large model output content suspected of being incorrect, obtaining the suspicious content in the answer data, which can effectively alleviate the hallucination problem of the large model. For this suspicious content, by retrieving external source knowledge information to assist in quality evaluation and comprehensively considering the authority of the external knowledge information source and the conflict between the external knowledge information and the answer data, the quality evaluation result of the Q&A data can be measured more accurately, improving the accuracy and reliability of data quality inspection.

[0150] In some possible application scenarios, the Q&A data quality inspection method of this embodiment can be used to conduct quality spot checks on the Q&A data in the Q&A training dataset, and determine the quality average score of the Q&A training dataset according to the spot check results. Furthermore, the sampling ratio of the Q&A training dataset can be set according to the quality average score, sampled according to the sampling ratio, and the sampled results can be used to train the Q&A model to be trained.

[0151] Of course, the Q&A data quality inspection method of this embodiment can also be applied to other scenarios, which will not be elaborated here one by one.

[0152] In some embodiments of this application, the target large model used in the foregoing embodiments is described.

[0153] As introduced above, the target large model can be obtained by training the initial large model. The initial large model can be a general large model or a domain large model in the field to which the Q&A data to be quality inspected belongs. By training the initial large model, the ability of the model to identify incorrect content in the input data can be enhanced.

[0154] Refer to Figure 3 As shown, the training process of the target large model provided in this embodiment can include the following steps:

[0155] Step S200: Invoke the initial large model through the third prompt instruction to instruct the initial large model to combine the input sample question data and the masked sample answer data to predict the predicted entity corresponding to the masked part.

[0156] By sending the training sample data into the initial large model, instruct the large model to combine the sample question data and the masked sample answer data in the training sample data to predict the entity corresponding to the mask in the sample answer data as the predicted entity.

[0157] Among them, the masking process is to mask the key entities in the sample answer data.

[0158] In a possible implementation, the following processing steps can also be added before this step:

[0159] S1. Obtain sample question-and-answer data, where the sample question-and-answer data includes sample question data and sample answer data.

[0160] Taking the medical Q&A scenario as an example, exemplarily, the sample question data Q is like "How does dopamine increase blood pressure?", and the corresponding sample answer data A is like "Dopamine can act on the α receptors of blood vessels, causing vasoconstriction, thereby achieving the effect of increasing blood pressure."

[0161] S2. Extract the key entities in the sample answer data.

[0162] The process of extracting the key entities in the sample answer data can be realized by means of natural language processing technology. For example, a pre-trained keyword recognition model can be used to identify the key entities in the sample answer data. Another example is that the ability of the initial large model can be called to extract the key entities in the sample answer data. Specifically, by calling the initial large model, the initial large model is instructed to extract the key entities in the sample answer data according to the input sample question data and sample answer data.

[0163] The following provides an optional example of a prompt instruction for calling the initial large model to extract key entities:

[0164] "The following is a pair of medical-related knowledge Q&A. Please, based on the question and answer, mark all the non-subject key information with bold (two asterisks) in Markdown format and output the corresponding marked answer.

[0165] The following is the question: Q: How does dopamine increase blood pressure?

[0166] The following is the answer: A: Dopamine can act on the α receptors of blood vessels, causing vasoconstriction, thereby achieving the effect of increasing blood pressure."

[0167] By sending the above prompt to the initial large model, the output of the model is as follows:

[0168] Dopamine can act on the α receptors of blood vessels , causing vasoconstriction , thereby achieving increasing blood pressure the effect.

[0169] Among them, the parts in bold in Markdown format are the key entities recognized by the model.

[0170] By invoking the capabilities of the initial large model to identify key entities, the natural language understanding and generation capabilities of the initial large model can be fully utilized to improve the accuracy of key entity identification, and there is no need to additionally train a key entity identification model.

[0171] S3. Mask at least one key entity in the sample answer data to obtain the masked sample answer data.

[0172] After identifying each key entity in the sample answer data in the above steps, at least one key entity in the sample answer data can be masked to obtain the masked sample answer data.

[0173] Optionally, in the masking process, one key entity in the sample answer data can be masked each time, and the remaining key entities are left unmasked. In this way, more than one masked sample answer data can be obtained, and the masked key entities in different masked sample answer data can be different. This facilitates the initial large model in step S200 to more accurately predict the predicted entity corresponding to the masked part.

[0174] The following exemplifies a masked sample answer data A1: "Dopamine can act on the α receptors of blood vessels, making [MASK], so as to achieve the effect of raising blood pressure."

[0175] After obtaining the masked sample answer data, the sample question data and the masked sample answer data can be assembled into the third prompt instruction, and the initial large model is called through the third prompt instruction to obtain the predicted entity corresponding to the masked part output by the model.

[0176] The following exemplifies an optional example of the third prompt instruction:

[0177] "The following is a piece of medical knowledge. Please give the correct information for the masked part according to your mastery of medical knowledge:

[0178] Dopamine can act on the α receptors of blood vessels, making [MASK], so as to achieve the effect of raising blood pressure."

[0179] Feeding the above third prompt instruction into the initial large model can obtain the predicted entity corresponding to [MASK], for example: "vasodilation".

[0180] For the sake of convenience of explanation, three key entities in the previous sample answer data A are defined as follows: E1: "α receptors of blood vessels", E2: "vasoconstriction", E3: "raising blood pressure".

[0181] For these three key entities, the entities predicted by the initial large model are respectively:

[0182] O1: "β receptors of blood vessels", O2: "vasodilation", O3: "raising blood pressure".

[0183] Step S210: For each of the key entities in the sample response data, determine whether the key entity and the corresponding predicted entity express the same meaning, and filter out the target predicted entities with different expressed meanings.

[0184] Among them, the predicted entity corresponding to the key entity refers to the predicted entity corresponding to the mask generated by the initial large model after masking the key entity in the sample response data.

[0185] In this step, for the predicted entity corresponding to the mask generated by the initial large model, it can be further determined whether the predicted entity and the key entity corresponding to the mask express the same meaning, that is, whether the meaning of the sentence remains unchanged after replacing the key entity in the sample response data with the corresponding predicted entity. Through this judgment process, the knowledge mastery of the initial large model can be understood. After judgment, the target predicted entities with different expressed meanings can be filtered out, that is, the target predicted entity and the corresponding key entity express different meanings.

[0186] In this step, the process of determining whether the key entity and the corresponding predicted entity express the same meaning can directly compare the semantic similarity between the key entity and the predicted entity. If the similarity exceeds the threshold, it can be considered that the two express the same meaning. In addition, this embodiment also provides a judgment method, that is, substituting the key entity and the predicted entity into the sample response data respectively, and judging whether the key entity and the predicted entity express the same meaning by comparing whether the two sample response data express the same meaning. Specifically as follows:

[0187] For any one of the key entities in the sample response data (defined as X1), use the corresponding predicted entity for replacement to obtain the replaced sample response data (defined as X2).

[0188] Call the initial large model to instruct the initial large model to compare whether the sample response data X1 before replacement and the replaced sample response data X2 express the same meaning.

[0189] If the comparison result output by the initial large model indicates the same meaning, it is determined that the predicted entity in the replaced sample response data X2 and the corresponding key entity express the same meaning; otherwise, it is determined that the predicted entity in the replaced sample response data X2 and the corresponding key entity express different meanings.

[0190] In this embodiment, the natural language understanding ability of the initial large model can be called to compare whether the meanings of the two sample response data X1 and X2 are the same, and a more accurate comparison result can be obtained.

[0191] It is understandable that the sample response data X1 may contain more than two key entities. When replacing the key entities therein, only one key entity can be replaced each time, which is convenient for accurately determining the key entities with different expressed meanings and the predicted entities subsequently.

[0192] Taking the replacement of the key entity E2: "vasoconstriction" with the predicted entity O2: "vasodilation" in the above example as an example, the following provides a prompt instruction prompt for calling the initial large model to compare meanings:

[0193] "The following are two pieces of medical knowledge. Please compare whether the contents of these two pieces of knowledge are exactly the same:

[0194] Dopamine can act on the α receptors of blood vessels, causing vasoconstriction, thereby achieving the effect of raising blood pressure.

[0195] Dopamine can act on the α receptors of blood vessels, causing vasodilation, thereby achieving the effect of raising blood pressure."

[0196] Sending the above prompt instruction prompt into the initial large model, the output result of the model is: different meanings. Therefore, it can be determined that the key entity E2 and the predicted entity O2 have different expressed meanings.

[0197] Similarly, it can be determined in turn that the key entity E1 and the predicted entity O1 have different expressed meanings; the key entity E3 and the predicted entity O3 have the same expressed meanings. The target predicted entities with different expressed meanings finally obtained include: the predicted entity O1 and the predicted entity O2.

[0198] Step S220: Replace each of the key entities in the sample response data with the corresponding predicted entity to obtain the first data.

[0199] Taking the above example for illustration, the first data is:

[0200] Dopamine can act on the β receptors of blood vessels, causing vasodilation, thereby achieving the effect of raising blood pressure.

[0201] Step S230: Determine the training input data based on the sample question data and the first data, and determine the training output data based on the target predicted entity, so as to train the initial large model with the task of indicating the content suspected of being incorrect in the input data for the initial large model, and obtain the trained target large model.

[0202] To train the initial large model and improve its ability to identify incorrect content from the input data, based on the processing results of the foregoing steps, training data is constructed in this step. The training data includes training input data and training output data. Among them, the training input data is determined based on the sample question data and the first data. The training output data is determined based on the target prediction entity. The training task is to instruct the initial large model to identify the content suspected of being incorrect in the input data.

[0203] In a possible implementation, the target prediction entity in the first data can be marked to obtain the second data.

[0204] The training input data is composed of the set task description text, the sample question data, and the first data, and the training output data is composed of the second data. The initial large model is trained using the training input data and the training output data. Among them, the set task description text is used to instruct the large model to mark and output the content suspected of being incorrect in the first data.

[0205] Exemplarily, the first data is: "Dopamine can act on the β receptors of blood vessels, dilate blood vessels, and thus achieve the effect of raising blood pressure." Mark the target prediction entity in the first data to obtain the second data: "Dopamine can act on the β receptors of blood vessels , causing blood vessels to dilate , and thus achieve the effect of raising blood pressure."

[0206] The set task description text can be: "The following is a piece of medical knowledge. Please underline (two ) the places you suspect are incorrect in bold using Markdown format."

[0207] Then the training input data composed of the set task description text, the sample question data, and the first data is:

[0208] "The following is a piece of medical knowledge. Please underline (two ) the places you suspect are incorrect in bold using Markdown format:

[0209] Title: How does dopamine raise blood pressure?

[0210] Content: Dopamine can act on the β receptors of blood vessels, dilate blood vessels, and thus achieve the effect of raising blood pressure."

[0211] The training output data composed of the second data is as follows:

[0212] "Dopamine can act on the β receptors of blood vessels , causing vasodilation , thus achieving the effect of raising blood pressure.

[0213] It can be understood that, according to different requirements of the task description text, the form of the training output data can also be different. For example, when the task description text is used to instruct the large model to only output the content suspected of being incorrect in the first data, the corresponding training output data can only include the target prediction entity.

[0214] In the training process of the target large model provided in this embodiment, by predicting the prediction entity corresponding to the mask in the input sample answer data through the initial large model and comparing whether the prediction entity and the corresponding key entity express the same meaning, the true mastery degree of the initial large model on relevant knowledge can be understood. Based on this, training input and output data are constructed to instruct the initial large model to identify the content suspected of being incorrect in the input data as the training task, and the initial large model is trained to obtain the trained target large model, which can further improve the ability of the target large model to identify the content suspected of being incorrect in the input data.

[0215] In an optional solution, after obtaining the first data in the above step S220 and before constructing the training input and output data in step S230, this embodiment further provides a data augmentation method, which can process the first data to obtain additional first data, expanding the data volume of the first data, thereby facilitating the construction of more model training data in the subsequent steps and improving the model training effect.

[0216] Specifically:

[0217] After obtaining the first data, if it is determined that the number of prediction entities included in the first data is more than two, any one of the prediction entities can be restored to the corresponding key entity according to a set probability to obtain additional first data.

[0218] Among them, the set probability value can be set by the user. For example, the value can be 0.1 or other values.

[0219] Combined with Figure 4 shown, a complete process of a question and answer data quality inspection method is introduced. It includes the training process of the target large model and the inference process based on the trained target large model, that is, the process of using the target large model for question data quality inspection.

[0220] For the training process:

[0221] 1. Mine key information. For the obtained training question and answer pairs (sample question and answer data, which can be obtained by manual annotation), mine the key information (key entities) through the initial large model.

[0222] 2-3. Mask key information and let the large model restore it. Mask the key information in the training Q&A pairs, and generate the restored information (predicted entities) corresponding to the mask through the initial large model.

[0223] 4. Compare the consistency. Compare whether the key information and the restored information are consistent through the initial large model.

[0224] 5. Reconstruct the training data based on the consistency and train the large model. Specifically, replace each key entity in the sample answer data with the corresponding predicted entity to obtain the first data. Determine the training input data based on the sample question data and the first data. Based on the consistency comparison result, filter out the inconsistent target predicted entities, and determine the training output data based on the target predicted entities. Use the above training input data and training output data to train the initial large model with the task of indicating the content suspected of being incorrect in the input data by the initial large model to obtain the trained target large model.

[0225] For the inference process:

[0226] 1. Mine the suspicious content. Specifically, for the Q&A pair to be tested, mine the suspicious content in the answer data through the target large model trained above.

[0227] 2. Generate retrieval questions. Specifically, for each suspicious content, generate the corresponding retrieval question through the target large model, and the retrieval question is used to retrieve the knowledge information for verifying the correctness of the suspicious content.

[0228] 3. Query relevant knowledge. Specifically, query the relevant knowledge (external knowledge information) in each information source through the retrieval question.

[0229] 4. Compare and give the authority and conflict scores. Specifically, through the target large model, determine the authority score of the information source of the relevant knowledge, and compare the conflict between the relevant knowledge and the answer data to give the conflict score.

[0230] 5. Update the quality score. Specifically, comprehensively determine the quality score of the Q&A pair based on the authority score and conflict score of the information source corresponding to each suspicious content.

[0231] The solution provided by the embodiment of the present application has the following advantages:

[0232] (1) By mining and constructing training data from the existing correctly labeled sample Q&A data, the large model can learn to find the parts (suspicious content) that may have problems in the Q&A pair, which is more intelligent than the existing rules or directly letting the large model judge the quality score of the Q&A data, and can alleviate the problem of inaccurate scoring caused by the large model hallucination.

[0233] (2) By generating retrieval questions corresponding to the doubtful content, the large model can effectively search for external knowledge information related to relevant parts, with higher efficiency and accuracy than extracting all key information for individual matching or performing full-text retrieval on Q&A data.

[0234] (3) Judging the authority of the retrieved knowledge and its conflict with the knowledge to be quality-checked to update the quality score of the Q&A pair can more accurately evaluate the quality of the Q&A data.

[0235] The Q&A data quality inspection device provided by the embodiments of the present application will be described below. The Q&A data quality inspection device described below can be correspondingly referred to the Q&A data quality inspection method described above.

[0236] See Figure 5 , Figure 5 which is a schematic structural diagram of a Q&A data quality inspection device disclosed in the embodiments of the present application.

[0237] As Figure 5 shown, the device may include:

[0238] A data acquisition unit 11, configured to acquire Q&A data to be quality-checked, where the Q&A data includes question data and answer data;

[0239] A doubtful content identification unit 12, configured to send the Q&A data into a configured target large model to instruct the target large model to output the content suspected of being incorrect in the answer data as doubtful content;

[0240] An external knowledge retrieval unit 13, configured to retrieve external knowledge information for verifying the correctness of the doubtful content in each information source for the doubtful content in the answer data;

[0241] An external knowledge analysis unit 14, configured to determine the source authority of the external knowledge information and the conflict between the external knowledge information and the answer data;

[0242] A quality evaluation unit 15, configured to determine the quality evaluation result of the Q&A data by combining the source authority and the conflict.

[0243] In a possible implementation, the process of the external knowledge retrieval unit retrieving external knowledge information for verifying the correctness of the doubtful content in each information source for the doubtful content in the answer data includes:

[0244] Invoking the target large model through a first prompt instruction to instruct the target large model to generate a retrieval question for the doubtful content in the answer data within the Q&A data, where the retrieval question is used to retrieve external knowledge information for verifying the correctness of the doubtful content;

[0245] Invoke a retrieval tool through the retrieval query sentence to retrieve the external knowledge information from each information source.

[0246] In a possible implementation, when the external knowledge retrieval unit fails to retrieve the external knowledge information by invoking the retrieval tool through the retrieval query sentence, it is further configured to:

[0247] Iteratively execute the step of invoking the target large model through the first prompt instruction, and add, in each invocation, the retrieval query sentence that the target large model has previously generated and is known to be unable to retrieve the external knowledge information to the first prompt instruction;

[0248] Until the external knowledge information is retrieved by the retrieval query sentence generated by the target large model.

[0249] In a possible implementation, the process by which the external knowledge analysis unit determines the source authority of the external knowledge information and the conflict between the external knowledge information and the answer data includes:

[0250] Invoke the target large model through the second prompt instruction to instruct the target large model to evaluate the source authority of the external knowledge information, where the external knowledge information is used to verify the correctness of the doubtful content in the question-and-answer data, and to evaluate the conflict between the external knowledge information and the answer data;

[0251] Obtain the source authority score of the external knowledge information output by the target large model and the conflict score between the external knowledge information and the answer data.

[0252] In a possible implementation, there is more than one piece of doubtful content. Correspondingly, the source authority includes the source authority of the external knowledge information corresponding to each piece of doubtful content, and the conflict includes the conflict between the external knowledge information corresponding to each piece of doubtful content and the answer data. The process by which the quality evaluation unit determines the quality evaluation result of the question-and-answer data by combining the source authority and the conflict includes:

[0253] For each piece of doubtful content, calculate a quality penalty score based on the corresponding source authority and conflict;

[0254] Combine the quality penalty scores of each piece of doubtful content to determine the quality score of the question-and-answer data.

[0255] In a possible implementation, the quality evaluation unit is further configured to:

[0256] If the doubtful content does not exist in the output of the target large model invoked by the doubtful content recognition unit, it is determined that the quality of the question-and-answer data is qualified.

[0257] In one possible implementation, the target large model is obtained by training an initial large model. The apparatus of the present application may further include:

[0258] A large model training unit, configured to train the initial large model to obtain the target large model. The training process includes:

[0259] Invoking the initial large model through a third prompt instruction to instruct the initial large model to combine the input sample question data and the masked sample answer data to predict the predicted entity corresponding to the masked part, where the masking process is to mask the key entities in the sample answer data;

[0260] For each of the key entities in the sample answer data, determine whether the key entity and the corresponding predicted entity express the same meaning, and filter out the target predicted entities with different expressed meanings;

[0261] Replace each of the key entities in the sample answer data with the corresponding predicted entity to obtain the first data;

[0262] Determine the training input data based on the sample question data and the first data, and determine the training output data based on the target predicted entity, so as to instruct the initial large model to identify the content suspected of being incorrect in the input data as a training task to train the initial large model and obtain the trained target large model.

[0263] In one possible implementation, before invoking the initial large model through the third prompt instruction, the large model training unit is further configured to:

[0264] Obtain sample Q&A data, where the sample Q&A data includes sample question data and sample answer data;

[0265] Extract the key entities in the sample answer data;

[0266] Perform a masking process on at least one of the key entities in the sample answer data to obtain the masked sample answer data.

[0267] In one possible implementation, the process by which the large model training unit extracts the key entities in the sample answer data includes:

[0268] Invoking the initial large model to instruct the initial large model to extract the key entities in the sample answer data according to the input sample question data and sample answer data.

[0269] In one possible implementation, the process by which the large model training unit determines whether each of the key entities in the sample answer data and the corresponding predicted entity express the same meaning includes:

[0270] For any of the key entities in the sample answer data, use the corresponding predicted entity for replacement to obtain the replaced sample answer data;

[0271] Invoke the initial large model to instruct the initial large model to compare whether the sample answer data and the replaced sample answer data express the same meaning;

[0272] If the comparison result output by the initial large model indicates the same meaning, determine that the predicted entity in the replaced sample answer data expresses the same meaning as the corresponding key entity; otherwise, determine that the predicted entity in the replaced sample answer data expresses a different meaning from the corresponding key entity.

[0273] In a possible implementation, the large model training unit determines training input data based on the sample question data and the first data, and determines training output data based on the target predicted entity, to instruct the initial large model to identify the content suspected of being incorrect in the input data as a training task to train the initial large model. The process includes:

[0274] Mark the target predicted entity in the first data to obtain the second data;

[0275] The training input data is composed of the set task description text, the sample question data, and the first data, and the training output data is composed of the second data. Use the training input data and the training output data to train the initial large model, where the set task description text is used to instruct the large model to mark and output the content suspected of being incorrect in the first data.

[0276] In a possible implementation, after obtaining the first data, the large model training unit is further configured to:

[0277] When the number of predicted entities included in the first data is more than two, restore any one of the predicted entities to the corresponding key entity according to a set probability to obtain the new first data.

[0278] An embodiment of the present application also provides an electronic device. Refer to Figure 6 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiment of the present application. The electronic device in the embodiment of the present application may include, but is not limited to, fixed terminals such as mobile phones, tablet computers, teaching large screens, wearable devices, and the like. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and usage scope of the embodiment of the present application.

[0279] As Figure 6As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603, so as to implement the question-and-answer data quality inspection method of the foregoing embodiments of the present application. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0280] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0281] In an embodiment of the present application, there is also provided a computer program product including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement any one of the question-and-answer data quality inspection methods provided in the embodiments of the present application.

[0282] In an embodiment of the present application, there is also provided a computer-readable storage medium carrying one or more computer programs, which, when executed by an electronic device, can enable the electronic device to implement any one of the question-and-answer data quality inspection methods provided in the embodiments of the present application.

[0283] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in the present application, the connection relationship between the modules indicates that they have a communication connection, which may be specifically implemented as one or more communication buses or signal lines.

[0284] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions accomplished by computer programs can easily be implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits. However, for the present application, in more cases, software program implementation is a better embodiment. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, and includes several instructions for causing a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0285] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0286] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0287] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

Claims

1. A method for quality inspection of question and answer data, characterized in that: include: Acquire question and answer data to be inspected, wherein the question and answer data includes question data and answer data; Sending the question and answer data to the configured target macro model to instruct the target macro model to output the content suspected of being erroneous in the answer data as questionable content; With respect to the questionable content in the answer data, searching for external knowledge information for verifying the correctness of the questionable content in each information source; Determining the authority of the source of the external knowledge information and the conflict between the external knowledge information and the answer data; The quality evaluation result of the question and answer data is determined in combination with the authority of the information source and the conflict.

2. The method according to claim 1, characterized in that The process of retrieving external knowledge information for verifying the correctness of the questionable content in the answer data from each information source includes: The target macromodel is called through a first prompt instruction to instruct the target macromodel to generate a search question for the questionable content in the answer data in the question-and-answer data, wherein the search question is used to retrieve external knowledge information to verify the correctness of the questionable content; The search tool is called through the search question to search for the external knowledge information in various information sources.

3. The method according to claim 2, characterized in that When the external knowledge information cannot be retrieved by calling the search tool through the search question, the method further includes: Iteratively executing the step of calling the target large model through the first prompt instruction, and adding a search question historically generated by the target large model and known to be unable to retrieve the external knowledge information to the first prompt instruction each time the call is made; Until the external knowledge information is retrieved through the search question generated by the target macro model.

4. The method according to claim 1, characterized in that: The process of determining the authority of the source of the external knowledge information and the conflict between the external knowledge information and the answer data includes: The target big model is called through a second prompt instruction to instruct the target big model to evaluate the source authority of the external knowledge information, the external knowledge information is used to verify the correctness of the questionable content in the question and answer data, and to evaluate the conflict between the external knowledge information and the answer data; The source authority score of the external knowledge information output by the target large model and the conflict score between the external knowledge information and the answer data are obtained.

5. The method according to claim 1, characterized in that The questionable content is more than one questionable content, and correspondingly, the source authority includes the source authority of the external knowledge information corresponding to each questionable content, and the conflict includes the conflict between the external knowledge information corresponding to each questionable content and the answer data; The process of determining the quality evaluation result of the question and answer data in combination with the authority of the information source and the conflict includes: For each of the questionable contents, a quality penalty score is calculated according to the corresponding source authority and the conflict; The quality score of the question and answer data is determined by combining the quality penalty points of each of the questionable contents.

6. The method according to any one of claims 1 to 5, characterized in that: The target large model is obtained by training the initial large model, and the training process of the target large model includes: The initial large model is called through a third prompt instruction to instruct the initial large model to combine the input sample question data and the sample answer data after masking to predict the predicted entity corresponding to the masked part, wherein the masking is to mask the key entities in the sample answer data; For each of the key entities in the sample answer data, determining whether the key entity and the corresponding prediction entity express the same meaning, and screening out target prediction entities that express different meanings; Replacing each of the key entities in the sample answer data with a corresponding predicted entity to obtain first data; The training input data is determined based on the sample question data and the first data, and the training output data is determined based on the target prediction entity, so as to instruct the initial large model to identify suspected erroneous content in the input data as a training task to train the initial large model and obtain a trained target large model.

7. The method according to claim 6, characterized in that Before calling the initial large model through the third prompt instruction, it also includes: Acquire sample question and answer data, wherein the sample question and answer data includes sample question data and sample answer data; Extracting key entities from the sample answer data; Masking is performed on at least one key entity in the sample answer data to obtain masked sample answer data.

8. The method according to claim 7, characterized in that The process of extracting key entities from the sample answer data includes: The initial large model is called to instruct the initial large model to extract key entities in the sample answer data according to the input sample question data and sample answer data.

9. The method according to claim 6, characterized in that For each of the key entities in the sample answer data, the process of determining whether the key entity and the corresponding predicted entity express the same meaning includes: For any one of the key entities in the sample answer data, replace it with the corresponding predicted entity to obtain the replaced sample answer data; Calling the initial large model to instruct the initial large model to compare whether the sample answer data and the replaced sample answer data express the same meaning; If the comparison result output by the initial large model indicates the same meaning, it is determined that the predicted entity in the replaced sample answer data and the corresponding key entity express the same meaning; otherwise, it is determined that the predicted entity in the replaced sample answer data and the corresponding key entity express different meanings.

10. The method according to claim 6, characterized in that The process of determining training input data based on the sample question data and the first data, and determining training output data based on the target prediction entity, so as to instruct the initial large model to identify suspected erroneous content in the input data as a training task to train the initial large model includes: Marking the target prediction entity in the first data to obtain second data; The training input data is composed of a set task description text, the sample question data and the first data, and the training output data is composed of the second data. The initial large model is trained using the training input data and the training output data, wherein the set task description text is used to instruct the large model to mark and output content suspected of being erroneous in the first data.

11. The method according to claim 6, characterized in that After obtaining the first data, the method further includes: When the number of predicted entities included in the first data is more than two, any one of the predicted entities is restored to a corresponding key entity according to a set probability to obtain newly added first data.

12. The method according to any one of claims 1 to 5, characterized in that: Also includes: If the questionable content does not exist in the output of the target large model, the quality of the question and answer data is determined to be qualified.

13. A question and answer data quality inspection device, characterized in that: include: A data acquisition unit, used to acquire question and answer data to be inspected, wherein the question and answer data includes question data and answer data; A questionable content identification unit, used for sending the question and answer data to a configured target macro model to instruct the target macro model to output the content suspected of being erroneous in the answer data as questionable content; An external knowledge retrieval unit, configured to retrieve external knowledge information for verifying the correctness of the questionable content in the answer data from various information sources; An external knowledge analysis unit, used to determine the authority of the source of the external knowledge information and the conflict between the external knowledge information and the answer data; A quality evaluation unit is used to determine a quality evaluation result of the question and answer data in combination with the authority of the information source and the conflict.

14. An electronic device, characterized in that: include: Memory and processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the question and answer data quality inspection method as described in any one of claims 1 to 12.

15. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the question and answer data quality inspection method according to any one of claims 1 to 12 is implemented.

16. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, each step of the question and answer data quality inspection method as described in any one of claims 1 to 12 is implemented.

Citation Information

Cited By

  • Artificial intelligence vertical large model training method focusing on petroleum coke industry

    CN120744027A

  • Data processing method and device

    CN121478938A