Attribution method and system
Through the automated attribution method combined with the large language model, the problems of low efficiency and poor accuracy of manual attribution analysis are solved, and efficient identification and optimization of wrong examples of the question-and-answer system are realized, improving the efficiency and accuracy of attribution analysis.
Patent Information
- Application Number
- CN202510408727.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, manual attribution analysis is inefficient when processing large-scale wrong case data sets and relies on subjective judgment of experts, resulting in inconsistent attribution results and poor accuracy, making it difficult to cope with the rapid iterative optimization of machine learning models and systems.
Through the automated attribution method, the evaluation data and module evaluation indicators of the Q&A system are used, and attribution analysis is performed in combination with large language models. The target modules that cause wrong examples are automatically identified, and the attribution analysis report is generated and optimization suggestions are provided.
It improves the efficiency and accuracy of attribution analysis, reduces the dependence of manual attribution, ensures the consistency and reliability of attribution results, and supports the rapid optimization of the Q&A system.
Smart Images

Figure CN120338083A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of artificial intelligence technology, and in particular, to an attribution method and system. Background Art
[0002] In the context of the rapid development of artificial intelligence technology, the training and deployment cycles of machine learning models have been significantly shortened, and the iteration frequency of machine learning models and systems based on machine learning models has also increased accordingly. Attribution analysis has become a key link in optimizing models and systems. Attribution analysis refers to analyzing the wrong cases (i.e., cases where the performance does not meet expectations) that occur in the actual application of a model or system, and identifying the reasons that lead to the wrong cases. Then, based on the results of the attribution analysis, the model or system is optimized and improved.
[0003] Currently, attribution analysis mainly relies on manual attribution. However, manual attribution requires a large amount of human resources and time, especially when dealing with large-scale wrong case data sets, the efficiency is low. And the attribution results of manual attribution depend on the subjective judgment of experts, which may lead to inconsistent opinions among different experts, affecting the accuracy and consistency of attribution. Due to relying on manual operations, it is difficult for manual attribution to cope with the complexity and dynamic changes of large-scale systems, and the scalability is poor. Therefore, manual attribution is time-consuming, laborious, subjective, and difficult to expand, seriously restricting the rapid optimization and deployment of models and systems.
[0004] The content in the background art section is only the information known to the inventor personally, and does not represent that the above information has entered the public domain before the filing date of this disclosure, nor does it represent that it can become the prior art of this disclosure. Summary of the Invention
[0005] This specification provides an attribution method and system. The attribution method can achieve automated attribution analysis, determine the module that causes the wrong case in the question-and-answer system, thereby improving the efficiency of attribution analysis, and enhancing the accuracy of attribution analysis while reducing the dependence on manual attribution.
[0006] In a first aspect, this specification provides an attribution method, including:
[0007] Obtain Q&A data, where the Q&A data includes: the question information received by the Q&A system and the answer information output by the Q&A system for the question information; obtain the evaluation data corresponding to the Q&A data, where the evaluation data includes the overall evaluation indicators corresponding to the Q&A system and the module evaluation indicators corresponding to multiple modules of the Q&A system, the overall evaluation indicator characterizes the degree of adaptation of the answer information to the question information, and the module evaluation indicator characterizes the degree of adaptation of the intermediate result output by the corresponding module in the Q&A system to the question information; based on the overall evaluation indicator, determine whether the Q&A data is a wrong example; and in the case where the Q&A data is a wrong example, determine a target module among the multiple modules based on the module evaluation indicators corresponding to the multiple modules, and the target module is the module that causes the Q&A data to be a wrong example.
[0008] In some embodiments, determining the target module among the multiple modules based on the module evaluation indicators corresponding to the multiple modules includes: obtaining the index ranges corresponding to the multiple modules; and based on the index ranges corresponding to the multiple modules, determining the modules that meet the target conditions among the multiple modules as the target modules, where the target conditions include: the module evaluation indicator corresponding to the module is not within the index range corresponding to the module.
[0009] In some embodiments, determining the target module among the multiple modules based on the module evaluation indicators corresponding to the multiple modules includes: obtaining the module operation data corresponding to the multiple modules, where the module operation data corresponding to a module is the data generated by the module during the process of the Q&A system generating the answer information; and determining the target module among the multiple modules based on the module evaluation indicators corresponding to the multiple modules and the module operation data corresponding to the multiple modules.
[0010] In some embodiments, determining the target module among the multiple modules based on the module evaluation indicators corresponding to the multiple modules and the module operation data corresponding to the multiple modules includes: providing the module evaluation indicators corresponding to the multiple modules and the module operation data corresponding to the multiple modules to a large model for attribution analysis, and determining the target module based on the attribution analysis result of the large model.
[0011] In some embodiments, providing the module evaluation metrics corresponding to the multiple modules and the module operation data corresponding to the multiple modules to a large model for attribution analysis, and determining the target module based on the attribution analysis result of the large model includes: obtaining preset attribution reference information, where the attribution reference information includes error description information corresponding to multiple preset errors, and each preset error has an association relationship with one of the multiple modules; providing the module evaluation metrics corresponding to the multiple modules, the module operation data corresponding to the multiple modules, and the attribution reference information to the large model, and guiding the large model to determine a target error hit by the question-and-answer system among the multiple preset errors; and determining the module associated with the target error as the target module.
[0012] In some embodiments, providing the module evaluation metrics corresponding to the multiple modules and the module operation data corresponding to the multiple modules to a large model for attribution analysis, and determining the target module based on the attribution analysis result of the large model includes: traversing the multiple modules in the execution order of the multiple modules, and for the current module in the traversal: providing the question information, the module evaluation metrics corresponding to the current module, and the module operation data corresponding to the current module to the large model, and guiding the large model to predict whether the current module will cause the adaptation degree between the answer information and the question information not to meet the expectation; and if so, determining the current module as the target module.
[0013] In some embodiments, for any target module among the multiple modules, the module operation data corresponding to the target module includes at least one of the following: intermediate results output by the target module; or log information generated by the target module.
[0014] In some embodiments, determining whether the question-and-answer data is a wrong example based on the overall evaluation metric includes: obtaining the metric range corresponding to the overall evaluation metric; and determining whether the question-and-answer data is a wrong example based on the metric range corresponding to the overall evaluation metric and the wrong example condition, where the wrong example condition includes: if the overall evaluation metric is not within the metric range corresponding to the overall evaluation metric, the question-and-answer data is a wrong example, or if the overall evaluation metric is within the metric range corresponding to the overall evaluation metric, the question-and-answer data is not a wrong example.
[0015] In some embodiments, after determining the target module, the method further includes: generating an attribution analysis report corresponding to the question-and-answer data based on the question-and-answer data, the evaluation data, and the target module; and displaying the attribution analysis report on an interaction interface or sending the attribution analysis report to a target device.
[0016] In some embodiments, the method further includes: providing the Q&A data, the evaluation data, and the target module to a large model, and guiding the large model to generate optimization suggestions for the target module; generating an attribution analysis report corresponding to the Q&A data based on the Q&A data, the evaluation data, and the target module, including: generating the attribution analysis report based on the Q&A data, the evaluation data, the target module, and the optimization suggestions.
[0017] In some embodiments, the method further includes: after detecting a change in the target module, inputting the question information into the changed Q&A system to obtain new Q&A data, where the new Q&A data includes the question information and new answer information; obtaining evaluation data corresponding to the new Q&A data, and determining whether the change in the target module meets the expectation based on the evaluation data corresponding to the new Q&A data.
[0018] In some embodiments, the Q&A system is a retrieval-enhanced Q&A system, and the multiple modules include at least two of: a question rewriting module, a knowledge retrieval module, a knowledge rearrangement module, a knowledge filtering module, and a generation module.
[0019] In some embodiments, the method further includes: obtaining the question information from a test dataset and sending the question information to the Q&A system.
[0020] In some embodiments, the question information is information input by a user to the Q&A system.
[0021] In a second aspect, this specification also provides an attribution system, including:
[0022] At least one storage medium storing at least one instruction set for performing attribution processing; and
[0023] At least one processor communicatively connected to the at least one storage medium, where when the attribution system runs, the at least one processor reads the at least one instruction set and executes the attribution method according to the instructions of the at least one instruction set as described in any item of the first aspect.
[0024] As can be seen from the above technical solutions, in the attribution method and system provided in the embodiments of this specification, it is possible to obtain question-and-answer data and the corresponding evaluation data for the question-and-answer data. Through the overall evaluation index, it is determined whether the question-and-answer data is an incorrect example; and in the case where the question-and-answer data is an incorrect example, based on the module evaluation indexes corresponding to the multiple modules, a target module is determined among the multiple modules, and the target module is the module that causes the question-and-answer data to be an incorrect example. Based on the evaluation data, automated attribution analysis is realized, and the target module that causes incorrect examples in the question-and-answer system is determined, thereby improving the efficiency of attribution analysis, reducing the dependence on manual attribution while enhancing the accuracy of attribution analysis, and providing reliable data support for the optimization of the question-and-answer system. Based on the analysis method of evaluation data, it is also possible to reduce the errors caused by subjectivity in manual attribution and improve the accuracy of attribution results.
[0025] Other functions of the attribution method and system provided in this specification will be partially listed in the following description. The creative aspects of the attribution method and system provided in this specification can be fully explained by practice or using the methods, devices, and combinations described in the detailed examples below. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0027] Figure 1 FIG. shows a schematic diagram of an application scenario of an attribution method provided according to an embodiment of this specification;
[0028] Figure 2 FIG. shows a hardware structure diagram of a computing system provided according to an embodiment of this specification;
[0029] Figure 3 FIG. shows a flowchart of an attribution method provided according to an embodiment of this specification; and
[0030] Figure 4 FIG. shows a schematic diagram of an attribution analysis process provided according to an embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] The following description provides specific application scenarios and requirements of this specification, aiming to enable those skilled in the art to manufacture and use the content in this specification. For those skilled in the art, various local modifications to the disclosed embodiments are obvious, and the general principles defined here can be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the illustrated embodiments, but has the broadest scope consistent with the claims.
[0032] The terms used herein are for the purpose of describing specific example embodiments only and are not restrictive. For example, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. When used in this specification, the terms "comprising", "including", and / or "containing" mean that the associated integers, steps, operations, elements, and / or components exist, but do not exclude the existence of one or more other features, integers, steps, operations, elements, components, and / or groups, or the addition of other features, integers, steps, operations, elements, components, and / or groups in the system / method.
[0033] In view of the following description, these features and other features of this specification, as well as the operations and functions of the related elements of the structure, and the economy of the combination and manufacture of the components can be significantly improved. Referring to the accompanying drawings, all of which form a part of this specification. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.
[0034] The flowcharts used in this specification illustrate the operations implemented by the system according to some embodiments in this specification. It should be clearly understood that the operations in the flowchart may not be implemented in sequence. On the contrary, the operations may be implemented in reverse order or simultaneously. In addition, one or more other operations may be added to the flowchart. One or more operations may be removed from the flowchart.
[0035] For the convenience of description, the terms that will appear in the following text of this specification are first explained.
[0036] Retrieval Augmented Generation (RAG): Retrieval Augmented Generation is a technology that combines retrieval and generation, aiming to enhance the generation ability of large language models by dynamically retrieving relevant information. Such methods usually rely on external knowledge bases or document collections, and can retrieve the context or facts related to the input from the knowledge base first when dealing with natural language tasks, and then use the large language model to generate accurate answers or texts based on the retrieval results.
[0037] Badcase: Refers to cases where the model's performance does not meet expectations, usually manifested as prediction errors, unreasonable outputs, or inconsistencies with user expectations.
[0038] In this specification, the Large Language Model (LLM) can also be abbreviated as the large model. A large language model is a natural language processing model based on deep learning technology, with the number of parameters usually reaching billions to hundreds of billions or even higher, and having powerful language understanding and generation capabilities. The large language model can adopt the Transformer architecture or its variants (such as GPT, BERT, etc.). This architecture uses the Attention Mechanism to achieve global modeling of sequential data, can efficiently handle long-range dependencies, and thus performs well in natural language tasks. The large language model learns the statistical features and semantic correlations of language by pre-training on a large-scale corpus, enabling it to have excellent generalization capabilities. The core capabilities of the large language model include but are not limited to: understanding context semantics, generating coherent and grammatically correct text, performing logical reasoning, and handling multi-task scenarios. Its usage methods usually include two modes: direct inference and fine-tuning. In the direct inference mode, the user guides the large language model to generate specific outputs by designing prompts. Prompts can be task descriptions or instructions in text form, used to stimulate the semantic understanding and generation capabilities of the large language model. In the fine-tuning mode, the large language model is further trained on a small-scale dataset in a specific domain to optimize its performance on specific tasks. The powerful generalization ability and flexibility of the large language model make it an important tool in the field of artificial intelligence technology, providing an efficient and accurate solution for automated text generation and understanding.
[0039] In some embodiments, the large language model can also have the ability to understand and generate data of other modalities (such as vision, audio, etc.). In this case, the large language model can also be called a Multimodal Large Language Model (MLLMs). MLLMs provide a richer and more natural interaction experience by integrating various types of inputs and outputs such as text, images, and sounds. The core advantage of MLLMs is their ability to process and understand information from different modalities and fuse this information to complete complex tasks. For example, MLLMs can analyze a picture and generate descriptive text, or generate corresponding images according to text descriptions. This cross-modal understanding and generation ability make MLLMs have broad application prospects in multiple fields.
[0040] It should be noted that the key technologies of large language models can be found in the detailed description in the paper "A Survey of Large Language Models" (paper number: arXiv:2303.18223v16, publication time: March 11, 2025, public link: https: / / doi.org / 10.48550 / arXiv.2303.18223), which will not be elaborated in this specification.
[0041] The application scenarios of this specification will be introduced below.
[0042] The solution provided in this specification can be used for attribution analysis, which includes locating wrong examples in the question-and-answer data and locating the root causes of the wrong examples. The attribution method in the embodiments of this specification can run through all stages of the model or system life cycle. For example, from the training and evaluation stage to the inference stage, the model or system can be optimized specifically based on the attribution method provided in this specification. For example, the solution provided in this specification can attribute the question-and-answer data of a question-and-answer system (such as a RAG system), locate the wrong examples in the question-and-answer system and the specific modules that cause the wrong examples. Taking the question-and-answer system as an example, the question-and-answer data includes: the question information received by the question-and-answer system and the answer information output by the question-and-answer system for the question information.
[0043] In some embodiments, the solution provided in this specification can be applied to the training and evaluation stage of the question-and-answer system. The question-and-answer data can include the question information in the test data set and the answer information generated by the question-and-answer system based on the question information. By analyzing the evaluation data of the above question-and-answer data, wrong examples and the specific modules that cause the wrong examples can be determined, and then the optimization direction of the question-and-answer system in the training and evaluation stage can be clarified.
[0044] In some embodiments, the solution provided in this specification can also be applied to the inference stage of the question-and-answer system. The question information in the question-and-answer data can be the information input by the user to the question-and-answer system, and the answer information can be the answer information generated by the question-and-answer system based on the question information. By analyzing the wrong examples generated during the actual use of the user, the weak links of the question-and-answer system in the real scenario can be located more accurately, so as to optimize the online question-and-answer system specifically.
[0045] Through the attribution analysis method provided in this specification, the reasons for the frequent occurrence of wrong examples in each stage of the question-and-answer system life cycle can also be identified, so as to provide a clear direction for the optimization of the question-and-answer system and further improve the robustness and practicality of the question-and-answer system.
[0046] It should be noted that the above application scenarios of the attribution analysis are only some examples among multiple usage scenarios. The attribution method provided in this specification can be applied not only to the scenarios listed above, but also to other scenarios that require misexample attribution. Those skilled in the art should understand that when the attribution method provided in this specification is applied to other usage scenarios, its implementation manner and technical effects are similar. This embodiment does not limit the object type of the attribution analysis, and is applicable to various question-and-answer data types and application scenarios such as text, images, videos, and audios generated by the question-and-answer system.
[0047] Figure 1 FIG. shows a schematic diagram of an application scenario of an attribution method provided according to an embodiment of this specification. As Figure 1 shown, the application scenario 100 may include an attribution system 110 and a question-and-answer system.
[0048] The attribution system 110 obtains question-and-answer data, where the question-and-answer data is the question information received by the question-and-answer system and the answer information output by the question-and-answer system in response to the question information. The attribution system 110 further obtains the evaluation data corresponding to the question-and-answer data. The attribution system 110 determines whether the question-and-answer data is a misexample based on the overall evaluation index in the evaluation data; and in the case of a misexample, determines the target module that causes the question-and-answer data to be a misexample among multiple modules of the question-and-answer system based on the module evaluation index.
[0049] The above application scenario 100 can be applied to multiple stages of the question-and-answer system, such as the training evaluation stage and the inference stage of the question-and-answer system.
[0050] In the training evaluation stage of the question-and-answer system, the question-and-answer system can generate answer information based on the question information in the test data set. The attribution system 110 can perform attribution analysis based on the answer data generated from the test data set and the corresponding evaluation data.
[0051] In the inference stage of the question-and-answer system, the question-and-answer system can generate answer information based on the user's question information. The attribution system 110 can perform attribution analysis based on the answer data generated from the user's question and the corresponding evaluation data.
[0052] In some embodiments, the attribution system 110 may store data and instructions for implementing the attribution method provided in this specification, and may execute or be used to execute the data and instructions. In some embodiments, the attribution system 110 may include a hardware device with data information processing capabilities and the necessary programs required to drive the hardware device to work.
[0053] It can be understood that the question-and-answer system and the attribution system 110 may correspond to the same system or different systems, and this specification does not limit this.
[0054] It should be noted that the attribution system 110 can correspond to a single device or a cluster of devices, and this specification places no restrictions on this. When the attribution system 110 corresponds to a single device, the attribution method can be executed entirely on that device. When the attribution system 110 corresponds to a cluster of devices, the attribution method can be executed in cooperation on multiple devices corresponding to the device cluster, and this specification places no restrictions on this.
[0055] It should be noted that all user data obtained in this specification has been authorized by the user and does not involve user privacy.
[0056] Figure 2 The hardware structure diagram of a computing system 200 provided according to an embodiment of this specification is shown. The computing system 200 can serve as Figure 1 the attribution system 110 in
[0057] As Figure 2 shown, the computing system 200 may include at least one storage medium 230 and at least one processor 220. In some embodiments, the computing system 200 may further include a communication port 250 and an internal communication bus 210. The computing system 200 may further include I / O components 260.
[0058] The internal communication bus 210 can connect different system components. For example, the internal communication bus 210 can connect the storage medium 230, the processor 220, the communication port 250, and the I / O components 260, etc.
[0059] The I / O components 260 support input / output between the computing system 200 and other components.
[0060] The communication port 250 is used for data communication between the computing system 200 and the outside world. For example, the communication port 250 can be used for data communication between the computing system 200 and the network 140. The communication port 250 can be a wired communication port or a wireless communication port.
[0061] The storage medium 230 may include a data storage device. The data storage device can be a non-transitory storage medium or a transitory storage medium. For example, the data storage device can include one or more of a magnetic disk 232, a read-only storage medium (ROM) 234, or a random access storage medium (RAM) 235. The storage medium 230 further includes at least one instruction set stored in the data storage device. The instruction set may include computer program code, and the computer program code may include programs, routines, objects, components, data structures, processes, modules, and so on.
[0062] At least one processor 220 may be communicatively coupled to at least one storage medium 230. When the computing system 200 is running, at least one processor 220 reads the at least one instruction set and, according to the instructions of the at least one instruction set, executes the attribution method provided in this specification. The processor 220 may execute the steps included in the attribution method. The processor 220 may be in the form of one or more processors. In some embodiments, the processor 220 may include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application specific integrated circuit (ASIC), an application specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physics processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of performing one or more functions, etc., or any combination thereof.
[0063] For illustrative purposes only, the computing system 200 in the drawings only shows one processor 220. However, it should be noted that the computing system 200 in this specification may also include multiple processors. Therefore, the operations and / or method steps disclosed in this specification may be executed by one processor or jointly executed by multiple processors. For example, if it is described in this specification that the processor 220 of the computing system 200 executes step A and step B, it should be understood that step A and step B may also be jointly or separately executed by two different processors 220 (for example, the first processor executes step A, the second processor executes step B, or the first and second processors jointly execute steps A and B).
[0064] Figure 3 A flowchart of an attribution method P300 provided according to an embodiment of this specification is shown. As before, the computing system 200 may execute the attribution method P300 of this specification. Specifically, the processor 220 in the computing system 200 may read the instruction set stored in its local storage medium and then, according to the provisions of the instruction set, execute the attribution method P300 of this specification. As Figure 3 shown, the method P300 may include steps S310 - S340.
[0065] S310: Obtain question - and - answer data, where the question - and - answer data includes: the question information received by the question - and - answer system and the answer information output by the question - and - answer system for the question information.
[0066] In some embodiments, the question-answering system has the ability to output answer information based on the question information. The question-answering system can be a system based on the RAG architecture, or a system based on a neural network or LLM model architecture. The attribution method in the embodiments of this specification can attribute the above-built system. It should be noted that the question-answering system mentioned in the embodiments of this specification is not limited to a hardware system or a software system, and can also refer to a model with the ability to output answer information based on the question information.
[0067] As mentioned above, the attribution method in the embodiments of this specification can run through all stages of the model or system life cycle. Based on different application scenarios, the sources of the question-and-answer data are different.
[0068] In some embodiments, the solution provided in this specification can be applied to the training and evaluation stage of the question-answering system. In this scenario, the computing system 200 can obtain the question information from the test data set and send the question information to the question-answering system. Specifically, the computing system 200 evaluates the question-answering system during the training and testing process through the question information samples in the test data set. The test data set can contain various types of question information, covering different fields and scenarios, to help train the question-answering system to answer questions better. The computing system 200 can evaluate the performance of the question-answering system based on the evaluation data. The computing system 200 can also compare the answers generated by the question-answering system with the sample answers in the test data set, so as to evaluate the performance of the question-answering system. Testers can adjust and optimize the question-answering system based on the attribution results of the computing system 200 for the question-answering system, so as to improve the accuracy and generation ability of the system. The embodiments of this specification use the computing system 200 with an automatic attribution function to quickly evaluate the question-answering system, significantly shortening the evaluation cycle.
[0069] In some embodiments, the solution provided in this specification can also be applied to the inference stage of the question-answering system. In this scenario, the question information in the question-and-answer data is the information input by the user to the question-answering system. That is to say, the computing system 200 can obtain the question-and-answer data from the interaction between the user and the question-answering system. The question-and-answer data can include the question information submitted by the user and the answer information generated by the question-answering system according to these questions during the actual application process of the question-answering system. The question-and-answer data can also cover multiple question-and-answer pairs exchanged between the user and the question-answering system in the current round of conversation. Based on the above question-and-answer data in the interaction process between the user and the answer system, the computing system 200 can attribute the wrong cases of the question-and-answer data of the online question-answering system, analyze and identify the errors or deficiencies that occur when the question-answering system generates answers. The computing system 200 can find the target module in the question-answering system that causes the wrong cases, and can continuously monitor the target module to realize the automatic evaluation of the online question-answering system.
[0070] S320: Obtain the evaluation data corresponding to the Q&A data. The evaluation data includes the overall evaluation indicators corresponding to the Q&A system and the module evaluation indicators corresponding to multiple modules of the Q&A system. The overall evaluation indicators represent the degree of adaptation of the answer information to the question information, and the module evaluation indicators represent the degree of adaptation of the intermediate results output by the corresponding modules in the Q&A system to the question information.
[0071] In some embodiments, the evaluation data can be obtained from an evaluation model or system specifically designed to evaluate the Q&A system, and the acquisition of the evaluation data can rely on an external evaluation model or system. In some embodiments, the Q&A system itself has the ability to generate or output evaluation data, or an evaluation module capable of generating evaluation data is configured. In this case, the evaluation data can be directly obtained from the Q&A system. Additionally, in some embodiments, the computing system 200 can also have the ability to evaluate the Q&A system to generate evaluation data and obtain the evaluation data through internal calculations.
[0072] In some embodiments, if the Q&A system is a Q&A system based on the RAG architecture, the multiple modules may include at least two of: a question rewriting module, a knowledge retrieval module, a knowledge rearrangement module, a knowledge filtering module, and a generation module. The above-mentioned question rewriting module is responsible for optimizing and reconstructing the question information to make it more accurate, standardized, or suitable for subsequent retrieval and generation processing. The knowledge retrieval module is responsible for retrieving documents or fragments related to the question from a pre-constructed knowledge base or external document library according to the question information or the rewritten question information. The knowledge rearrangement module is responsible for sorting multiple relevant documents or document blocks to select relevant knowledge to generate an answer. The knowledge filtering module is responsible for filtering the retrieved documents to remove low-quality, repetitive, or irrelevant content to the question information to ensure that the information finally provided by the system has high quality and accuracy. The generation module generates the final answer by inputting the question information and knowledge into a generation model. It should be noted that the modules in the Q&A system described in this specification are not limited, and those skilled in the art can select appropriate modules to construct the Q&A system based on different usage scenarios.
[0073] In some embodiments, the module evaluation metrics corresponding to different modules may be different. For example, the question rewriting module may include metrics such as rewriting accuracy and semantic consistency. The evaluation metrics of the knowledge retrieval module may include metrics such as context precision and context recall. The evaluation metrics of the knowledge rearrangement module may include metrics such as rearrangement precision and rearrangement relevance. The evaluation metrics of the knowledge filtering module may include metrics such as filtering accuracy and noise filtering ability. The evaluation metrics of the generation module may include metrics such as language fluency and semantic consistency. It should be noted that the above module evaluation metrics are not used to limit the capabilities or evaluation directions of each module in this specification. Those skilled in the art can select the module evaluation metrics of different modules based on different usage scenarios to evaluate the adaptation degree of the intermediate results output by the corresponding module to the question information. In some embodiments, different modules may also have the same module evaluation metrics, and this specification does not limit the selection of module evaluation metrics either.
[0074] S330: Determine whether the Q&A data is a wrong example based on the overall evaluation metric.
[0075] In some embodiments, the overall evaluation metric may be represented by metrics such as the accuracy, precision, recall, or adaptability of the answer information, or may also be represented by the satisfaction of users with the answer information collected during the use of the Q&A system.
[0076] In some embodiments, the computing system 200 may also obtain the metric range corresponding to the overall evaluation metric; and determine whether the Q&A data is a wrong example based on the metric range corresponding to the overall evaluation metric and the wrong example condition. Wherein, the wrong example condition includes: if the overall evaluation metric is not within the metric range corresponding to the overall evaluation metric, the Q&A data is a wrong example, or if the overall evaluation metric is within the metric range corresponding to the overall evaluation metric, the Q&A data is not a wrong example.
[0077] For example, if the metric range corresponding to the overall evaluation metric is [0.6, 1], and the overall evaluation metric of the current Q&A data is 0.45, then it is considered that the Q&A data is a wrong example.
[0078] For another example, the evaluation results can be divided into three cases according to the metric range of the overall evaluation metric: correct answer (metric range: [0.8, 1]), wrong answer (metric range: [0, 0.4]), and poor answer (that is, the answer is not good enough, metric range: (0.4, 0.8)). In the case of a wrong answer or a poor answer, it will be determined as a wrong example. If the overall evaluation metric of the current Q&A data is 0.45, then the answer is considered to be a "poor answer" and belongs to the wrong example category. Through the above embodiments, the computing system 200 can directly identify wrong examples based on the overall evaluation metric.
[0079] S340: When the Q&A data is a wrong example, determine a target module from the multiple modules based on the module evaluation metrics corresponding to the multiple modules, where the target module is the module that causes the Q&A data to be a wrong example.
[0080] Multiple embodiments for determining the target module are disclosed in this specification. For example, attribution analysis can be performed through index quantification, or attribution can be performed through comprehensive analysis.
[0081] In some embodiments, attribution analysis can be performed through index quantification. The computing system 200 obtains the index ranges corresponding to the multiple modules; and based on the index ranges corresponding to the multiple modules, determines the modules that meet the target conditions in the multiple modules as the target modules, where the target conditions include: the module evaluation metrics corresponding to the module are not within the index range corresponding to the module.
[0082] In some embodiments, if a module has multiple module evaluation metrics, each module evaluation metric and its corresponding index range can be compared and judged according to the target conditions. The target conditions can be: there is an index among the multiple module evaluation metrics corresponding to the module that is not within the index range corresponding to the module. For example, taking the knowledge retrieval module as an example, the evaluation metrics of the knowledge retrieval module can include metrics such as context precision and context recall rate, and the index ranges are [0.8, 1] and [0.7, 1] respectively. If the context precision and context recall rate metrics of the Q&A data are 0.7 and 0.75 respectively, since the context precision metric (0.7) is not within the index range [0.8, 1], it can be considered that the knowledge retrieval module is the target module.
[0083] In some embodiments, if a module has multiple module evaluation metrics, the weighted values of the multiple evaluation metrics and the index ranges can also be comprehensively compared to judge the target module. The target conditions can be: the weighted value of the multiple module evaluation metrics corresponding to the module is not within the index range corresponding to the module. Taking the knowledge retrieval module as an example, assuming the index range of the context precision is [0.8, 1], the context precision and context recall rate in the Q&A data are 0.7 and 0.75 respectively, and the weights are 0.5. At this time, the weighted value can be calculated as: (0.7 * 0.5) + (0.75 * 0.5) = 0.725. Since 0.725 is not within the index range [0.8, 1], the weighted average value of the knowledge retrieval module does not meet the expectation, so it can be considered that the knowledge retrieval module is the target module.
[0084] By comparing the overall evaluation metrics and the actual values of the module evaluation metrics with the preset index ranges in the above embodiments, the computing system 200 can quickly determine the wrong example and the corresponding target module.
[0085] In some embodiments, there may be multiple modules whose module evaluation indicators are not within the indicator range corresponding to the module. In the embodiments of this specification, a variety of solutions are provided. For example, the module with the execution order at the front can be used as the target module according to the execution order of the multiple modules. This is because if there is a problem with the previous module, it may cause the output results of the subsequent modules to be less compatible with the question information, affecting the performance of the entire system. For example, in a RAG question-answering system, if the knowledge retrieval module does not correctly extract information highly related to the question information in the initial stage, the subsequent knowledge rearrangement module, generation module, etc. may be sorted and generated based on inaccurate or irrelevant contexts, resulting in a decrease in the quality of the final answer. In this way, although the evaluation indicators of the subsequent modules may not meet expectations, the root cause of the problem is often in the front module. Therefore, by using the module with the front execution order as the target module, the computing system 200 can quickly locate the source of the problem. In actual operation, the computing system 200 can check the module evaluation indicators corresponding to the front module according to the execution order of each module.
[0086] In some embodiments, the target module can be further determined from multiple modules that are not within the scope of the indicator by comprehensively analyzing the module evaluation indicators and the module operation data. The comprehensive analysis method can also be directly used to determine the target module from all modules. Specifically, the comprehensive analysis method includes: the computing system 200 obtains the module operation data corresponding to the multiple modules, wherein the module operation data corresponding to a module is the data generated by the module in the process of the question-and-answer system generating the answer information; and based on the module evaluation indicators corresponding to the multiple modules and the module operation data corresponding to the multiple modules, the target module is determined from the multiple modules.
[0087] In some embodiments, taking the target module as an example, the module operation data corresponding to the target module includes at least one of the following: an intermediate result output by the target module; or log information generated by the target module.
[0088] In some embodiments, the module operation data may include intermediate results generated during the execution of the module, log information, variable states during the calculation process, etc. By analyzing this module operation data, the computing system 200 can more accurately and quickly locate the target module. For example, the intermediate results include data or temporary results generated by the module during the processing, which can help the computing system 200 locate the target module in the question-answering system. For example, in the knowledge retrieval module, the output retrieval results can be used as intermediate results, and it can be analyzed based on the module evaluation metrics of the knowledge retrieval module and the retrieval results whether the inaccurate retrieval leads to the unsatisfactory quality of the finally generated answer. Log information usually contains the input, output, and execution status of each operation. By analyzing the log information, the computing system 200 can trace possible abnormal behaviors or incorrect operations during the module execution process, helping to locate the target module that causes the wrong case.
[0089] In some embodiments, the computing system 200 can also perform comprehensive analysis through a large model. The computing system 200 can provide the module evaluation metrics corresponding to the multiple modules and the module operation data corresponding to the multiple modules to the large model for attribution analysis, and determine the target module based on the attribution analysis result of the large model.
[0090] This specification provides multiple embodiments of using a large model for attribution analysis. For example, the large model can be guided to perform a classification task based on preset attribution reference information to locate the target module corresponding to the wrong case, thereby achieving the purpose of attribution analysis. The large model can also be guided to perform a generation task based on the module evaluation metrics and module operation data to identify the target module corresponding to the wrong case to implement attribution analysis.
[0091] In some embodiments, the large model performs a classification task to determine the target module. The computing system 200 obtains preset attribution reference information, where the attribution reference information includes error description information corresponding to multiple preset errors, and each preset error has an association relationship with one of the multiple modules; provides the module evaluation metrics corresponding to the multiple modules, the module operation data corresponding to the multiple modules, and the attribution reference information to the large model, and guides the large model to determine the target error hit by the question-answering system among the multiple preset errors; and determines the module associated with the target error as the target module.
[0092] In some embodiments, a prompt can be generated based on the module evaluation metrics corresponding to the multiple modules, the module operation data corresponding to the multiple modules, and the attribution reference information, and the prompt is input into the large model to guide the large model to determine the target error hit by the question-answering system among the multiple preset errors; and determines the module associated with the target error as the target module.
[0093] In some embodiments, the attribution reference information may be stored in a tabular form, or other feasible storage methods may be used. Please refer to Table 1, which is used to store attribution reference information, and systematically lists some common preset error points (Failure Points, FP) in the RAG system and their corresponding descriptions. Please continue to refer to Table 2, which is used to store the mapping relationship between error points and corresponding modules, showing the association between each preset error and its corresponding module, and establishing a mapping relationship between error points and functional modules. The large model can accurately classify the error examples into corresponding error points based on the description in Table 1, and then map the error points to specific functional modules according to Table 2 to determine the target module that causes the error example.
[0094] Table 1
[0095]
[0096] Table 2
[0097]
[0098] It should be noted that only some error points and their corresponding modules are shown in Table 1 and Table 2. Those skilled in the art can further establish and expand the preset attribution reference information according to actual application scenarios to meet different usage requirements.
[0099] In some embodiments, the large model can also perform a generation task to determine a target module. The computing system 200 traverses the multiple modules in the execution order of the multiple modules, and for the current module in the traversal: provides the question information, the module evaluation index corresponding to the current module, and the module operation data corresponding to the current module to the large model, and guides the large model to predict whether the current module will cause the adaptation degree between the answer information and the question information to not meet expectations; and if so, determines the current module as the target module.
[0100] Based on the same reasons as above, in the above embodiment, the computing system 200 traverses these modules one by one in the order of module execution, and makes judgments starting from the front module. The large model can perform a preliminary screening of each module according to the module evaluation index, and quickly locate the modules that may have problems. Then, the large model can further identify the target module by comparing and analyzing the module operation data of these modules. In this embodiment, the large model can not only perform a rapid evaluation based on the module evaluation index, but also accurately locate the target module through in-depth analysis of the module operation data.
[0101] In the above embodiments of this specification, multiple methods for determining the target module are disclosed. These methods can be executed in parallel or serially. For example, the computing system 200 can perform attribution through both metric quantification analysis and comprehensive analysis. These two methods can be executed simultaneously and complement each other, thereby improving the efficiency and accuracy of attribution analysis. On the other hand, these two methods can also be executed serially, that is, first perform attribution through metric quantification analysis. If the target module cannot be determined, then further use comprehensive analysis for attribution to ensure that the target module can be accurately identified.
[0102] In some embodiments, the comprehensive analysis method can also be implemented by guiding the large model in multiple ways. These guiding methods can be executed in parallel or serially. Taking serial execution as an example, the large model can be first guided to perform a classification task based on the preset attribution reference information, so as to perform attribution processing on the defined wrong case types. For the wrong cases not defined in the preset attribution reference information, the computing system 200 can further guide the large model to perform a generation task, analyze the wrong cases and module performance through the generation task, and determine the target module. The computing system 200 can automatically complete the qualitative classification of wrong cases and module-level attribution based on the preset attribution reference information and in combination with the analysis ability of the large language model, greatly improving the attribution efficiency.
[0103] This specification does not limit the combination of different attribution analysis methods. The computing system 200 can flexibly select or combine these methods according to actual needs to achieve attribution processing that meets the requirements.
[0104] Traditional attribution methods rely on manual annotation and analysis, which are time-consuming and laborious, and are difficult to meet the requirements of rapid model iteration. Moreover, the results of manual attribution are often limited by the subjective judgment of experts, which may lead to inconsistent attribution conclusions. Through the attribution method of this specification, based on multiple attribution methods, the powerful reasoning capabilities of metric quantification analysis and large language models can be used to reduce the interference of human subjective judgment, improving the accuracy and consistency of attribution results. Through the above embodiments, the computing system 200 can also efficiently locate the target module, identify the error source, and thus provide accurate attribution suggestions for subsequent system optimization.
[0105] After determining the target module, the computing system 200 can also generate an attribution analysis report corresponding to the Q&A data based on the Q&A data, the evaluation data, and the target module; display the attribution analysis report on the interaction interface, or send the attribution analysis report to the target device.
[0106] In some embodiments, the computing system 200 further has an intelligent dispatching function, which can accurately assign each error case to the corresponding handler. The attribution analysis report can be displayed on the interaction interface or sent to the target device to help technicians comprehensively understand the root cause of the error cases. In some embodiments, the target device can be the device of the person in charge corresponding to the target module, and the interaction interface can be the interaction interface of the device of the person in charge corresponding to the target module. By using the intelligent dispatching function to assign error cases to the corresponding handlers, the efficient processing of error cases is achieved through automated assignment on the basis of accurate positioning of error cases.
[0107] In some embodiments, the attribution analysis report generated based on the Q&A data, the evaluation data, and the target module may include the following content.
[0108] 1. Basic information of error cases: including basic information of error cases such as error case ID, occurrence time, and Q&A data. Other basic information such as system version and execution environment may also be included to evaluate the impact of system status and configuration on the occurrence of problems.
[0109] 2. Error type analysis: Classify the error cases in combination with preset attribution reference information and classify the error cases to specific error points.
[0110] 3. Module-level attribution result: including the target module determined based on the attribution method of the embodiments of this specification. The aforementioned evaluation results such as correct answer, wrong answer (index range [0, 0.4]), and poor answer may also be included.
[0111] 4. Module operation data: Display the intermediate output status of each module to help technicians further understand the performance of each module of the Q&A system.
[0112] In some embodiments, the attribution report may further include optimization suggestions for the target module to provide specific guidance for optimization. The optimization suggestions can be generated by guiding the large model, thereby helping to improve the performance of the target module. Specifically: the computing system 200 can also provide the Q&A data, the evaluation data, and the target module to the large model and guide the large model to generate optimization suggestions for the target module; generating the attribution analysis report corresponding to the Q&A data based on the Q&A data, the evaluation data, and the target module includes: generating the attribution analysis report based on the Q&A data, the evaluation data, the target module, and the optimization suggestions.
[0113] In some embodiments, the optimization suggestions may be improvement methods proposed for the problems existing in the target module, such as adjusting the parameters of the target module, improving the algorithm design of the target module, or optimizing the data processing flow, etc. Through these optimization suggestions, the computing system 200 can provide developers with a clearer improvement direction and specific operation steps, so that the question answering system can achieve better performance in training evaluation and actual application.
[0114] In some embodiments, the computing system 200 can also continuously monitor data to form a complete closed-loop learning link, providing data-driven support for the rapid iteration and precise optimization of the model. In some embodiments, it is also possible to perform automated testing on the changed or repaired system to verify whether the processing results of the same question information meet the expectations, and continuously monitor the changed question answering system. After detecting a change in the target module, the computing system 200 can input the question information into the changed question answering system to obtain new question answering data, where the new question answering data includes the question information and new answer information; obtain the evaluation data corresponding to the new question answering data, and determine whether the change in the target module meets the expectations based on the evaluation data corresponding to the new question answering data.
[0115] In some embodiments, it is possible to directly determine whether the change in the target module meets the expectations based on whether the new evaluation data meets the preset requirements. For example, the index range of the overall evaluation index is [0.6, 1]. If the overall evaluation index corresponding to the same question information after the change of the question answering system becomes 0.7, it can be considered that the change meets the expectations. It is also possible to determine whether the change in the target module meets the expectations by comparing the changes in the evaluation indexes corresponding to the same question information before and after the change of the question answering system. For example, the index range of the overall evaluation index is [0.6, 1]. The overall evaluation index before the change is 0.5, and after the change is 0.7, then it can be considered that the change meets the expectations.
[0116] It should be noted that the evaluation data in actual applications is more complex and may involve evaluation data in multiple dimensions. The above embodiments for judging based on the overall evaluation index are only examples provided in this specification. If the overall evaluation index after the change still does not meet the expectations, it is necessary to further perform attribution analysis on the new question answering data to determine whether the change in the target module conforms to the adjustment direction, or to identify a new target module when the change direction of the target module is correct.
[0117] In some implementations, the computing system 200 can also track the Q&A data of the same type as the wrong example for a long time, and monitor the evaluation data of this type of Q&A data to ensure that the changed Q&A system can effectively solve similar problems, thus ensuring the robustness of the Q&A system. Through long-term tracking, the computing system 200 can timely detect potential fallback risks and prevent the performance of the Q&A system from degrading or experiencing other adverse changes after optimization. If it is monitored that a certain indicator shows a fallback or abnormal change, the computing system 200 can also timely issue a prompt to remind relevant personnel to take necessary optimization measures to ensure that the system always maintains a high performance level.
[0118] Please refer to Figure 4 , through the embodiments in this specification, the computing system 200 can establish a complete attribution analysis closed-loop: including wrong example identification, wrong example attribution, attribution result generation, and wrong example backtesting, specifically.
[0119] Wrong example identification: The computing system 200 identifies wrong examples in the Q&A system through overall evaluation indicators.
[0120] Wrong example attribution: For module evaluation indicators, conduct attribution analysis based on indicator quantification, or comprehensive attribution analysis based on large models, etc., to locate the cause of the wrong example.
[0121] Attribution result generation: After clarifying the target module, generate an attribution report, and can also propose optimization suggestions based on the attribution results to provide technical support for technicians' repair work.
[0122] Wrong example backtesting: After the problem is repaired, the computing system 200 will also automatically backtest the wrong example and continuously monitor the performance of the repaired Q&A system to verify the optimization effect.
[0123] Through the above complete closed-loop process, the computing system 200 can continuously discover wrong examples and locate the source, helping technicians continuously optimize the performance of the Q&A system and ensuring that the Q&A system maintains efficient, accurate, and stable performance during long-term operation.
[0124] In summary, in the attribution method and system provided in this specification, the computing system 200 can obtain question-and-answer data and the corresponding evaluation data for the question-and-answer data, and determine whether the question-and-answer data is a wrong example through the overall evaluation index; and in the case where the question-and-answer data is a wrong example, determine a target module among the multiple modules based on the module evaluation indexes corresponding to the multiple modules, where the target module is the module that causes the question-and-answer data to be a wrong example. Based on the evaluation data, automated attribution analysis is realized, and the target module that causes wrong examples in the question-and-answer system is determined, thereby improving the efficiency of attribution analysis, reducing the dependence on manual attribution while improving the accuracy of attribution analysis, and providing reliable data support for the optimization of the question-and-answer system. Based on the analysis method of the evaluation data, it is also possible to reduce the errors caused by subjectivity in manual attribution and improve the accuracy of the attribution results.
[0125] On the other hand, this specification provides a computer-readable non-transitory storage medium storing at least one instruction set for performing attribution processing. When the at least one instruction set is executed by a processor, the at least one instruction set directs the processor to implement the steps of the attribution method P300 described in this specification. In some possible implementation manners, various aspects of this specification can also be implemented in the form of a program product, which includes program code. When the program product runs on the computing system 200, the program code is used to cause the computing system 200 to execute the steps of the attribution method P300 described in this specification. The program product for implementing the above method can be a portable compact disc read-only memory (CD-ROM) including program code and can run on the computing system 200. However, the program product of this specification is not limited to this. In this specification, the readable storage medium can be any tangible medium that contains or stores a program that can be used by or in combination with an instruction execution system. The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium include: a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer-readable storage medium can include a data signal propagated as part of a carrier wave in a baseband, where the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium that can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above. The program code for performing the operations of this specification can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the computing system 200, partially on the computing system 200, executed as an independent software package, partially on the computing system 200 and partially on a remote computing system, or entirely on a remote computing system.
[0126] The above description is of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require a particular order or a sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0127] In summary, after reading this detailed disclosure, those skilled in the art will appreciate that the foregoing detailed disclosure may be presented by way of example only and is not necessarily limiting. Although not explicitly stated herein, those skilled in the art will understand that this specification is intended to embrace various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be proposed by this specification and are within the spirit and scope of the exemplary embodiments of this specification.
[0128] In addition, certain terms in this specification have been used to describe the embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" mean that the specific features, structures, or characteristics described in connection with that embodiment may be included in at least one embodiment of this specification. Thus, it should be emphasized and understood that two or more references to "an embodiment" or "one embodiment" or "alternative embodiments" in various parts of this specification do not necessarily all refer to the same embodiment. Additionally, the specific features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.
[0129] It should be understood that in the foregoing description of the embodiments of this specification, for the purpose of helping to understand a feature, and for the purpose of simplifying this specification, this specification combines various features in a single embodiment, drawing, or its description. However, this does not mean that the combination of these features is necessary, and those skilled in the art may very well mark out some of the devices as separate embodiments when reading this specification. That is to say, the embodiments in this specification can also be understood as the integration of multiple sub - embodiments. And the content of each sub - embodiment is also valid when it has fewer features than all the features of a single foregoing disclosed embodiment.
[0130] Each patent, patent application, publication of patent application, and other materials cited herein, such as articles, books, specifications, publications, documents, items, etc., except for those that are inconsistent with or conflict with this document, or those that have a limiting effect on the broadest scope of the claims, may be incorporated herein by reference and used for all purposes now or hereafter associated with this document. In addition, in the event of any inconsistency or conflict between the description, definition, and / or use of relevant terms in any material and the description, definition, and / or use of relevant terms in this document, the terms in this document shall prevail.
[0131] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can adopt alternative configurations based on the embodiments in this specification to implement the application in this specification. Therefore, the embodiments of this specification are not limited to the embodiments precisely described in the application.
Claims
1. An attribution method, comprising: Obtaining Q&A data, where the Q&A data includes: question information received by a Q&A system and answer information output by the Q&A system for the question information; Obtaining evaluation data corresponding to the Q&A data, where the evaluation data includes an overall evaluation index corresponding to the Q&A system and module evaluation indexes corresponding to multiple modules of the Q&A system. The overall evaluation index represents the degree of adaptation of the answer information to the question information, and the module evaluation index represents the degree of adaptation of the intermediate result output by the corresponding module in the Q&A system to the question information; Based on the overall evaluation index, determining whether the Q&A data is a wrong example; and In the case where the Q&A data is a wrong example, determining a target module from the multiple modules based on the module evaluation indexes corresponding to the multiple modules, where the target module is the module that causes the Q&A data to be a wrong example.
2. The method according to claim 1, wherein, The determining a target module from the multiple modules based on the module evaluation indexes corresponding to the multiple modules includes: Obtaining the index ranges corresponding to the multiple modules; and Based on the index ranges corresponding to the multiple modules, determining the modules that meet the target conditions in the multiple modules as the target modules, where the target conditions include: the module evaluation index corresponding to a module is not within the index range corresponding to the module.
3. The method according to claim 1, wherein, The determining a target module from the multiple modules based on the module evaluation indexes corresponding to the multiple modules includes: Obtaining the module operation data corresponding to the multiple modules, where the module operation data corresponding to a module is data generated by the module during the process of the Q&A system generating the answer information; and Based on the module evaluation indexes corresponding to the multiple modules and the module operation data corresponding to the multiple modules, determining the target module from the multiple modules.
4. The method according to claim 3, wherein, The determining the target module from the multiple modules based on the module evaluation indexes corresponding to the multiple modules and the module operation data corresponding to the multiple modules includes: Providing the module evaluation indexes corresponding to the multiple modules and the module operation data corresponding to the multiple modules to a large model for attribution analysis, and determining the target module based on the attribution analysis result of the large model.
5. The method according to claim 4, wherein The providing the module evaluation indexes corresponding to the multiple modules and the module operation data corresponding to the multiple modules to a large model for attribution analysis, and determining the target module based on the attribution analysis result of the large model includes: Obtaining preset attribution reference information, where the attribution reference information includes error description information corresponding to multiple preset errors, and each preset error has an association relationship with one of the multiple modules; Providing the module evaluation indexes corresponding to the multiple modules, the module operation data corresponding to the multiple modules, and the attribution reference information to the large model, and guiding the large model to determine a target error hit by the Q&A system from the multiple preset errors; and Determining the module associated with the target error as the target module.
6. The method according to claim 4, wherein Providing the module evaluation metrics corresponding to the multiple modules and the module operation data corresponding to the multiple modules to a large model for causal analysis, and determining the target module based on the causal analysis result of the large model, includes: Traversing the multiple modules in the execution order of the multiple modules, and for the current module in the traversal: Providing the question information, the module evaluation metrics corresponding to the current module, and the module operation data corresponding to the current module to the large model, and guiding the large model to predict whether the current module will cause the adaptation degree between the answer information and the question information not to meet the expectation; and If so, determining the current module as the target module.
7. The method according to claim 3, wherein For any target module among the multiple modules, the module operation data corresponding to the target module includes at least one of the following: The intermediate result output by the target module; or The log information generated by the target module.
8. The method according to claim 1, wherein Based on the overall evaluation metrics, determining whether the Q&A data is a wrong example, includes: Obtaining the index range corresponding to the overall evaluation metrics; and Based on the index range corresponding to the overall evaluation metrics and the wrong example condition, determining whether the Q&A data is a wrong example, where the wrong example condition includes: If the overall evaluation metrics are not within the index range corresponding to the overall evaluation metrics, the Q&A data is a wrong example, or If the overall evaluation metrics are within the index range corresponding to the overall evaluation metrics, the Q&A data is not a wrong example.
9. The method according to claim 1, wherein After determining the target module, the method further includes: Generating an attribution analysis report corresponding to the Q&A data based on the Q&A data, the evaluation data, and the target module; Displaying the attribution analysis report on the interaction interface, or sending the attribution analysis report to the target device.
10. The method according to claim 9, wherein The method further includes: Providing the Q&A data, the evaluation data, and the target module to the large model, and guiding the large model to generate optimization suggestions for the target module; The generating an attribution analysis report corresponding to the Q&A data based on the Q&A data, the evaluation data, and the target module includes: Generating the attribution analysis report based on the Q&A data, the evaluation data, the target module, and the optimization suggestions.
11. The method according to claim 9, wherein, The method further includes: After detecting a change in the target module, inputting the question information into the changed Q&A system to obtain new Q&A data, where the new Q&A data includes the question information and new answer information; Obtaining the evaluation data corresponding to the new Q&A data, and determining whether the change of the target module meets the expectation based on the evaluation data corresponding to the new Q&A data.
12. The method according to claim 1, wherein, The Q&A system is a Q&A system based on retrieval augmented generation, and the multiple modules include at least two of: a question rewriting module, a knowledge retrieval module, a knowledge rearrangement module, a knowledge filtering module, and a generation module.
13. The method according to claim 1, wherein, The method further includes: Obtaining the question information from the test dataset, and sending the question information to the Q&A system.
14. The method according to claim 1, wherein, The question information is the information input by the user to the Q&A system.
15. An attribution system, includes: At least one storage medium storing at least one instruction set for performing attribution processing; And At least one processor communicatively connected to the at least one storage medium, wherein when the attribution system runs, the at least one processor reads the at least one instruction set and executes the attribution method according to any one of claims 1-14 according to the instructions of the at least one instruction set.
Citation Information
Patent Citations
Question and answer system evaluation method and device, computing equipment and storage medium
CN117290694A
Medical question and answer text quality evaluation system and method based on pre-training language model
CN117766160A
Evaluation method and device of large language model, electronic equipment and storage medium
CN118093824A
Question answering method and device based on large model, training method and device, intelligent agent, equipment and medium
CN119106123A
Artificial intelligence model evaluation method and device and storage medium
CN119271550A