Recall Accuracy Evaluation Method, System, Device, Medium and Product

By dividing the knowledge base into knowledge blocks and using the large language model to generate problems and traceability results, the automatic evaluation of the accuracy of the large language model recall is achieved, the problem of high cost of human evaluation is solved, and the evaluation accuracy and efficiency are improved.

CN119149687BActive Publication Date: 2025-06-20BEIJING QINGCHENG JIZHI TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411134516.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2025-06-20
Estimated Expiration
2044-08-19

AI Technical Summary

Technical Problem

When injecting data search capabilities into large language model applications through retrieval enhancement generation technology, it is necessary to measure the accuracy of the searched related recall content, but the existing technology relies on human resources and is costly.

Method used

By dividing the preset knowledge base into various knowledge blocks, and generating the corresponding problems of each knowledge block through a large language model, obtaining the traceability results of the knowledge blocks, and automatically evaluating the accuracy of the recall.

Benefits of technology

It realizes automated assessment of recall accuracy without human participation, reduces assessment costs, and improves the accuracy of recall accuracy assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119149687B_ABST
    Figure CN119149687B_ABST
Patent Text Reader

Abstract

The present application discloses a recall accuracy evaluation method, system, device, medium and product, relating to the technical field of large model applications. The disclosed recall accuracy evaluation method includes: dividing a preset knowledge base into respective knowledge chunks, and generating respective questions corresponding to each of the knowledge chunks through a large language model; inputting each of the questions into the large language model respectively to obtain respective knowledge chunk traceability results corresponding to each of the questions, wherein the knowledge chunk traceability result is the knowledge chunk recalled by the large language model based on the input question; and automatically evaluating the recall accuracy according to the knowledge chunk corresponding to each of the questions and the respective knowledge chunk traceability result corresponding thereto. The technical solution of the present application aims to solve the technical problem of how to reduce the cost required for evaluation when evaluating the recall accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of large model applications, and particularly to a recall accuracy evaluation method, system, device, medium and product. Background Art

[0002] Large language models are trained with massive amounts of data and use billions of parameters to generate raw outputs for tasks such as answering questions, translating languages, and completing sentences. Retrieval-augmented generation technology is used to optimize the output of large language models, enabling them to reference knowledge bases outside of the training data sources before generating answers. Based on the already powerful capabilities of large language models, retrieval-augmented generation technology can expand them to access internal knowledge bases of specific domains or organizations. Specifically, the first step of retrieval-augmented generation technology is to read external data, i.e., new data outside of the large language model's original training dataset. The new data can come from multiple data sources, such as APIs, databases, or document repositories. The second step is to convert the data into a digital representation form through embedding language model technology and store it in a vector database. This process creates a knowledge base that a generative AI model can understand. The third step is to perform a relevance search, convert the user query into a vector representation form, and match it with the vector database. For example, in a retrieval-augmented generation technology application that can answer organizational human resources questions, if an employee searches: "How many annual leave days do I have?", the system will retrieve the annual leave policy document and the employee's personal past leave records. These specific documents will be recalled because they are highly relevant to what the employee entered. Next, the retrieval-augmented generation technology model enhances the user input by adding the retrieved relevant data to the context. Finally, the enhanced content is provided to the large language model together to generate a more relevant answer for the user query.

[0003] However, when injecting data search capabilities into large language model applications through retrieval-augmented generation technology, it is also necessary to measure the accuracy of the relevant recall content retrieved. Nowadays, the accuracy of the recall content is usually evaluated manually. However, manual evaluation requires multiple manual interventions and incurs relatively high costs.

[0004] Therefore, how to reduce the cost required for evaluation when evaluating recall accuracy is a technical problem that those skilled in the art still need to solve. Summary of the Invention

[0005] The main purpose of this application is to provide a recall accuracy evaluation method, system, device, medium and product, aiming to solve the technical problem of how to reduce the cost required for evaluation when evaluating recall accuracy.

[0006] To achieve the above object, this application proposes a recall accuracy evaluation method, and the method includes:

[0007] Divide a preset knowledge base into individual knowledge chunks, and use a large language model to generate questions corresponding to each of the knowledge chunks.

[0008] Input each of the questions into the large language model to obtain a knowledge chunk traceability result corresponding to each of the questions, where the knowledge chunk traceability result is the knowledge chunk recalled by the large language model based on the input question.

[0009] Automatically evaluate the recall accuracy according to the knowledge chunk corresponding to each question and the knowledge chunk traceability result corresponding to each question.

[0010] In one embodiment, the step of automatically evaluating the recall accuracy according to the knowledge chunk corresponding to each question and the knowledge chunk traceability result corresponding to each question includes:

[0011] For each of the questions, compare the knowledge chunk and the knowledge chunk traceability result corresponding to the same question to obtain a knowledge chunk comparison result corresponding to each of the questions.

[0012] Calculate the hit rate and / or the reciprocal of the average rank according to each of the knowledge chunk comparison results, and automatically evaluate the recall accuracy according to the hit rate and / or the reciprocal of the average rank.

[0013] In one embodiment, the step of dividing the preset knowledge base into individual knowledge chunks includes:

[0014] Identify the data types in the preset knowledge base, where the data type is one or more.

[0015] Divide the knowledge base into individual knowledge chunks according to the data types.

[0016] In one embodiment, when it is detected that the number of data types is one, the step of dividing the knowledge base into individual knowledge chunks according to the data types includes:

[0017] When it is detected that the data type is text, divide the knowledge base into individual knowledge chunks according to chapters.

[0018] When it is detected that the data type is a table, use each cell as a knowledge chunk to obtain individual knowledge chunks.

[0019] In one embodiment, when it is detected that the number of data types is multiple, the step of dividing the knowledge base into individual knowledge chunks according to the data types includes:

[0020] Divide the knowledge base into multiple knowledge regions according to each of the data types.

[0021] Obtain the division rules corresponding to each of the data types, and based on each of the division rules, divide each of the knowledge areas into individual knowledge blocks.

[0022] In one embodiment, the step of inputting each of the questions into the large language model to obtain the knowledge block traceability results corresponding to each of the questions includes:

[0023] Convert each of the knowledge blocks into vectors and store each of the vectors in the RAG application;

[0024] Input each of the questions into the large language model, and through the large language model, determine the target vectors corresponding to each of the questions in the vectors stored in the RAG application;

[0025] Use the knowledge blocks corresponding to each of the target vectors as the knowledge block traceability results corresponding to each of the questions, and output each of the knowledge block traceability results.

[0026] In addition, to achieve the above object, the present application also proposes a recall accuracy evaluation system, and the recall accuracy evaluation system includes:

[0027] A question generation module, configured to divide a preset knowledge base into individual knowledge blocks, and through a large language model, generate questions corresponding to each of the knowledge blocks;

[0028] A data recall module, configured to input each of the questions into the large language model to obtain the knowledge block traceability results corresponding to each of the questions, where the knowledge block traceability result is the knowledge block recalled by the large language model based on the input question;

[0029] An automatic evaluation module, configured to automatically evaluate the recall accuracy according to the knowledge blocks corresponding to each of the questions and the knowledge block traceability results corresponding to each of the questions.

[0030] In addition, to achieve the above object, the present application also proposes a recall accuracy evaluation device, and the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the recall accuracy evaluation method as described above.

[0031] In addition, to achieve the above object, the present application also proposes a storage medium, and the storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the recall accuracy evaluation method as described above are implemented.

[0032] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the recall accuracy evaluation method described above.

[0033] In the present application, by dividing a preset knowledge base into individual knowledge chunks and generating respective questions corresponding to each of the knowledge chunks through a large language model, a correspondence relationship between the knowledge chunks and the questions can be established; then, each of the questions is input into the large language model to obtain respective knowledge chunk traceability results corresponding to each of the questions, where the knowledge chunk traceability result is the knowledge chunk recalled by the large language model based on the input question, and a correspondence relationship between the question and the recalled knowledge chunk can be obtained based on the large language model; finally, according to each of the knowledge chunks corresponding to each of the questions and the respective knowledge chunk traceability results, the recall accuracy is automatically evaluated, and the recall accuracy of the large language model can be evaluated based on the correspondence relationship between the question and the knowledge chunk, as well as the correspondence relationship between the question and the recalled knowledge chunk.

[0034] In this way, the present application can automatically evaluate the recall accuracy without manual participation, thereby reducing the cost required for evaluating the recall accuracy. Moreover, the present application also divides the knowledge base into chunks, making the correspondence relationship between the questions and the knowledge chunks clear, and then, by comparing the knowledge chunks with the knowledge chunk traceability results, the obtained recall result must be either accurate or inaccurate, thereby improving the accuracy of the recall accuracy evaluation as a whole. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0036] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0037] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the recall accuracy evaluation method of the present application;

[0038] Figure 2 It is a schematic flowchart provided for a specific embodiment of the recall accuracy evaluation method of the present application;

[0039] Figure 3 It is a schematic flowchart provided for another embodiment of the recall accuracy evaluation method of the present application;

[0040] Figure 4 It is a schematic diagram of the module structure of the recall accuracy evaluation system according to an embodiment of the present application;

[0041] Figure 5 It is a schematic diagram of the device structure of the hardware operating environment involved in the recall accuracy evaluation method according to an embodiment of the present application.

[0042] The implementation, functional characteristics and advantages of the present application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. Specific embodiments

[0043] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0044] For a better understanding of the technical solutions of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific embodiments.

[0045] The main solution of the embodiment of the present application is: dividing a preset knowledge base into each knowledge block, and generating respective corresponding questions for each of the knowledge blocks through a large language model; inputting each of the questions into the large language model respectively to obtain respective corresponding knowledge block traceability results for each of the questions, wherein the knowledge block traceability result is the knowledge block recalled by the large language model based on the input question; automatically evaluating the recall accuracy according to each of the knowledge blocks corresponding to each of the questions and the respective corresponding knowledge block traceability results.

[0046] In the following embodiments, for the convenience of description, the recall accuracy evaluation device is used as the execution subject for elaboration. Specifically, the recall accuracy evaluation device may be a terminal such as a server or a personal computer. The present application does not limit this.

[0047] It is understandable that large language models are trained with massive amounts of data and use billions of parameters to generate raw outputs for tasks such as answering questions, translating languages, and completing sentences. Retrieval-Augmented Generation (RAG) technology is used to optimize the output of large language models, enabling them to reference knowledge bases outside of their training data sources before generating answers. Based on the already powerful capabilities of large language models, RAG technology can expand them to access internal knowledge bases of specific domains or organizations. Specifically, the first step of RAG technology is to read external data, i.e., new data outside of the large language model's original training dataset. The new data can come from multiple data sources, such as APIs, databases, or document repositories. The second step is to convert the data into a digital representation through embedding language model technology and store it in a vector database. This process creates a knowledge base that can be understood by the generative AI model. The third step is to perform a relevance search, convert the user query into a vector representation, and match it with the vector database. For example, in an RAG technology application that can answer organizational human resources questions, if an employee searches: "How many annual leave days do I have?", the system will retrieve the annual leave policy document and the employee's past leave records. These specific documents will be recalled because they are highly relevant to what the employee entered. Next, the RAG technology model enhances the user input by adding the retrieved relevant data in context. Finally, the enhanced content is provided to the large language model together to generate a more relevant answer to the user's query.

[0048] However, when injecting data search capabilities into large language model applications through RAG technology, it is also necessary to measure the accuracy of the relevant retrieved content. Currently, the accuracy of the retrieved content is usually evaluated manually. However, manual evaluation requires multiple human interventions and incurs relatively high costs.

[0049] Therefore, how to reduce the cost required for evaluation when assessing retrieval accuracy is a technical problem that those skilled in the art still need to solve.

[0050] To solve the above problems, this application proposes a method for evaluating retrieval accuracy.

[0051] Please refer to Figure 1 , Figure 1 , which is a schematic flowchart of the first embodiment of the method for evaluating retrieval accuracy of this application.

[0052] In this embodiment, the method for evaluating retrieval accuracy includes steps S10 to S30:

[0053] Step S10, divide a preset knowledge base into each knowledge block, and generate questions corresponding to each of the knowledge blocks through a large language model;

[0054] It should be noted that in this embodiment, the knowledge base is only used to evaluate the recall accuracy. Therefore, the content of the knowledge base is not restricted in this application. It should also be noted that a knowledge block refers to a block obtained after dividing the knowledge base.

[0055] In this embodiment, to establish the correspondence between questions and knowledge blocks, the knowledge base can be divided into multiple blocks, that is, multiple knowledge blocks are obtained. Then, for each knowledge block, one or more questions can be generated. Among them, the questions are used to input into the large language model to evaluate the recall accuracy.

[0056] It can be understood that the technical solution of generating questions for the content in each knowledge block is a mature existing technology, and this application does not limit it.

[0057] Step S20: Input each of the questions into the large language model respectively to obtain the knowledge block traceability results corresponding to each of the questions, where the knowledge block traceability result is the knowledge block recalled by the large language model based on the input question.

[0058] It can be understood that after inputting the question into the large language model, the large language model can infer the knowledge block associated with the input question among the divided knowledge blocks and return the knowledge block. In this embodiment, the knowledge block traceability result is the returned knowledge block. Thus, based on the large language model, the knowledge block traceability results corresponding to each question can be obtained.

[0059] Step S30: Automatically evaluate the recall accuracy according to the knowledge blocks and the knowledge block traceability results corresponding to each of the questions.

[0060] In this embodiment, after obtaining the knowledge block traceability results corresponding to each question, the knowledge block traceability results can be compared with the knowledge blocks when generating the questions, so as to evaluate the recall accuracy.

[0061] In a feasible implementation manner, the above step S20 further includes:

[0062] Step S201: Convert each of the knowledge blocks into vectors and store each of the vectors in the RAG application.

[0063] It should be noted that since the large language model cannot understand the knowledge base or knowledge blocks, after obtaining the knowledge blocks, it is also necessary to convert the knowledge blocks into a data format that the large language model can understand. In this embodiment, the knowledge blocks can be converted into vectors and then the vectors are stored. In addition, the RAG application is an application developed by the RAG technology. In this embodiment, the RAG application is used to store the vectors obtained by converting the knowledge blocks.

[0064] Step S202: Input each of the said problems into the large language model, and through the large language model, determine the target vector corresponding to each of the said problems respectively from the vectors stored in the RAG application;

[0065] It can be understood that the target vector refers to the vector corresponding to the currently input problem among each vector.

[0066] Exemplarily, after the input problem is converted into a vector, it will be compared with the vectors stored in the RAG application, and then the vector with the highest similarity can be used as the target vector.

[0067] In a feasible implementation manner, the number of target vectors can be one or multiple. When the number of target vectors is one, the target vector has the highest similarity with the vector converted from the input problem. When the number of target vectors is multiple, the vectors corresponding to the knowledge chunks can be sorted according to the similarity magnitude, so that for one problem, multiple target vectors can be output.

[0068] Step S203: Take the knowledge chunks corresponding to each of the said target vectors as the knowledge chunk traceability results corresponding to each of the said problems respectively, and output each of the knowledge chunk traceability results.

[0069] In this embodiment, by converting the knowledge chunks into vectors, it is not only convenient for the large language model to understand the problems and knowledge chunks, but also the target vector can be quickly determined based on the similarity, thus reducing the time required to determine the target vector, and further improving the efficiency of evaluating the recall accuracy.

[0070] In a feasible implementation manner, the above-mentioned step S30 further includes:

[0071] Step S301: For each of the said problems, compare the knowledge chunks and the knowledge chunk traceability results corresponding to the same problem to obtain the knowledge chunk comparison results corresponding to each of the said problems respectively;

[0072] It can be understood that when there is one target vector, the knowledge block comparison result is one of "the knowledge block traceability result is the same as the knowledge block that generated the question" and "the knowledge block traceability result is different from the knowledge block that generated the question". When there are multiple target vectors, the knowledge block comparison result is one of "the knowledge block traceability result contains the knowledge block that generated the question" and "the knowledge block traceability result does not contain the knowledge block that generated the question". Among them, when the knowledge block comparison result is "the knowledge block traceability result contains the knowledge block that generated the question", the ranking of each knowledge block in the knowledge block traceability result can also be determined based on the similarity, and then the ranking of the knowledge block that generated the question in the knowledge traceability result can be obtained. It can be understood that when there are multiple target vectors, determining the ranking of the knowledge block that generated the question in the knowledge traceability result can comprehensively evaluate the recall accuracy of the large language model, and can avoid the problem that the evaluated accuracy is relatively one-sided when there is one target vector.

[0073] Step S302, according to each of the knowledge block comparison results, calculate the hit rate and / or the mean reciprocal rank, and automatically evaluate the recall accuracy according to the hit rate and / or the mean reciprocal rank.

[0074] It should be noted that the recall accuracy is generally measured by two indicators: Hit Rate (HR) and Mean Reciprocal Rank (MRR). Among them, the hit rate calculates the proportion of questions in the top k recalled knowledge blocks that contain the correct answer, that is, how many times the top k retrieved items in the Q&A hit the correct answer. The mean reciprocal rank evaluates the performance of the retrieval system through the ranking of the correct retrieval result value in the retrieval result. For example, if the first relevant document is retrieved first, the reciprocal rank is 1; if it is retrieved second, the reciprocal rank is 1 / 2, and so on. Finally, the reciprocal ranks of different Q&As are averaged.

[0075] In this embodiment, by calculating the hit rate and the mean reciprocal rank through the knowledge block comparison result, the recall accuracy can be intuitively presented through data, improving the user experience.

[0076] In an embodiment of the present application, by dividing a preset knowledge base into various knowledge chunks, and using a large language model to generate questions corresponding to each of the knowledge chunks, a correspondence relationship between the knowledge chunks and the questions can be established; then, each of the questions is input into the large language model to obtain a knowledge chunk traceability result corresponding to each of the questions, where the knowledge chunk traceability result is the knowledge chunk recalled by the large language model based on the input question, and a correspondence relationship between the question and the recalled knowledge chunk can be obtained based on the large language model; finally, according to the knowledge chunk corresponding to each of the questions and the knowledge chunk traceability result corresponding to each of the questions, the recall accuracy is automatically evaluated, and the recall accuracy of the large language model can be evaluated based on the correspondence relationship between the question and the knowledge chunk, and between the question and the recalled knowledge chunk.

[0077] In this way, the present application can automatically evaluate the recall accuracy without manual participation, thereby reducing the cost required for evaluating the recall accuracy. Moreover, the present application also makes the correspondence relationship between the question and the knowledge chunk clear by dividing the knowledge base into chunks, and then, by comparing the knowledge chunk with the knowledge chunk traceability result, the obtained recall result must be either accurate or inaccurate, which overall improves the accuracy of calculating the recall accuracy.

[0078] Further, based on the first embodiment of the recall accuracy evaluation method of the present application, a second embodiment of the recall accuracy evaluation method of the present application is proposed.

[0079] In this embodiment, the above step S10 further includes:

[0080] Step S101, identifying the data types in the preset knowledge base, where the data types are one or more;

[0081] It can be understood that after obtaining the knowledge base, in order to evaluate the recall accuracy based on the limited content in the knowledge base, the data types of the content in the knowledge base can be identified for division according to the data types. Specifically, the data types can be text, table, image, etc.

[0082] Step S102, dividing the knowledge base into various knowledge chunks according to the data types.

[0083] It can be understood that different knowledge bases usually store different contents, and a knowledge base usually stores multiple data types. For example, a knowledge base can simultaneously have text, table, and image contents. Therefore, in order to fully utilize the content in the knowledge base, the content belonging to the same data type can be divided into one knowledge chunk, and thus, various knowledge chunks can be obtained.

[0084] Specifically, the above step S102 includes:

[0085] Step S1021: According to each of the data types, divide the knowledge base into multiple knowledge areas;

[0086] It should be noted that, in addition to directly dividing the content of the same data type into one area and then using this area as a knowledge block, it is also possible to, after dividing the content of the knowledge blocks belonging to the same data type into the same area according to the data type, further divide the content in the same area to obtain knowledge blocks.

[0087] Step S1022: Obtain the respective division rules corresponding to each of the data types, and based on each of the division rules, divide each of the knowledge areas into respective knowledge blocks.

[0088] It can be understood that for different data types, different division rules can be preset in advance, so as to fully split the content of different data types.

[0089] In a feasible implementation manner, when it is detected that the number of the data types is one, the above-mentioned step S102 further includes:

[0090] Step S1023: When it is detected that the data type is text, divide the knowledge base into respective knowledge blocks according to chapters;

[0091] Step S1024: When it is detected that the data type is a table, use each cell as a knowledge block to obtain respective knowledge blocks.

[0092] It is worth mentioning that in this application, for a table, the cell division method can be adopted to divide it into multiple knowledge blocks, so as to evaluate the recall accuracy based on the limited content in the table.

[0093] Exemplarily, for a table, the cells can be traversed with code to generate questions such as "what is the number in the nth row and mth column", and then the recall accuracy can be evaluated by assessing whether the correct value is returned.

[0094] In this embodiment, by dividing the text content according to chapters and dividing the table according to cells in this application, the number of generated questions and the amount of test data used to evaluate the recall accuracy can be increased, and thus the recall accuracy can be evaluated more accurately.

[0095] Furthermore, based on the first embodiment and the second embodiment of the recall accuracy evaluation method of this application, two specific embodiments of the recall accuracy evaluation method of this application are proposed.

[0096] Please refer to Figure 2, when the content in the knowledge base is a table, each cell can be obtained by traversal. Then, for each cell, a question such as "what is the number in the nth row and mth column" can be generated. Then, the question can be input into the large language model, and based on the answer of the large language model, the recall accuracy can be evaluated.

[0097] Please refer to Figure 3 , when the content in the knowledge base is text, the text can be chunked, and then the large language model is used to generate questions corresponding to each text chunk. At the same time, the text chunks are converted into vectors and stored in the RAG application for recall. Thus, based on a knowledge base with only text, the evaluation of recall accuracy can be achieved at a reduced cost.

[0098] It should also be noted that when answering questions with the large language model, a data indicating whether the recall result is relevant to the question will be generated. By answering questions multiple times, various indicators for evaluating search recall can be calculated. For a table, the total number of tests is the product of the number of rows and columns; if the test unit is extended from cells to row and column ranges, the model capabilities such as calculating the average value, finding the maximum and minimum values, and data analysis and comparison can also be evaluated. For text, the number of tests can be adjusted by defining different lengths of chunks and the number of questions generated for each chunk. Thus, the present application can be applied to different application scenarios.

[0099] In addition, since the present application can generate test data without introducing additional manual and model links, it is very suitable for the cold start of the early evaluation system of the RAG application. And after being put into the production environment, the RAG application will face a constantly changing knowledge base. The present application also supports flexible sampling in the knowledge base through code and conveniently tracks the parameters used in each evaluation, so as to adjust the evaluation process, thereby further improving the evaluation efficiency.

[0100] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the recall accuracy evaluation method of the present application. Based on this technical concept, more forms of simple transformation are within the protection scope of the present application.

[0101] The present application also provides a recall accuracy evaluation system. Please refer to Figure 4 , the recall accuracy evaluation system includes:

[0102] A question generation module 10, configured to divide a preset knowledge base into respective knowledge chunks, and generate questions corresponding to each of the knowledge chunks through a large language model;

[0103] A data recall module 20, configured to input each of the questions into the large language model respectively to obtain a knowledge chunk traceability result corresponding to each of the questions, where the knowledge chunk traceability result is the knowledge chunk recalled by the large language model based on the input question;

[0104] An automatic evaluation module 30 for automatically evaluating the recall accuracy according to the knowledge blocks corresponding to each of the questions and the knowledge block traceability results corresponding to each of them.

[0105] In one embodiment, the automatic evaluation module 30 is further configured to:

[0106] For each of the questions, compare the knowledge blocks corresponding to the same question with the knowledge block traceability results to obtain the knowledge block comparison results corresponding to each of the questions;

[0107] Calculate the hit rate and / or the reciprocal of the average rank according to the knowledge block comparison results, and automatically evaluate the recall accuracy according to the hit rate and / or the reciprocal of the average rank.

[0108] In one embodiment, the question generation module 10 is further configured to:

[0109] Identify the data types in a preset knowledge base, where the data types are one or more;

[0110] Divide the knowledge base into respective knowledge blocks according to the data types.

[0111] In one embodiment, when it is detected that the number of the data types is one, the question generation module 10 is further configured to:

[0112] When it is detected that the data type is text, divide the knowledge base into respective knowledge blocks according to chapters;

[0113] When it is detected that the data type is a table, regard each cell as a knowledge block to obtain respective knowledge blocks.

[0114] In one embodiment, when it is detected that the number of the data types is multiple, the question generation module 10 is further configured to:

[0115] Divide the knowledge base into multiple knowledge areas according to the data types;

[0116] Obtain the division rules corresponding to the data types respectively, and divide each of the knowledge areas into respective knowledge blocks based on the division rules.

[0117] In one embodiment, the data recall module 20 is further configured to:

[0118] Convert each of the knowledge blocks into a vector respectively, and store each of the vectors in a RAG application;

[0119] Input each of the above problems into the large language model, and through the large language model, determine the target vector corresponding to each of the problems in the vectors stored in the RAG application;

[0120] Use the knowledge block corresponding to each of the target vectors as the knowledge block traceability result corresponding to each of the problems, and output each of the knowledge block traceability results.

[0121] The recall accuracy evaluation system provided by this application adopts the recall accuracy evaluation method in the above embodiment, and can solve the technical problem of how to reduce the cost required for evaluation when evaluating recall accuracy. Compared with the prior art, the beneficial effects of the recall accuracy evaluation system provided by this application are the same as those of the recall accuracy evaluation method provided by the above embodiment, and other technical features in the recall accuracy evaluation system are the same as the features disclosed in the method of the above embodiment, and will not be elaborated here.

[0122] This application provides a recall accuracy evaluation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the recall accuracy evaluation method in the first embodiment above.

[0123] Next, refer to Figure 5 , which shows a schematic structural diagram of a recall accuracy evaluation device suitable for implementing the embodiments of this application. Figure 5 The shown recall accuracy evaluation device is only an example, and should not bring any limitation to the functions and usage scope of the embodiments of this application.

[0124] As Figure 5As shown, the recall accuracy evaluation device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. In the RAM 1004, various programs and data required for the operation of the recall accuracy evaluation device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the recall accuracy evaluation device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a recall accuracy evaluation device having various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be implemented or had alternatively.

[0125] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.

[0126] The recall accuracy evaluation device provided by the present application adopts the recall accuracy evaluation method in the above-mentioned embodiment, and can solve the technical problem of how to reduce the cost required for evaluation when evaluating recall accuracy. Compared with the prior art, the beneficial effects of the recall accuracy evaluation device provided by the present application are the same as those of the recall accuracy evaluation method provided by the above-mentioned embodiment, and other technical features in the recall accuracy evaluation device are the same as the features disclosed in the method of the previous embodiment, and will not be elaborated here.

[0127] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0128] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0129] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the recall accuracy evaluation method in the above embodiments.

[0130] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0131] The above computer-readable storage medium can be included in the recall accuracy evaluation device; it can also exist separately and not be assembled into the recall accuracy evaluation device.

[0132] The above computer-readable storage medium carries one or more programs, which, when executed by the recall accuracy evaluation device, cause the recall accuracy evaluation device to: divide a preset knowledge base into respective knowledge chunks, and generate questions corresponding to each of the knowledge chunks through a large language model; input each of the questions into the large language model to obtain a knowledge chunk traceability result corresponding to each of the questions, where the knowledge chunk traceability result is the knowledge chunk recalled by the large language model based on the input question; and automatically evaluate the recall accuracy according to the knowledge chunk corresponding to each of the questions and the knowledge chunk traceability result corresponding to each of the questions.

[0133] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0134] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0135] The modules involved in the embodiments of the present application can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.

[0136] The readable storage medium provided by the present application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned recall accuracy evaluation method, and can solve the technical problem of how to reduce the cost required for evaluation when evaluating recall accuracy. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the recall accuracy evaluation method provided by the above embodiments, and will not be elaborated here.

[0137] The present application also provides a computer program product, including a computer program, and the steps of the recall accuracy evaluation method as described above are implemented when the computer program is executed by a processor.

[0138] The computer program product provided by the present application can solve the technical problem of how to reduce the cost required for evaluation when evaluating recall accuracy. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the recall accuracy evaluation method provided by the above embodiments, and will not be elaborated here.

[0139] The above are only some embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or any direct / indirect application in other related technical fields is included in the patent protection scope of the present application.

Claims

1. A recall accuracy evaluation method, characterized in that: The recall accuracy evaluation method includes: Divide the preset knowledge base into knowledge blocks, and generate questions corresponding to each of the knowledge blocks through a large language model, wherein the knowledge base is divided into knowledge blocks in a manner including dividing according to data types, the data types including text and table; for a text type knowledge base, dividing according to chapters; for a table type knowledge base, taking each cell as a knowledge block; Convert each of the knowledge blocks into a vector, and store each of the vectors in the RAG application; Inputting each of the questions into the large language model respectively, and determining the target vector corresponding to each of the questions from the vectors stored in the RAG application through the large language model; Taking the knowledge blocks corresponding to the target vectors as the knowledge block tracing results corresponding to the questions, and outputting the knowledge block tracing results, wherein the knowledge block tracing results are the knowledge blocks recalled by the large language model based on the input questions; For each of the above-mentioned questions, the knowledge block and the knowledge block tracing result corresponding to the same question are compared to obtain the knowledge block comparison result corresponding to each of the above-mentioned questions. When each of the above-mentioned questions corresponds to one of the above-mentioned target vectors, the knowledge block comparison result is one of "the knowledge block tracing result is the same as the knowledge block" and "the knowledge block tracing result is different from the knowledge block"; when each of the above-mentioned questions corresponds to multiple target vectors, the knowledge block comparison result is one of "the knowledge block tracing result contains the knowledge block" and "the knowledge block tracing result does not contain the knowledge block"; According to the comparison results of each of the knowledge blocks, the hit rate and / or the average reciprocal ranking are calculated, and the recall accuracy is automatically evaluated according to the hit rate and / or the average reciprocal ranking.

2. The recall accuracy evaluation method according to claim 1, characterized in that: The step of dividing the preset knowledge base into various knowledge blocks includes: Identify a data type in a preset knowledge base, wherein the data type is one or more; According to the data type, the knowledge base is divided into various knowledge blocks.

3. The recall accuracy evaluation method according to claim 2, characterized in that: When it is detected that the number of the data types is multiple, the step of dividing the knowledge base into knowledge blocks according to the data types includes: According to each of the data types, the knowledge base is divided into a plurality of knowledge areas; The division rules corresponding to each of the data types are obtained, and based on each of the division rules, each of the knowledge areas is divided into various knowledge blocks.

4. A recall accuracy evaluation system, characterized in that: The recall accuracy evaluation system comprises: A question generation module is used to divide the preset knowledge base into various knowledge blocks, and generate questions corresponding to each of the knowledge blocks through a large language model, wherein the knowledge base is divided into knowledge blocks in a manner including dividing according to data types, the data types including text and table; for a text type knowledge base, dividing according to chapters; for a table type knowledge base, taking each cell as a knowledge block; A data recall module is used to convert each of the knowledge blocks into a vector, and store each of the vectors in the RAG application; input each of the questions into the large language model, and determine the target vector corresponding to each of the questions in the vectors stored in the RAG application through the large language model; use the knowledge block corresponding to each of the target vectors as the knowledge block tracing result corresponding to each of the questions, and output the knowledge block tracing result, wherein the knowledge block tracing result is the knowledge block recalled by the large language model based on the input question; An automatic evaluation module is used to compare the knowledge blocks and knowledge block tracing results corresponding to the same question for each of the above-mentioned questions, and obtain the knowledge block comparison results corresponding to each of the above-mentioned questions. When each question corresponds to one of the above-mentioned target vectors, the knowledge block comparison result is one of "the knowledge block tracing result is the same as the knowledge block" and "the knowledge block tracing result is different from the knowledge block"; when each question corresponds to multiple target vectors, the knowledge block comparison result is one of "the knowledge block tracing result contains the knowledge block" and "the knowledge block tracing result does not contain the knowledge block"; based on the comparison results of each of the above-mentioned knowledge blocks, the hit rate and / or the average ranking inverse are calculated, and the recall accuracy is automatically evaluated based on the hit rate and / or the average ranking inverse.

5. A recall accuracy assessment device, characterized in that: The recall accuracy assessment device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the recall accuracy assessment method according to any one of claims 1 to 3.

6. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the recall accuracy evaluation method according to any one of claims 1 to 3 are implemented.

7. A computer program product, characterized in that The computer program product comprises a computer program, which implements the steps of the recall accuracy assessment method according to any one of claims 1 to 3 when executed by a processor.

Citation Information

Patent Citations

  • Large language model question and answer optimization method and device, electronic equipment and storage medium

    CN117851575A