Quality detection method for search enhancement generation system, electronic device, and program product
Patent Information
- Application Number
- CN202610903430.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-09-25
AI Technical Summary
但是上述方法存在颗粒度较粗以及方法单一的问题,从而导致质量检测的精度较差
Smart Images

Figure CN122817367A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to a quality inspection method, electronic device, and program product for a search enhancement generation system. Background Technology
[0002] Currently, when using Retrieval-Augmented Generation (RAG) systems to answer user questions, it's typically necessary to understand the user's question to analyze their intent, retrieve relevant data from the database based on that intent, and then generate an answer based on the relevant data and the user's question. To ensure the accuracy of the answer, related technologies perform quality checks before providing feedback to the user. For example, these methods use test sets, manual sampling, and relevance algorithms to perform quality checks on the RAG system. However, these methods suffer from coarse granularity and a lack of diversity, resulting in poor accuracy in quality checks. Furthermore, these methods cannot monitor answer quality in real time, exhibiting significant lag and greatly degrading the user experience. Summary of the Invention
[0003] This disclosure provides a quality inspection method, electronic device, and program product for a retrieval enhancement generation system.
[0004] According to one aspect of this disclosure, a quality detection method for a retrieval enhancement generation system is provided, comprising: The process involves obtaining the answer to be tested for quality. This answer is generated by the retrieval enhancement generation system based on user search intent data and multiple related data sets. Search intent data is generated based on user question data. Related data is data associated with the user search intent data. The relevance between user question data and user search intent data is analyzed to obtain relevant results. These relevant results indicate the accuracy of the user search intent data. The correlation between each related data set and the user question data is analyzed to obtain correlation results. These correlation results indicate the accuracy of the correlation data. The matching between the answer to be tested for quality and the correlation data is analyzed to obtain matching results. These matching results indicate the authenticity of the answer to be tested for quality. If the relevant results, correlation results, and matching results all meet the detection conditions, the answer to be tested for quality passes the detection.
[0005] According to at least one embodiment of this disclosure, the correlation between user question data and user search intent data is analyzed to obtain relevant results, including: Using the first-largest language model of the retrieval enhancement generation system, the topic of user query data is obtained; the relevance between the topic of user query data and user search intent data is analyzed to obtain topic-related results; using the first-largest language model, key information of user query data is obtained; the coverage of key information by user search intent data is analyzed to obtain coverage results; using the first-largest language model, redundant information of user query data is obtained; the distribution of key information and redundant information in user search intent data is analyzed to obtain redundancy results; using the first-largest language model, the mean of topic-related results, coverage results, and redundant results is calculated to obtain relevance results.
[0006] According to at least one embodiment of this disclosure, the correlation between the topic of user question data and user search intent data is analyzed to obtain topic-related results, including: The topic of user query data is converted into topic vector, and the user search intent data is converted into user search intent vector; similarity is calculated on the topic vector and the user search intent vector to obtain topic relevance results.
[0007] According to at least one embodiment of this disclosure, the distribution of key information and redundant information in user search intent data is analyzed to obtain redundancy results, including: Obtain the redundancy quantity of redundant information; calculate the sum of the redundancy quantity and the critical quantity; divide the sum of the redundancy quantity and the critical quantity by the redundancy quantity to obtain the redundancy result.
[0008] According to at least one embodiment of this disclosure, analyzing the correlation between each piece of related data and user question data to obtain a correlation result includes: using a first large language model to obtain multiple inference question data corresponding to each piece of related data; for each piece of related data, using the first large language model to calculate multiple correlation values corresponding to the related data; the correlation value is obtained by performing correlation calculation on the inference question data and user question data; the correlation value corresponds one-to-one with the inference question data; and using the first large language model to calculate the average of the multiple correlation values corresponding to the related data to obtain the correlation result.
[0009] According to at least one embodiment of this disclosure, a first large language model is used to calculate multiple association values corresponding to the association data, including: converting each inference question data corresponding to the association data into an inference vector, and converting the user question data into a user question data vector; performing association calculation on each inference vector and the user question data vector to obtain the association value corresponding to each inference vector.
[0010] According to at least one embodiment of this disclosure, analyzing the matching between the answer to be quality tested and the associated data to obtain a matching result includes: using a first large language model to obtain multiple factual data in the answer to be quality tested; using the first large language model to compare the multiple factual data with the associated data to obtain factual data that matches the associated data; and using the first large language model to perform a division calculation on the number of factual data and the number of factual data that matches the associated data to obtain a matching result.
[0011] According to at least one embodiment of this disclosure, if any one of the relevant results, association results, and matching results fails to meet the detection conditions, the answer to be tested fails the detection. When the relevant results fail to meet the detection conditions, the second language model of the retrieval enhancement generation system is used to generate updated user search intent data based on the user question data. The second language model is used to search the database based on the updated user search intent data to obtain updated association data. The generative model of the retrieval enhancement generation system is used to generate an updated answer based on the updated association data and user question data. Alternatively, when the relevant results meet the detection conditions, but the association results do not meet the detection conditions, the second language model is used to generate updated association data based on the user search intent data. The generative model is used to generate an updated answer based on the updated association data and user question data. Alternatively, when both the relevant results and association results meet the detection conditions, but the matching results do not meet the detection conditions, the generative model is used to generate an updated answer based on the association data and user question data.
[0012] According to at least one embodiment of this disclosure, if the relevant result, the association result, and the matching result all meet the detection conditions, then the answer to be tested for quality control passes the detection, including: using a first large language model to calculate the relevant level corresponding to the relevant result, the association level corresponding to the association result, and the matching level corresponding to the matching result; if the relevant level, the association level, and the matching level are all not lower than the level threshold, then the answer to be tested for quality control passes the detection; wherein, a relevant level not lower than the level threshold indicates that the user's search intent data is accurate; an association level not lower than the level threshold indicates that the association data is accurate; and a matching level not lower than the level threshold indicates that the answer to be tested for quality control is authentic.
[0013] According to at least one embodiment of this disclosure, if any one of the relevance level, association level, and matching level is lower than a level threshold, the answer to be tested fails the quality inspection; wherein, a relevance level lower than the level threshold indicates that the user's search intent data is inaccurate; an association level lower than the level threshold indicates that the association data is inaccurate; and a matching level lower than the level threshold indicates that the answer to be tested is not genuine.
[0014] According to another aspect of this disclosure, an electronic device is provided, comprising: a memory storing execution instructions; and a processor executing the execution instructions stored in the memory, causing the processor to perform a quality detection method of a retrieval enhancement generation system according to any embodiment of this disclosure.
[0015] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements a quality detection method for a retrieval enhancement generation system according to any embodiment of this disclosure. Attached Figure Description
[0016] The accompanying drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, serve to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification.
[0017] Figure 1 This is a schematic diagram of the overall process of a quality detection method for a retrieval enhancement generation system according to one embodiment of the present disclosure.
[0018] Figure 2 This is a flowchart illustrating step S110 provided in an embodiment of this disclosure.
[0019] Figure 3 This is a flowchart illustrating step S120 provided in an embodiment of this disclosure.
[0020] Figure 4 This is a flowchart illustrating step S1201 provided in an embodiment of this disclosure.
[0021] Figure 5 This is a flowchart illustrating step S1202 provided in an embodiment of this disclosure.
[0022] Figure 6 This is a flowchart illustrating step S1203 provided in an embodiment of this disclosure.
[0023] Figure 7 This is a flowchart illustrating step S130 provided in an embodiment of this disclosure.
[0024] Figure 8 This is a flowchart illustrating step S1302 provided in an embodiment of this disclosure.
[0025] Figure 9 This is a flowchart illustrating step S140 provided in an embodiment of this disclosure.
[0026] Figure 10 This is a flowchart illustrating step S150 provided in an embodiment of this disclosure.
[0027] Figure 11This is a schematic flowchart of a method for correcting a failed quality inspection result provided in an embodiment of this disclosure.
[0028] Figure 12 This is a schematic structural block diagram of a quality inspection device for a search enhancement generation system according to one embodiment of the present disclosure.
[0029] Figure 13 This is a schematic structural block diagram of an electronic device equipped with a quality inspection device for a retrieval enhancement generation system, according to one embodiment of the present disclosure. Detailed Implementation
[0030] The present disclosure will now be described in further detail with reference to the accompanying drawings and examples. It should be understood that the specific examples described herein are for illustrative purposes only and are not intended to limit the scope of the disclosure. Furthermore, it should be noted that, for ease of description, only the parts relevant to the present disclosure are shown in the accompanying drawings.
[0031] It should be noted that, where there is no conflict, the embodiments and features described in this disclosure can be combined with each other. The technical solutions of this disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0032] In related technologies, retrieval enhancement generation systems generate corresponding answers to user questions by performing multiple steps. However, because the answer generation process involves multiple steps, and the more steps there are, the greater the risk of error propagation, making quality control of the retrieval enhancement generation system particularly important. For example, an error in the first step can cause all subsequent steps to output irrelevant answers.
[0033] In some examples, related technologies use test sets, manual sampling, and relevance algorithms to perform quality checks on the retrieval enhancement generation system. Let's take using a test set to perform quality checks on the retrieval enhancement generation system as an example. Specifically, multiple sets of user question data and correct answers are set up; the user question data is input into the retrieval enhancement generation system, which outputs the actual answer; the actual answer and the correct answer are compared to determine if the retrieval enhancement generation system is abnormal. However, this approach has the following problems: First, even if the actual answer and the correct answer are inconsistent, it can only roughly determine that the retrieval enhancement generation system is abnormal, but it cannot accurately determine which step in the above-mentioned process went wrong, that is, it cannot determine the specific reason why the retrieval enhancement generation system generated an incorrect answer, thus making it impossible to optimize the retrieval enhancement generation system in a targeted manner. Second, the above method performs quality checks on the retrieval enhancement generation system periodically, rather than in real time. Therefore, the above method has a lag; incorrect answers are still fed back to the user, and incorrect answers cannot be corrected in real time, greatly reducing the user experience. Third, in the above method, the retrieval enhancement generation system cannot perform quality checks and corrections on its own, resulting in poor intelligence and reliability.
[0034] To address this, this disclosure proposes the following technical solution: In this solution, the retrieval enhancement generation system automatically performs quality checks on the output of each step involved, obtaining corresponding quality check results. Multiple quality check results are then combined to determine whether the answer to be tested passes the quality check. If the answer fails the quality check, the multiple quality check results are combined to accurately determine which step caused the failure. The retrieval enhancement generation system then re-executes the problematic step to generate an updated answer, thus achieving real-time correction of answers that fail the quality check. This not only improves the reliability, intelligence, and fault tolerance of the retrieval enhancement generation system but also ensures a better user experience.
[0035] To facilitate description and make the technical solutions of this disclosure easier to understand, the terminology of this disclosure will be explained before describing the technical solutions of this disclosure.
[0036] 1. Retrieval Enhancement Generation System: First, retrieve factual data and document fragments related to the user's question from the database. Then, use the retrieved content as constraints and basis to input into the generation model, and generate accurate answers or text based on the retrieval results.
[0037] 2. Manual sampling inspection: For the large number of answers output by the search enhancement generation system, a certain proportion of samples are randomly selected by humans for inspection and verification, rather than inspecting all objects one by one, in order to determine whether the overall results are qualified and compliant.
[0038] 3. Relevance Algorithm: This refers to comparing and calculating two or more sets of text, data, and feature information to quantitatively determine the degree of content association, similarity, or matching.
[0039] 4. Cosine similarity: A mathematical metric used to measure the degree of similarity between two vectors in a direction. It characterizes the similarity between the two vectors by calculating the cosine of the angle between them.
[0040] Figure 1 A schematic flowchart illustrating the overall process of a quality inspection method for a retrieval enhancement generation system according to one embodiment of this disclosure is shown. Figure 1 The method shown includes steps S110 to S150. This method can be executed by electronic devices such as mobile phones and tablets.
[0041] In step S110, the answer to be tested for quality is obtained.
[0042] Among them, the answer to be tested for quality is generated by the retrieval enhancement generation system based on the user's search intent data and multiple related data; the search intent data is generated based on the user's question data; and the related data is data associated with the user's search intent data.
[0043] In some examples, the retrieval enhancement generation system obtains user question data and sequentially executes query analysis, knowledge retrieval, and enhancement generation steps based on the user question data to obtain the corresponding answer. After obtaining the answer, it is not directly fed back to the user, but the answer is subjected to quality inspection, that is, steps S120 to S150 are executed.
[0044] In step S120, the correlation between user question data and user search intent data is analyzed to obtain relevant results.
[0045] Among them, the relevant results characterize whether the user search intent data is accurate.
[0046] Step S120 is used to perform quality checks on the output of the query analysis step. Since the query analysis step outputs user search intent data, it is necessary to analyze this data to detect any problems with the query analysis step. Based on the analysis of step S110, it is known that the user search intent data is obtained from user query data. Therefore, this disclosure determines whether the query analysis step is abnormal by analyzing the correlation between the user query data and the user search intent data.
[0047] Specifically, if the correlation results show a high correlation between the user's question data and the user's search intent data, it indicates that the user's search intent data is accurate, and the query analysis steps are normal. If the correlation results show a weak correlation between the user's question data and the user's search intent data, it indicates that the user's search intent data is inaccurate, and the query analysis steps are abnormal.
[0048] In step S130, the correlation between each piece of related data and the user's question data is analyzed to obtain the correlation results.
[0049] Among them, the association result characterizes whether the associated data is accurate.
[0050] Step S130 is used to perform quality checks on the output of the knowledge retrieval step. Since the knowledge retrieval step outputs related data, it is necessary to analyze this related data to detect any problems with the knowledge retrieval step. Based on the analysis of step S110, it is known that the related data is associated with the user's search intent data, which is obtained based on the user's question data. Therefore, by analyzing the correlation between the user's question data and the related data, it is determined whether the knowledge retrieval step is abnormal.
[0051] Specifically, if the association results show a high correlation between the associated data and the user's question data, it indicates that the associated data is accurate, and the knowledge retrieval process proceeds normally. If the association results show a weak correlation between the associated data and the user's question data, it indicates that the associated data is inaccurate, and the knowledge retrieval process is abnormal.
[0052] In step S140, the matching between the answer to be tested and the associated data is analyzed to obtain the matching result.
[0053] The matching result indicates whether the answer to be tested is true.
[0054] Step S140 is used to perform quality checks on the output of the enhancement generation step. Based on the analysis of step S110, it is known that the enhancement generation step generates the answer to be tested for quality based on associated data and user question data. In other words, the answer to be tested for quality comes from the associated data, is a part of the associated data, or the associated data should cover the entire content of the answer to be tested for quality. Therefore, it is necessary to analyze the matching between the answer to be tested for quality and the associated data to determine whether the enhancement generation step is abnormal.
[0055] In this scenario, if the matching results show that the associated data completely covers all the content of the answer to be tested, it indicates that there is no unfounded content in the answer, meaning the answer and associated data are a perfect match and the answer is authentic, then the enhancement generation step proceeds normally. If the matching results show that the associated data does not completely cover all the content of the answer to be tested, it indicates that there is some unfounded content in the answer, which can also be understood as the retrieval enhancement generation system arbitrarily generating some data without basis. In other words, the answer and associated data do not match, the answer is not authentic, and the enhancement generation step is abnormal.
[0056] In step S150, if the relevant results, association results, and matching results all meet the detection conditions, then the answer to be tested for quality control passes the test.
[0057] Based on the analysis of steps S120 to S140 above, the following statements indicate that the relevant results meet the detection conditions: the user's question data and the user's search intent data are highly correlated, the user's search intent data is accurate, and therefore the query analysis step is normal. The association results meet the detection conditions: the association data and the user's question data are highly correlated, the association data is accurate, and therefore the knowledge retrieval step is normal. The matching results meet the detection conditions: the association data completely covers all the content of the answer to be tested for quality, the answer to be tested does not contain any unfounded content, that is, the answer to be tested is true and does not deviate from the facts, and therefore the enhancement generation step is normal. In other words, all steps performed by the retrieval enhancement generation system in generating the answer to be tested for quality are normal, and therefore the answer to be tested passes the test. This means the answer to be tested is correct.
[0058] It should be noted that the execution order of steps S120 to S140 described above is only one implementation method. In other implementation methods, steps S120 to S140 can also be performed in parallel.
[0059] After performing step S150, the retrieved enhanced generation system outputs the answer that passed the test.
[0060] It is understood that the first language model, the second language model, and the generative model involved in the embodiments of this disclosure are all pre-trained models. The first language model, the second language model, and the generative model are used to process different data according to different training data in order to achieve different functions.
[0061] This embodiment provides a quality detection method for a retrieval enhancement generation system. It determines the accuracy of user search intent data by analyzing the correlation between user question data and user search intent data; it determines the accuracy of related data by analyzing the correlation between related data and user question data; it determines the authenticity of the answer to be tested by analyzing the matching between the answer and related data; and it determines whether the answer to be tested passes the quality detection by combining the above correlation results, association results, and matching results. This technical solution achieves precise dynamic quality detection at each step involved in answer generation, thus not only enabling real-time monitoring of answer quality but also ensuring quality detection accuracy from multiple dimensions, greatly improving the user experience.
[0062] In one or more embodiments of this disclosure, Figure 2 This is a flowchart illustrating step S110 provided in an embodiment of this disclosure. Figure 2 As shown, step S110 may include steps S1101 to S1104.
[0063] In step S1101, the retrieval enhancement system obtains user question data.
[0064] In some examples, users can interact with the search enhancement generation system in multiple ways. For instance, users can input questions into the system via voice or natural language interaction. For example, a user might ask, "What renovation packages does the company offer?"
[0065] In step S1102, the second language model of the retrieval enhancement generation system is used to analyze the user's question data to obtain the user's search intent data.
[0066] In some examples, because user query data may contain redundant or incomplete information, to address this issue, embodiments of this disclosure use a second language model to accurately analyze the intent corresponding to the user query data, obtaining user search intent data. This allows for more precise database searches using the user search intent data. For example, the user query data "What renovation packages does the company offer?" corresponds to the user search intent data "Company renovation packages in 2025".
[0067] In step S1103, the user search intent data is converted into a user search intent vector using the second language model; the database is searched based on the user search intent vector to obtain at least one related data.
[0068] The database contains multiple vectors. Each vector corresponds one-to-one with a document related to the user. It is understood that the associated data in this embodiment includes at least one vector associated with the user's search intent vector.
[0069] In step S1104, the generative model of the retrieval enhancement generation system is used to generate answers to be quality tested based on the associated data and user question data.
[0070] In one or more embodiments of this disclosure, Figure 3 This is a flowchart illustrating step S120 provided in an embodiment of this disclosure. Figure 3 As shown, step S120 may include steps S1201 to S1204.
[0071] In step S1201, the first language model of the retrieval enhancement generation system is used to obtain the topic of the user question data; the correlation between the topic of the user question data and the user search intent data is analyzed to obtain topic-related results.
[0072] In some examples, in order to detect whether user search intent data accurately understands user question data, embodiments of this disclosure perform quality checks on user search intent data from three dimensions: relevance, completeness, and conciseness, thereby improving the accuracy of quality checks.
[0073] For example, step S1201 performs quality inspection on the user search intent data from the dimension of relevance. Specifically, the first language model extracts the topic of the user question data based on prompt words; the topic is converted into a topic vector, and the cosine similarity between the topic vector and the user search intent data vector is calculated.
[0074] In step S1202, the first language model is used to obtain key information from the user's question data; the coverage of the key information by the user's search intent data is analyzed to obtain the coverage result.
[0075] For example, step S1202 performs quality checks on the user search intent data from the perspective of completeness. Specifically, the first language model extracts key information from the user's question data based on prompt words. This key information includes at least one keyword. For example, the keywords are "2025" and "renovation package".
[0076] In step S1203, the first language model is used to obtain redundant information from the user's question data; the key information and the distribution of redundant information in the user's search intent data are analyzed to obtain the redundancy results.
[0077] For example, step S1203 performs quality inspection on the user search intent data from the perspective of conciseness. Specifically, the first language model extracts redundant information from the user question data based on prompt words. The redundant information includes at least one redundant word. The reason this embodiment performs quality inspection on the user search intent data from the perspective of conciseness is that the user question data may contain some information unrelated to the intent. If the user search intent data also includes redundant information, it will result in the user search intent data being less concise, reducing the accuracy of the user search intent data, and thus affecting the accuracy of subsequent steps. For example, redundant information may include modal particles, etc.
[0078] In step S1204, the first large language model is used to calculate the mean of the topic-related results, coverage results, and redundancy results to obtain the relevant results.
[0079] For example, the relevant results, coverage results, and redundancy results are all percentage scores.
[0080] In some examples, without considering the context, the first language model is used to directly calculate the mean of topic-related results, coverage results, and redundant results to obtain the relevant results.
[0081] In some examples, considering different scenarios, the importance of topic-related results, coverage results, and redundant results varies across different scenarios. Therefore, embodiments of this disclosure assign different weights to topic-related results, coverage results, and redundant results in different scenarios, thereby enabling the relevant results to more accurately characterize whether the user's search intent data is correct.
[0082] For example, the first large language model is used to calculate the product of topic-related results and the weights of topic-related results, the product of coverage results and the weights of coverage results, and the product of redundancy results and the weights of redundancy results. The average of the product of topic-related results and the weights of topic-related results, the product of coverage results and the weights of coverage results, and the product of redundancy results and the weights of redundancy results is calculated to obtain the relevant results.
[0083] It should be noted that the prompt words involved in the embodiments of this disclosure are pre-set, and the embodiments of this application do not specifically limit the prompt words.
[0084] Based on the above analysis of steps S1201 to S1204, it can be seen that this disclosure performs quality detection on user search intent data from three dimensions: relevance, completeness, and simplicity. This not only improves the accuracy of judging the correctness of user search intent data, but also improves the accuracy of judging anomalies in the query analysis steps.
[0085] In one or more embodiments of this disclosure, Figure 4 This is a flowchart illustrating step S1201 provided in an embodiment of this disclosure. Figure 4 As shown, step S1201 may include steps S12011 to S12012.
[0086] In step S12011, the topic of the user's question data is converted into a topic vector, and the user's search intent data is converted into a user search intent vector.
[0087] In step S12012, similarity calculation is performed on the topic vector and the user search intent vector to obtain topic relevance results.
[0088] For example, the cosine similarity between the topic vector and the user search intent vector is calculated, and this cosine similarity is the topic relevance result.
[0089] In one or more embodiments of this disclosure, Figure 5 This is a flowchart illustrating step S1202 provided in an embodiment of this disclosure. Figure 5 As shown, step S1202 may include steps S12021 to S12022.
[0090] In step S12021, the number of key information items and the coverage of key information items included in the user search intent data are obtained.
[0091] In step S12022, the coverage quantity and the key quantity are calculated by division to obtain the coverage result.
[0092] For example, coverage result = number of coverage / number of keys × 100%.
[0093] In one or more embodiments of this disclosure, Figure 6 This is a flowchart illustrating step S1203 provided in an embodiment of this disclosure. Figure 6 As shown, step S1203 may include steps S12031 to S12033.
[0094] In step S12031, the redundancy quantity of the redundancy information is obtained.
[0095] In step S12032, the sum of the redundancy quantity and the critical quantity is calculated.
[0096] In step S12033, the redundancy result is obtained by dividing the sum of the redundancy quantity and the critical quantity by the redundancy quantity.
[0097] For example, the redundancy result = number of redundancies / (number of redundancies + number of criticalities) × 100%.
[0098] In one or more embodiments of this disclosure, Figure 7 This is a flowchart illustrating step S130 provided in an embodiment of this disclosure. Figure 7 As shown, step S130 may include steps S1301 to S1303.
[0099] In step S1301, the first large language model is used to obtain multiple reasoning question data corresponding to each associated data.
[0100] In some examples, for each piece of related data, the first language model uses the prompt word and the related data to infer all possible inference questions corresponding to that related data.
[0101] In step S1302, for each piece of associated data, multiple associated values corresponding to the associated data are calculated using the first large language model.
[0102] The correlation value is calculated by performing correlation analysis on the reasoning question data and the user question data. Each correlation value corresponds one-to-one with the reasoning question data.
[0103] In step S1303, the first large language model is used to calculate the average of multiple associated values corresponding to the associated data to obtain the association result.
[0104] The larger the value of the association result, the stronger the correlation and the better the match between the associated data and the user's question data. The smaller the value of the association result, the weaker the correlation and the worse the match between the associated data and the user's question data.
[0105] In one or more embodiments of this disclosure, Figure 8 This is a flowchart illustrating step S1302 provided in an embodiment of this disclosure. Figure 8 As shown, step S1302 may include steps S13021 to S13022.
[0106] In step S13021, each inference question data corresponding to the associated data is converted into an inference vector, and the user question data is converted into a user question data vector.
[0107] In step S13022, the correlation is calculated for each inference vector and the user question data vector to obtain the correlation value corresponding to each inference vector.
[0108] For example, the cosine similarity between each inference vector and the user question data vector is calculated, and this cosine similarity is the association value.
[0109] The above-disclosed technical solution quantifies the correlation between related data and user query data through reverse reasoning, thereby determining whether the related data is correct.
[0110] In one or more embodiments of this disclosure, Figure 9 This is a flowchart illustrating step S140 provided in an embodiment of this disclosure. Figure 9 As shown, step S140 may include steps S1401 to S1403.
[0111] In step S1401, the first large language model is used to obtain multiple factual data from the answer to be quality checked.
[0112] In some examples, the first language model breaks down the answer to be quality checked into multiple factual data based on the prompt words. For instance, the factual data includes categorical factual data and price factual data. For example, the categorical factual data includes renovation packages A, B, C, and D. The price factual data are the prices of renovation packages A, B, C, and D, respectively. That is, this example includes five factual data points.
[0113] In step S1402, the first large language model is used to compare multiple factual data with related data to obtain factual data that matches the related data.
[0114] Determining whether correlated data matches factual data can be understood as determining whether the correlated data includes factual data, or whether the factual data is part of the correlated data. Specifically, for each piece of factual data, if it is part of the correlated data, it means that the factual data is real and not generated out of thin air by the first language model. If all or part of the factual data is not part of the correlated data, it means that all or part of the factual data is not real, or that all or part of the factual data was generated out of thin air by the first language model. For example, after searching for correlated data based on type factual data, if only renovation packages A, B, and C are found, but renovation package D is not found, it indicates a weak match between the type factual data and the correlated data, and that the type factual data contains inaccurate content, i.e., it deviates from the facts.
[0115] In step S1403, the first large language model is used to perform a division calculation on the number of fact data and the number of fact data that match the associated data to obtain the matching result.
[0116] For example, the matching result = number of matches / number of fact data × 100%.
[0117] Steps S1401 to S1403 above determine the authenticity of the answer to be tested by comparing the related data and the content of the answer to be tested, thus preventing the output answer from contradicting the user's question data.
[0118] In one or more embodiments of this disclosure, Figure 10 This is a flowchart illustrating step S150 provided in an embodiment of this disclosure. Figure 10 As shown, step S150 may include the following steps S1501 to S1502.
[0119] In step S1501, the first large language model is used to calculate the relevance level corresponding to the relevant results, the association level corresponding to the associated results, and the matching level corresponding to the matching results.
[0120] In some examples, grading rules are set for relevant, associated, and matched results. Specifically, percentage scores are mapped to multiple grades, each corresponding to a numerical range. For example, 0–20% is mapped to grade 1, 21–40% to grade 2, 41–60% to grade 3, 61–80% to grade 4, and 81–100% to grade 5. If the relevant result is 45%, then the relevant grade is grade 3.
[0121] For example, using the first major language model, based on the above-mentioned level classification rules, the relevance level corresponding to the relevant results, the association level corresponding to the associated results, and the matching level corresponding to the matching results are automatically calculated in order to accurately measure whether the answer to be tested for quality is true.
[0122] In step S1502, if the relevant level, association level, and matching level are all not lower than the level threshold, then the answer to be tested for quality will pass the test.
[0123] Among them, a relevance level not lower than the level threshold indicates that the user's search intent data is accurate.
[0124] A correlation level not lower than the level threshold indicates that the correlation data is accurate.
[0125] A matching level not lower than the level threshold indicates that the answer to be tested is genuine.
[0126] It is understandable that the level threshold is set according to actual needs. For example, in the embodiments of this disclosure, the level threshold is five levels. If the relevant level, the associated level, and the matching level are all not lower than five levels, then the answer to be tested will pass the test.
[0127] In one or more embodiments of this disclosure, if the answer to be tested for quality fails the test after step S1501 is performed, the answer to be tested for quality is corrected. Figure 11 This is a schematic flowchart illustrating a method for correcting failed quality inspection results provided in this embodiment of the disclosure. Figure 11 As shown, the method may include the following steps S1601 to S1609.
[0128] In step S1601, if any of the relevant results, association results, and matching results does not meet the detection conditions, the answer to be tested for quality will fail the test.
[0129] For example, if any of the relevant level, associated level, and matching level is lower than the level threshold, the answer to be tested for quality will fail the test.
[0130] Among them, a relevance level below the level threshold indicates that the user's search intent data is inaccurate.
[0131] A correlation level below the threshold indicates that the correlation data is inaccurate.
[0132] A matching level below the level threshold indicates that the answer to be tested for quality is not genuine.
[0133] Understandably, when an answer to be tested fails the quality check, the retrieval enhancement generation system intercepts that answer to avoid providing incorrect answers to the user. Simultaneously, the retrieval enhancement generation system automatically generates an updated answer.
[0134] Step S1601 of this disclosure adopts a "one-vote veto" rule to determine whether the answer to be tested for quality inspection passes the test, thereby greatly reducing the risk of outputting an incorrect answer.
[0135] Understandably, the query analysis, knowledge retrieval, and enhancement generation steps are executed sequentially. Therefore, the retrieval and enhancement generation system identifies the step most likely to experience an anomaly based on the execution order. When a step malfunctions, but the preceding steps are normal, not only does the malfunctioning step need to be re-executed, but all subsequent steps also need to be re-executed. The preceding steps do not need to be regenerated. This is done to reduce latency and improve user experience, and also to prevent the introduction of new errors after re-executing the normal steps.
[0136] In step S1602, it is determined whether the relevant result does not meet the detection conditions; if the relevant result does not meet the detection conditions, step S1603 is executed; if the relevant result meets the detection conditions, step S1606 is executed.
[0137] In step S1603, the second language model of the retrieval enhancement generation system is used to generate updated user search intent data based on the user question data.
[0138] In step S1604, the second language model is used to search the database based on the updated user search intent data to obtain the updated related data.
[0139] In step S1605, the generative model of the retrieval enhancement generation system is used to generate an updated answer based on the updated related data and user question data.
[0140] For step S1605, if the updated answer still fails the test, steps S1602 to S1605 are repeated until the updated answer passes the test. If the number of times steps S1603 to S1605 are repeated reaches the upper limit, the answer corresponding to the largest relevant result in the aforementioned process is output to avoid excessive delay.
[0141] In step S1606, it is determined whether the association result does not meet the detection conditions. If the association result does not meet the detection conditions, step S1607 is executed; if the association result meets the detection conditions, step S1609 is executed.
[0142] In step S1607, the second language model is used to generate updated related data based on the user search intent data.
[0143] In step S1608, the generative model is used to generate an updated answer based on the updated association data and user question data.
[0144] In step S1609, the generative model is used to generate an updated answer based on the associated data and user question data.
[0145] In some embodiments of this disclosure, to continuously optimize the retrieval enhancement generation system, embodiments of this disclosure store detection data corresponding to each answer to be quality checked. This detection data includes, but is not limited to, user question data, user search intent data, related data, answers, relevant results, associated results, and matching results. The detection data corresponding to answers that fail quality check are periodically analyzed to identify steps requiring optimization.
[0146] For example, when a large number of cases have a relevance level below the threshold, the query analysis step needs optimization. In this case, the prompts should be optimized to more accurately identify user intent. When a large number of cases have an association level below the threshold, the knowledge retrieval step needs optimization. In this case, the database content should be supplemented based on the actual situation to find more relevant related data. When a large number of cases have a matching level below the threshold, the enhancement generation step needs optimization. In this case, the enhancement generation algorithm should be optimized.
[0147] Based on the above analysis of steps S1601 to S1609, it can be seen that the retrieval enhancement generation system of this disclosure possesses a certain degree of self-diagnosis and correction capabilities. The retrieval enhancement generation system intercepts answers that fail the quality check to avoid providing incorrect answers to the user. Simultaneously, the retrieval enhancement generation system can automatically diagnose which step caused the answer to fail the quality check, and then re-execute that abnormal step to accurately correct the failed answer. This not only significantly reduces the risk of factual errors caused by problems in a certain step, improving the system's fault tolerance and reliability, but also enhances the user experience.
[0148] Based on the above analysis, this disclosure achieves precise dynamic quality checks at each step involved in generating the answer. This allows for accurate identification of the steps that cause the answer to fail the quality check, and further precise correction of the failed answer based on the abnormal steps. This significantly reduces the risk of factual errors caused by problems in a single step, improving the system's fault tolerance and reliability, and enhancing the user experience. Furthermore, the retrieval enhancement generation system can automatically diagnose which step caused the answer to fail the quality check, and then re-execute that abnormal step to accurately correct the failed answer. This not only significantly reduces the risk of factual errors caused by problems in a single step, improving the system's fault tolerance and reliability, but also enhances the user experience.
[0149] This disclosure also provides a quality detection device for a retrieval enhancement generation system. Figure 12 This is a schematic structural block diagram of a quality inspection device for a search enhancement generation system according to one embodiment of this disclosure. Figure 12 As shown, the device includes: The generation module 10 is used to obtain the answer to be tested for quality; the answer to be tested for quality is generated by the retrieval enhancement generation system based on the user's search intent data and multiple related data; the search intent data is generated based on the user's question data; the related data is the data associated with the user's search intent data.
[0150] Analysis module 20 is used to analyze the correlation between user question data and user search intent data to obtain relevant results; the relevant results indicate whether the user search intent data is accurate. Analysis module 20 is also used to analyze the correlation between each piece of related data and the user's question data to obtain the correlation results; the correlation results indicate whether the related data is accurate. Analysis module 20 is also used to analyze the matching between the answer to be tested and the associated data, and obtain the matching results; the matching results indicate whether the answer to be tested is true. The results module 30 is also used to ensure that the answer to be tested passes the test if the relevant results, associated results and matching results all meet the detection conditions.
[0151] The specific implementation process of the functions and roles of each module in the above-mentioned device is detailed in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here. The execution subject of the quality detection method of the retrieval enhancement generation system in the specific embodiments of this disclosure can be an electronic device such as a server (including a local server or a cloud server).
[0152] Therefore, based on any of the above embodiments, this disclosure also provides an electronic device that can execute the quality detection method of the retrieval enhancement generation system of any of the embodiments described above.
[0153] Figure 13 This is a schematic structural block diagram of an electronic device equipped with a quality inspection device for a retrieval enhancement generation system, according to one embodiment of the present disclosure.
[0154] The hardware architecture of electronic devices / devices can be implemented using a bus architecture. A bus architecture can include any number of interconnect buses and bridges, depending on the specific application and overall design constraints of the hardware. Bus 1100 connects various circuits, including one or more processors 1200, memory 1300, and / or hardware modules. Bus 1100 can also connect various other circuits 1400, such as peripherals, voltage regulators, power management circuits, external antennas, etc. Bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Component (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one connection line is used in this diagram, but this does not indicate that there is only one bus or one type of bus.
[0155] For ease of explanation, certain steps of the above method are described in relation to modules. It should be understood that the corresponding module performing one or more steps of the above method may be one or more hardware modules specifically configured to perform the corresponding step, or implemented by a processor configured to perform the corresponding step, or stored in a computer-readable medium for implementation by a processor, or implemented by some combination thereof.
[0156] This disclosure also provides a readable storage medium storing a computer program that, when executed by a processor, is used to implement the methods described above. A "readable storage medium" can be any means capable of containing, storing, communicating, propagating, or transmitting a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples of a readable storage medium include: an electrical connection with one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable read-only memory (CDROM), etc.
[0157] This disclosure also provides a computer program product, the methods of which can be implemented wholly or partially through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented wholly or partially as a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, all or part of the processes or functions of this disclosure are performed.
[0158] Computer programs or instructions can be stored in a readable storage medium or transferred from one readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The readable storage medium can be any available medium capable of access, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; an optical medium, such as a digital video optical disc; or a semiconductor medium, such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or it can include both volatile and non-volatile types of storage media.
[0159] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0160] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and computer program products according to this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0161] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0162] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0163] Those skilled in the art should understand that the above embodiments are merely for illustrating the present disclosure and are not intended to limit the scope of the disclosure. Those skilled in the art can make other changes or modifications based on the above disclosure, and these changes or modifications still fall within the scope of the present disclosure.
Claims
1. A quality detection method for a retrieval enhancement generation system, characterized in that, include: Obtain the answer to the quality inspection; The answer to be tested for quality is generated by the retrieval enhancement generation system based on user search intent data and multiple related data. The search intent data is generated based on user query data; The associated data is data that is associated with the user's search intent data; Analyze the correlation between the user question data and the user search intent data to obtain relevant results; The relevant results indicate whether the user search intent data is accurate; Analyze the correlation between each of the associated data and the user question data to obtain the correlation results; The correlation result indicates whether the correlation data is accurate; Analyze the matching between the answer to be tested and the associated data to obtain the matching results; The matching result indicates whether the answer to be tested for quality is true; If the relevant results, the association results, and the matching results all meet the detection conditions, then the answer to be tested for quality will pass the detection.
2. The method as described in claim 1, characterized in that, Analyzing the correlation between the user question data and the user search intent data yields relevant results, including: Using the first language model of the retrieval enhancement generation system, the topic of the user question data is obtained; the relevance between the topic of the user question data and the user search intent data is analyzed to obtain topic-related results. Using the first large language model, key information from the user's question data is obtained; the coverage of the user's search intent data with the key information is analyzed to obtain the coverage result. Using the first large language model, redundant information of the user question data is obtained; the distribution of key information and redundant information in the user search intent data is analyzed to obtain redundancy results. Using the first large language model, the average of the topic-related results, the coverage results, and the redundancy results is calculated to obtain the related results.
3. The method as described in claim 2, characterized in that, Analyze the coverage of the key information by the user search intent data to obtain the coverage results, including: The number of key information items obtained, and the coverage of key information items included in the user search intent data; The coverage result is obtained by dividing the coverage quantity and the key quantity.
4. The method as described in claim 1, characterized in that, Analyze the correlation between each of the associated data and the user question data to obtain the correlation results, including: The first major language model is used to obtain multiple reasoning question data corresponding to each of the aforementioned associated data; For each piece of associated data, multiple association values are calculated using the first large language model; the association values are obtained by performing association calculations on the inference question data and the user question data; the association values correspond one-to-one with the inference question data. The first large language model is used to calculate the average of multiple associated values corresponding to the associated data to obtain the associated result.
5. The method as described in claim 4, characterized in that, Analyze the matching between the answer to be tested and the associated data to obtain the matching results, including: Use the first large language model to obtain multiple factual data from the answer to be quality tested; The first large language model is used to compare multiple factual data with the associated data to obtain factual data that matches the associated data. The first large language model is used to divide the number of factual data and the number of factual data that match the associated data to obtain the matching result.
6. The method as described in claim 1, characterized in that, The method further includes: if any one of the relevant results, the association results, and the matching results does not meet the detection conditions, then the answer to be tested for quality fails the detection; If the relevant results do not meet the detection conditions, the second language model of the retrieval enhancement generation system is used to generate updated user search intent data based on the user question data. Using the second major language model, the database is searched based on the updated user search intent data to obtain the updated related data; Using the generative model of the retrieval enhancement generation system, an updated answer is generated based on the updated associated data and the user question data; Alternatively, if the relevant results meet the detection conditions, but the association results do not meet the detection conditions, then the second large language model is used to generate the updated association data based on the user search intent data; Using the generative model, an updated answer is generated based on the updated association data and the user question data; Alternatively, if both the relevant results and the association results satisfy the detection conditions, and the matching results do not satisfy the detection conditions, then the generation model is used to generate an updated answer based on the association data and the user question data.
7. The method as described in claim 6, characterized in that, If the relevant results, the association results, and the matching results all meet the detection conditions, then the answer to be tested for quality control passes the detection, including: The first major language model is used to calculate the relevance level of the relevant results, the association level of the association results, and the matching level of the matching results. If the relevant level, the association level, and the matching level are all not lower than the level threshold, then the answer to be tested for quality will pass the test. Wherein, the relevance level being no lower than the level threshold indicates that the user's search intent data is accurate; The association level being no lower than the association threshold indicates that the association data is accurate. The matching level being no lower than the level threshold indicates that the answer to be tested is genuine.
8. The method as described in claim 7, characterized in that, If any one of the relevant results, the association results, and the matching results does not meet the detection conditions, then the answer to be tested for quality fails the detection, including: If any one of the relevant level, the association level, and the matching level is lower than the level threshold, then the answer to be tested for quality fails the test. Wherein, a relevance level below the level threshold indicates that the user's search intent data is inaccurate; The association level being lower than the association threshold indicates that the association data is inaccurate. The matching level being lower than the level threshold indicates that the answer to be tested for quality is not genuine.
9. An electronic device, characterized in that, include: The memory stores execution instructions; as well as A processor that executes the execution instructions stored in the memory, causing the processor to perform the quality detection method of the retrieval enhancement generation system according to any one of claims 1 to 8.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the quality detection method of the retrieval enhancement generation system according to any one of claims 1 to 8.