Question and answer processing method, system and equipment based on three-dimensional entropy evaluation and medium
By constructing a semantic graph and calculating a three-dimensional evaluation system of semantic entropy, available entropy, and sample entropy, the problems of illusion recognition and parameter updating in large-scale language model question-answering processing are solved, achieving more efficient and reliable question-answering processing.
Patent Information
- Application Number
- CN202511536093.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-10-27
AI Technical Summary
Existing question-answering techniques for large-scale language models mostly employ single-dimensional evaluation, which fails to identify high-frequency and highly harmful hallucinations and cannot adapt to parameter updates across various domains in a timely manner, thus limiting the reliability and practical application value of model question-answering.
A question-answering processing method based on three-dimensional entropy evaluation is constructed. By building a semantic graph and calculating semantic entropy, available entropy and sample entropy, and combining dynamic threshold comparison and feedback adjustment mechanism, semantic contradictions, prediction anomalies and sample deviations are comprehensively detected, thereby improving the reliability of the model question-answering system.
Through a multi-dimensional evaluation mechanism, semantic structure confusion, lack of prediction stability, and deviation from domain knowledge are effectively identified, reducing the risk of dangerous suggestion output and improving the reliability and security of model question answering.
Smart Images

Figure CN121009182A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a question-answering processing method, system, device and medium based on three-dimensional entropy assessment. Background Technology
[0002] With the widespread application of large language models, model-driven question answering has emerged in various fields. Leveraging its ability to rapidly generate natural language responses, it has become an important tool for improving service efficiency across these sectors. However, responses generated by large language models based on questions often suffer from the illusion problem, which not only affects the accuracy of user decisions but may also pose security risks in critical areas such as healthcare.
[0003] Current model-based question answering technologies mostly employ single-dimensional evaluation, and the real sample database supporting hallucination recognition lacks a regular update and maintenance mechanism. This results in the inability to identify high-frequency and harmful hallucinations in various fields, and the inability to adapt to changes in parameters across different fields in a timely manner. Consequently, the hallucination problems in model-based question answering cannot be accurately identified, and the processing efficiency is low, which in turn restricts the overall reliability of model-based question answering and its practical application value in various fields. Summary of the Invention
[0004] The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.
[0005] The main objective of this disclosure is to propose a question-answering processing method, system, device, and storage medium based on three-dimensional entropy evaluation. This method can comprehensively detect and handle semantic contradictions, prediction anomalies, and sample deviations by constructing a semantic graph and calculating a three-dimensional evaluation system of semantic entropy, available entropy, and sample entropy, combined with dynamic threshold comparison and feedback adjustment mechanisms, thereby improving the reliability of the model question-answering system.
[0006] A first aspect of this application provides a question-answering processing method based on three-dimensional entropy evaluation for a central controller, the method comprising: A question-and-answer set to be processed is constructed based on the target question, the target answer, and the historical context; the target answer is generated by the large language model based on the target question; the historical context is the contextual content associated with the target question. Based on the question-and-answer set to be processed, a corresponding semantic graph is constructed; each node in the semantic graph is obtained based on the target entities in the question-and-answer set to be processed, and each edge in the semantic graph is obtained based on the target attributes and target associations in the question-and-answer set to be processed; the target associations represent the logical relationships between the target entities in the question-and-answer set to be processed. Multiple combinations to be verified are extracted from the question and answer set to be processed; wherein each combination to be verified consists of a target entity in the question and answer set to be processed, a target attribute in the question and answer set to be processed, and a target attribute value in the question and answer set to be processed. Based on the semantic graph and each of the combinations to be verified, the semantic entropy, usable entropy, and sample entropy of the target response are calculated respectively; the semantic entropy is used to quantify the logical consistency of the target response, the usable entropy is used to quantify the predictive stability of the large language model when generating the target response based on the target question, and the sample entropy is used to quantify the matching degree between the target response and the real sample library of the domain corresponding to the target question; Based on the semantic entropy, the available entropy, and the sample entropy, a three-dimensional entropy result of the target response is generated, and the question-answering processing scheme corresponding to the three-dimensional entropy result is executed.
[0007] In some embodiments of this application, constructing a corresponding semantic graph based on the question-and-answer set to be processed includes: Extract the target entity, the target attribute, and the target relationship from the question-and-answer set to be processed; the target entity represents a domain-specific object in the question-and-answer set to be processed; the target attribute represents the feature of the target entity; the target relationship includes subject-verb-object relationship, causal relationship, temporal relationship, and attribute relationship between the target entities; The target entity and the target attribute are used as nodes in the semantic graph; The target association is used as the edge of the semantic graph, and the confidence of each edge in the semantic graph is calculated based on the labeled confidence of the target entity and the labeled confidence of the target attribute. The semantic graph is obtained by integrating the nodes, edges, and confidence scores of each edge in the semantic graph.
[0008] In some embodiments of this application, the semantic entropy of the target response is calculated through the following steps: Based on the confidence level of each edge in the semantic graph, the target association relationship of each edge in the semantic graph is analyzed according to the preset semantic analysis model to obtain the rationality score of each target association relationship; Based on the reasonableness score of each target association, the degree of confusion between each node and each edge in the semantic graph is calculated; the degree of confusion represents the degree of semantic and logical consistency between the target response, the target question, and the historical context. Based on the degree of confusion, the semantic entropy of the target response is calculated; the semantic entropy represents the probability that the target response contains a semantic contradiction.
[0009] In some embodiments of this application, the available entropy of the target response is calculated through the following steps: Retrieve the probability distribution of candidate content for each of the combinations to be verified during the process of generating the target response by the large language model; the probability distribution of candidate content is the probability distribution of multiple alternative contents when the large language model generates the combination to be verified. The dispersion of the probability distribution of candidate content corresponding to each of the aforementioned combinations to be verified is statistically analyzed. Based on the probability distribution dispersion of each of the combinations to be verified, the available entropy of the target response is calculated; the available entropy characterizes the probability degree of prediction anomalies in the target response.
[0010] In some embodiments of this application, the sample entropy of the target response is calculated through the following steps: Retrieve the real sample library corresponding to the target problem; Calculate the feature distance between the target response and all samples in the real sample library of the domain; Based on the aforementioned feature distances, reference samples are extracted from the samples; the reference samples are those samples that meet preset sample conditions; the preset sample conditions include the feature distance being less than a preset feature distance threshold. Statistically analyze the probability distribution of feature distances between the target response and each of the reference samples; Based on the feature distance probability distribution, the sample entropy of the target response is calculated; the sample entropy represents the degree of matching between the target response and the real sample library of the corresponding domain of the target question.
[0011] In some embodiments of this application, generating a three-dimensional entropy result of the target response based on the semantic entropy, the available entropy, and the sample entropy, and then executing the question-answering processing scheme corresponding to the three-dimensional entropy result, includes: The semantic entropy, the available entropy, and the sample entropy are weighted and fused to obtain the three-dimensional entropy result of the target response; wherein, the weights of the semantic entropy, the available entropy, and the sample entropy in the weighted fusion process are determined according to the needs of the domain corresponding to the target response. Retrieve the hallucination threshold of the domain corresponding to the question-and-answer set to be processed; The three-dimensional entropy result of the target response is compared with the illusion threshold of the corresponding domain of the question-and-answer set to be processed to obtain the comparison result; Execute the corresponding question-and-answer processing scheme based on the comparison results; If the three-dimensional entropy result of the target response is lower than the illusion threshold of the domain corresponding to the question-and-answer set to be processed, the target response is output; if the three-dimensional entropy result of the target response is higher than or equal to the illusion threshold of the domain corresponding to the question-and-answer set to be processed, the semantic entropy, the available entropy, and the sample entropy of the target response are analyzed to determine the illusion type of the target response; based on the illusion type of the target response, a question-and-answer anomaly prompt is generated and the output of the target response is prevented.
[0012] In some embodiments of this application, after generating a three-dimensional entropy result of the target response based on the semantic entropy, the available entropy, and the sample entropy, and executing the question-answering processing scheme corresponding to the three-dimensional entropy result, the method further includes: In response to error messages from abnormal user feedback, the hallucination type of the error message is obtained; When the hallucination type of the erroneous information is based on the semantic entropy anomaly, the semantic error between the semantic entropy of the target response and the semantic entropy threshold of the domain corresponding to the target response is calculated, so as to adjust the calculation steps of the semantic entropy according to the semantic error; the adjustment of the calculation steps of the semantic entropy includes adjusting the judgment threshold of the reasonableness score output by the preset semantic analysis model in the semantic entropy; When the illusion type of the error message is based on the anomaly of the available entropy, the prediction error between the available entropy of the target response and the available entropy threshold of the corresponding domain of the target response is calculated, so as to adjust the calculation steps of the available entropy according to the prediction error; the adjustment of the calculation steps of the available entropy includes the adjustment of the statistical standard of the dispersion in the available entropy; When the illusion type of the error message is based on the abnormality of the sample entropy, the true error between the sample entropy of the target response and the sample entropy threshold of the corresponding domain of the target response is calculated, so as to adjust the calculation steps of the sample entropy according to the true error; the adjustment of the calculation steps of the sample entropy includes updating the domain true sample library in the sample entropy.
[0013] To achieve the above objectives, a second aspect of the present invention provides a question-answering processing system based on three-dimensional entropy assessment, the system comprising: The first module is used to construct a set of questions and answers to be processed based on the target question, the target answer, and the historical context; the target answer is generated by the large language model based on the target question; the historical context is the context content associated with the target question. The second module is used to construct a corresponding semantic graph based on the question-and-answer set to be processed; each node in the semantic graph is obtained based on the target entities in the question-and-answer set to be processed, and each edge in the semantic graph is obtained based on the target attributes and target associations in the question-and-answer set to be processed; the target associations represent the logical relationships between the target entities in the question-and-answer set to be processed. An extraction module is used to extract multiple combinations to be verified from the question-and-answer set to be processed; wherein each combination to be verified consists of a target entity in the question-and-answer set to be processed, a target attribute in the question-and-answer set to be processed, and a target attribute value in the question-and-answer set to be processed; The calculation module is used to calculate the semantic entropy, usable entropy, and sample entropy of the target response based on the semantic graph and each of the combinations to be verified. The semantic entropy is used to quantify the logical consistency of the target response, the usable entropy is used to quantify the predictive stability of the large language model when generating the target response based on the target question, and the sample entropy is used to quantify the matching degree between the target response and the real sample library of the domain corresponding to the target question. The execution module is used to generate a three-dimensional entropy result of the target response based on the semantic entropy, the available entropy, and the sample entropy, so as to execute the question-answering processing scheme corresponding to the three-dimensional entropy result.
[0014] To achieve the above objectives, a third aspect of the present invention provides an electronic device, comprising: at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, the instructions being executed by the at least one control processor to enable the at least one control processor to perform the above-described question-answering method based on three-dimensional entropy evaluation.
[0015] To achieve the above objectives, a fourth aspect of the present invention provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-described question-answering method based on three-dimensional entropy assessment.
[0016] This application provides a question-answering processing method based on three-dimensional entropy evaluation. The method constructs a question-and-answer set to be processed based on the target question, target answer, and historical context. Based on this set, a corresponding semantic graph is constructed, extracting multiple combinations to be verified. The semantic entropy, usable entropy, and sample entropy of the target answer are calculated based on the semantic graph and each combination to be verified. A three-dimensional entropy result for the target answer is generated based on the semantic entropy, usable entropy, and sample entropy, and the question-answering processing scheme corresponding to the three-dimensional entropy result is executed. This method, by constructing a semantic graph and calculating the three-dimensional evaluation system of semantic entropy, usable entropy, and sample entropy, combined with dynamic threshold comparison and feedback adjustment mechanisms, comprehensively detects and handles semantic contradictions, prediction anomalies, and sample deviations, thereby improving the reliability of the model-based question-answering system.
[0017] It is understood that the beneficial effects of the second to fourth aspects compared with the related technologies are the same as the beneficial effects of the first aspect compared with the related technologies. Please refer to the relevant description in the first aspect above, which will not be repeated here. Attached Figure Description
[0018] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is a flowchart illustrating a question-answering processing method based on three-dimensional entropy evaluation provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a question-answering processing system based on three-dimensional entropy evaluation provided in an embodiment of this application; Figure 3 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0019] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0020] In the description of this application, the use of terms such as "first," "second," etc., is for the purpose of distinguishing technical features only and should not be construed as indicating or implying relative importance or implicitly indicating the number of technical features indicated or the order of the technical features indicated.
[0021] In the description of this application, it should be understood that the orientation descriptions, such as up, down, etc., are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.
[0022] In the description of this application, it should be noted that, unless otherwise explicitly defined, terms such as "setup," "installation," and "connection" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this application in conjunction with the specific content of the technical solution.
[0023] With the widespread application of large language models, model-driven question answering has emerged in various fields. Leveraging its ability to rapidly generate natural language responses, it has become an important tool for improving service efficiency across these sectors. However, responses generated by large language models based on questions often suffer from the illusion problem, which not only affects the accuracy of user decisions but may also pose security risks in critical areas such as healthcare.
[0024] Current model-based question answering technologies mostly employ single-dimensional evaluation, and the real sample database supporting hallucination recognition lacks a regular update and maintenance mechanism. This results in the inability to identify high-frequency and harmful hallucinations in various fields, and the inability to adapt to changes in parameters across different fields in a timely manner. Consequently, the hallucination problems in model-based question answering cannot be accurately identified, and the processing efficiency is low, which in turn restricts the overall reliability of model-based question answering and its practical application value in various fields.
[0025] Based on this, embodiments of this application provide a question-answering processing method, system, electronic device, and medium based on three-dimensional entropy evaluation. The aim is to improve the reliability of the model question-answering system by constructing a semantic graph and calculating a three-dimensional evaluation system of semantic entropy, available entropy, and sample entropy, combined with dynamic threshold comparison and feedback adjustment mechanisms, to comprehensively detect and handle semantic contradictions, prediction anomalies, and sample deviations.
[0026] The question-answering method, system, electronic device, and medium based on three-dimensional entropy assessment provided in this application are specifically described through the following embodiments. First, the question-answering method based on three-dimensional entropy assessment in this application embodiment is described.
[0027] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0028] Foundational artificial intelligence technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0029] The question-answering method based on three-dimensional entropy evaluation provided in this application relates to the field of artificial intelligence technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the question-answering method based on three-dimensional entropy evaluation, but is not limited to the above forms.
[0030] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0031] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirects to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0032] Therefore, referring to Figure 1 This application provides a question-answering processing method based on three-dimensional entropy evaluation. This method is applied to a central controller, which can be a server, an electronic device, or a mobile terminal, etc. There are no specific limitations here. The method includes the following steps S110 to S150.
[0033] Step S110: Construct a question-and-answer set to be processed based on the target question, target answer, and historical context; the target answer is generated by the large language model based on the target question; the historical context is the contextual content associated with the target question; Step S120: Based on the question-and-answer set to be processed, construct the corresponding semantic graph; each node in the semantic graph is obtained based on the target entities in the question-and-answer set to be processed, and each edge in the semantic graph is obtained based on the target attributes and target associations in the question-and-answer set to be processed; the target associations represent the logical relationships between the target entities in the question-and-answer set to be processed. Step S130: Extract multiple combinations to be verified from the question and answer set to be processed; wherein each combination to be verified consists of the target entity in the question and answer set to be processed, the target attribute in the question and answer set to be processed, and the target attribute value in the question and answer set to be processed. Step S140: Based on the semantic graph and each combination to be verified, calculate the semantic entropy, available entropy and sample entropy of the target response respectively; the semantic entropy is used to quantify the logical consistency of the target response, the available entropy is used to quantify the predictive stability of the large language model when generating the target response based on the target question, and the sample entropy is used to quantify the matching degree between the target response and the real sample library of the corresponding domain of the target question. Step S150: Generate a three-dimensional entropy result of the target response based on semantic entropy, available entropy, and sample entropy, and execute the question-and-answer processing scheme corresponding to the three-dimensional entropy result.
[0034] In this step, a question-and-answer set to be processed is first constructed based on the target question, target answer, and historical context. The target answer is generated by a large language model based on the target question, the historical context is the contextual content associated with the target question, and the question-and-answer set to be processed refers to the data set formed by integrating the target question, target answer, and historical context. Specifically, natural language processing techniques can be used to extract question entities, answer elements, and contextual association information, providing a structured data foundation for subsequent semantic graph construction and entropy calculation.
[0035] Furthermore, a corresponding semantic graph is constructed based on the question-and-answer set to be processed. Each node in the semantic graph is obtained from the target entities in the question-and-answer set to be processed, and each edge is obtained from the target attributes and target relationships in the question-and-answer set to be processed. Each target relationship represents the logical relationship between the target entities in the question-and-answer set to be processed.
[0036] Specifically, a semantic graph refers to a knowledge network that expresses semantic relationships through nodes and edges. Specifically, it can be implemented by using named entity recognition technology to extract target entities as nodes and dependency parsing technology to extract subject-verb-object relationships as edges. Through the visual expression of logical relationships between entities, it provides a structured analysis framework for semantic contradiction detection, overcoming the shortcomings of traditional text matching methods that cannot capture deep logical contradictions.
[0037] Furthermore, multiple verification combinations are extracted from the question-and-answer set to be processed. Each verification combination consists of a triplet structure containing the target entity, target attribute, and target attribute value from the question-and-answer set to be processed. Specifically, it can be extracted from the question-and-answer text using regular expression matching or sequence labeling models. By decomposing the response content into the smallest quantifiable and verifiable semantic unit, it supports fine-grained analysis for subsequent entropy value calculation, breaking through the limitation of traditional methods that only evaluate the credibility of the overall statement.
[0038] Furthermore, based on the semantic graph and each combination to be verified, the semantic entropy, usable entropy, and sample entropy of the target response are calculated respectively. Among them, the semantic entropy is used to quantify the logical consistency of the target response, the usable entropy is used to quantify the predictive stability of the large language model when generating the target response based on the target question, and the sample entropy is used to quantify the matching degree between the target response and the real sample library of the corresponding domain of the target question.
[0039] Preferably, semantic entropy is a quantitative indicator representing the logical consistency of the response content. Specifically, it can be achieved by calculating the dispersion of the rationality score of the relationship between nodes in the semantic graph, which is used to detect semantic confusion caused by contradictory relationships between entities. Usable entropy is a quantitative indicator representing the stability of the model generation process. Specifically, it can be achieved by calculating the variance of the probability distribution of candidate content generated by the statistical language model, which is used to identify abnormal outputs caused by fluctuations in model predictions. Sample entropy is a quantitative indicator representing the domain knowledge matching degree. Specifically, it can be achieved by calculating the cosine similarity distribution between the response feature vector and the real sample library, which is used to detect illusory content that deviates from domain common sense.
[0040] Furthermore, based on semantic entropy, available entropy, and sample entropy, a three-dimensional entropy result for the target response is generated to execute the question-answering processing scheme corresponding to the three-dimensional entropy result. The three-dimensional entropy result refers to a comprehensive quantitative value that integrates semantics, generation stability, and domain matching. Specifically, a dynamic weighted algorithm can be used to adjust the entropy weights according to domain requirements. This multi-dimensional evaluation mechanism improves the comprehensiveness of illusion recognition and overcomes the high misjudgment rate problem associated with single-indicator evaluation.
[0041] Therefore, this step enhances the ability to identify deep semantic contradictions by constructing a semantic graph and integrating entity relationship parsing. Combined with a multi-dimensional entropy calculation mechanism, it improves the accuracy of detecting prediction anomalies and knowledge deviations, and realizes multi-dimensional evaluation of responses generated by large language models. It effectively identifies complex illusion features such as semantic structure confusion, lack of prediction stability, and domain knowledge deviation, thereby improving the reliability of model question-answering processing, reducing the risk of dangerous suggestion output, and laying the foundation for safe applications in various fields.
[0042] In some embodiments, in step S120, a corresponding semantic graph is constructed based on the question-and-answer set to be processed, including the following steps S210 to S240: Step S210: Extract target entities, target attributes, and target relationships from the question-and-answer set to be processed; target entities represent domain-specific objects in the question-and-answer set to be processed; target attributes represent the characteristics of target entities; target relationships include subject-verb-object relationships, causal relationships, temporal relationships, and attribute relationships between target entities; Step S220: Use the target entity and target attributes as nodes in the semantic graph; Step S230: Treat the target association as edges of the semantic graph, and calculate the confidence of each edge in the semantic graph based on the label confidence of the target entity and the label confidence of the target attribute. Step S240: Integrate the nodes, edges, and confidence scores of each edge in the semantic graph to obtain the semantic graph.
[0043] In this embodiment, target entities, target attributes, and target relationships are first extracted from the question-and-answer set to be processed. Target entities represent domain-specific objects in the question-and-answer set to be processed, target attributes represent the characteristics of target entities, and target relationships include subject-verb-object relationships, causal relationships, temporal relationships, and attribute relationships between target entities.
[0044] Specifically, target entities are extracted from the question-and-answer set to be processed and defined as domain-specific objects, such as disease names or drug names in a medical scenario; target attributes are extracted from the question-and-answer set to be processed and defined as features of target entities, such as drug dosage or disease symptoms; target relationships include subject-verb-object relationships, causal relationships, temporal relationships, and attribute relationships, such as the causal relationship between disease and symptoms.
[0045] Furthermore, target entities and target attributes are used as nodes in the semantic graph, and target relationships are used as edges in the semantic graph. The confidence scores of each edge in the semantic graph are calculated based on the labeled confidence scores of the target entities and target attributes. The labeled confidence scores of the target entities and target attributes are pre-assigned during the extraction process.
[0046] Preferably, the confidence of an edge is calculated by weighting the confidence of the target entity and the target attribute, for example, by using the arithmetic mean or geometric mean of the confidence of the attribute. The integration process of the semantic graph includes connecting nodes and edges topologically according to the graph structure and storing the confidence of the edge as an attribute value.
[0047] Therefore, this embodiment can construct a structured semantic graph, effectively capturing entity, attribute, and relation information in the question-answer set, providing a foundation for subsequent semantic analysis and entropy calculation, and helping to more accurately assess the semantic consistency and reliability of the responses. Simultaneously, by introducing confidence level calculation, the credibility of each element in the graph can be quantified, further improving the accuracy of semantic analysis.
[0048] In some embodiments, the semantic entropy of the target response in step S140 is calculated through the following steps, including steps S310 to S330: Step S310: Based on the confidence of each edge in the semantic graph, analyze the target association relationship of each edge in the semantic graph according to the preset semantic analysis model, and obtain the rationality score of each target association relationship. Step S320: Based on the reasonableness score of the relationship between each target, calculate the degree of confusion between each node and each edge in the semantic graph; the degree of confusion represents the degree of semantic logical consistency between the target response and the target question and historical context; Step S330: Calculate the semantic entropy of the target response based on the degree of confusion; the semantic entropy represents the probability that the target response contains a semantic contradiction.
[0049] In this embodiment, firstly, based on the confidence level of each edge in the semantic graph, the target association relationships of each edge in the semantic graph are analyzed according to a preset semantic analysis model to obtain a reasonableness score for each target association relationship. The reasonableness score quantifies the logical reasonableness of the target association relationship through the preset semantic analysis model, such as whether the subject-verb-object relationship conforms to grammatical rules or whether the causal relationship conforms to domain common sense.
[0050] Furthermore, based on the reasonableness scores of the relationships between each target, the degree of confusion between each node and each edge in the semantic graph is statistically analyzed. The degree of confusion characterizes the semantic logical consistency between the target response, the target question, and the historical context. This degree of confusion can be determined by calculating the variance of the reasonableness scores to statistically analyze the frequency of logical conflicts between nodes and edges.
[0051] Furthermore, the semantic entropy of the target response is calculated based on the degree of confusion. Semantic entropy represents the probability of a semantic contradiction in the target response. It can be calculated by normalizing the degree of confusion and then weighting it with confidence levels. For example, the degree of confusion can be mapped to a probability distribution, and then the semantic entropy can be calculated using the information entropy formula. A higher entropy value indicates a greater probability of semantic contradiction.
[0052] Therefore, this embodiment can quantitatively evaluate the quality of the target response from the perspective of semantic logical consistency, effectively identify possible semantic contradictions in the response, and improve the reliability of the question-answering system. Moreover, by introducing semantic graphs and pre-set semantic analysis models, it achieves in-depth analysis of the semantic relationships in the response content, which can more accurately capture semantic-level issues compared to methods that simply rely on keyword matching.
[0053] In some embodiments, the available entropy of the target response in step S140 is calculated by the following steps, including steps S410 to S430: Step S410: Retrieve the probability distribution of candidate content for each combination to be verified during the generation of the target response by the large language model; the probability distribution of candidate content is the probability distribution of multiple alternative contents when the large language model generates the combination to be verified. Step S420: Calculate the dispersion of the probability distribution of candidate content corresponding to each combination to be verified; Step S430: Calculate the available entropy of the target response based on the probability distribution dispersion of each combination to be verified; the available entropy characterizes the probability of an anomaly in the target response.
[0054] In this embodiment, firstly, the probability distribution of candidate content for each combination to be verified is retrieved during the generation of the target response by the large language model. This candidate content probability distribution is obtained by calling the model inference log, representing the probability distribution of multiple alternative contents when the large language model generates the combination to be verified, including the top-k candidate words and their probability values for each combination to be verified during the generation process.
[0055] Furthermore, the dispersion of the probability distribution of candidate content corresponding to each combination to be verified is statistically analyzed. Specifically, firstly, target units are screened from the question-and-answer set to be processed. For example, entities with non-empty NERs in token_info (such as "aspirin" and "5mg" in the medical scenario), attribute values in entity attributes (such as "dosage: 5mg" and "interest rate: 3.5%), and factual statement phrases (such as "hypertension requires a low-salt diet") are extracted first, while function words without factual verification value (such as "of" and "oh") are excluded to ensure that all units are information that can be verified through external knowledge or rules.
[0056] Furthermore, the probabilities of the top-k candidate tokens for the target unit are extracted. Specifically, the LLM generation logs captured during the preprocessing stage (including the logits output when each token is generated) are called. If the target unit is a single token (e.g., "5mg"), the top-k (k-5 by default in the document, adjustable in the domain) candidate tokens and their probabilities at the time of token generation are directly extracted (e.g., the top-5 candidates for "5mg" are "5mg", "10mg", "5q", "3mg", and "2mq7", with probabilities of 0.82, 0.1, 0.05, 0.02, and 0.01, respectively). If the target unit is multiple tokens (e.g., "aspirin"), the intersection probability distribution of the top-k probabilities of each token is taken to ensure the overall prediction stability of the covered unit.
[0057] Furthermore, the dispersion of a single target unit is calculated. Specifically, variance quantifies the dispersion of the top-k probability distribution (the larger the variance, the higher the dispersion, and the more unstable the model prediction). The probability of the top-k candidate tokens for a single target unit (denoted as...) ,..., ,in, The target unit number, For the first When the target unit is generated, the first To calculate the probability of each candidate token, first calculate the mean probability, as shown in the following formula: ; in, For the first The average probability of each target unit. The total number of target units, For the first When the target unit is generated, the first The probability of each candidate token. Then substituting into the variance formula: ; in, For the first The variance of each target unit, The total number of target units, For the first When the target unit is generated, the first The probability of each candidate token. For the first The mean probability of each target unit is calculated. Then, the dispersion of the current target unit is calculated (e.g., the top-5 probability variance of "5mg" is 0.12, indicating that the probability is concentrated; the top-5 probability variance of "XYZ antihypertensive tablets" is 0.02, indicating that the probability is dispersed).
[0058] Furthermore, the dispersion of multiple target units is aggregated. Specifically, the average variance of all selected target units is calculated and substituted into the formula: ; in, The total number of target units selected. For the first The variance of each target unit is used to calculate the overall top-k probability distribution dispersion, which is the core input of the usable entropy and directly reflects the model's predictive stability for verifiable information.
[0059] Furthermore, based on the probability distribution dispersion of each combination to be verified, the available entropy of the target response is calculated. Specifically, the available entropy calculation integrates the dispersion weighted average of all combinations to be verified, with the weights dynamically adjusted according to the node level of the combination to be verified in the semantic graph, so as to characterize the probability of the target response having prediction anomalies through available entropy.
[0060] Therefore, by analyzing the probability distribution of candidate content, this embodiment can assess the model's certainty about different options, quantify the predictive stability of the large language model when generating the target response, and reflect the reliability of the target response through available entropy, thus providing an important basis for identifying and handling illusion problems in model question answering.
[0061] In some embodiments, the sample entropy of the target response in step S140 is calculated by the following steps, including steps S510 to S550: Step S510: Retrieve the real sample library corresponding to the target problem's domain; Step S520: Calculate the feature distance between the target response and all samples in the real sample library of the domain; Step S530: Extract reference samples from the samples based on each feature distance; the reference samples are those that meet the preset sample conditions; the preset sample conditions include feature distances less than preset feature distance thresholds. Step S540: Calculate the probability distribution of the feature distance between the target response and each reference sample; Step S550: Based on the feature distance probability distribution, calculate the sample entropy of the target response; the sample entropy represents the matching degree between the target response and the real sample library of the corresponding domain of the target question.
[0062] In this embodiment, a real sample library corresponding to the target question is retrieved. For example, for a question-answering system in the medical field, a medical database containing a large number of real case records is retrieved as the real sample library. Then, the feature distance between the target answer and all samples in the real sample library is calculated. Preferably, the feature distance calculation uses a cosine similarity algorithm to measure the directional difference between the target answer and each sample in the sample library in the vector space.
[0063] Furthermore, reference samples are extracted from the sample based on each feature distance. The reference samples are those that meet preset sample conditions, including feature distances less than a preset feature distance threshold. For example, the feature distance threshold is set to 0.8, and samples with a feature distance less than 0.8 from the target response are extracted as reference samples.
[0064] Furthermore, the probability distribution of the feature distance between the target response and each reference sample is calculated. Specifically, the feature distance probability distribution is constructed using a kernel density estimation method to create a continuous probability density function. This involves dividing the feature distance into multiple intervals and counting the number of reference samples falling into each interval to obtain the probability distribution.
[0065] Furthermore, based on the feature distance probability distribution, the sample entropy of the target response is calculated. Preferably, the sample entropy calculation uses the Shannon entropy formula to integrate the probability density function, so as to characterize the matching degree between the target response and the real sample library corresponding to the target question through the sample entropy.
[0066] Therefore, this embodiment can promptly detect anomalous responses that do not conform to domain knowledge, improving the accuracy and reliability of the question-answering system. It can effectively quantify the degree of matching between the target response and real samples, providing an important basis for evaluating the reliability of responses generated by large language models. Moreover, by introducing a real sample library as a reference benchmark, it can adapt to the characteristics of different domains, improving the targeting and effectiveness of illusion recognition.
[0067] In some embodiments, in step S150, a three-dimensional entropy result of the target response is generated based on semantic entropy, available entropy, and sample entropy, so as to execute the question-answering processing scheme corresponding to the three-dimensional entropy result, including the following steps S610 to S650: Step S610: The semantic entropy, available entropy, and sample entropy are weighted and fused to obtain the three-dimensional entropy result of the target response; wherein, the weights of semantic entropy, available entropy, and sample entropy in the weighted fusion process are determined according to the needs of the domain corresponding to the target response. Step S620: Retrieve the hallucination threshold of the domain corresponding to the question-and-answer set to be processed; Step S630: Compare the three-dimensional entropy result of the target response with the illusion threshold of the corresponding domain of the question-and-answer set to be processed, and obtain the comparison result; Step S640: Execute the corresponding question-and-answer processing scheme based on the comparison results; Step S650: If the three-dimensional entropy result of the target response is lower than the illusion threshold of the corresponding domain of the question and answer set to be processed, output the target response; if the three-dimensional entropy result of the target response is higher than or equal to the illusion threshold of the corresponding domain of the question and answer set to be processed, analyze the semantic entropy, available entropy, and sample entropy of the target response to determine the illusion type of the target response; based on the illusion type of the target response, generate a question and answer anomaly prompt and prevent the output of the target response.
[0068] In this embodiment, semantic entropy, available entropy, and sample entropy are first weighted and fused to obtain the three-dimensional entropy result of the target response. The weights of semantic entropy, available entropy, and sample entropy during the weighted fusion process are determined according to the needs of the domain corresponding to the target response. For example, a higher sample entropy weight can be assigned to the medical domain to enhance the matching degree of real samples, while a higher semantic entropy weight can be assigned to the education domain to enhance logical consistency.
[0069] Further, the illusion threshold for the domain corresponding to the question-and-answer set to be processed is retrieved. When retrieving the illusion threshold, a dynamic parameter table in the domain knowledge base needs to be referenced. This parameter table automatically adjusts the threshold baseline value based on the domain data update frequency. Then, the three-dimensional entropy result of the target response is compared with the illusion threshold for the domain corresponding to the question-and-answer set to be processed, obtaining the comparison result. The comparison result triggers a differentiated processing branch. When the three-dimensional entropy result exceeds the threshold, entropy component analysis is used to locate the main source of anomalies. For example, semantic entropy anomalies correspond to logical contradiction-type illusions, which can be predicted using an entropy anomaly correspondence model; fluctuation-type illusions correspond to sample entropy anomalies, which correspond to data deviation-type illusions.
[0070] Specifically, if the three-dimensional entropy result of the target response is lower than the hallucination threshold of the corresponding domain of the question-and-answer set to be processed, the target response is output. For example, if the three-dimensional entropy result is 0.6, which is lower than the hallucination threshold of 0.7 in the medical domain, the target response is directly output. If the three-dimensional entropy result of the target response is higher than or equal to the hallucination threshold of the corresponding domain of the question-and-answer set to be processed, the semantic entropy, available entropy, and sample entropy of the target response are analyzed to determine the hallucination type of the target response. Based on the hallucination type of the target response, a question-and-answer anomaly prompt is generated and the output of the target response is prevented. For example, if the three-dimensional entropy result is 0.8, which is higher than the hallucination threshold of 0.7 in the medical domain, the entropy values of each dimension are further analyzed. If the semantic entropy is 0.9, the available entropy is 0.7, and the sample entropy is 0.8, it is determined to be a hallucination of the semantic contradiction type, an anomaly prompt of "The response has a semantic contradiction, please use with caution" is generated, and the output of the target response is prevented.
[0071] Therefore, this embodiment, by setting a domain-specific illusion threshold, can effectively identify and filter illusory responses from different domains. It can flexibly adjust the weights of entropy in each dimension according to the characteristics of different domains, achieving accurate evaluation of the target response. Furthermore, for responses exhibiting illusion, it can further analyze the type of illusion and generate targeted anomaly prompts, helping users understand potential risks. This significantly improves the reliability and security of the question-answering system, making it particularly suitable for domains with high accuracy requirements.
[0072] In some embodiments, after generating a three-dimensional entropy result of the target response based on semantic entropy, available entropy, and sample entropy in step S150, and executing the question-answering processing scheme corresponding to the three-dimensional entropy result, the following steps S710 to S740 are included: Step S710: In response to the error message from the user's abnormal feedback, obtain the illusion type of the error message; Step S720: When the hallucination type of the error message is based on semantic entropy anomaly, calculate the semantic error between the semantic entropy of the target response and the semantic entropy threshold of the domain corresponding to the target response, so as to adjust the semantic entropy calculation steps according to the semantic error; the adjustment of the semantic entropy calculation steps includes adjusting the judgment threshold of the reasonableness score output by the preset semantic analysis model in the semantic entropy. Step S730: When the illusion type of the error message is based on an anomaly in available entropy, calculate the prediction error between the available entropy of the target response and the available entropy threshold of the corresponding domain of the target response, so as to adjust the calculation steps of available entropy according to the prediction error; the adjustment of the calculation steps of available entropy includes the adjustment of the statistical standard of the dispersion in available entropy. Step S740: When the illusion type of the error message is based on sample entropy anomaly, calculate the true error between the sample entropy of the target response and the sample entropy threshold of the corresponding domain of the target response, so as to adjust the calculation steps of sample entropy according to the true error; the adjustment of the calculation steps of sample entropy includes updating the domain true sample library in the sample entropy.
[0073] In this embodiment, in response to the user's abnormal feedback error information, the hallucination type of the error information is first obtained. Specifically, when the hallucination type of the error information is based on semantic entropy anomaly, the semantic error between the semantic entropy of the target response and the semantic entropy threshold of the corresponding domain of the target response is calculated, so as to adjust the semantic entropy calculation steps according to the semantic error. The semantic error is determined by comparing the difference between the semantic entropy and the semantic entropy threshold, and the adjustment of the semantic entropy calculation steps includes dynamic correction of the reasonableness scoring judgment threshold in the preset semantic analysis model.
[0074] Furthermore, when the illusion type of the erroneous information is based on anomalies in available entropy, the prediction error between the available entropy of the target response and the available entropy threshold of the corresponding domain is calculated, and the calculation steps of available entropy are adjusted according to the prediction error. The prediction error is calculated as the deviation between available entropy and the available entropy threshold, and the adjustment of the available entropy calculation steps includes optimization of the statistical standard for the dispersion of the candidate content probability distribution.
[0075] Furthermore, when the illusion type of the erroneous information is based on sample entropy anomalies, the true error between the sample entropy of the target response and the sample entropy threshold of the corresponding domain is calculated, and the sample entropy calculation steps are adjusted according to the true error. The true error is obtained through the difference between the sample entropy and the sample entropy threshold. The adjustment of the sample entropy calculation steps includes incremental updates to the domain's real sample library, such as adding recent clinical case data in the medical field.
[0076] Therefore, this embodiment, by adjusting the calculation steps of semantic entropy, available entropy, and sample entropy for different types of hallucination anomalies, can more accurately identify and handle various hallucination problems. It can dynamically adjust the parameters of the three-dimensional entropy evaluation model based on user feedback error information, improving the accuracy and adaptability of the model's question-answering processing. Furthermore, by updating the domain-specific real-world sample library, this embodiment can adapt to parameter changes in various domains in a timely manner, enhancing the model's application value in different fields, thereby improving the overall reliability of the model's question-answering processing and reducing potential security risks in critical domains.
[0077] In some implementations, large language model (LLM) problem processing is achieved by constructing a three-dimensional entropy evaluation model. Specifically, the initial question received by the LLM and the preliminary answer generated by the LLM based on the initial question are first preprocessed to prepare for subsequent hallucination recognition. The core objective of preprocessing is to transform the raw, unstructured initial question and the preliminary answer generated by the LLM into clean, ordered, and semantically rich intermediate data, providing accurate input for subsequent semantic entropy calculation, usable entropy calculation, and sample entropy calculation. For example, the first step is text cleaning to remove noise and standardize the format. The second step is text decomposition and key information extraction. Finally, standardized intermediate data is output.
[0078] Furthermore, based on the preprocessed data, a three-dimensional entropy evaluation model of "semantic entropy + usable entropy + sample entropy" is adopted to ultimately determine whether the results output by the large model are credible. Specifically, "semantic entropy" quantifies semantic logical consistency, "usable entropy" quantifies model prediction confidence, and "sample entropy" quantifies real-world scenario adaptability, thereby achieving efficient identification of factual errors, logical contradictions, and irrelevant content in the large language model.
[0079] Specifically, the first step is the calculation of three-dimensional entropy. The three entropy values are independent of each other and can be calculated in parallel. Because each of them depends on different fields of the structured data (semantic entropy uses entity_attribute and logical connectors, usable entropy uses the verifiable unit probability of token_info, and sample entropy uses scene and sample library), there is no computational dependency. The definitions and calculation methods for each dimension are as follows: For semantic entropy, a semantic dependency graph (such as subject-verb-object relations, causal relations, and temporal relations) is constructed, and the degree of disorder between nodes (entities / concepts) and edges (relationships) in the graph is calculated. Semantic entropy is used to quantify the semantic logical consistency between the LLM-generated content and the input question and context.
[0080] Specifically, semantic entropy calculation extracts "entity-attribute-relationship" triples (such as "nifedipine tablets-dosage-5mg") from the entity_attribute of structured data and logical conjunctions (such as causal and adversative relationships) from token_info to construct a semantic dependency graph; the reasonableness score of each semantic relationship is calculated using the pre-trained model BERT-REL. Substituting into the formula, we obtain the semantic entropy (the higher the value, the more chaotic the logic), and its formula is as follows: ; in, Score the reasonableness of semantic relationships (calculated using a pre-trained semantic model such as BERT-REL). This represents the probability distribution for the reasonableness score. Higher semantic entropy indicates more chaotic semantic logic and a higher probability of logical contradictions or thematic deviation. Semantic entropy is calculated for each iteration. This represents the first element to be evaluated in the current semantic dependency graph. A semantic relation, and All of them rely on real-time data generation.
[0081] Specifically, the semantic dependency graph needs to be reconstructed before each semantic entropy calculation, rather than constructing a general graph once. This is because semantic entropy needs to quantify the semantic logical consistency between the current LLM output and the corresponding input question and context. Since the content of each LLM output (e.g., different disease descriptions in a medical scenario, different product parameters in an e-commerce scenario), the topic of the input question, and the contextual relationships are all different, a general graph cannot adapt to the personalized semantic logic verification requirements of a single calculation. Therefore, it needs to be dynamically constructed based on the dedicated structured data after each preprocessing. The specific construction method of the semantic dependency graph is as follows: 1. Node Definition and Extraction: Nodes are defined as "entities / concepts" and extracted from standardized intermediate data. For example, "core entities" (such as diseases, drugs, and products) and "attribute concepts" (such as dosage and volume) are obtained from the `entity_attribute` field, and high-semantic-value words (such as nouns and technical terms) are extracted from the `token_info` field. For instance, in a medical scenario, "hypertension (disease entity)," "nifedipine tablets (drug entity)," and "5mg (dosage concept)" are extracted as nodes.
[0082] 2. Edge (Relationship) Definition and Extraction: Using "semantic association" as the graph edge, it covers three types of relationships. For example, "subject-verb-object relationship" (e.g., "patient-take-nifedipine tablets") is extracted from the part-of-speech tagging and dependency parsing of token_info; "causal / temporal relationship" (e.g., "high blood pressure-therefore-take antihypertensive drugs") is extracted from the logic_conjunction field; and "attribute relationship" (e.g., "nifedipine tablets-dosage-5mg") is extracted from the entity_attribute field.
[0083] 3. Graph structure integration: The graph is organized using a "node-edge-attribute" ternary structure, and each edge is labeled with "relationship type" (such as subject-predicate, causality, attribute) and "initial rationality score".
[0084] Specifically, core entity / concept data is extracted directly from the `entity_attribute` (entity-attribute-relationship list) and `token_info` (entities and their confidence scores labeled with Named Entity Recognition (NER)) fields in structured data. For example, entries with non-empty NER values are selected from `token_info` (e.g., "hypertension (disease entity, confidence score 0.97)"). Semantic relationship data is extracted from `logic_conjunction` (logic conjunction position and type), dependency syntax annotations in `token_info` (implicit subject-verb-object relationship), and the `relation` field of `entity_attribute` (e.g., "treatment" "has"). For example, "therefore (causal conjunction)" is obtained from `logic_conjunction`, and entities in the preceding and following sentences are linked to form "causal relationship edges". Initial rationality score data is assigned to the graph edges based on entity label confidence and preliminary scores from the semantic association model, serving as the basis for subsequent semantic entropy calculations. The basic input.
[0085] For available entropy, verifiable units (such as entities, values, and factual statements) are first extracted from structured data. Then, the top-k probability distribution dispersion of these units is statistically analyzed. Here, available entropy is used to quantify the "effective confidence" of the predicted probability distribution during the LLM generation process. Unlike traditional single token probabilities, it focuses on the predictive stability of "verifiable information".
[0086] Specifically, usable entropy is calculated by filtering verifiable units (such as the entity "nifedipine tablets" and the numerical value "5mg") from the token_info of structured data, extracting the top-k candidate token probabilities (the LLM generation probabilities associated with the preprocessing stage) at the time of generation of each unit, and substituting the variance of the candidate probability of each unit into the formula to obtain the usable entropy (the higher the value, the less stable the model's prediction of the facts). The calculation formula is as follows: ; in, For the first When the verifiable unit is generated, the first verifiable unit is... The probability of each candidate token (normalized intermediate data). Variance. The higher the available entropy, the less stable the model's predictions of verifiable information are, and the higher the probability of factual errors. Usable entropy is a core indicator for quantifying the stability of LLM predictions based on verifiable information; a higher value indicates a more unstable prediction (higher risk of hallucination); in the formula, " "of The total number of verifiable units is calculated from the extracted entities, values, factual statements, and other units. It is not preset and changes dynamically with each LLM output.
[0087] To address sample entropy, a domain-specific real-world sample library is constructed (e.g., disease diagnosis cases in the medical field, compliance documents in the financial field). The feature distance distribution entropy between the generated text and similar texts in the sample library is calculated. Sample entropy is used to quantify the matching degree between the LLM-generated content and the real-world sample library, resolving the illusion of "model confidence but disconnect from reality."
[0088] Specifically, the sample entropy calculation uses the scene identifier (e.g., "medical") of the structured data to call the corresponding real sample library in the domain. Then, it uses Sentence-BERT to calculate the feature distance between the generated text and the top-5 similar samples in the library, calculates the probability distribution of the distance, and substitutes it into the formula to obtain the sample entropy (the higher the value, the greater the deviation from the real scene). The calculation formula is as follows: ; in, To generate text and the first The feature distance between similar samples (calculated using Sentence-BERT). This represents the probability distribution of the feature distance. Higher sample entropy indicates a greater deviation between the generated content and the real scene, and a higher probability of the presence of fictitious information. Sample entropy is a core metric for quantifying the matching degree between LLM-generated content and a real-world sample library. A higher value indicates a more dispersed distribution of feature distances between the generated text and similar texts in the sample library, resulting in a lower matching degree and a higher risk of the illusion that the model is "confident but out of touch with reality." Conversely, a lower value indicates a higher matching degree and stronger content credibility. The term refers to the "Top-m similarity number" used in feature distance calculation, which is the number of samples with the highest semantic similarity to the generated text selected from a domain-specific real-world sample library. One sample (document default) =5, which can be adjusted according to the needs of different fields. For example, in the medical field, it can be set to ensure accuracy. =8). The parameters need to be dynamically determined in conjunction with the size of the sample library to ensure that enough reference samples are covered and to avoid the influence of single sample bias on the calculation results. These parameters are the basis for constructing the feature distance distribution.
[0089] In this context, "similar samples" refers to the real data within the "domain-specific real sample library" mentioned earlier in the sample entropy calculation logic. Specifically, it refers to the samples selected from this library that have the highest semantic relevance to the currently generated LLM text. This sample library is a dedicated real dataset pre-built for scenarios such as education, healthcare, and finance. For example, it includes disease diagnosis cases and drug instructions in the healthcare scenario, and compliance documents and product descriptions in the financial scenario. All of these have undergone authoritative verification (e.g., medical samples conform to clinical guidelines, and financial samples come from regulatory agencies). During the calculation, the LLM-generated text and all samples in the sample library are first converted into semantic feature vectors using the Sentence-BERT model. Then, the samples are sorted by indicators such as cosine similarity to select the top-m samples (m=5 by default for documents, but can be adjusted). These highly relevant real samples selected are the "similar samples." That is, generating text and the first one of them indivual( Values range from 1 to The feature distance of the sample.
[0090] Furthermore, the entropy value fusion decision is made by first standardizing the entropy values and assigning weights. Specifically, since the three-dimensional entropy values are calculated in different dimensions (semantic entropy is based on logical association, usable entropy is based on probability concentration, and sample entropy is based on scene matching degree), the three must first be standardized to the [0,1] interval to eliminate the difference in dimensions.
[0091] For example, using min-max normalization, the formula is as follows: ; in, To standardize the processing results, This corresponds to the maximum value of the entropy in the historical data. This corresponds to the minimum entropy value in historical data.
[0092] Furthermore, differentiated weights are assigned to the three-dimensional entropy values based on the needs of different application scenarios. For example, in the medical scenario, "sample entropy" (the degree of matching with real medical knowledge) has the highest weight (0.4) because errors in medical facts can endanger lives; in the e-commerce scenario, "semantic entropy" (the logical coherence of customer service scripts) has a slightly higher weight (0.35) to avoid affecting the user experience due to logical confusion; in the education scenario, "usable entropy" (the model's confidence in knowledge points) has an increased weight (0.35) to help identify "seemingly correct but actually incorrect" knowledge point outputs from the model.
[0093] Furthermore, the fusion decision formula and threshold determination are determined. Specifically, the "comprehensive reliability score" output by the large model is calculated through weighted summation, and its formula is as follows: ; in, To calculate the overall reliability score, The weights of semantic entropy, The weights are the available entropy. The weights are the sample entropy, and , This is the standardized semantic entropy. This represents the standardized usable entropy. The standardized sample entropy indicates that the higher the overall reliability score, the stronger the reliability and the lower the probability of hallucination.
[0094] Furthermore, based on historical labeled data (text already labeled "hallucination / non-hallucination"), ROC curve analysis is used to set "hallucination judgment thresholds" for different scenarios. For example, in the medical scenario, a score < 0.6 indicates the presence of hallucination; in the e-commerce scenario, a score < 0.5 indicates the presence of hallucination. Simultaneously, the "abnormal fluctuations" of single-dimensional entropy values are used to assist in the judgment: if the overall score of a text segment does not reach the threshold, and the entropy value of a certain dimension suddenly increases (e.g., semantic entropy increases by 2 times compared to the previous segment), the type of hallucination (e.g., logical contradiction) can be accurately located.
[0095] Furthermore, the real sample database needs to be updated and maintained regularly. This real sample database is the core support for calculating "sample entropy," and its sample quality (authenticity, domain suitability, and timeliness) directly determines the accuracy of sample entropy in identifying "factual errors / irrelevant content illusions." The specific update steps are as follows: (1) Based on the data source list planned in the early stage, the raw data is obtained in a targeted manner by combining "automatic collection + manual collection + API connection". Specifically, the data is collected only by "file MD5 value comparison" or "update timestamp" (such as replacing old parameter samples after e-commerce product parameters are updated, without duplicate entry into the database).
[0096] (2) The raw data often has problems such as “disordered format, redundant information, and incorrect expression” (such as headers and footers in PDF guidelines and emoticons in UGC evaluations). It needs to be cleaned through three steps: “format standardization → redundancy filtering → error correction” to output “clean raw text”.
[0097] (3) Verification is the core step in "eliminating falsehoods and retaining truth" of samples. It is necessary to combine "authoritative cross-validation + domain rule verification + manual sampling review" to ensure that the samples meet the scene authenticity standards 100%.
[0098] (4) Store the structured sample units in the sample management library and build a "multi-dimensional index" to ensure that the sample entropy calculation can be matched quickly (e.g., input "nifedipine tablets 5mg", and the corresponding drug sample can be retrieved within 100ms). Moreover, the update cycle can be set according to the scenario.
[0099] Furthermore, the parameters of the 3D entropy calculation and entropy fusion parts are dynamically iterated and maintained based on user feedback (misjudgment cases, etc.). User feedback (especially misjudgment cases) is the core basis for exposing the "parameter bias" and "insufficient scene adaptation" of the 3D entropy model. For example, in the medical scenario, the misjudgment of "LLM outputting the wrong drug dosage but not being recognized" may be due to the imbalance of similarity weights of sample entropy; in the e-commerce scenario, "logical contradictions are judged as normal" may be due to the excessively narrow temporal correlation window of semantic entropy.
[0100] Specifically, the first step is to pinpoint the root cause of misjudgments. This is done through "case reproduction → dimensional decomposition → cross-validation" to accurately determine whether the misjudgment stems from "parameter deviation in the 3D entropy calculation module" or "weight imbalance in the fusion module," thus avoiding blind parameter tuning. The steps are as follows: I. Case Reproduction and Data Backtracking. First, reproduce the misjudgment scenario and backtrack the key intermediate data in the system's calculation process to identify "which step went wrong." Specifically, first extract the intermediate results of semantic entropy, available entropy, and sample entropy calculations for the misjudged case—for example, in a medical missed case, backtrack the "similarity calculation log" of the sample entropy to see if the system matched samples related to "aspirin dosage," and what the "cosine semantic similarity" and "BM25 keyword similarity" were at the time of matching (if only old samples from 3 years ago were matched, and the dosage was 3g, it may be that the timeliness weight was not in effect).
[0101] Furthermore, examine the "weights of each dimension" used in calculating misjudged cases (e.g., the weights in the medical scenario were: semantic entropy 0.25, available entropy 0.35, and sample entropy 0.4) to determine whether the key errors were not amplified due to the low weight of a certain dimension (e.g., a sample entropy weight of 0.4 is still insufficient to cover dosage errors and needs to be further increased).
[0102] II. Three-Dimensional Entropy Dimensional Decomposition and Analysis. For each entropy dimension, we verify one by one whether there are parameter biases leading to misjudgments. For example, the core analysis logic and case are as follows: Regarding semantic entropy, the analysis focused on whether the semantic association probability threshold was too low (e.g., a normal association probability ≥ 0.8, but the actual setting was 0.6, leading to unidentified logical contradictions), whether the temporal association window was too narrow (e.g., cross-round logical association only considered 3 tokens, not covering long text logic), and whether the domain terminology association model was suitable (e.g., the medical term "overdose" was not associated with "dosage error"). The analysis revealed that the association probability of "aspirin-5g dose" in the semantic entropy calculation was 0.75 (higher than the then-current threshold of 0.7), thus it was judged as "logically coherent" and had no issues; the temporal association window covered the complete sentence, with no cross-round omissions—eliminating the possibility of semantic entropy parameter bias.
[0103] Regarding the available entropy, the analysis focused on whether the probability concentration threshold was too high (e.g., a concentration ≥ 0.6 is considered confidence, but a setting of 0.5 resulted in the failure to identify "false confidence"), whether the sliding window size was too small (e.g., window = 3, failing to capture the trend changes in token probabilities), and whether the candidate probability and distribution entropy thresholds were too low (e.g., distribution entropy ≥ 0.6 is marked as anomaly, but a setting of 0.7 resulted in the failure to identify uniform distributions). Analysis revealed that when LLM generated "5g," the probability concentration was 0.82 (above the threshold of 0.6), and the candidate probability distribution entropy was 0.35 (below the threshold of 0.7). The system judged this as "model confidence"—the available entropy was normal, thus eliminating the possibility of available entropy parameter bias.
[0104] For sample entropy, analyze whether the similarity calculation weights are unbalanced (e.g., overemphasis on semantic similarity and neglect of keyword / numerical matching), whether the weight of sample timeliness is too low (e.g., old sample weight = 0.8, which does not significantly reduce the influence of old samples), and whether the sample library coverage is insufficient (e.g., no latest sample of aspirin "4g maximum dose"). Analysis revealed that during sample entropy matching, the cosine semantic similarity was 0.9 (high, due to semantic matching of "aspirin-dosage"), and the BM25 keyword similarity was 0.3 (low, due to the mismatch between "5g" and the "4g" value in the sample library). However, the BM25 weight was only 0.3, resulting in a total similarity of 0.9 × 0.7 + 0.3 × 0.3 = 0.72 (above the threshold of 0.6, thus considered a match). Furthermore, the matched sample was from the 2021 guidelines (timeliness weight = 0.9, which did not reduce the impact), and there were no updated "4g maximum dose" samples from 2023. The root cause was identified as the "BM25 weight being too low" and the "sample timeliness weight being too high" in the sample entropy.
[0105] III. Cross-validation of fusion module weights. If there is no significant deviation in the 3D entropy calculation, it is necessary to verify whether the anomalies in key dimensions have not been amplified due to "fusion weight imbalance": 1. Weight Sensitivity Analysis: In misjudgment cases, the three-dimensional entropy values are fixed, and the weights of each dimension are adjusted to observe changes in the overall score. For example, in a medical case, the sample entropy weight is increased from 0.4 to 0.5, and the overall score is recalculated. The original score was 1. (0.25×0.32+0.35×0.28+0.4×0.35)=0.72 (normal); Adjusted score: 1 (0.25×0.32+0.35×0.28+0.5×0.35)=0.685 (still higher than the threshold of 0.6, unresolved); further increase the sample entropy weight to 0.6, score: 1 (0.25×0.32+0.35×0.28+0.6×0.35)=0.65 (close to the threshold); if the BM25 weight of the sample entropy is adjusted simultaneously (increased from 0.3 to 0.5), the sample entropy value increases from 0.35 to 0.6, at which point the score is 1. (0.25×0.32+0.35×0.28+0.6×0.6)=0.52 (below the threshold of 0.6, judged as hallucination), thus the misjudgment can be resolved by verifying "sample entropy parameter adjustment + weight fine-tuning".
[0106] 2. Cross-scenario weight comparison: Compare the weight settings of similar cases in other scenarios (e.g., the sample entropy weight of "product parameter value error" in e-commerce scenario is 0.55) to determine whether the current scenario weight is too low (the sample entropy weight of 0.4 for "dosage value error" in medical scenario is indeed too low).
[0107] IV. Parameter Iteration of the 3D Entropy Calculation and Fusion Module. Based on the root cause localization results, targeted parameter tuning was performed on the "3D Entropy Calculation Module" and the "Fusion Module". Specifically, for the 3D entropy calculation module, adjustments were made to the calculation parameters of semantic entropy, available entropy, and sample entropy. For the fusion module, the basic scene weights were adjusted based on the "frequency of misjudgment in each dimension" in the feedback. Semantic entropy weight, Entropy weights can be used. =Sample entropy weight), to ensure that the weight of dimensions with high misjudgment is increased.
[0108] Furthermore, regarding the semantic association probability threshold used in semantic entropy calculation to determine whether the "association of semantic units is reasonable", the value is adjusted when the threshold is too low, resulting in the failure to identify logical contradictions, or the threshold is too high, resulting in normal text being misjudged as contradictions.
[0109] like Figure 2 As shown in some embodiments of this application, a question-answering processing system based on three-dimensional entropy evaluation is provided. The system includes a first module 210, a second module 220, an extraction module 230, a calculation module 240, and an execution module 250. Specifically: The first module 210 is used to construct a set of questions and answers to be processed based on the target question, the target answer, and the historical context; the target answer is generated by the large language model based on the target question; the historical context is the contextual content associated with the target question; The second module 220 is used to construct a corresponding semantic graph based on the question-and-answer set to be processed; each node in the semantic graph is obtained based on the target entities in the question-and-answer set to be processed, and each edge in the semantic graph is obtained based on the target attributes and target relationships in the question-and-answer set to be processed; the target relationships represent the logical relationships between the target entities in the question-and-answer set to be processed. Extraction module 230 is used to extract multiple combinations to be verified from the question and answer set to be processed; wherein each combination to be verified consists of a target entity, a target attribute, and a target attribute value in the question and answer set to be processed; The computation module 240 is used to calculate the semantic entropy, available entropy, and sample entropy of the target response based on the semantic graph and each combination to be verified. The semantic entropy is used to quantify the logical consistency of the target response, the available entropy is used to quantify the predictive stability of the large language model when generating the target response based on the target question, and the sample entropy is used to quantify the matching degree between the target response and the real sample library of the corresponding domain of the target question. The execution module 250 is used to generate a three-dimensional entropy result of the target response based on semantic entropy, available entropy and sample entropy, so as to execute the question-answering processing scheme corresponding to the three-dimensional entropy result.
[0110] It should be noted that the question-answering system based on three-dimensional entropy assessment provided in this embodiment is based on the same inventive concept as the question-answering method based on three-dimensional entropy assessment described above. Therefore, the relevant content of the question-answering method based on three-dimensional entropy assessment described above is also applicable to the content of the question-answering system based on three-dimensional entropy assessment. Therefore, it will not be repeated here.
[0111] To address this, the system constructs a question-and-answer set to be processed based on the target question, target answer, and historical context. Based on this set, a corresponding semantic graph is built, extracting multiple combinations to be verified. Using the semantic graph and each combination to be verified, the semantic entropy, usable entropy, and sample entropy of the target answer are calculated. Finally, a three-dimensional entropy result for the target answer is generated based on these entropies, and the question-and-answer processing scheme corresponding to the three-dimensional entropy result is executed. In this way, a three-dimensional evaluation system can be implemented by constructing a semantic graph and calculating semantic entropy, usable entropy, and sample entropy. Combined with dynamic threshold comparison and feedback adjustment mechanisms, this system comprehensively detects and addresses semantic contradictions, prediction anomalies, and sample deviations, thereby improving the reliability of the model-based question-and-answer system.
[0112] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described question-and-answer processing method based on three-dimensional entropy evaluation.
[0113] like Figure 3 , Figure 3 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. The electronic device includes: At least one battery; At least one memory; At least one processor; At least one program; The program is stored in memory, and the processor executes at least one program to implement the question-answering processing method based on three-dimensional entropy evaluation described above in this disclosure.
[0114] This electronic device can be any smart terminal, including mobile phones, tablets, personal digital assistants (PDAs), and in-vehicle computers.
[0115] The electronic devices according to embodiments of this application will now be described in detail.
[0116] The processor 1600 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this disclosure. The memory 1700 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1700 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1700 and is called and executed by the processor 1600 to execute a question-answering processing method based on three-dimensional entropy evaluation according to an embodiment of this disclosure.
[0117] The input / output interface 1800 is used to implement information input and output. The communication interface 1900 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 2000 transmits information between various components of the device (e.g., processor 1600, memory 1700, input / output interface 1800, and communication interface 1900); The processor 1600, memory 1700, input / output interface 1800 and communication interface 1900 are connected to each other within the device via bus 2000.
[0118] This disclosure also provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-described question-answering processing method based on three-dimensional entropy assessment.
[0119] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0120] The embodiments described in this disclosure are for the purpose of more clearly illustrating the technical solutions of this disclosure and do not constitute a limitation on the technical solutions provided by this disclosure. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by this disclosure are also applicable to similar technical problems.
[0121] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this disclosure, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0122] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0123] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0124] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any related variations, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0125] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0126] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0127] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0128] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0129] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0130] The above is a detailed description of the preferred embodiments of this application. However, the embodiments of this application are not limited to the above-described implementation methods. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the embodiments of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of the embodiments of this application.
[0131] The embodiments of this application have been described in detail above with reference to the accompanying drawings. However, this application is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of this application.
Claims
1. A question-answering processing method based on three-dimensional entropy evaluation, characterized in that, The method includes: A question-and-answer set to be processed is constructed based on the target question, the target answer, and the historical context; the target answer is generated by the large language model based on the target question; the historical context is the contextual content associated with the target question. Based on the question-and-answer set to be processed, a corresponding semantic graph is constructed; each node in the semantic graph is obtained based on the target entities in the question-and-answer set to be processed, and each edge in the semantic graph is obtained based on the target attributes and target associations in the question-and-answer set to be processed; the target associations represent the logical relationships between the target entities in the question-and-answer set to be processed. Multiple combinations to be verified are extracted from the question and answer set to be processed; wherein each combination to be verified consists of a target entity in the question and answer set to be processed, a target attribute in the question and answer set to be processed, and a target attribute value in the question and answer set to be processed. Based on the semantic graph and each of the combinations to be verified, the semantic entropy, usable entropy, and sample entropy of the target response are calculated respectively; the semantic entropy is used to quantify the logical consistency of the target response, the usable entropy is used to quantify the predictive stability of the large language model when generating the target response based on the target question, and the sample entropy is used to quantify the matching degree between the target response and the real sample library of the domain corresponding to the target question; Based on the semantic entropy, the available entropy, and the sample entropy, a three-dimensional entropy result of the target response is generated, and the question-answering processing scheme corresponding to the three-dimensional entropy result is executed.
2. The question-answering processing method based on three-dimensional entropy evaluation according to claim 1, characterized in that, The step of constructing a corresponding semantic graph based on the question-and-answer set to be processed includes: Extract the target entity, the target attribute, and the target relationship from the question-and-answer set to be processed; the target entity represents a domain-specific object in the question-and-answer set to be processed; the target attribute represents the feature of the target entity; the target relationship includes subject-verb-object relationship, causal relationship, temporal relationship, and attribute relationship between the target entities; The target entity and the target attribute are used as nodes in the semantic graph; The target association is used as the edge of the semantic graph, and the confidence of each edge in the semantic graph is calculated based on the labeled confidence of the target entity and the labeled confidence of the target attribute. The semantic graph is obtained by integrating the nodes, edges, and confidence scores of each edge in the semantic graph.
3. The question-answering processing method based on three-dimensional entropy evaluation according to claim 2, characterized in that, The semantic entropy of the target response is calculated through the following steps: Based on the confidence level of each edge in the semantic graph, the target association relationship of each edge in the semantic graph is analyzed according to the preset semantic analysis model to obtain the rationality score of each target association relationship; Based on the reasonableness score of each target association, the degree of confusion between each node and each edge in the semantic graph is calculated; the degree of confusion represents the degree of semantic and logical consistency between the target response, the target question, and the historical context. Based on the degree of confusion, the semantic entropy of the target response is calculated; the semantic entropy represents the probability that the target response contains a semantic contradiction.
4. The question-answering processing method based on three-dimensional entropy evaluation according to claim 1, characterized in that, The available entropy of the target response is calculated through the following steps: Retrieve the probability distribution of candidate content for each of the combinations to be verified during the process of generating the target response by the large language model; the probability distribution of candidate content is the probability distribution of multiple alternative contents when the large language model generates the combination to be verified. The dispersion of the probability distribution of candidate content corresponding to each of the aforementioned combinations to be verified is statistically analyzed. Based on the probability distribution dispersion of each of the combinations to be verified, the available entropy of the target response is calculated; the available entropy characterizes the probability degree of prediction anomalies in the target response.
5. The question-answering processing method based on three-dimensional entropy evaluation according to claim 1, characterized in that, The sample entropy of the target response is calculated through the following steps: Retrieve the real sample library corresponding to the target problem; Calculate the feature distance between the target response and all samples in the real sample library of the domain; Based on the aforementioned feature distances, reference samples are extracted from the samples; The reference sample is one of the samples that meets the preset sample conditions; the preset sample conditions include the feature distance being less than a preset feature distance threshold. Calculate the probability distribution of the feature distances between the target response and each of the reference samples; Based on the feature distance probability distribution, the sample entropy of the target response is calculated; the sample entropy represents the degree of matching between the target response and the real sample library of the corresponding domain of the target question.
6. The question-answering processing method based on three-dimensional entropy evaluation according to claim 1, characterized in that, The step of generating a three-dimensional entropy result of the target response based on the semantic entropy, the available entropy, and the sample entropy, and executing the question-answering processing scheme corresponding to the three-dimensional entropy result, includes: The semantic entropy, the available entropy, and the sample entropy are weighted and fused to obtain the three-dimensional entropy result of the target response; wherein, the weights of the semantic entropy, the available entropy, and the sample entropy in the weighted fusion process are determined according to the needs of the domain corresponding to the target response. Retrieve the hallucination threshold of the domain corresponding to the question-and-answer set to be processed; The three-dimensional entropy result of the target response is compared with the illusion threshold of the corresponding domain of the question-and-answer set to be processed to obtain the comparison result; Execute the corresponding question-and-answer processing scheme based on the comparison results; If the three-dimensional entropy result of the target response is lower than the illusion threshold of the domain corresponding to the question-and-answer set to be processed, the target response is output; if the three-dimensional entropy result of the target response is higher than or equal to the illusion threshold of the domain corresponding to the question-and-answer set to be processed, the semantic entropy, the available entropy, and the sample entropy of the target response are analyzed to determine the illusion type of the target response; based on the illusion type of the target response, a question-and-answer anomaly prompt is generated and the output of the target response is prevented.
7. The question-answering processing method based on three-dimensional entropy evaluation according to claim 6, characterized in that, After generating a three-dimensional entropy result of the target response based on the semantic entropy, the available entropy, and the sample entropy, and executing the question-answering processing scheme corresponding to the three-dimensional entropy result, the method further includes: In response to error messages from abnormal user feedback, the hallucination type of the error message is obtained; When the hallucination type of the erroneous information is based on the semantic entropy anomaly, the semantic error between the semantic entropy of the target response and the semantic entropy threshold of the domain corresponding to the target response is calculated, so as to adjust the calculation steps of the semantic entropy according to the semantic error; the adjustment of the calculation steps of the semantic entropy includes adjusting the judgment threshold of the reasonableness score output by the preset semantic analysis model in the semantic entropy; When the illusion type of the error message is based on the anomaly of the available entropy, the prediction error between the available entropy of the target response and the available entropy threshold of the corresponding domain of the target response is calculated, so as to adjust the calculation steps of the available entropy according to the prediction error; the adjustment of the calculation steps of the available entropy includes the adjustment of the statistical standard of the dispersion in the available entropy; When the illusion type of the error message is based on the abnormality of the sample entropy, the true error between the sample entropy of the target response and the sample entropy threshold of the corresponding domain of the target response is calculated, so as to adjust the calculation steps of the sample entropy according to the true error; the adjustment of the calculation steps of the sample entropy includes updating the domain true sample library in the sample entropy.
8. A question-answering processing system based on three-dimensional entropy evaluation, characterized in that, The system includes: The first module is used to construct a set of questions and answers to be processed based on the target question, the target answer, and the historical context; the target answer is generated by the large language model based on the target question; the historical context is the context content associated with the target question. The second module is used to construct a corresponding semantic graph based on the question-and-answer set to be processed; each node in the semantic graph is obtained based on the target entities in the question-and-answer set to be processed, and each edge in the semantic graph is obtained based on the target attributes and target associations in the question-and-answer set to be processed; the target associations represent the logical relationships between the target entities in the question-and-answer set to be processed. An extraction module is used to extract multiple combinations to be verified from the question-and-answer set to be processed; wherein each combination to be verified consists of a target entity in the question-and-answer set to be processed, a target attribute in the question-and-answer set to be processed, and a target attribute value in the question-and-answer set to be processed; The calculation module is used to calculate the semantic entropy, usable entropy, and sample entropy of the target response based on the semantic graph and each of the combinations to be verified. The semantic entropy is used to quantify the logical consistency of the target response, the usable entropy is used to quantify the predictive stability of the large language model when generating the target response based on the target question, and the sample entropy is used to quantify the matching degree between the target response and the real sample library of the domain corresponding to the target question. The execution module is used to generate a three-dimensional entropy result of the target response based on the semantic entropy, the available entropy, and the sample entropy, so as to execute the question-answering processing scheme corresponding to the three-dimensional entropy result.
9. An electronic device, characterized in that, It includes at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, which, when executed by the at least one control processor, enables the at least one control processor to perform a question-answering processing method based on three-dimensional entropy assessment as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform a question-answering method based on three-dimensional entropy assessment as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Three-dimensional back projection imaging method and device of array interference SAR (Synthetic Aperture Radar)
CN119247361A
Online medical question and answer dynamic retrieval enhancement generation method
CN120045662A
Medical question and answer method based on big language model illusion detection
CN120144701A
Intelligent operation and maintenance question-answering system for cable manufacturing equipment
CN120561253A
Semantic map production system and method
US20220156536A1