A Method, Device, Equipment and Storage Medium for Eliminating Hallucinations in Large Language Models

By combining semantic analysis and probability statistics, a target noun association network in a large language model is constructed and the entropy difference value is calculated, the hallucination risk type is determined and optimized, and the dependence and generalization ability of the hallucination problem of the large language model in the existing technology is solved, and more reliable and generalized hallucination processing is achieved.

CN119849509BActive Publication Date: 2025-06-27SHANDONG INSPUR SCI RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510337458.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-27
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

Existing large language models have hallucinations problems in text generation and processing, such as factual errors, logical contradictions, semantic biases and non-fidelity, and rely on external knowledge bases for hallucination detection, resulting in unreliability in the process and weak generalization ability.

Method used

By combining semantic analysis with probability statistics, a correlation network between target nouns is constructed, entropy difference values ​​are calculated to determine the type of hallucination risk, and the text is optimized using corresponding optimization strategies to eliminate hallucinations.

Benefits of technology

It realizes the generalization ability and reliability of hallucination processing methods without relying on external knowledge base, and fundamentally eliminates the hallucination risks in the analysis results of large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119849509B_ABST
    Figure CN119849509B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, equipment and storage medium for eliminating hallucinations in large language models, which relates to the technical field of natural language processing and includes: constructing an association network among target nouns according to the association relationship among the target nouns, and obtaining the association paths corresponding to the target nouns according to the association network; determining a reference path from each association path, and obtaining the entropy difference value between the first target entropy corresponding to each association path and the second target entropy corresponding to the reference path; determining the hallucination risk type corresponding to the initial text to be analyzed based on the entropy difference value, and optimizing the initial text to be analyzed by using an optimization strategy corresponding to the hallucination risk type, so as to eliminate the hallucinations in the analysis result after analyzing the initial text to be analyzed by a preset large language model. By combining semantic analysis with probability statistics, it is ensured that no external knowledge base is relied on during the hallucination analysis process, and the generalization ability and reliability of the hallucination processing method are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly relates to a method, device, equipment and storage medium for eliminating hallucinations in large language models. Background Art

[0002] Large language models perform well in tasks such as text generation, question answering, summarization, code generation, etc., but there are still some hallucination problems, including factual errors (the generated content does not match the real-world facts), logical contradictions (the internal information of the text is inconsistent), semantic deviations (the generated text fails to accurately express the user's intention), and non-faithfulness (in summarization or translation tasks, the generated content does not match the source text).

[0003] Currently, hallucination detection methods mainly rely on comparison with external knowledge bases. This method has problems such as unreliable hallucination detection processes, being limited by the coverage of the knowledge base, and weak generalization ability. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method, device, equipment and storage medium for eliminating hallucinations in large language models, which can combine semantic analysis with probability statistics, so that there is no need to rely on external knowledge bases during the hallucination analysis process, ensuring the generalization ability and reliability of the hallucination processing method. The specific solutions are as follows:

[0005] In a first aspect, the present application provides a method for eliminating hallucinations in large language models, including:

[0006] Analyze the initial text to be analyzed to obtain target nouns, and encode the feature maps corresponding to each of the target nouns into target vectors; the feature maps include noun concepts, semantic labels, and context information corresponding to each of the target nouns;

[0007] According to the association relationships between the target nouns determined based on the target vectors, construct an association network between the target nouns, and obtain the association paths corresponding to each of the target nouns according to the association network; the nodes in the association network represent the target nouns, the directed edges between the nodes represent the association relationships between the target nouns, and the weights of the directed edges represent the association strengths between the nodes;

[0008] Determine a reference path from each of the associated paths, and obtain the entropy difference value between the first target entropy corresponding to each of the associated paths and the second target entropy corresponding to the reference path; the target entropy includes a feature activation entropy, a path selection entropy, and an output distribution entropy. The feature activation entropy characterizes the uncertainty of the activation distribution of the features corresponding to the target noun in the associated path. The path selection entropy characterizes the uncertainty of the path selection from the path start point to the path end point in the associated path. The output distribution entropy characterizes the uncertainty of the distribution of the output results generated by the preset large language model based on the associated path.

[0009] Based on the entropy difference value, determine the hallucination risk type corresponding to the initial text to be analyzed, and use an optimization strategy corresponding to the hallucination risk type to optimize the initial text to be analyzed, so as to eliminate the hallucination in the analysis result after the preset large language model analyzes the initial text to be analyzed.

[0010] Optionally, the analyzing the initial text to be analyzed to obtain the target noun includes:

[0011] Use a preset syntax analysis technique to identify the initial text to be analyzed, so as to obtain the first initial nouns in the initial text to be analyzed, and evaluate the semantic importance of the first initial nouns, so as to determine the second initial nouns from each of the first initial nouns;

[0012] Judge whether the second initial noun is a complete noun. If the second initial noun is a complete noun, determine the second initial noun as the target noun;

[0013] If the second initial noun is not a complete noun, merge each of the second initial nouns to obtain a merged noun, and determine the merged noun as the target noun.

[0014] Optionally, before encoding the feature maps corresponding to each of the target nouns into target vectors, it further includes:

[0015] Input different text segments including the same target noun into the preset large language model, so that the preset large language model generates different first output results according to different text segments;

[0016] Obtain each of the first output results and obtain the first output change situation of the preset large language model according to each of the first output results;

[0017] Modify any target noun in the initial text to be analyzed, and input the corresponding modified text into the preset large language model to obtain the second output result output by the preset large language model, and obtain the second output change situation according to the second output result;

[0018] Input text fragments including the same target noun in the context into the preset large language model, and according to the third output result of the preset large language model;

[0019] Obtain the understanding degree of the preset large language model for each target noun according to the first output change situation, the second output change situation and the third output result, and generate respective feature maps corresponding to each target noun according to the understanding degree.

[0020] Optionally, the method for eliminating large language model hallucinations further includes: obtaining any two mutually related target nouns according to the association network, and obtaining the association strength between any two associated target nouns according to the feature overlap degree, context distance and co-occurrence frequency between the two target nouns; wherein, the feature overlap degree is the quotient of the target difference and the total number of features of any two target nouns, and the target difference is the difference between the first number of common features and the second number of mutually exclusive features between any two target nouns.

[0021] Optionally, the hallucination risk types include direct conflict, indirect conflict, context conflict and weight imbalance. The direct conflict indicates that the feature overlap degree between any two directly related target nouns is negative. The indirect conflict indicates that the semantics of any two target nouns connected by the same intermediate node are contradictory. The context conflict indicates that the difference in weight values between the edges formed by the same target noun in different positions in the association network is greater than the preset weight value difference threshold. The weight imbalance indicates that the weight value of the edge corresponding to the preset secondary noun is higher than the weight value of the edge corresponding to the preset primary noun.

[0022] Optionally, optimizing the initial text to be analyzed by using an optimization strategy corresponding to the hallucination risk type includes:

[0023] If the hallucination risk type is the context conflict, increase the usage times of the first target noun in the initial text to be analyzed and increase the association between the first target noun and the corresponding context; wherein, the first target noun is the target noun corresponding to the context conflict.

[0024] If the hallucination risk type is the direct conflict, add a target modifier before the third target noun; wherein, the third target noun is the target noun corresponding to the direct conflict, and the target modifier is a modifier that limits the third target noun.

[0025] If the hallucination risk type is the indirect conflict, adjust the position of the third target noun in the initial text to be analyzed, and add transitional text between each of the third target nouns; wherein, the third target noun is the target noun corresponding to the indirect conflict.

[0026] If the hallucination risk type is the weight imbalance, weaken the first representation intensity of the preset secondary noun in the initial text to be analyzed, and strengthen the second representation intensity of the preset primary noun in the initial text to be analyzed.

[0027] Optionally, after optimizing the initial text to be analyzed by using the optimization strategy corresponding to the hallucination risk type, the method further includes:

[0028] Input the optimized text to be analyzed into the preset large language model, obtain the optimization effect corresponding to the optimized text to be analyzed according to the fourth output result of the preset large language model, and adjust the optimization parameters corresponding to the optimization strategy according to the optimization effect.

[0029] In a second aspect, the present application provides a large language model hallucination elimination device, including:

[0030] A feature map acquisition module, configured to analyze the initial text to be analyzed to obtain target nouns, and encode the feature maps corresponding to each of the target nouns into target vectors; the feature maps include noun concepts, semantic labels, and context information corresponding to each of the target nouns.

[0031] An association path acquisition module, configured to construct an association network between each of the target nouns according to the association relationship between each of the target nouns determined based on the target vectors, and obtain the association path corresponding to each of the target nouns according to the association network; the nodes in the association network represent each of the target nouns, the directed edges between each node represent the association relationship between each of the target nouns, and the weights of each directed edge represent the association strength between the nodes.

[0032] An entropy difference acquisition module, configured to determine a reference path from each of the association paths, and obtain the entropy difference value between the first target entropy corresponding to each of the association paths and the second target entropy corresponding to the reference path; the target entropy includes a feature activation entropy, a path selection entropy, and an output distribution entropy, the feature activation entropy characterizes the uncertainty of the activation distribution of the features corresponding to the target noun in the association path, the path selection entropy characterizes the uncertainty of the path selection from the path start point to the path end point in the association path, and the output distribution entropy characterizes the uncertainty of the distribution of the output results generated by the preset large language model based on the association path.

[0033] A text optimization module, configured to determine the hallucination risk type corresponding to the initial text to be analyzed based on the entropy difference value, and optimize the initial text to be analyzed by using an optimization strategy corresponding to the hallucination risk type, so as to eliminate hallucinations in the analysis result after the preset large language model analyzes the initial text to be analyzed.

[0034] In a third aspect, the present application provides an electronic device, including:

[0035] A memory, configured to store a computer program;

[0036] A processor, configured to execute the computer program to implement the foregoing large language model hallucination elimination method.

[0037] In a fourth aspect, the present application provides a computer-readable storage medium, configured to store a computer program, where the computer program, when executed by a processor, implements the foregoing large language model hallucination elimination method.

[0038] In this application, first, the initial text to be analyzed is analyzed to obtain target nouns, and the feature maps corresponding to each of the target nouns are encoded as target vectors; the feature maps include noun concepts, semantic tags, and context information corresponding to each of the target nouns. Then, based on the association relationships between the target nouns determined based on the target vectors, an association network between the target nouns is constructed, and association paths corresponding to each of the target nouns are obtained according to the association network; the nodes in the association network represent the target nouns, the directed edges between the nodes represent the association relationships between the target nouns, and the weights of the directed edges represent the association strengths between the nodes. After that, a reference path is determined from each of the association paths, and the entropy difference value between the first target entropy corresponding to each of the association paths and the second target entropy corresponding to the reference path is obtained; the target entropy includes feature activation entropy, path selection entropy, and output distribution entropy. The feature activation entropy characterizes the uncertainty of the activation distribution of the features corresponding to the target nouns in the association path, the path selection entropy characterizes the uncertainty of the path selection from the start point to the end point of the path in the association path, and the output distribution entropy characterizes the uncertainty of the distribution of the output results generated by the preset large language model based on the association path. Finally, based on the entropy difference value, the hallucination risk type corresponding to the initial text to be analyzed is determined, and the initial text to be analyzed is optimized using an optimization strategy corresponding to the hallucination risk type, so as to eliminate the hallucinations in the analysis results after the preset large language model analyzes the initial text to be analyzed. It can be seen that through semantic analysis of the text to be analyzed, this application can obtain the association network corresponding to the nouns in the text. By analyzing the association network between the nouns, the types of large language model hallucination risks can be obtained, and the text to be analyzed can be optimized using corresponding processing methods, fundamentally eliminating the risk of hallucinations in the analysis results of the large language model for the text, and ensuring the reliability of the hallucination processing process; by combining semantic analysis with probability statistics to determine the hallucination risk type in the text, it is not necessary to rely on an external knowledge base during the hallucination analysis process, improving the generalization ability of the hallucination processing method. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0040] Figure 1 It is a flowchart of a method for eliminating hallucinations in a large language model disclosed in this application;

[0041] Figure 2Schematic diagram of a specific method for eliminating hallucinations in large language models disclosed in this application;

[0042] Figure 3 Flowchart of a method for obtaining entropy difference value disclosed in this application;

[0043] Figure 4 Schematic diagram of the structure of a device for eliminating hallucinations in large language models disclosed in this application;

[0044] Figure 5 Structural diagram of an electronic device disclosed in this application. Detailed implementation manners

[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0046] Currently, hallucination detection methods mainly rely on comparison with external knowledge bases. This method has problems such as unreliable hallucination detection processes, being limited by the coverage of the knowledge base, and weak generalization ability. For this reason, this application provides a method for eliminating hallucinations in large language models. By combining semantic analysis and probability statistics, it is ensured that no external knowledge base is required during the hallucination analysis process, and the generalization ability and reliability of the hallucination processing method are guaranteed.

[0047] See Figure 1 As shown, an embodiment of the present invention discloses a method for eliminating hallucinations in large language models, including:

[0048] Step S11: Analyze the initial text to be analyzed to obtain target nouns, and encode the feature maps corresponding to each of the target nouns into target vectors; the feature maps include noun concepts, semantic tags, and context information corresponding to each of the target nouns.

[0049] In this embodiment, the main process of the method for eliminating hallucinations in large language models is as shown in Figure 2As shown, first, obtain the association paths between nouns in the text to be analyzed, and use entropy difference calculation to quantify the association paths to evaluate the hallucination risk distribution of the original text; select the most suitable combination of optimization strategies according to the risk type; gradually apply the selected strategies and re-evaluate after each application; verify and optimize the effect through small-scale sampling tests; adjust the optimization parameters according to the verification results; finally, complete the optimization by applying the best parameter configuration. In addition, the system also creates an optimization strategy library, continuously updates the strategy effectiveness score according to the historical application effect, and realizes the self-optimization of strategy selection. It should be noted that the system for detecting and eliminating hallucinations in large language models in this embodiment mainly includes a noun feature mapper, a feature path analyzer, an entropy difference calculator, and a noun optimization engine; therefore, the main process of the above method for eliminating hallucinations in large language models is also as follows: use the noun feature mapper to identify and analyze the core nouns in the input text (i.e., the initial text to be analyzed), where the noun feature mapper is mainly responsible for identifying the key semantic units (i.e., target nouns) in the input text and establishing the mapping relationship between the key semantic units and the internal representations of the model; use the feature path analyzer to construct the association network between noun phrases. The feature path analyzer mainly constructs a higher-level association network and performs structural and semantic analysis on the network; use the entropy difference calculator to calculate the uncertainty differences of each path to provide an accurate mathematical quantification basis; use the noun optimization engine to optimize the structure of the input text and eliminate the hallucination risk from the root. Finally, output the hallucination-free text after optimization.

[0050] It should be noted that the noun feature mapper is based on the following key findings: during the training process of large language models, their internal neurons will self-organize to form functional units for processing specific concepts. These functional units are mainly formed around noun concepts, rather than other parts of speech such as verbs or adjectives. For example, when the model processes nouns such as "computer" and "Einstein", it will activate specific neuron clusters, and these clusters represent the knowledge and associations related to these concepts.

[0051] The noun feature mapper is a fundamental component of the present invention, responsible for identifying key nouns in the input text and establishing mapping relationships between them and the internal representations of the model. The core function of this mapper is to understand the concept network activated within the large language model when processing specific nouns, thereby providing basic data for subsequent hallucination detection. Correspondingly, the process of analyzing the initial text to be analyzed to obtain target nouns specifically includes: using preset grammar analysis techniques to identify the initial text to be analyzed to obtain the first initial nouns in the initial text to be analyzed, and evaluating the semantic importance of the first initial nouns to determine the second initial nouns from each of the first initial nouns; determining whether the second initial nouns are complete nouns. If the second initial nouns are complete nouns, then the second initial nouns are determined as target nouns; if the second initial nouns are not complete nouns, then each of the second initial nouns is merged to obtain a merged noun, and the merged noun is determined as the target noun. Specifically, the noun feature mapper first needs to identify key nouns from the input text. The system adopts multi-level analysis, including basic grammar analysis (using grammar analysis techniques to identify noun components in the text, that is, obtaining the first initial nouns in the initial text to be analyzed), semantic importance evaluation (determining which nouns have core semantic value in the current context, that is, determining the second initial nouns with core semantic value), and concept integration (integrating related nouns into a complete semantic unit, such as "artificial intelligence" instead of separate "artificial" and "intelligence", that is, merging the second initial nouns with incomplete semantics).

[0052] In this embodiment, each recognized noun is assigned multi-dimensional tags (i.e., semantic tags) to form a rich semantic description: including entity type (person, place, organization, time, etc.), degree of abstraction (a continuous spectrum from concrete entities to completely abstract concepts), professional field (the knowledge field it belongs to, such as general, technology, medicine, etc.), and information density (the amount of information required to express the noun). A noun usually has multiple tags at the same time. For example, "artificial intelligence" may be marked as "technology" (entity type), "semi-abstract concept" (degree of abstraction), "computer science" (professional field), and "high information content" (information density). The noun feature mapper uses a non-invasive mapping method to observe the model's response to the input to infer its internal processing mechanism, just like a psychologist infers the thinking process by observing behavior. Correspondingly, before encoding the feature maps corresponding to each target noun into target vectors, it also includes: inputting different text segments including the same target noun into a preset large language model so that the preset large language model generates different first output results according to different text segments; obtaining each first output result and obtaining the first output change situation of the preset large language model according to each first output result; modifying any target noun in the initial text to be analyzed and inputting the corresponding modified text into the preset large language model to obtain the second output result output by the preset large language model and obtaining the second output change situation according to the second output result; inputting the text segments including the same target noun in the context into the preset large language model and according to the third output result of the preset large language model; obtaining the understanding degree of the preset large language model for the target noun according to the first output change situation, the second output change situation, and the third output result, and generating the feature maps corresponding to each target noun according to the understanding degree; specifically, the process of obtaining the understanding degree of the large language model for each noun includes:

[0053] Differential input test: The system provides the model with multiple variant texts containing the target noun and observes the output changes. For example, provide "Einstein invented the theory of relativity" and "Einstein liked the violin", and analyze the different outputs caused by the two inputs (i.e., the first output change situation).

[0054] Noun substitution analysis: By replacing, deleting, or modifying the target noun, measure the degree of output change. For example, replace "Einstein" with "Newton" and observe the degree of change in the model output (i.e., the second output change situation) to infer the degree of differentiation of the model for these two concepts.

[0055] Context sensitivity test: Use the same noun in different contexts and analyze how the model adjusts its understanding of the noun according to the context based on the model output results (i.e., the third output result).

[0056] Through these analyses, the system generates a feature activation map for each noun, mainly recording the set of concepts, semantic tags, and context information associated with the noun. Finally, the content contained in the above map is encoded into high-dimensional vectors and stored in a vector database to support efficient similarity search and semantic matching. By inputting different texts and observing the output changes of the large language model, it is possible to more accurately understand the understanding of each noun by the large language model, thus ensuring the correctness of the generated feature images and further ensuring the reliability of the subsequent generated association paths.

[0057] Step S12: Based on the association relationships between the target nouns determined based on the target vectors, construct an association network between the target nouns, and obtain the association paths corresponding to the target nouns according to the association network. The nodes in the association network represent the target nouns, the directed edges between the nodes represent the association relationships between the target nouns, and the weights of the directed edges represent the association strengths between the nodes.

[0058] The feature path analyzer is the second core component of the system, responsible for analyzing the semantic association paths (i.e., the association network) between nouns and identifying potential inconsistencies. Based on the basic data generated by the noun feature mapper, the analyzer constructs a higher-level association network and performs structural and semantic analyses on the network.

[0059] The feature path analyzer regards the noun concepts in the input text as nodes in the feature activation map and constructs the association paths between them. The system represents each noun as a node, and the association relationship between nouns is represented as a directed weighted edge (referring to the line connecting two nodes, which has directionality and a weight value. Directionality indicates that the relationship points from one node to another, and the weight value represents the strength or importance of this relationship). The attributes of the node include the concept set, semantic label, and context information obtained from the noun feature mapper (all of which are obtained from the noun feature mapper), while the weight of the edge is calculated based on the feature set similarity, context distance, and historical co-occurrence frequency (the calculation uses context information, which measures the frequency of two nouns appearing together in similar contexts. When calculating, the system analyzes the previously processed text samples and records the number of times the noun pair appears together in the same or similar contexts. This calculation may also refer to the domain information in the semantic labels of the nouns and give higher weights to co-occurrences within the same domain); specifically, any two target nouns that are mutually associated are obtained according to the association network, and the association strength between any two associated target nouns is obtained according to the feature overlap degree, context distance, and co-occurrence frequency between the two target nouns; among them, the feature overlap degree is the quotient of the target difference and the total number of features of any two target nouns. The target difference is the difference between the first number of common features and the second number of mutually exclusive features between any two target nouns. That is, when the system uses the path analysis algorithm to calculate the weight of the edge, it uses three main factors: feature overlap degree ((number of common features - number of mutually exclusive features) / total number of features, where the number of mutually exclusive features is summarized by the system, such as "vegetarianism" and "meat consumption"), context distance, and co-occurrence frequency. These weights directly reflect the strength of the association between two nouns, and through this calculation method, the connection between nouns can be efficiently evaluated. It can not only find the "strongest path" (i.e., the reference path) with the highest weight sum, but also identify sub-optimal paths and compare the weight differences between them. When the weight gap between the strongest path and the sub-optimal path is large, it indicates that this association is more reliable; while when the weights of multiple paths are similar, it indicates the existence of semantic ambiguity.

[0060] It should be noted that conflict detection is the core function of the path analyzer, and the system can initially identify four typical conflicts included in the hallucination risk type:

[0061] Direct conflict: When calculating the edge weight between two directly connected nouns, if the system detects that their feature overlap degree is negative (mutual exclusivity), it will be initially marked as a direct conflict. For example, the comparison of the feature sets of "vegetarian" and "eating steak" will result in a negative overlap degree.

[0062] Indirect conflict: When the system analyzes multiple paths, it detects conflicting results generated by paths connected through different intermediate nodes. For example, there is a conflict in the time characteristics between the path of "Mozart" → "background of the times" and the path of "background of the times" → "electronic synthesizer" (some nouns will contain time characteristics).

[0063] Context conflict: When the edges formed by the same noun at different positions in the text have significantly different weight values, the system will initially identify it as a context conflict. For example, the weight of the edge connecting "apple" with food-related nouns is contrasted with the weight of the edge connecting it with technology-related nouns.

[0064] Strength imbalance: When calculating the paths between nodes, if the edge weight of a secondary noun is abnormally higher than that of a topic noun, the system will detect a deviation in focus. For example, the edge weight of "Einstein" → "patent office clerk" exceeds the edge weight of "Einstein" → "physicist".

[0065] That is, the types of hallucination risks in this embodiment include direct conflict, indirect conflict, context conflict, and weight imbalance. A direct conflict indicates that the feature overlap degree between any two directly associated target nouns is negative. An indirect conflict indicates that there is a semantic contradiction between any two target nouns connected through the same intermediate node. A context conflict indicates that the difference in weight values between the edges formed by the same target noun at different positions in the association network is greater than a preset weight value difference threshold. A weight imbalance indicates that the weight value of the edge corresponding to a preset secondary noun is higher than the weight value of the edge corresponding to a preset primary noun.

[0066] Step S13: Determine a reference path from each of the association paths, and obtain the entropy difference value between the first target entropy corresponding to each association path and the second target entropy corresponding to the reference path; the target entropy includes feature activation entropy, path selection entropy, and output distribution entropy. The feature activation entropy represents the uncertainty of the activation distribution of the features corresponding to the target noun in the association path. The path selection entropy represents the uncertainty of path selection from the start point to the end point of the path in the association path. The output distribution entropy represents the uncertainty of the distribution of the output results generated by a preset large language model based on the association path.

[0067] The entropy difference calculator is used to quantify the uncertainty difference of different noun paths, providing a mathematical basis for hallucination risk assessment. It should be noted that in this embodiment, the path analyzer can only initially identify the above four typical conflicts, and the specific conflict types need to be finally determined according to the results of the entropy difference calculator.

[0068] For each identified association path, the system calculates three types of entropy: feature activation entropy, path selection entropy, and output distribution entropy. Among them, the feature activation entropy characterizes the uncertainty of the activation distribution of the features corresponding to the target noun in the association path, the path selection entropy characterizes the uncertainty of the path selection from the start point to the end point of the path in the association path, and the output distribution entropy characterizes the uncertainty of the distribution of the output results generated by the preset large language model based on the association path.

[0069] Step S14: Determine the hallucination risk type corresponding to the initial text to be analyzed based on the entropy difference value, and use the optimization strategy corresponding to the hallucination risk type to optimize the initial text to be analyzed, so as to eliminate the hallucination in the analysis result after the preset large language model analyzes the initial text to be analyzed.

[0070] In this embodiment, the process of optimizing the initial text to be analyzed by using the optimization strategy corresponding to the hallucination risk type may specifically include: if the hallucination risk type is the context conflict, increase the usage times of the first target noun in the initial text to be analyzed, and increase the association between the first target noun and the corresponding context; where the first target noun is the target noun corresponding to the context conflict; if the hallucination risk type is a direct conflict, add a target modifier before the third target noun; where the third target noun is the target noun corresponding to the direct conflict, and the target modifier is a modifier that limits the third target noun; if the hallucination risk type is an indirect conflict, adjust the position of the third target noun in the initial text to be analyzed, and add transitional text between the third target nouns; where the third target noun is the target noun corresponding to the indirect conflict; if the hallucination risk type is a weight imbalance, weaken the first representation intensity of the preset secondary noun in the initial text to be analyzed, and strengthen the second representation intensity of the preset primary noun in the initial text to be analyzed. Specifically, the system adopts five optimization strategies and dynamically selects according to the hallucination type:

[0071] 1. Noun strengthening: Enhance the representation intensity of key nouns in the text, mainly to solve the context conflict in the third part. The methods include: repeating key nouns, adding modificatory descriptions, and increasing context associations.

[0072] 2. Noun precisification: Improve the semantic precision of nouns, mainly to solve the direct conflict in the third part. The methods include: replacing fuzzy nouns with specific expressions, adding restrictive modifiers, and decomposing complex concepts into basic components.

[0073] 3. Conflict path isolation: Reduce the mutual interference between conflicting nouns, mainly addressing the indirect conflicts in the third part. The methods include: adjusting the order of noun appearance, adding transitional expressions, and marking different contexts (here it means separating unclear contexts and adding some descriptive words to reduce noun conflicts).

[0074] 4. Feature balance adjustment: Balance the representation strengths of different nouns, mainly addressing the strength imbalance in the third part. The methods include: weakening secondary nouns and strengthening primary nouns. For example, in the original text: "Turing designed the computer and developed the artificial intelligence system", here the computer is the primary noun (because this is Turing's main contribution), and the artificial intelligence system is the secondary noun. Finally, it is changed to Turing proposed the theoretical basis of the universal computer (Turing machine) in 1936. And in 1950, he published a thought experiment on machine intelligence, providing an early conceptual framework for the field of artificial intelligence.

[0075] 5. Entropy stability optimization: Reduce the uncertainty of high-entropy paths, applicable to all situations. The methods include: adding redundant information, providing auxiliary contexts, and enhancing path connection strength.

[0076] In this embodiment, after optimizing the initial text to be analyzed using the optimization strategy corresponding to the hallucination risk type, it further includes: inputting the optimized text to be analyzed into a preset large language model, obtaining the optimization effect corresponding to the optimized text to be analyzed according to the fourth output result of the preset large language model, and adjusting the optimization parameters corresponding to the optimization strategy according to the optimization effect. That is, applying the selected strategy and re-evaluating after each application; verifying and optimizing the effect through small-scale sampling tests; adjusting the optimization parameters according to the verification results; and finally completing the optimization by applying the best parameter configuration. By optimizing the text according to the optimization strategy corresponding to the hallucination risk type, the pertinence and reliability of the optimization process are ensured.

[0077] Thus, through semantic analysis of the text to be analyzed, this application can obtain the association network corresponding to the nouns in the text. By analyzing the association network between nouns, the types of large language model hallucination risks can be obtained, and the text to be analyzed is optimized using corresponding processing methods, fundamentally eliminating the risk of hallucinations in the analysis results of the large language model for the text, ensuring the reliability of the hallucination processing process; by combining semantic analysis with probability statistics to determine the hallucination risk types in the text, it is not necessary to rely on an external knowledge base during the hallucination analysis process, improving the generalization ability of the hallucination processing method.

[0078] Based on the foregoing embodiments, this application describes a process for detecting and eliminating large language model hallucinations. To make the technical solutions in this application more complete, next, this application will elaborate in detail on how to select the reference path and how to calculate the entropy difference value. SeeFigure 3 As shown in Figure 3 , an embodiment of the present invention discloses a calculation process of entropy difference value, including:

[0079] Step S21, calculate the target entropy of each associated path to obtain a target calculation result; wherein, the target entropy includes feature activation entropy, path selection entropy and output distribution entropy, the feature activation entropy characterizes the uncertainty of the activation distribution of the features corresponding to the target noun in the associated path, the path selection entropy characterizes the uncertainty of path selection from the path start point to the path end point in the associated path, and the output distribution entropy characterizes the uncertainty of the distribution of the output results generated by the preset large language model based on the associated path.

[0080] In this embodiment, it is necessary to calculate the feature activation entropy, path selection entropy and output distribution entropy corresponding to each of the associated paths respectively.

[0081] 1. Feature activation entropy: used to quantify the uncertainty of feature activation distribution on the path, and its calculation process is as follows:

[0082] First, it is necessary to calculate the activation intensity of the feature corresponding to the target noun, and its formula is as follows:

[0083] ;

[0084] Wherein, represents the activation intensity of feature i, is the weight of feature i on noun n, and the feature weight represents the importance of a certain feature to the noun concept, usually extracted from the knowledge base or obtained through pre-training, is the activation frequency of feature i, indicating the frequency of activation of feature i in the current context, obtained after model training. Then, an operation of normalization processing is required, and its calculation formula is as follows:

[0085] ;

[0086] Wherein, is the probability distribution corresponding to the above feature activation intensity, is the activation intensity of feature j, and finally the feature activation entropy is calculated from the probability, and its calculation formula is as follows:

[0087] ;

[0088] Wherein, is the above-mentioned feature activation entropy.

[0089] 2. Path selection entropy: used to quantify the uncertainty of possible path selection from the start point to the end point, and its calculation process is as follows:

[0090] First, it is necessary to calculate the weight value of the edges in the path, and its calculation formula is as follows:

[0091] ;

[0092] Among them, represents the weight value of the k-th edge in path j, and the weight value is calculated from the feature overlap degree, context distance, and co-occurrence frequency between different nouns on the path, and represent different nouns respectively, represents the feature overlap degree between nouns, represents the context distance between nouns (the number of tokens between words), represents the co-occurrence frequency of two nouns (counted from the preset corpus), , , are the coefficients of the weights, which are obtained by training with the data in the corpus and satisfy .

[0093] After that, it is necessary to calculate the comprehensive strength of path j, which is obtained by multiplying the weight values of each edge in the path, and its calculation formula is as follows:

[0094] ;

[0095] Normalize all path strengths using the Softmax function to obtain the probability distribution , and its calculation formula is as follows:

[0096] ;

[0097] Among them, is the comprehensive strength of path m.

[0098] Finally, calculate the path selection entropy from the probability, and its calculation formula is as follows:

[0099] ;

[0100] 3. Output distribution entropy: used to quantify the uncertainty of the output distribution that may be generated based on the current path.

[0101] First, it is necessary to calculate the probability distribution corresponding to the frequency of the model generating output k, and its calculation formula is as follows:

[0102] ;

[0103] Among them, Let \(k\) be the frequency of the model generating output \(k\) from historical data (which can be statistically calculated based on training data). Let \(m\) be the frequency of the model generating output \(m\) from historical data.

[0104] Finally, calculate the output distribution entropy , and its calculation formula is as follows:

[0105] ;

[0106] Step S22: Determine a reference path from each of the associated paths according to the target calculation result, obtain the first target entropy corresponding to the reference path and the second target entropies corresponding to each of the associated paths, and calculate the entropy difference value between the first target entropy and the second target entropies.

[0107] In this embodiment, by calculating the feature activation entropy, path selection entropy, and output distribution entropy for each path, and analyzing the paths according to the target calculation result, the system selects a most reliable path as the reference path , and then calculates the entropy difference values between other candidate paths and the reference path . Its specific calculation formula is as follows:

[0108] ;

[0109] Among them, , , are weight coefficients dynamically adjusted according to the specific scenario. Thus, an entropy difference calculator can be used to calculate four types of conflicts. Among them, direct conflict: the feature activation entropy is extremely high, indicating a large uncertainty in the feature distribution; indirect conflict: the path selection entropy is extremely high, indicating a contradiction between multiple paths; context conflict: the feature activation entropies of the same noun in different paths are significantly different; intensity imbalance: the entropy difference value between the reference path and the candidate path exceeds the threshold. In addition, the entropy difference calculator is also used for risk assessment. Set the entropy difference threshold . When , it is marked as a high-risk path. For high-risk paths, the system can reduce the weight of this path in the final inference, or require the model to provide alternative explanations or evidence, or mark relevant content as "low confidence" information in the output. This method of integrating path analysis and entropy difference calculation enables the system to efficiently identify potential hallucination risk points, while providing an accurate mathematical quantification basis. By calculating the three types of entropy corresponding to each associated path, the system can combine semantic analysis with probability statistics, thereby obtaining the hallucination risk types in the text, and then optimizing the text to fundamentally eliminate the hallucination risk.

[0110] It can be seen that by performing semantic analysis on the text to be analyzed, the present application can obtain the association network corresponding to the nouns in the text, and can obtain the types of risks of large language model hallucinations by analyzing the association network between nouns, and use corresponding processing methods to optimize the text to be analyzed, fundamentally eliminating the risk of hallucinations in the analysis results of the large language model for the text, ensuring the reliability of the hallucination processing process; by combining semantic analysis with probability statistics to determine the types of hallucination risks in the text, it is not necessary to rely on an external knowledge base during the hallucination analysis process, improving the generalization ability of the hallucination processing method.

[0111] See Figure 4 As shown, an embodiment of the present invention discloses a large language model hallucination elimination device, including:

[0112] A feature map acquisition module 11, configured to analyze the initial text to be analyzed to obtain target nouns, and encode the feature maps corresponding to each of the target nouns into target vectors; the feature maps include noun concepts, semantic labels, and context information corresponding to each of the target nouns;

[0113] An association path acquisition module 12, configured to construct an association network between each of the target nouns according to the association relationship between each of the target nouns determined based on the target vectors, and obtain the association path corresponding to each of the target nouns according to the association network; the nodes in the association network represent each of the target nouns, the directed edges between the nodes represent the association relationships between each of the target nouns, and the weights of the directed edges represent the association strength between the nodes;

[0114] An entropy difference acquisition module 13, configured to determine a reference path from each of the association paths, and obtain the entropy difference value between the first target entropy corresponding to each of the association paths and the second target entropy corresponding to the reference path; the target entropy includes a feature activation entropy, a path selection entropy, and an output distribution entropy, the feature activation entropy characterizes the uncertainty of the activation distribution of the features corresponding to the target nouns in the association path, the path selection entropy characterizes the uncertainty of the path selection from the start point to the end point of the path in the association path, and the output distribution entropy characterizes the uncertainty of the output result distribution generated by the preset large language model based on the association path;

[0115] A text optimization module 14, configured to determine the type of hallucination risk corresponding to the initial text to be analyzed based on the entropy difference value, and use an optimization strategy corresponding to the type of hallucination risk to optimize the initial text to be analyzed, so as to eliminate the hallucinations in the analysis results after the preset large language model analyzes the initial text to be analyzed.

[0116] It can be seen that by performing semantic analysis on the text to be analyzed, the present application can obtain the associated network corresponding to the nouns in the text. By analyzing the associated network among the nouns, the types of risks of large language model hallucinations can be obtained, and the text to be analyzed can be optimized using corresponding processing methods, fundamentally eliminating the risk of hallucinations in the analysis results of the large language model for the text and ensuring the reliability of the hallucination processing process. By combining semantic analysis with probability statistics to determine the types of hallucination risks in the text, it is not necessary to rely on an external knowledge base during the hallucination analysis process, improving the generalization ability of the hallucination processing method.

[0117] In some specific embodiments, the feature map acquisition module 11 may specifically include:

[0118] A text recognition unit, configured to recognize the initial text to be analyzed by using a preset syntax analysis technique to obtain the first initial nouns in the initial text to be analyzed, and perform a semantic importance evaluation on the first initial nouns to determine the second initial nouns from among the first initial nouns;

[0119] A noun judgment unit, configured to judge whether the second initial noun is a complete noun. If the second initial noun is a complete noun, the second initial noun is determined as the target noun;

[0120] A noun merging unit, configured to, if the second initial noun is not a complete noun, merge the second initial nouns to obtain a merged noun, and determine the merged noun as the target noun.

[0121] In some specific embodiments, the feature map acquisition module 11 may further include:

[0122] A text input unit, configured to input different text segments including the same target noun into the preset large language model, so that the preset large language model generates different first output results according to the different text segments;

[0123] A first output result acquisition unit, configured to acquire each of the first output results and obtain the first output change situation of the preset large language model according to each of the first output results;

[0124] A second output result acquisition unit, configured to modify any target noun in the initial text to be analyzed, input the corresponding modified text into the preset large language model to obtain the second output result output by the preset large language model, and obtain the second output change situation according to the second output result;

[0125] A third output result acquisition unit, configured to input text fragments including the same target noun in the context into the preset large language model, and according to the third output result of the preset large language model;

[0126] A feature map generation unit, configured to obtain the understanding degree of the preset large language model for each of the target nouns according to the first output change situation, the second output change situation, and the third output result, and generate each of the feature maps corresponding to each of the target nouns according to the understanding degree.

[0127] In some specific embodiments, the large language model hallucination elimination device further includes:

[0128] An association strength acquisition module, configured to obtain any two mutually associated target nouns according to the association network, and obtain the association strength between any two associated target nouns according to the feature overlap degree, context distance, and co-occurrence frequency between the two target nouns; wherein, the feature overlap degree is the quotient of the target difference and the total number of features of any two target nouns, and the target difference is the difference between the first number of common features and the second number of mutually exclusive features between any two target nouns.

[0129] In some specific embodiments, the text optimization module 14 may specifically include:

[0130] An association increase unit, configured to increase the usage times of the first target noun in the initial text to be analyzed and increase the association between the first target noun and the corresponding context if the hallucination risk type is the context conflict; wherein, the first target noun is the target noun corresponding to the context conflict;

[0131] A modifier addition unit, configured to add a target modifier before the third target noun if the hallucination risk type is the direct conflict; wherein, the third target noun is the target noun corresponding to the direct conflict, and the target modifier is a modifier that limits the third target noun;

[0132] A position adjustment unit, configured to adjust the position of the third target noun in the initial text to be analyzed and add transitional text between each of the third target nouns if the hallucination risk type is the indirect conflict; wherein, the third target noun is the target noun corresponding to the indirect conflict;

[0133] A representation strength adjustment unit, configured to weaken the first representation strength of the preset secondary noun in the initial text to be analyzed and strengthen the second representation strength of the preset primary noun in the initial text to be analyzed if the hallucination risk type is the weight imbalance.

[0134] In some specific embodiments, the text optimization module 14 may further include:

[0135] A parameter adjustment unit, configured to input the text to be analyzed after optimization into the preset large language model, obtain the optimization effect corresponding to the text to be analyzed after optimization according to the fourth output result of the preset large language model, and adjust the optimization parameters corresponding to the optimization strategy according to the optimization effect.

[0136] Furthermore, an embodiment of the present application also discloses an electronic device. Figure 5 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be regarded as any limitation to the scope of use of the present application.

[0137] Figure 5 This is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store computer programs, and the computer programs are loaded and executed by the processor 21 to implement the relevant steps in the large language model hallucination elimination method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0138] In this embodiment, the power supply 23 is used to provide working voltages for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and specific limitations are not imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and specific limitations are not imposed here.

[0139] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be short-term storage or permanent storage.

[0140] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it may be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the large language model hallucination elimination method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks.

[0141] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the method for eliminating large language model hallucinations disclosed above is implemented. For the specific steps of this method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated herein.

[0142] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the method part for relevant details.

[0143] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0144] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0145] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0146] The above has introduced the technical solution provided by this application in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A large language model hallucination elimination method, characterized in that: include: Analyze the initial text to be analyzed to obtain target nouns, and encode the feature graphs corresponding to each target noun into a target vector; The feature graph includes noun concepts, semantic labels and context information corresponding to each of the target nouns; According to the association relationship between the target nouns determined based on the target vector, an association network between the target nouns is constructed, and an association path corresponding to each target noun is obtained according to the association network; the nodes in the association network represent the target nouns, the directed edges between the nodes represent the association relationship between the target nouns, and the weight of each directed edge represents the association strength between the nodes; Determine a reference path from each of the associated paths, and obtain an entropy difference value between a first target entropy corresponding to each of the associated paths and a second target entropy corresponding to the reference path; the target entropy includes feature activation entropy, path selection entropy and output distribution entropy, the feature activation entropy represents the uncertainty of activation distribution of the feature corresponding to the target noun in the associated path, the path selection entropy represents the uncertainty of path selection from the path starting point to the path end point in the associated path, and the output distribution entropy represents the uncertainty of distribution of output results generated by a preset large language model based on the associated path; Determining the hallucination risk type corresponding to the initial text to be analyzed based on the entropy difference value, and optimizing the initial text to be analyzed using an optimization strategy corresponding to the hallucination risk type, so as to eliminate hallucinations in the analysis result after the preset large language model analyzes the initial text to be analyzed; Among them, the hallucination risk types include direct conflict, indirect conflict, context conflict and weight imbalance. The direct conflict indicates that the feature overlap between any two directly related target nouns is a negative value. The indirect conflict indicates that there is a contradiction in the semantics of any two target nouns connected by the same intermediate node. The context conflict indicates that the difference in weight values ​​between the edges formed by the same target noun at different positions in the association network is greater than a preset weight value difference threshold. The weight imbalance indicates that the weight value of the edge corresponding to the preset secondary noun is higher than the weight value of the edge corresponding to the preset primary noun. The method of determining the hallucination risk type corresponding to the initial text to be analyzed based on the entropy difference value, and optimizing the initial text to be analyzed using an optimization strategy corresponding to the hallucination risk type, includes: if the hallucination risk type is the context conflict, increasing the number of times a first target noun is used in the initial text to be analyzed, and increasing the association between the first target noun and the corresponding context; wherein the first target noun is a target noun corresponding to the context conflict; If the hallucination risk type is the direct conflict, a target modifier is added before the third target noun; wherein the third target noun is the target noun corresponding to the direct conflict, and the target modifier is a modifier that limits the third target noun; If the hallucination risk type is the indirect conflict, adjusting the position of the third target noun in the initial text to be analyzed, and adding transition text between each of the third target nouns; wherein the third target noun is the target noun corresponding to the indirect conflict; If the hallucination risk type is the weight imbalance, the first representation strength of the preset secondary noun in the initial text to be analyzed is weakened, and the second representation strength of the preset primary noun in the initial text to be analyzed is strengthened.

2. The large language model hallucination elimination method according to claim 1, characterized in that: The step of analyzing the initial text to be analyzed to obtain the target noun includes: Using a preset grammar analysis technology to identify the initial text to be analyzed to obtain first initial nouns in the initial text to be analyzed, and performing semantic importance evaluation on the first initial nouns to determine second initial nouns from each of the first initial nouns; determining whether the second initial noun is a complete noun, and if the second initial noun is a complete noun, determining the second initial noun as a target noun; If the second initial noun is not a complete noun, each of the second initial nouns is merged to obtain a merged noun, and the merged noun is determined as a target noun.

3. The large language model hallucination elimination method according to claim 1, characterized in that: Before encoding the feature graphs corresponding to the target nouns into target vectors, the method further includes: Inputting different text segments including the same target noun into the preset large language model, so that the preset large language model generates different first output results according to the different text segments; Obtain each of the first output results and obtain a first output change of the preset large language model according to each of the first output results; Modify any target noun in the initial text to be analyzed, and input the corresponding modified text into the preset large language model to obtain a second output result output by the preset large language model, and obtain a second output change according to the second output result; Inputting a text segment including the same target noun in the context into the preset large language model, and according to a third output result of the preset large language model; The degree of understanding of each of the target nouns by the preset large language model is obtained according to the first output change, the second output change and the third output result, and the feature graphs corresponding to each of the target nouns are generated according to the degree of understanding.

4. The large language model hallucination elimination method according to claim 1, characterized in that: Also includes: According to the association network, any two target nouns that are associated with each other are obtained, and according to the feature overlap, context distance and co-occurrence frequency between the any two target nouns, the association strength between any two associated target nouns is obtained; wherein the feature overlap is the quotient of the target difference and the total number of features of any two target nouns, and the target difference is the difference between the first number of common features and the second number of mutually exclusive features between any two target nouns.

5. The large language model hallucination elimination method according to claim 1, characterized in that: After optimizing the initial text to be analyzed by using the optimization strategy corresponding to the hallucination risk type, the method further includes: The optimized text to be analyzed is input into the preset large language model, an optimization effect corresponding to the optimized text to be analyzed is obtained according to a fourth output result of the preset large language model, and optimization parameters corresponding to the optimization strategy are adjusted according to the optimization effect.

6. A large language model hallucination elimination device, characterized in that: include: A feature map acquisition module is used to analyze the initial text to be analyzed to obtain target nouns, and encode the feature maps corresponding to each target noun into a target vector; the feature maps include noun concepts, semantic labels and context information corresponding to each target noun; an association path acquisition module, configured to construct an association network between the target nouns according to the association relationship between the target nouns determined based on the target vector, and acquire an association path corresponding to each target noun according to the association network; the nodes in the association network represent the target nouns, the directed edges between the nodes represent the association relationship between the target nouns, and the weight of each directed edge represents the association strength between the nodes; An entropy difference acquisition module is used to determine a reference path from each of the associated paths, and obtain an entropy difference value between a first target entropy corresponding to each of the associated paths and a second target entropy corresponding to the reference path; the target entropy includes feature activation entropy, path selection entropy and output distribution entropy, the feature activation entropy represents the uncertainty of the activation distribution of the feature corresponding to the target noun in the associated path, the path selection entropy represents the uncertainty of the path selection from the path start point to the path end point in the associated path, and the output distribution entropy represents the uncertainty of the distribution of the output result generated by the preset large language model based on the associated path; a text optimization module, configured to determine the hallucination risk type corresponding to the initial text to be analyzed based on the entropy difference value, and optimize the initial text to be analyzed using an optimization strategy corresponding to the hallucination risk type, so as to eliminate hallucinations in an analysis result after the preset large language model analyzes the initial text to be analyzed; Among them, the hallucination risk types include direct conflict, indirect conflict, context conflict and weight imbalance. The direct conflict indicates that the feature overlap between any two directly related target nouns is a negative value. The indirect conflict indicates that there is a contradiction in the semantics of any two target nouns connected by the same intermediate node. The context conflict indicates that the difference in weight values ​​between the edges formed by the same target noun at different positions in the association network is greater than a preset weight value difference threshold. The weight imbalance indicates that the weight value of the edge corresponding to the preset secondary noun is higher than the weight value of the edge corresponding to the preset primary noun. The text optimization module may specifically include: an association adding unit, configured to increase the number of times a first target noun is used in the initial text to be analyzed, and increase the association between the first target noun and the corresponding context if the hallucination risk type is the context conflict; wherein the first target noun is the target noun corresponding to the context conflict; a modifier adding unit, configured to add a target modifier before the third target noun if the hallucination risk type is the direct conflict; wherein the third target noun is the target noun corresponding to the direct conflict, and the target modifier is a modifier that limits the third target noun; a position adjustment unit, configured to adjust the position of the third target noun in the initial text to be analyzed and add transition text between each of the third target nouns if the hallucination risk type is the indirect conflict; wherein the third target noun is the target noun corresponding to the indirect conflict; A representation strength adjustment unit is used to weaken the first representation strength of the preset minor noun in the initial text to be analyzed and to strengthen the second representation strength of the preset major noun in the initial text to be analyzed if the hallucination risk type is the weight imbalance.

7. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the large language model hallucination elimination method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: Used to store a computer program, which, when executed by a processor, implements the large language model hallucination elimination method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • High and new technology enterprise evaluation method and system, computer equipment and storage medium

    CN112766788A

  • Knowledge boundary identification method based on interpretability of large language model and related device

    CN119599136A