A model hallucination suppression method and system based on two-way vectorization retrieval

By constructing a parametric belief field and a fact reference field, combining local Riemannian metrics and curvature singularity maps, and utilizing a neural Jacobian solver and a geodesic repair force network, the illusion problem in generative large language models is solved, achieving accuracy and stability of generated content.

CN122633892APending Publication Date: 2026-08-25WEIMAI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611125905.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-28
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing generative large language models suffer from illusions when generating content, making it difficult to accurately locate and suppress these illusions. This results in generated content that does not match objective facts, affecting their application in professional fields.

Method used

A method based on dual-path vectorization retrieval is adopted to construct a parametric belief field and a fact reference field. The hallucination source point is located by combining local Riemannian metric and curvature singularity map. The purity manifold is generated by variational information bottleneck and geodesic distance clustering. The model hallucination is suppressed by using a neural Jacobian solver and a geodesic repair force network.

Benefits of technology

It achieves precise location and effective filtering of hallucination sources, alleviates the inaccuracy and instability of generated content, and ensures real-time correction and hallucination-free output of generated content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122633892A_ABST
    Figure CN122633892A_ABST
Patent Text Reader

Abstract

The application provides a model hallucination inhibition method and system based on two-way vector quantization retrieval, comprising: constructing and constructing a local Riemann metric according to a parameter belief field and a fact reference field, locating a hallucination source point based on the local Riemann metric, generating a curvature singularity graph, synchronously combining the local Riemann metric and the curvature singularity graph to adaptively purify the retrieved content based on a variational information bottleneck, obtaining a fact core set, and clustering the fact core set using geodesic distance to generate a purity manifold, predicting a hallucination bias sequence using a neural Jacobian solver, and generating a geometric correction force vector according to a geodesic repair force network, at each decoding time step of a large language model, simulating a network in real time to repair the local Riemann metric according to a Ricci flow, and based on the repaired local Riemann metric, the purity manifold constraint and the geometric correction force vector, the model hallucination is inhibited to generate hallucination-free retrieval content, thereby realizing efficient and accurate inhibition of model hallucination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and more specifically, to a method and system for suppressing model illusions based on dual-path vectorized retrieval. Background Technology

[0002] With the rapid development of natural language processing technology, generative large language models have been widely used in various text generation tasks. However, they often suffer from the phenomenon that the generated content does not match the objective facts, contains false information, or has logical contradictions. This "illusion" problem seriously limits the application of large language models in professional fields.

[0003] Existing hallucination suppression methods are mainly divided into two categories: one is based on external knowledge retrieval, which obtains factual information by retrieving external knowledge bases and constrains model generation. However, most of these methods rely on Euclidean distance for similarity matching, which is often difficult to adapt to the nonlinear curvature of semantic space and is prone to fact matching bias. The other category is based on internal model correction, which optimizes the generation logic by modifying the model structure or training strategy. However, these methods often lack quantitative modeling of the difference between the model's internal cognition and external facts, making it difficult to locate and suppress "hallucinations". Summary of the Invention

[0004] To address the shortcomings of existing technologies, embodiments of the present invention provide a model illusion suppression method and system based on dual-path vectorized retrieval, in order to solve at least one of the aforementioned technical problems in the prior art.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a model hallucination suppression method based on dual-path vectorized retrieval, comprising: S1: Construct a parametric belief field and a factual reference field based on the user query text and a preset external vector knowledge base, and construct a local Riemannian metric based on the parametric belief field and the factual reference field. Simultaneously combine the local Riemannian metric to locate the illusion source point and generate a curvature singularity map. S2: Based on the variational information bottleneck, the retrieved content is adaptively purified by simultaneously combining local Riemannian metrics and curvature singularity graphs to obtain a set of fact kernels. The fact kernels within the set are then clustered using geodesic distance to generate a purity manifold. S3: Based on the hidden state sequence, local Riemannian metric, curvature singularity graph and fact kernel set in the decoding process of the large language model, the hallucination bias sequence is predicted simultaneously by combining the neural Jacobian solver, and the geometric correction force vector is generated according to the geodesic repair force network. S4: At each decoding time step, the local Riemann metric is repaired in real time according to the Ricci flow simulation network. Based on the repaired local Riemann metric, purity manifold constraints and geometric correction force vector, model illusion suppression is performed to generate illusion-free search content.

[0006] Secondly, embodiments of the present invention also provide a model hallucination suppression system based on dual-path vectorized retrieval, comprising: The data processing module is used to construct a parametric belief field and a fact reference field, and to construct a local Riemannian metric based on the parametric belief field and the fact reference field, and to construct a curvature singularity map based on the local Riemannian metric. The manifold construction module is used to adaptively purify the retrieved content by combining local Riemannian metrics and curvature singularity graphs, obtain a set of fact kernels, and cluster the fact kernels within the set of fact kernels to generate a purity manifold. The hallucination prediction module is used to predict hallucinations by combining the hidden state sequence in the decoding process of the large language model with the local Riemann metric, curvature singularity graph and fact kernel set according to the neural Jacobian solver, and to obtain the predicted hallucination bias sequence. A correction generation module is used to generate a geometric correction force vector based on the geodesic repair force network and the predicted hallucination bias sequence. The illusion correction module is used to suppress model illusions based on local Riemannian metrics, purity manifold constraints, and geometric correction force vectors, and generate illusion-free search content.

[0007] Compared with the prior art, the present invention has the following beneficial effects: Step S1 constructs a dual-path semantic vector field consisting of a parametric belief field and a factual reference field, and simultaneously combines the local Riemannian metric to quantify the difference between the internal cognition of the model and the external facts, thereby achieving accurate localization of the hallucination source point and solving the problem of inaccurate hallucination localization in the existing technology. Step S2 employs variational information bottleneck technology and scalar curvature-based perception to adaptively purify the retrieved content of the model. It also combines geodesic distance clustering to generate a purity manifold, thereby effectively filtering out non-factual noise and alleviating the problem of incomplete fact purification in existing technologies.

[0008] Step S3 predicts the hallucination bias using a neural Jacobian solver and generates a correction force vector using a geodesic repair force network. Step S4 combines Ricci flow to repair the local Riemann metric in real time. The repaired local Riemann metric, purity manifold constraints, and geometric correction force vector are combined to suppress the hallucination in the model, thereby achieving real-time correction of the generation process and alleviating the problems of untimely correction of generated content and unstable suppression effect in the prior art. Attached Figure Description

[0009] Figure 1 This is a flowchart of the steps of a model illusion suppression method based on dual-path vectorized retrieval according to the present invention; Figure 2 This is a schematic diagram of a model hallucination suppression system based on dual-path vectorized retrieval according to the present invention. Detailed Implementation

[0010] To make the technical solution of the present invention clearer and its technical advantages more apparent, the technical solution of the present invention will be clearly and completely described below in conjunction with specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of the present invention.

[0011] It should be noted that, in this document, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the present invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will understand, explicitly and implicitly, that the embodiments described herein can be combined with other embodiments.

[0012] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart of the steps of a model hallucination suppression method based on dual-path vectorized retrieval according to the present invention. The following is a detailed introduction to this model hallucination suppression method based on dual-path vectorized retrieval.

[0013] Step S1: Construct a parametric belief field and a factual reference field based on the user query text and a preset external vector knowledge base, and construct a local Riemannian metric based on the parametric belief field and the factual reference field. Simultaneously combine the local Riemannian metric to locate the illusion source point and generate a curvature singularity map.

[0014] Specifically, step S1 includes steps S11-S15: Step S11, construct the parametric belief map, specifically: First, during the inference process of the Large Language Model (LLM), standard forward propagation is performed on the query text input by the user until the hidden state is output by the last Transformer block of the LLM.

[0015] Among them, LLM can be selected from mainstream generative models such as GPT-3.5 and LLaMA-2. In this embodiment, the LLaMA-2-7B model is used as an example. It contains 32 Transformer blocks, each of which contains a multi-head attention mechanism and a feedforward neural network. The model parameters are frozen during inference and no parameter updates are performed.

[0016] Next, at the residual flow output position of each Transformer block in the LLM, a pre-trained intra-layer relation probe is inserted. This probe consists of a linear layer and a sigmoid activation function, with the input dimension consistent with the hidden state dimension of the corresponding layer and the output dimension being the total number of predefined relation types. After being mapped by the sigmoid activation function, a predicted probability value for the corresponding relation type is generated.

[0017] In this embodiment of the LLaMA-2-7B model, assuming that the hidden state dimension of each Transformer block is 4096, the input dimension of the intra-layer relation probe needs to be set to 4096; the predefined relation types include at least 10 common relation types: belonging, containing, causality, parallel, hierarchical, association, attribute, action, condition, and time, so the output dimension of the probe needs to be set to 10.

[0018] The intra-layer relation probe is trained using a manually labeled entity relation dataset. The training objective is to minimize the cross-entropy loss between the probe's predicted value and the corresponding true value in the entity relation dataset, until the cross-entropy loss reaches a preset threshold and the fluctuation amplitude of the loss value within multiple consecutive periods is lower than the preset threshold, or the preset maximum training step size is reached.

[0019] Next, named entity recognition is performed on the user query text to extract all entity sets. For example, if the user query text is "Introduce the main products and founder of xx company", the entity set E={xx company, main products, founder} can be extracted by the BERT-based named entity recognition model.

[0020] At the same time, for the hidden state of each Transformer block Extract the hidden state vector corresponding to each entity, where, The number of layers in the Transformer block. To query the length of the text sequence, For the hidden state dimension, Let be the set of real numbers; for each pair of entities The corresponding magnitude hidden state is concatenated into a vector. Input the intra-layer relation probe to obtain the predicted probability of the layer for various types of relations.

[0021] Continuing with the previous example, assuming it is split into 12 tokens by LLM, then its sequence "xx company" corresponds to the hidden state vector. Similarly, the hidden state vector corresponding to the entity "main product" The hidden state vector corresponding to the entity "founder" The concatenated vector corresponding to entity pairs (company xx, founder) Input the intra-layer relation probe and output the predicted probabilities of 10 relation types. Assume the output results are: {belong to: 0.05, contain: 0.03, causal: 0.02, parallel: 0.01, superordinate: 0.85, ..., time: 0.01}.

[0022] Finally, the outputs of intra-layer relation probes from all layers are aggregated, the final probability of each type of relation between entities is calculated, relations with probabilities higher than a preset threshold are selected, and a weighted directed multigraph is constructed as a parametric belief graph. At the same time, a pre-trained graph attention network is used to encode the parametric belief graph to obtain a high-dimensional dense vector of each entity node.

[0023] The final probability is quantified by taking the average of the output probabilities of the intra-layer relation probes of all layers.

[0024] The parameter belief graph uses entities as nodes, predefined relationships as edges, and weights as the corresponding prediction probabilities.

[0025] In this embodiment, the graph attention network contains two layers of attention heads, with eight attention heads in each layer, and the output dimension is set to 512. After inputting the parameter belief graph into the network, high-dimensional dense vectors of each entity can be obtained. Its training data uses a manually labeled entity graph dataset, which contains entity graphs related to the query text domain, and each sample is labeled with a real high-dimensional semantic vector of the entity node. The training objective is to minimize the mean squared error between the network's output entity node vector and the real semantic vector. The graph attention network in training is periodically validated. Training stops when the mean squared error on the validation set no longer decreases for several consecutive rounds.

[0026] Step S12, construct the retrieval fact graph, specifically: First, an encoder compatible with the current large language model is used to encode the user query text into a dense vector. Then, the top-K document fragments with the highest cosine similarity to the dense vector are retrieved from the preset external vector knowledge base through maximum inner product search, and a set of document fragments is generated.

[0027] Subsequently, a BERT-based triple extraction model is used to extract fact triples from each document fragment, obtaining fact triples in the form of (subject, predicate, object). For example, document fragment 1 is "Company xx was founded in 1976 by person A and person B", and its corresponding fact triples include: (Company xx, founder, person A), (Company xx, founder, person B), (Company xx, founding time, 1976).

[0028] Next, the extracted fact triples are subjected to coreference resolution processing. Coreference resolution means using coreference resolution models such as SpanBERT to merge different references to the same real-world entity into a unified entity identifier.

[0029] Next, the processed fact triples are linked together to form a chain of evidence, that is, for each fact triple... and ,like Then establish a connection. .

[0030] For example, for the two triples (Company xx, Founder, Person A) and (Person A, Leading R&D, Machine 1), since the object "Person A" of the former is the same as the subject "Person A" of the latter, a connection can be established between the two to form the following chain of evidence: (Company xx, Founder, Person A) → (Person A, Leading R&D, Machine 1).

[0031] Finally, a weighted directed graph with entities as nodes, predicates as edges, and comprehensive scores as weights is constructed based on the evidence chain as the retrieval fact graph. The comprehensive score is represented as the product of the normalized value of the frequency of occurrence of a predefined relation and the document retrieval ranking correction value. The normalized value of the frequency of occurrence of a relation is represented as the frequency of occurrence of the relation in all document fragments divided by the maximum frequency of occurrence. The document retrieval ranking correction value is represented as 1 - document fragment ranking / total number of extracted document fragments. Simultaneously, the pre-trained graph attention network from step S11 is used to encode the retrieval fact graph, resulting in high-dimensional dense vectors for each entity node.

[0032] Step S13: Based on the parametric belief map and the retrieved fact map, generate the parametric belief field and the fact reference field, specifically: First, the entity node vectors of the parameter belief graph and the retrieval fact graph are respectively regarded as semantic spaces. Field values ​​at discrete sampling points, where The semantic space is represented by the entity node vector dimension, and is the space composed of all entity semantic vectors.

[0033] Taking the entity node vectors of a parametric belief graph as an example, the vectors of entity "xx company" and entity "main products" are... , Corresponding to discrete sampling points x1 and x2 in the semantic space, their field values ​​are respectively , Similarly, the entity node vectors retrieved from the fact graph correspond to discrete sampling points and field values ​​in the semantic space.

[0034] Next, a thermonuclear diffusion process is used to smoothly propagate the discrete field values ​​to the entire semantic space, generating a parametric belief field and a fact reference field.

[0035] It should be noted that for any point in the semantic space, the corresponding parametric belief field vector and fact reference field vector can be obtained through the thermonuclear diffusion operator. The thermonuclear diffusion operator for the thermonuclear diffusion process is defined as follows: ; in, This is the diffusion time hyperparameter; The dimension of the entity node vector; The coordinates are any position in the semantic space; For entity nodes The coordinates; For entity nodes The corresponding vector; It is the Euclidean norm.

[0036] Step S14: Construct a local Riemannian metric based on the parametric belief field and the factual reference field. Specifically: First, obtain the difference vector field between the parametric belief field and the factual reference field: ; in For the parameter belief field; As a field of factual reference; Difference vector field The dimension of the vector is consistent with that of the parameter belief field and the fact reference field, and is used to characterize the degree of difference between the internal cognition of the model and the external facts at a certain point in the semantic space.

[0037] Next, at each point in the semantic space At that point, a local Riemannian metric tensor is constructed by combining the aforementioned difference vector field. The local Riemannian metric tensor is represented as: ; in, It is the identity matrix; These are preset hyperparameters; Let the difference vector field be the parameter belief field and the fact reference field; This is represented as a vector outer product operation.

[0038] Step S15: Construct a curvature singularity map based on local Riemannian metrics. Specifically: First, based on the local Riemann metric tensor, an automatic differentiation technique is used to obtain a Riemann dataset consisting of Christofel symbols, Riemann curvature tensors, Ricci curvature tensors, and scalar curvature.

[0039] It should be noted that the Christofel notation used is of the first type, which can be represented as: ; in, For the first class of Christofel symbols, subscript , , The range of values ​​is It is a tensor that depends on the local Riemannian metric; For local Riemannian metric tensors The Line 1 Column elements represent the first element in the semantic space. peacekeeping The degree of correlation between semantic coordinates; Represented as local Riemannian metric tensor elements For semantic coordinates The partial derivatives of each component are used to reflect the rate at which the metric tensor changes with the semantic coordinates.

[0040] The Riemann curvature tensor can be expressed as: ; in, Let Riemann curvature tensor be the superscript. and subscript , , The range of values ​​is , used to characterize the bending properties in different directions in the semantic space; It is a second type of Christofel symbol, derived from the first type of Christofel symbol. Elements of the inverse matrix of the local Riemannian metric tensor Summation yields the result; For local Riemannian metric tensors The inverse matrix of the first Line 1 Column elements are used to convert first-class Christofel symbols to second-class symbols; Indicates to From 1 to Summation is performed to cover all dimensions of the semantic space.

[0041] The Ricci curvature tensor can be expressed as: ; in, Let Ricci curvature tensor be the index. , The range of values ​​is , used to represent the semantic space along the first line at a certain point. peacekeeping Average curvature in the dimensional direction; Let be the value of the Riemann curvature tensor under the corresponding subscript.

[0042] The scalar curvature can be expressed as: ; in, For a point in semantic space The scalar curvature at a point is a scalar value used to characterize the overall curvature at that point; Indicates to and From 1 to Perform a double summation.

[0043] Next, combined with the preset neighborhood radius Local extremum detection is performed in the semantic space to filter out local maxima and local minima of scalar curvature, and the local maxima and local minima are defined as positive and negative curvature singularities, respectively.

[0044] Among them, for point If its corresponding neighborhood radius Within the formed neighborhood, the scalar curvature values ​​of all points are less than 1. scalar curvature of a point ,but It is a singularity with positive curvature; Similarly, for point If its corresponding neighborhood radius Within the formed neighborhood, the scalar curvature values ​​of all points are greater than 1. scalar curvature of a point ,but It is a singularity with negative curvature.

[0045] Finally, the location, scalar curvature value, and corresponding type of all curvature singularities are recorded to construct a curvature singularity graph.

[0046] Step S2: Based on the variational information bottleneck, the retrieved content is adaptively purified by simultaneously combining local Riemannian metrics and curvature singularity graphs to obtain a set of fact kernels. Then, the fact kernels within the set of fact kernels are clustered using geodesic distance to generate a purity manifold.

[0047] Specifically, step S2 includes steps S21-S25: Step S21: Map the retrieved document fragments to the semantic space, obtain their semantic coordinates, and query the absolute value of the scalar curvature at those coordinates. Specifically: First, common semantic encoders such as the Sentence-BERT cross-encoder are used to encode the retrieved document fragments to generate corresponding semantic feature vectors.

[0048] Next, the semantic feature vector of each document fragment is mapped to the semantic space through a linear projection layer. Get the semantic coordinates corresponding to the document fragment.

[0049] The linear projection layer is represented as a pre-trained linear layer, whose input dimension is consistent with the semantic feature vector dimension, and whose output dimension is consistent with the corresponding vector dimension in the semantic space; its training objective is to minimize the Euclidean distance between the projected vector and the corresponding entity node vector in the semantic space.

[0050] Subsequently, for each semantic coordinate, the curvature singularity graph generated in step S1 is queried in combination with the preset query rules to obtain the corresponding scalar curvature absolute value.

[0051] The preset query rules include: If the semantic coordinates are within the neighborhood of a singularity in the curvature singularity graph, the scalar absolute value of curvature of the corresponding singularity is directly obtained. Otherwise, select the coordinate closest to that coordinate. The system identifies a singularity and obtains the corresponding absolute value of scalar curvature through local linear interpolation.

[0052] It should be noted that in this embodiment, Euclidean distance is used as the distance metric between the semantic coordinate position and the singular position within the curvature singularity graph.

[0053] For example, suppose a certain semantic coordinate is not within the neighborhood of any singularity in the curvature singularity graph, and by performing Euclidean distance calculation, the three singularities Q1, Q2, and Q3 that are closest to the current semantic coordinate in terms of Euclidean distance are obtained, and the Euclidean distances are 0.012, 0.015, and 0.011, respectively, and the corresponding absolute values ​​of scalar curvature are 0.8, 0.6, and 0.2, respectively; The corresponding interpolation weights are obtained using a Gaussian weighted interpolation method based on Euclidean distance. The calculation formula can be expressed as: ; in, For the current semantic coordinates and the first The Euclidean distance between the singularities, and , Let be the neighborhood radius, and be the denominator. The sum of the exponential terms corresponding to each singularity is used to normalize the weights, ensuring that the sum of all weights is 1. Assuming the radius of the domain in this example If the value is 0.01, then the difference weights of Q1, Q2, and Q3 in this example are 0.37, 0.17, and 0.47, respectively. Thus, the absolute value of the scalar curvature corresponding to this semantic coordinate is 0.492. After retaining two decimal places, the final absolute value of the scalar curvature is 0.49.

[0054] Step S22: Based on the variational information bottleneck and combined with the absolute value of scalar curvature, the feature vector of the document fragment is compressed and purified to obtain the corresponding fact kernel. Specifically: First, a fact-refining encoder consisting of an encoder and a decoder is constructed, where the encoder input is the semantic feature vector of a document fragment. The output is the mean of the latent variable distribution. With log variance The decoder input is a latent variable. The output is the reconstructed semantic feature vector of the document fragment. .

[0055] In this embodiment, both the encoder and decoder are composed of a 3-layer multilayer sensing architecture, with ReLU as the activation function. The loss function during pre-training is defined as follows: ; in, The total loss value during the pre-training process of the encoder is used to refine facts; Represented as latent variables Follows the posterior distribution Under the premise that the expression within the parentheses is expected; Represented as posterior distribution With prior distribution KL divergence between them; Represented as latent variables The posterior probability distribution is used to describe the probability distribution in a given document fragment. semantic feature vector Under the condition of latent variables The probability distribution of all possible values, which is derived from the mean of the encoder output of the fact-refined encoder. Sum of logarithmic variance Sure; Represented as given latent variables Under these conditions, document fragments semantic feature vector The reconstructed log-likelihood value; The prior is the standard normal distribution; This is a predefined KL divergence penalty coefficient.

[0056] Next, during inference, the KL divergence penalty coefficient is dynamically set based on the absolute value of the scalar curvature of the semantic coordinates of the document fragment. The specific formula is as follows: ; in, Pre-set baseline coefficients for the specific domain; Preset curvature sensitivity for the specific domain; semantic coordinates corresponding to document fragments The absolute value of the scalar curvature at a point is used to characterize the degree of curvature in the semantic spacetime of location.

[0057] For example, taking document fragment 1 as an example, its absolute value of scalar curvature Then the corresponding This indicates that the document fragment is located in a high curvature region and requires stronger compression to filter noise; taking document fragment 4 as an example, its scalar curvature absolute value Then the corresponding This indicates that the document fragment is located in a low curvature region, and the compression level can be appropriately reduced.

[0058] Finally, the semantic feature vector of the document fragment is input into the aforementioned fact-refining encoder, and the latent variable vector is sampled from the latent variable distribution output by the encoder. This latent variable vector is the fact kernel. The semantic coordinates of the fact kernel are consistent with the semantic coordinates of the corresponding document fragment.

[0059] Taking the feature vector v1 of document fragment 1 as an example, after inputting it into the fact refinement encoder, the mean is obtained. Sum of logarithmic variance Sampling yields fact kernel ,in, , It is the identity matrix. Its semantic coordinates are x1; after inputting the feature vector v2 of document fragment 2, the fact kernel is obtained. The semantic coordinate is x2; similarly, the fact cores of the remaining document fragments are obtained in the same way. , , This forms a fact core set. .

[0060] Step S23, obtain the geodesic distance between each pair of fact kernels in the fact kernel set, specifically: Obtain the local Riemannian metric of the current semantic space generated in step S1. The fast traversal algorithm is used to obtain the geodesic distance between each pair of fact kernels.

[0061] Based on facts and For example, The semantic coordinates are , The semantic coordinates are , and Between Geodesic distance below The quantification process is as follows: A uniform grid is constructed in the semantic space, assuming a grid resolution of 0.001, covering... and The semantic region in which it is located; by Using the source point, solve the equation of the function. ; in, Let be the distance function, which can be expressed as: ; Among them, the integration path For the semantic space from the source point To the target point geodesic lines, For any node in the grid; Let be the parametric expression for the geodesic, and ,For example, ,Right now When, corresponding , When, corresponding ; geodesic In parameters The tangent vector at the point; Represented as a local Riemannian metric tensor exist Tangent vector at point The result of its effect can be expressed as: ; Since the magnitude of the tangent vector is given, the result of the above integral can characterize the geodesic. The length of the source point To the target point geodesic distance .

[0062] Step S24: Based on geodesic distance, the DBSCAN clustering algorithm is used to cluster the fact kernel set to generate fact kernel clusters. Specifically: First, using geodesic distance as a metric, the DBSCAN clustering algorithm is used to cluster all fact kernels.

[0063] It should be noted that the neighborhood radius of the DBSCAN clustering algorithm... Set as an adaptive value, it can be adaptively adjusted using the following formula: ; in, A pre-defined basic neighborhood radius for domain experts; Curvature adjustment coefficients pre-set for domain experts; The absolute value of the local average scalar curvature is given. The local range corresponding to this parameter is a spherical region with a radius of 1 / 2 of the maximum geodesic distance between all pairs of fact kernels. The minimum number of samples for the DBSCAN clustering algorithm is set to 3, that is, when a certain fact kernel... When a neighborhood contains at least three fact kernels, it is used as the core point to form a cluster.

[0064] Next, the geodesic compactness of each cluster is obtained, which is expressed as the reciprocal of the average geodesic distance between all pairs of fact kernels within the cluster; The min-max normalization method was used to normalize the geodesic compactness of all clusters, and the normalized geodesic compactness was used as the confidence score of each fact kernel in the corresponding cluster.

[0065] Ultimately, fact kernels with confidence scores higher than a preset threshold are retained.

[0066] Step S25: Based on the filtered fact kernels, a purity manifold is constructed by combining their semantic coordinates, confidence scores, and local Riemannian metrics. Specifically: Collect the filtered set of fact kernels, each fact kernel containing a corresponding vector, semantic coordinates and confidence score; By integrating the set of fact kernels with local Riemannian metrics, a purity manifold is constructed, for example, for a set containing... , , The set of fact kernels, whose corresponding purity manifold can be represented as {set of fact kernels: [{fact kernels: Semantic coordinates: x1, confidence level: 0.94}, {fact kernel: Semantic coordinates: x2, confidence level: 0.94}, {fact kernel: Semantic coordinates: x3, confidence level: 0.94, Local Riemannian metric: }

[0067] Step S3: Based on the hidden state sequence, local Riemannian metric, curvature singularity graph and fact kernel set in the decoding process of the large language model, the hallucination bias sequence is predicted simultaneously by combining the neural Jacobian solver, and the geometric correction force vector is generated according to the geodesic repair force network.

[0068] Specifically, step S3 includes steps S31-S33: Step S31: The token-by-token generation process of the large language model is regarded as a continuous motion process on a Riemannian manifold. At each decoding time step, the current hidden state is projected onto the semantic space to obtain semantic coordinates. Specifically: First, we clarify the definition of a Riemannian manifold, whose base space is a semantic space. The metric is a local Riemannian metric. Each point on the manifold corresponds to a semantic state, and the path on the manifold corresponds to the model generation process.

[0069] Next, at each decoding time step in the large language model decoding process Based on the pre-trained projection matrix The current hidden state Projecting onto the semantic space yields the corresponding semantic coordinates. .

[0070] The projection matrix The training objective is to minimize the geodesic distance between the projected semantic coordinates and the corresponding fact kernel coordinates.

[0071] For example, using decoding time steps For example, the current hidden state Through projection matrix Projection yields semantic coordinates This coordinate corresponds to the semantic state after the first token is generated; similarly, the time step Hidden state After projection, semantic coordinates are obtained. This corresponds to the semantic state after generating the second token, and so on, forming a semantic coordinate sequence. , The length of the generated sequence.

[0072] Finally, the ideal hallucination-free generation path is defined as the connection query semantic point. semantic points of the answer A geodesic that satisfies the following equation: ; in, Represented as the semantic coordinate of the semantic space There are 1 component, of which The range of values ​​is , For semantic space dimension; For decoding time steps; Represented as semantic coordinates Each component corresponds to a decoding time step. The second derivative; For Christofel symbols of the first class, it is dependent on local Riemannian metrics. and semantic coordinates The tensor, its subscript , The range of values ​​is Corresponding to the dimension of semantic coordinates, it is used to characterize the influence of the curvature of semantic space on geodesics; Represented as semantic coordinates Each component corresponds to a decoding time step. The first derivative; Represented as semantic coordinates Each component corresponds to a decoding time step. The first derivative.

[0073] When the above equation holds true, the model generation path is the ideal path without illusions; If the equation does not hold, it indicates that the generated path deviates from the geodesic and there is a risk of illusion. It will be corrected by geometrically correcting the force vector.

[0074] Step S32: Input the Ricci curvature tensor at the current semantic coordinates into the neural Jacobian solver to obtain the hallucination bias sequence for multiple future time steps. Specifically: First, at each decoding time step Obtain the motion tangent vector This is used to characterize the direction of motion generated by the model, and the specific formula can be expressed as: ; in, Let be the projection matrix from the latent state space to the tangent space of the semantic space; This is the current hidden state. This is the hidden state from the previous moment.

[0075] Next, obtain the current semantic coordinates. Ricci curvature tensor at the location This is used to characterize the influence of the curvature of the semantic space at that point on the motion path; The and The input is processed by a neural Jacobian solver, which outputs a Jacobian field sequence for multiple future time steps.

[0076] It should be noted that the neural Jacobian solver is represented as a neural differential equation network, consisting of multiple ODE blocks. Each ODE block contains three layers of multilayer perceptrons with ReLU activation function and a tangent vector as the input dimension. With Ritchie curvature tensor The sum of the dimensions, and the output dimension is , To predict the number of steps; The neural Jacobian solver simulates the classical Jacobian equation, which is expressed as: ; in, Along the tangent vector The second covariant derivative; For Jacobi; This is the bilinear form of the Riemann curvature tensor, used to characterize the effect of curvature on the evolution of the Jacobian field.

[0077] The neural Jacobian solver simulates the evolution of the Jacobian field in the semantic space through iterative computation of ODE blocks, thereby obtaining the future... A Jacobian field sequence at each time step.

[0078] Next, the Jacobian field at each time step is projected onto a space with the same dimension as the semantic coordinates through a linear projection layer to obtain the illusion bias. This allows us to obtain the illusion bias sequence.

[0079] The projection matrix of the linear projection layer is a pre-trained matrix, and the training objective of this matrix is ​​to minimize the mean square error between the illusion bias and the actual deviation.

[0080] The illusion bias Represented as at time step The deviation vector between the model's current generated path and the ideal geodesic. The larger the modulus, the greater the degree of deviation, and the higher the probability of hallucination. The direction is used to characterize the specific semantic direction of the deviation.

[0081] Finally, the hallucination bias sequence is normalized, and the bias vector at each time step is normalized to the interval [-1, 1]. The normalization formula is as follows: ; in, for The model; This is a local minimum value, used to avoid cases where the denominator is zero.

[0082] Step S33: Input the current hidden state and the illusion bias sequence into the geodesic repair force network, and output the geometric correction force vector opposite to the bias direction. Specifically: The geodesic restoration force network consists of a fact attention layer, a curvature attention layer, and a correction force generation layer. Its specific input is a normalized illusion bias sequence. Current semantic coordinates Local Riemannian metric at [location] Curvature singularity diagram Singularity information within the neighborhood and distance in the fact kernel set Recent The vector and confidence level of each fact kernel are used to output a geometric correction force vector that is opposite to the bias direction.

[0083] It should be noted that the fact attention layer is used to mine the correlation between the fact kernel and the current semantic state, and simultaneously generate the contribution weights of the nearest fact kernels to the geometric correction force. The formula for the contribution weights can be expressed as: ; in, Represented as current semantic coordinates With the A close neighbor fact core semantic coordinates Geodesic distance between them; For the fact The confidence score.

[0084] The curvature attention layer dynamically adjusts the strength of the correction force based on the scalar curvature values ​​and types of neighboring singularities. Positive curvature singularities correspond to a stronger compressive correction force, while negative curvature singularities correspond to a stronger complementary correction force. Specifically, curvature weights are set based on the type and absolute value of the scalar curvature values ​​of the neighboring singularities, with a weight of [value missing] for positive curvature singularities. The weight of the negative curvature singularity is , The scaling factor is preset for domain experts, and ; For example, suppose that the neighborhood of a certain semantic coordinate in a curvature singularity graph contains three singularities: U1{type: positive curvature singularity, scalar curvature value: 0.8}, U2{type: negative curvature singularity, scalar curvature value: -0.6}, and U3{type: positive curvature singularity, scalar curvature value: 0.3}. In this example, 'a' is set to 1.2. Then, the weights corresponding to the above three singularities are 0.8, |-0.6|*1.2=0.72, and 0.3, respectively. After L1 normalization, the final weights are approximately 0.44, 0.40, and 0.16, respectively.

[0085] The correction force generation layer employs a two-layer multilayer perceptron, combining the weights output from the fact attention layer and the curvature attention layer to output the final geometric correction force vector. Specifically: First, select Each nearest neighbor fact kernel vector is multiplied by its corresponding fact attention weight to obtain a weighted fact kernel vector. Then, these three weighted vectors are summed to obtain a fact guidance vector, which is used to ensure that the direction of the correction force is aligned with the high-confidence facts. Next, the curvature weights of multiple singularities in the neighborhood are multiplied by the label curvature values ​​of the corresponding singularities and then summed to obtain the curvature intensity coefficient. This coefficient is used to dynamically adjust the overall intensity of the correction force. The larger the coefficient, the stronger the correction force. Subsequently, the fact-guided vector, curvature intensity coefficient and the feature vector of the current local Riemann metric are concatenated and used as the input of the two-layer multilayer perceptron. The first layer multilayer perceptron fuses and extracts the input features. The activation function of this layer is ReLU. The second layer multilayer perceptron maps the fused features output by the first layer into a vector with the same dimension as the semantic coordinates, generating a preliminary corrected force vector. Finally, the initial corrected force vector is multiplied by the corresponding normalized illusion bias vector to adjust the direction of the corrected force vector and generate the geometric corrected force vector.

[0086] For each decoding time step, the corresponding geometric correction force vector is generated according to the above process, forming a sequence of correction force vectors.

[0087] In step S4, at each decoding time step, the local Riemann metric is repaired in real time according to the Ricci flow simulation network, and based on the repaired local Riemann metric, purity manifold constraint and geometric correction force vector, model illusion suppression is performed to generate illusion-free search content.

[0088] Specifically, step S4 includes steps S41-S44: Step S41: Real-time repair of the local Riemann metric for each decoding time step based on the Ricci flow simulation network, specifically: The Ricci flow simulation network uses a U-Net structure, with the current decoding time step as input. Local Riemannian metric Current semantic coordinates scalar curvature value at and curvature singularity diagram The output is the restored local Riemannian metric, which contains information on singularities within the neighborhood. .

[0089] The Ricci flow simulation network comprises four encoding blocks and four decoding blocks. Each encoding block and decoding block contains a convolutional layer, a batch normalization layer, and a ReLU activation function. The training objective is to minimize the difference between the extreme value of the scalar curvature of the repaired local Riemannian metric and a preset threshold, until the difference is lower than the preset threshold and the fluctuation amplitude within multiple consecutive periods is lower than the preset threshold, or the preset maximum training step size is reached.

[0090] After pre-training, the Ricci flow simulation network extracts features from the input data through encoding blocks, captures the differences between diagonal and off-diagonal elements in the local Riemannian metric tensor, and simultaneously extracts scalar curvature. The numerical features and structured features such as the location coordinates, curvature values, and types of neighborhood singularities are extracted and fused to form a multi-dimensional feature map. The decoding block reduces the number of channels in the multi-dimensional feature map by using a transposed convolutional layer, and fuses the shallow features of the encoding block and the deep features of the decoding block by skip connections. During the decoding and reconstruction process, the Ricci flow simulation network dynamically allocates weights through an attention mechanism, focusing on correcting the metric elements corresponding to high curvature regions: for positive curvature regions, it reduces the values ​​of the diagonal elements of the metric tensor in that region, narrowing the difference between them and the surrounding elements; for negative curvature regions, it increases the values ​​of the diagonal elements of the metric tensor in that region.

[0091] Local Riemannian measurement after restoration Perform a symmetric positive definite check. If the check fails, that is, there are asymmetric elements or negative eigenvalues, then project them to the symmetric positive definite matrix space using the projection operator.

[0092] by Taking elements with medium to high curvature as an example, let's assume... A certain element Assuming the element's value is 0.63 higher than the average of its surrounding elements, corresponding to a positive curvature region, the Ricci flow simulation network adjusts this element to 0.65 through the decoding block, making the curvature of this region more uniform; for elements corresponding to negative curvature regions... If the value is lower than the mean of surrounding elements, the network adjusts it to 0.60 to eliminate excessive bending.

[0093] For each decoding time step, the local Riemann metric is repaired according to the above process to obtain the repaired local Riemann metric sequence.

[0094] Step S42, based on the purity manifold, the semantic coordinates of the current decoding time step. Imposing constraints, specifically: First, obtain the semantic coordinates of the current decoding time step. Geodesic distance to the purity manifold The distance of the geodesic line Represented as The minimum geodesic distance to the semantic coordinates of all fact kernels within the purity manifold.

[0095] Next, the average geodesic distance between all pairs of fact kernels within the purity manifold is defined as the purity constraint threshold. ; like This indicates that the current semantic state is close to the objective facts and requires no additional constraints; like This indicates that the current semantic state deviates from the factual constraints and needs to be adjusted.

[0096] The constraint adjustment is expressed as based on Soon The updated semantic coordinates are generated by weighting the semantic coordinates of each fact kernel. ; in, Represented as the fact kernel in the purity manifold semantic coordinates; The constraint weights for each fact kernel can be expressed as: ; Represented as and The geodesic distance between them.

[0097] It should be noted that the constraint weights L1 normalization must be performed before weighted fusion can be performed.

[0098] Step S43: Adjust the hidden state of the current decoding time step based on the geometric correction force vector. Specifically: First, the geometric correction force vector is projected onto the hidden state space using a preset projection matrix to generate the hidden state correction vector. The preset projection matrix and the projection matrix used in step S31 are transposes of each other.

[0099] Next, based on the corrected strength coefficient The hidden state correction vector is adjusted to obtain the final hidden state adjustment amount. ,Right now ; The corrected intensity coefficient Based on the current semantic coordinates To purity manifold The distance is dynamically adjusted, that is: ; The Represented as current semantic coordinates To purity manifold The geodesic distance.

[0100] Finally, the hidden state adjustment amount Superimposed on the current hidden state Obtain the corrected hidden state .

[0101] Step S44, based on the corrected hidden state Local Riemannian metric after restoration Including purity manifold constraints, the system decodes and generates illusion-free search content through a large language model. Specifically: First, the corrected hidden state Input the decoder of the large language model and perform token generation. In the attention mechanism of the decoder, a purity manifold constraint is introduced, that is, the calculation of attention weights needs to combine the similarity between the current hidden state and the fact kernel vector. The higher the similarity, the greater the attention weight, thereby ensuring that the generated tokens are consistent with objective facts.

[0102] The quantization function of the similarity can be expressed as: ; in, Represented as a fact kernel vector The projection vector into the hidden state space.

[0103] Next, during the token generation process, an illusion detection mechanism is introduced to determine in real time whether the generated token has the risk of illusion.

[0104] The hallucination detection mechanism is described as obtaining the geodesic distance between the semantic coordinates corresponding to the generated token and the purity manifold. If the geodesic distance is greater than... If the distance is less than or equal to the distance, the token is deemed to have a hallucination risk, the token is rejected, and a new token conforming to the factual constraints is generated by resampling; if the distance is less than or equal to the distance, the token is considered to have a hallucination risk. If the token is deemed to have no hallucination risk, the generation result is retained.

[0105] Finally, the above process is repeated, with the hidden state corrected and the generated token subjected to illusion detection at each decoding time step, until the complete search content is generated.

[0106] Figure 2 This is a schematic diagram of a model hallucination suppression system based on dual-path vectorized retrieval according to the present invention.

[0107] Specifically, a model-based hallucination suppression system based on dual-path vectorized retrieval includes: The data processing module is used to construct a parametric belief field and a fact reference field, and to construct a local Riemannian metric based on the parametric belief field and the fact reference field, and to construct a curvature singularity map based on the local Riemannian metric. The manifold construction module is used to adaptively purify the retrieved content by combining local Riemannian metrics and curvature singularity graphs, obtain a set of fact kernels, and cluster the fact kernels within the set of fact kernels to generate a purity manifold. The hallucination prediction module is used to predict hallucinations by combining the hidden state sequence in the decoding process of the large language model with the local Riemann metric, curvature singularity graph and fact kernel set according to the neural Jacobian solver, and to obtain the predicted hallucination bias sequence. A correction generation module is used to generate a geometric correction force vector based on the geodesic repair force network and the predicted hallucination bias sequence. The illusion correction module is used to suppress model illusions based on local Riemannian metrics, purity manifold constraints, and geometric correction force vectors, and generate illusion-free search content.

[0108] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0109] It should be understood that, in the embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0110] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A model hallucination suppression method based on dual-path vectorized retrieval, characterized in that, It includes the following steps: S1: Construct a parametric belief field and a factual reference field based on the user query text and a preset external vector knowledge base, and construct a local Riemannian metric based on the parametric belief field and the factual reference field. Simultaneously combine the local Riemannian metric to locate the illusion source point and generate a curvature singularity map. S2: Based on the variational information bottleneck, the retrieved content is adaptively purified by simultaneously combining local Riemannian metrics and curvature singularity graphs to obtain a set of fact kernels. The fact kernels within the set are then clustered using geodesic distance to generate a purity manifold. S3: Based on the hidden state sequence, local Riemannian metric, curvature singularity graph and fact kernel set in the decoding process of the large language model, the hallucination bias sequence is predicted simultaneously by combining the neural Jacobian solver, and the geometric correction force vector is generated according to the geodesic repair force network. S4: At each decoding time step, the local Riemann metric is repaired in real time according to the Ricci flow simulation network. Based on the repaired local Riemann metric, purity manifold constraints and geometric correction force vector, model illusion suppression is performed to generate illusion-free search content.

2. The model hallucination suppression method based on dual-path vectorized retrieval according to claim 1, characterized in that, Step S1 includes: Based on the intra-layer relation probe, named entity recognition and inter-entity relation prediction are performed on the user query text to obtain the entity set and the corresponding prediction probability. The intra-layer relation probe consists of a linear layer and a sigmoid activation function. The input dimension matches the hidden state of the corresponding layer, and the output dimension is the total number of predefined relation types. Based on the entity set and the predicted probability, a weighted directed graph with entities as nodes, predefined relationships as edges and predicted probabilities as weights is constructed as a parametric belief graph. The user query text is encoded into a dense vector. The top-K document fragments with the highest similarity to the dense vector are retrieved in the preset external vector knowledge base by maximum inner product search, and a set of document fragments is obtained. For each document fragment, a lightweight fact extraction model is used to extract fact triples. Coreference resolution is performed on all extracted triples, and the processed fact triples are connected into a chain of evidence. Based on the aforementioned evidence chain, a weighted directed graph with entities as nodes, predicates as edges, and comprehensive scores as weights is constructed as the retrieval fact graph; The overall score is represented as the product of the normalized value of the frequency of occurrence of predefined relationships and the document retrieval ranking correction value.

3. The model hallucination suppression method based on dual-path vectorized retrieval according to claim 2, characterized in that, The method further includes: The entity node corresponding vectors in the parameter belief graph and the retrieval fact graph are respectively regarded as field values ​​on discrete sampling points in the semantic space; The thermonuclear diffusion process is used to smoothly propagate discrete field values ​​to the entire semantic space, generating a parametric belief field and a fact reference field. The thermonuclear diffusion operator for the thermonuclear diffusion process is defined as: ; in, This is the diffusion time hyperparameter; The dimension of the entity node vector; The coordinates are any position in the semantic space; For entity nodes The coordinates; For entity nodes The corresponding vector; It is the Euclidean norm; For any point in the semantic space, the corresponding parameter belief field vector and fact reference field vector are obtained through the thermal kernel diffusion operator.

4. The model hallucination suppression method based on dual-path vectorized retrieval according to claim 3, characterized in that, The method further includes: Obtain the difference vector field between the parametric belief field and the factual reference field at each point in the semantic space. At that point, a local Riemannian metric tensor is constructed based on the difference vector field; The local Riemannian metric tensor is represented as: ; in, It is the identity matrix; These are preset hyperparameters; Let the difference vector field be the parameter belief field and the fact reference field; This is represented as a vector outer product operation; Based on the local Riemann metric tensor, a Riemann dataset consisting of Christofer symbols, Riemann curvature tensor, Ricci curvature tensor, and scalar curvature is obtained; Perform local extremum detection in the semantic space to filter out local maxima and local minima of scalar curvature; The local maxima and local minima are defined as positive and negative curvature singularities, respectively. Simultaneously record the position, scalar curvature value, and corresponding type of all curvature singularities to construct a curvature singularity graph.

5. The model hallucination suppression method based on dual-path vectorized retrieval according to claim 1, characterized in that, Step S2 includes: A linear projection layer is used to map each document fragment in the retrieved content into a semantic space to obtain the semantic coordinates corresponding to the document fragment; For each semantic coordinate, query the curvature singularity graph; If the semantic coordinates are located within the neighborhood of the singularity in the curvature singularity graph, the absolute value of the scalar curvature corresponding to the singularity is directly obtained. Otherwise, the absolute value of the scalar curvature is obtained through local linear interpolation; A fact-refining encoder based on the absolute value of scalar curvature is used to process any of the document fragments to obtain fact kernels, and all fact kernels constitute a fact kernel set.

6. The model hallucination suppression method based on dual-path vectorized retrieval according to claim 5, characterized in that, Fact-refining encoder, including: The fact-refining encoder operates based on the variational information bottleneck principle. For document fragments The KL divergence penalty coefficient is dynamically set as follows: ; in, Preset base coefficients; Preset curvature sensitivity; semantic coordinates corresponding to document fragments The absolute value of the scalar curvature at a point is used to characterize the degree of curvature in the semantic spacetime of location; The fact refinement encoder outputs a compressed latent variable vector, which is the fact kernel.

7. The model hallucination suppression method based on dual-path vectorized retrieval according to claim 5, characterized in that, The method further includes: Based on the local Riemannian metric corresponding to the current semantic space, a fast marching algorithm is used to obtain the geodesic distance between each pair of fact kernels. Using geodesic distance as a metric, the DBSCAN clustering algorithm is used to cluster all fact kernels, and the geodesic compactness of each cluster is obtained simultaneously. The normalized geodesic compactness is used as the confidence level of each fact kernel within the cluster; We retain fact kernels with confidence scores higher than a preset threshold and construct a purity manifold by combining the corresponding semantic coordinates, confidence scores, and local Riemannian metrics.

8. The model hallucination suppression method based on dual-path vectorized retrieval according to claim 1, characterized in that, Step S3 includes: The token generation process of the large language model is regarded as a continuous motion process on the Riemannian manifold. At each decoding time step, the current hidden state is projected onto the semantic space to obtain the semantic coordinates. Input the Ricci curvature tensor at the current semantic coordinates into the neural Jacobian solver to obtain the illusion bias sequence for multiple future time steps; Input the current hidden state and the illusion bias sequence into the geodesic repair force network, and output the geometric correction force vector opposite to the bias direction.

9. A model hallucination suppression method based on dual-path vectorized retrieval according to claim 1, characterized in that, Step S4 includes: At each decoding time step, obtain the scalar curvature and Ricci curvature tensors at the corresponding semantic coordinates during the decoding process of the large language model; The scalar curvature and Ricci curvature tensors are input into a pre-trained Ricci flow simulation network to obtain the metric correction increment, and the local Riemann metric at the current position is updated based on the metric correction increment. The large language model suppresses model illusions by using the repaired local Riemannian metric, purity manifold constraints, and geometrically corrected force vectors to generate illusion-free search content.

10. A model hallucination suppression system based on dual-path vectorized retrieval, used to implement the method described in any one of claims 1 to 9, characterized in that, include: The data processing module is used to construct a parametric belief field and a fact reference field, and to construct a local Riemannian metric based on the parametric belief field and the fact reference field, and to construct a curvature singularity map based on the local Riemannian metric. The manifold construction module is used to adaptively purify the retrieved content by combining local Riemannian metrics and curvature singularity graphs, obtain a set of fact kernels, and cluster the fact kernels within the set of fact kernels to generate a purity manifold. The hallucination prediction module is used to predict hallucinations by combining the hidden state sequence in the decoding process of the large language model with the local Riemann metric, curvature singularity graph and fact kernel set according to the neural Jacobian solver, and to obtain the predicted hallucination bias sequence. A correction generation module is used to generate a geometric correction force vector based on the geodesic repair force network and the predicted hallucination bias sequence. The illusion correction module is used to suppress model illusions based on local Riemannian metrics, purity manifold constraints, and geometric correction force vectors, and generate illusion-free search content.