A large model output reliability enhancement method based on a heterogeneous double model

By employing heterogeneous dual-model deep learning and reinforcement learning methods, the problem of output inaccuracy in large language models when faced with implicit semantic drift and dynamic changes is solved, achieving logically coherent and timely reliable output, thereby enhancing user trust.

CN122264113APending Publication Date: 2026-06-23SMO CLINPLUS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SMO CLINPLUS CO LTD
Filing Date
2026-03-24
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing large language models struggle to generate outputs that accurately align with reality when faced with implicit semantic drift and dynamic changes caused by symptoms or time differences. Furthermore, they lack timeliness verification and backtracking mechanisms for the reasoning chain, resulting in output content that lacks logical interpretability and user trust.

Method used

A heterogeneous dual-model approach is adopted, which uses deep learning semantic parsing and fruit fly algorithm to jointly optimize and accurately identify implicit semantic drift and dynamically activate localized knowledge subgraphs. A reinforcement learning mechanism is introduced to repair inference breakpoints, and a temporal logic scoring card is constructed to verify nodes, forming a traceable and reliable output.

Benefits of technology

It significantly improves the reliability and user trust of large model output, ensuring that the generated content is logically coherent, timely, accurately matches user needs, and provides verifiable output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122264113A_ABST
    Figure CN122264113A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of natural language processing of clinical trial medical texts, and specifically discloses a method for enhancing the reliability of large model output based on a heterogeneous double model, which comprises: performing deep learning semantic analysis on user input, fusing symptom and time feature encoding, combining fruit fly algorithm collaborative optimization, accurately positioning implicit semantic drift, and identifying implicit symptoms and time features; according to the identified symptoms and time features, dynamically activating and matching the corresponding localized knowledge sub-graph in the knowledge graph; the present application introduces time perception coding and symptom entity vector matrix in the semantic feature space through deep learning semantic analysis and fruit fly algorithm collaborative optimization, can accurately identify the implicit semantic drift in the user input due to the difference in symptoms or time, the two-stage search mechanism of the fruit fly algorithm gradually converges the semantic analysis range to the global optimal position, locates the semantic center of the user's real intention, and provides a semantic benchmark for knowledge retrieval and reasoning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing technology for clinical trial medical texts, and in particular to a method for enhancing the reliability of large model output based on heterogeneous dual models. Background Technology

[0002] As the scale of models expands and the application scenarios become more complex, relying solely on the results generated by the model itself is insufficient to guarantee the accuracy, consistency, and interpretability of the information. Therefore, improving the reliability of large model outputs has become an important direction for artificial intelligence research and application.

[0003] For example, Chinese Patent Publication No. CN119443230A describes a method and system for enhancing the reliability of large language models based on knowledge graphs. When using knowledge graphs to guide or constrain the generation of output text in a large language model, the generated output text can be injected into the large model as new knowledge, enabling the large model to use new knowledge in a timely manner. This effectively solves the problem of knowledge obsolescence in current large language models and further improves the reliability and accuracy of the natural language generated by the large language model.

[0004] In existing technologies, user input often contains implicit semantic drift due to differences in symptoms or time. Static knowledge graphs struggle to capture such dynamic changes, causing the model to fail to accurately align its generated answers with the local context. Even after initially anchoring the semantic scope through localized subgraphs, temporary relationships specific to particular scenarios within the graph are prone to loss, resulting in breaks in the model's reasoning chain at critical nodes. Furthermore, when the model output involves complex long-path reasoning, although the generated content appears complete, the lack of verification and backtracking mechanisms for the timeliness of each reasoning node leads to a lack of interpretability in the actual logical chain of the overall output, making it difficult for users to trust and adopt it. To address these issues, we propose a method for enhancing the reliability of large model output based on heterogeneous dual models. Summary of the Invention

[0005] To overcome the shortcomings of the prior art, the present invention provides a method for enhancing the reliability of large model output based on heterogeneous dual models, which can effectively solve the problems involved in the prior art.

[0006] The objective of this invention can be achieved through the following technical solution: This invention provides a method for enhancing the reliability of large model output based on heterogeneous dual models, comprising the following steps:

[0007] Step 1: Perform deep learning semantic parsing on user input, integrate symptom and time feature encoding, and combine fruit fly algorithm for collaborative optimization to accurately locate implicit semantic drift and identify implicit symptom and time features, ensuring that the model accurately understands the user's intent and eliminates semantic ambiguity caused by differences in symptom time.

[0008] Step 2: Based on the identified symptoms and time features, dynamically activate and match the corresponding localized knowledge subgraphs in the knowledge graph to narrow the semantic range and complete the initial anchoring of the semantic range, so that the model initially focuses on the knowledge area closely related to the user input.

[0009] Step 3: Based on real-time interactive feedback, a reinforcement learning mechanism is introduced to drive the graph neural network to perform dynamic relationship reasoning on the localized subgraph, automatically mining and completing the missing temporary relationship paths in the graph, repairing reasoning breakpoints, ensuring the integrity and timeliness of the knowledge chain, and avoiding interruption of the reasoning process due to missing relationships.

[0010] Step 4: Based on the completed dynamic knowledge subgraph, conduct long-path association reasoning, construct a complete reasoning link with time-sensitive markers, and form a logically coherent and time-sequential complete reasoning path.

[0011] Step 5: Use deep learning to construct a time-series logic scoring card, and perform timeliness verification and credibility scoring on each inference node. Filter out invalid logic, retain highly reliable inference paths, and eliminate invalid or untrustworthy logic nodes to ensure that every piece of information used in the final basis is true and valid.

[0012] Step 6: Based on the multi-dimensional logical verification results, optimize and adjust the generated content to form reliable output text with traceable anchor points, thereby increasing users' trust and acceptance of the model output and realizing the verifiability and traceability of the generated content.

[0013] Preferably, step 1 includes the following steps:

[0014] By using an encoder based on the Transformer architecture to perform deep semantic parsing of user input, and by introducing time-aware positional encoding and symptom entity vector matrix, the implicit time and symptom features in the text are mapped to a high-dimensional semantic space, thereby achieving deep semantic alignment and high-dimensional expression of spatiotemporal information in the text.

[0015] A symptom-time joint feature embedding layer is constructed, and the mapped features are deeply fused and encoded to generate a semantic feature tensor carrying spatiotemporal context information, which serves as the basic input for subsequent optimization processing, forming a spatiotemporally coupled joint semantic representation for subsequent optimization.

[0016] By introducing the olfactory search mechanism of the fruit fly algorithm into the feature space, and through swarm intelligence iterative optimization, the symptom weights and time thresholds in the semantic feature tensor are calibrated collaboratively to lock the precise location of implicit semantic drift. Adaptive collaborative calibration of spatiotemporal weights is achieved through swarm intelligence.

[0017] Preferably, step 1 further includes:

[0018] The position of the fruit fly swarm in the semantic feature space is initialized. The Euclidean distance of the semantic feature tensor is used as the taste concentration determination value to drive the swarm to perform a large-scale coarse olfactory search to identify potential drift regions. Potential drift regions are quickly located through a large-scale random search.

[0019] Based on the odor concentration value feedback from the olfactory search, the flight direction and stride of the fruit fly swarm are dynamically adjusted, and a visual fine search mechanism is adopted to accurately define the boundaries of symptoms and time features in the candidate drift area. The search direction is dynamically adjusted according to the feedback to achieve fine definition of feature boundaries.

[0020] The visual search is performed iteratively until the population converges, and the globally optimal individual position coordinates are output. The semantic center point corresponding to these coordinates is taken as the final localization result of the implicit semantic drift. The globally optimal semantic drift localization result is obtained through iterative convergence.

[0021] Preferably, step 2 includes the following steps:

[0022] The identified symptoms and time features are transformed into knowledge graph query commands. The symptoms nodes and time attribute edges in the graph are traversed to filter out the candidate knowledge subgraph set that meets the feature conditions, so as to achieve accurate mapping from user intent to graph query and significantly reduce the range of data to be processed.

[0023] Calculate the semantic similarity between the semantic center of the user input and the set of candidate knowledge subgraphs. Select the knowledge subgraph with the highest similarity as the activation object through a graph matching algorithm to complete the dynamic activation of the localized knowledge subgraph. This ensures that the activated subgraph is highly matched with the user's current needs and avoids the introduction of interference from irrelevant knowledge.

[0024] The activated localized knowledge subgraph is loaded into the working memory, severing the generalized connection channel between the large model and the global knowledge base. This ensures that the model's reasoning scope is precisely anchored within the semantic region defined by the knowledge subgraph, physically isolating global noise and fundamentally eliminating semantic drift during the reasoning process.

[0025] Preferably, step 3 includes the following steps:

[0026] The initially anchored localized knowledge subgraph is input into the graph neural network for initial state encoding. At the same time, a real-time interactive feedback receiving channel is constructed to transform the user's click stream and follow-up questions into reward signals for reinforcement learning. This realizes the spatial vectorization representation of the subgraph state and the quantitative transformation of the user's implicit feedback, laying the foundation for subsequent inference optimization.

[0027] The reward signal is input into the reinforcement learning model, which drives the graph neural network to explore paths between nodes in the knowledge subgraph with the goal of maximizing the cumulative reward. It identifies the breakpoints in the reasoning chain and actively discovers the weak links in the reasoning path through the reward-driven mechanism, ensuring the accuracy of breakpoint location.

[0028] For the identified breakpoints, a temporary relationship mining thread is initiated to extract evidence from an external corpus that is updated in real time. Temporary relationship connection edges are dynamically generated between the breakpoints to complete the missing relationships in real time, effectively repairing the connectivity of the reasoning chain to support complete logical deduction.

[0029] Preferably, step 3 further includes:

[0030] Feature extraction is performed on the node entities at both ends of the breakpoint to generate query vectors containing entity types and attributes. Semantic retrieval is then initiated against the real-time updated local news and announcement corpus to ensure that the retrieval content closely revolves around the breakpoint entity and improve the targeting of temporary relationship discovery.

[0031] Relations are extracted from the candidate text returned by the retrieval, the entity pairs and predicate logic contained therein are identified, and high-confidence temporary relations connecting entities at breakpoints are selected to ensure that the mined temporary relations are real and reliable, providing high-quality candidates for repairing breakpoints.

[0032] The selected high-confidence temporary relationships are dynamically written into the localized knowledge subgraph as edges with time-sensitive labels to complete the graph structure and form a complete dynamic knowledge subgraph. This enables the graph structure to self-repair in real time and ensures the physical connectivity of the reasoning chain.

[0033] Preferably, step 4 includes the following steps:

[0034] Starting with the core entities in the user input, a graph traversal algorithm is launched on the completed dynamic knowledge subgraph. A multi-hop path search combining breadth-first and depth-first approaches is performed along the relation edges to ensure broad coverage of candidate answers and in-depth mining of complex clues.

[0035] During the path search process, a current timestamp is attached to each traversed relation edge, and the effective period and data version number of each node along the path are recorded to achieve end-to-end timeliness marking and data traceability of the reasoning process;

[0036] All the searched paths are assembled and spliced ​​according to the number of jumps to generate a complete reasoning link that runs directly from the starting entity to the target answer entity, with each jump marked with a time limit, forming a transparent reasoning chain with complete temporal attributes and logical coherence.

[0037] Preferably, step 5 includes the following steps:

[0038] A temporal logic scoring card model is constructed, which takes each node entity in the inference link and its attached temporal attributes as input parameters and loads the feature evaluation dimensions of the scoring card to achieve standardized measurement and quantitative evaluation of the features of the inference nodes.

[0039] A deep learning classifier is used to perform binary classification on each node to identify whether its timestamp is within the time window of the user's question, and the timeliness verification result of each node is output to accurately determine the valid status of the node information at the time of the user's question.

[0040] By combining the verification results with the semantic relevance weights of the nodes, a comprehensive credibility score for each node is calculated. Nodes with scores higher than a preset threshold are selected to form a highly reliable inference path, ensuring that all nodes involved in the final generation have high timeliness and strong relevance.

[0041] Preferably, step 5 further includes:

[0042] Extract the timestamp attribute of each node in the inference chain and the implicit time context in the user's question. Input the two into the time series comparison encoder to calculate the relative time series relationship, so as to accurately quantify the degree of alignment between the timeliness of the node and the user's time requirement.

[0043] The calculated relative temporal relationship vector is input into the pre-trained failure logic detection model, and the probability value of whether the node has failed in the current question context is output, thus completing the reliable probability determination of whether the reasoning node has logically failed.

[0044] If the probability value exceeds the preset failure threshold, the node is determined to be logically invalid and is removed from the inference chain to ensure that invalid nodes do not enter the subsequent generation stage and pollute the output results; if it is lower than the preset failure threshold, it is determined to be valid and retained in the candidate path to ensure that the retained nodes meet the time context requirements of the user's question.

[0045] Preferably, step 6 includes the following steps:

[0046] The output highly reliable reasoning path is parsed into structured logical chain data, and key node entities on the chain are extracted as fact anchors for content generation, ensuring that the generated content is strictly based on the verified fact path, thus eliminating the risk of information illusion from the source.

[0047] By inputting factual anchors and logical relationships into the text generator of the large model, the generated content is forced to cover the anchor entities and their associated logic during the decoding process, thereby achieving forced alignment between the generated text and the factual skeleton and ensuring the logical integrity and accuracy of the output content.

[0048] In the final generated text, each key conclusion is marked with its corresponding reasoning anchor source through superscript markers or hyperlinks, forming a traceable and verifiable reliable output text. This empowers users to trace the source of the conclusions and fundamentally enhances their trust in the generated results.

[0049] Compared with the prior art, the beneficial effects of the present invention are:

[0050] 1. This method for enhancing the reliability of large model output based on heterogeneous dual models, through the collaborative optimization of deep learning semantic parsing and the fruit fly algorithm, introduces time-aware coding and symptom entity vector matrix into the semantic feature space. It can accurately identify implicit semantic drift caused by symptom or time differences in user input. The two-stage search mechanism of the fruit fly algorithm gradually converges the semantic parsing range to the global optimum, accurately locates the semantic center of the user's true intent, and provides an accurate semantic benchmark for knowledge retrieval and reasoning.

[0051] 2. This method for enhancing the reliability of large model output based on heterogeneous dual models can dynamically activate the corresponding localized knowledge subgraphs in the global knowledge graph based on the identified symptoms and time features. It can accurately anchor the model's reasoning scope within the relevant semantic region and cut off the generalization connection channel with the global knowledge base by modifying the attention mask matrix. This significantly reduces the scale of knowledge that the model needs to process, effectively avoids interference from irrelevant information, and ensures that the retrieved knowledge is highly matched with user needs in the time and space dimensions, thus significantly improving the accuracy of knowledge retrieval and reasoning efficiency.

[0052] 3. This method for enhancing the reliability of large model output based on heterogeneous dual models addresses the problem of inference breakpoints caused by the lack of temporary relationships in specific scenarios in localized subgraphs. It introduces reinforcement learning and graph neural networks to automatically identify inference breakpoints based on real-time user interaction feedback. Evidence is extracted from an external corpus that is updated in real time through asynchronous mining threads, and temporary relationship edges with time-sensitive labels are dynamically generated to ensure the timely update and completeness of the knowledge graph in the face of dynamically changing scenarios. This effectively repairs key breakpoints in the inference chain and ensures the continuity and integrity of knowledge inference.

[0053] 4. This method for enhancing the reliability of large model output based on heterogeneous dual models constructs a time-series comparison encoder and a failure logic detection model to perform timeliness verification and credibility scoring on each node in the inference chain. A deep learning classifier combined with a scoring card model comprehensively evaluates nodes from three dimensions: timeliness, semantic relevance, and structural importance. Based on dynamic thresholds, highly reliable nodes are selected to form a complete inference path, effectively filtering out failed or low-confidence logic nodes, ensuring that the final output inference chain has timeliness and credibility in real logic. Attached Figure Description

[0054] Figure 1 This is a schematic diagram illustrating the workflow of a method for enhancing the reliability of large model output based on heterogeneous dual models according to the present invention.

[0055] Figure 2 This is a schematic diagram of the method flow for a method to enhance the reliability of large model output based on heterogeneous dual models according to the present invention. Detailed Implementation

[0056] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.

[0057] Example 1, please refer to Figure 1 , Figure 2 This invention provides a technical solution: a method for enhancing the reliability of large model output based on heterogeneous dual models, comprising the following steps:

[0058] Step 1: Perform deep learning semantic parsing on user input, integrate symptom and temporal feature encoding, and combine fruit fly algorithm for collaborative optimization to accurately locate implicit semantic drift and identify implicit symptom and temporal features. This ensures that the model accurately understands the user's intent and eliminates semantic ambiguity caused by differences in symptom time. A Transformer-based encoder is used to perform deep semantic parsing on user input. By introducing time-aware positional encoding and symptom entity vector matrices, implicit temporal and symptom features in the text are mapped to a high-dimensional semantic space, achieving deep semantic alignment and high-dimensional expression of spatiotemporal information in the text. A symptom-time joint feature embedding layer is constructed, and the mapped features are deeply fused and encoded to generate a semantic feature tensor carrying spatiotemporal context information. This tensor serves as the basic input for subsequent optimization, forming a spatiotemporally coupled joint semantic representation for subsequent optimization. The olfactory search mechanism of the fruit fly algorithm is introduced into the feature space. Through swarm intelligence iterative optimization, the symptom weights and time thresholds in the semantic feature tensor are collaboratively calibrated to lock the precise location of implicit semantic drift. Adaptive collaborative calibration of spatiotemporal weights is achieved through swarm intelligence.

[0059] Furthermore, step 1 also includes: initializing the position of the fruit fly swarm in the semantic feature space, using the Euclidean distance of the semantic feature tensor as the taste concentration determination value, driving the swarm to perform a large-scale coarse olfactory search to identify potential drift regions, quickly locating potential drift regions through a large-scale random search, dynamically adjusting the flight direction and stride of the fruit fly swarm based on the taste concentration value fed back by the olfactory search, switching to a visual fine search mechanism to accurately define the boundaries of symptoms and time features within the candidate drift regions, dynamically adjusting the search direction based on feedback to achieve fine definition of feature boundaries, iteratively executing the visual search until the swarm converges, outputting the globally optimal individual position coordinates, using the semantic center point corresponding to these coordinates as the final localization result of the latent semantic drift, and obtaining the globally optimal semantic drift localization result through iterative convergence;

[0060] It should be noted that the semantic parsing module based on Transformer is constructed. Specifically, in the input layer of the encoder, time-aware encoding is superimposed on traditional positional encoding. This encoding is generated by combining a sine function with a learnable temporal bias, where the time granularity is set to the day level, and the encoding dimension is aligned with the word embedding dimension to 768 dimensions. Simultaneously, a symptom entity vector matrix is ​​constructed, with a dimension of N×768, where N is the number of predefined symptom nodes in the knowledge graph. These symptom nodes cover three levels of administrative divisions: country, province, and city. During the model's forward propagation, when a word in the input text is identified as a symptom entity or a time expression, the model automatically associates its corresponding word embedding with the symptom vector matrix. The corresponding row vectors in the semantic vector are weighted and fused, with the weight coefficients dynamically calculated through an attention mechanism. The fused semantic vector is then mapped to a high-dimensional semantic space. In the deep feature fusion stage, a symptom-time joint feature embedding layer is constructed. This layer receives the hidden state sequence output by the semantic parsing module. Two parallel gated linear units are used to process the symptom feature stream and the time feature stream, respectively. Specific parameters are set as follows: the hidden dimension of the gated linear unit is 512, and the activation function uses the Sigmoid gate mechanism. After independent transformation, the two feature streams interact through element-wise multiplication, and are then input into a two-layer fully connected fusion network. This network has an input dimension of 512, and the first layer outputs 256 dimensions, which are then processed by a layer... erNorm normalization restores the second-layer output dimension to 512, ultimately generating a semantic feature tensor carrying spatiotemporal context information. This tensor has a shape of [sequence length, 512], serving as the basic input data for subsequent fruit fly algorithm optimization and is also stored in a cache for use in the decoding stage. The fruit fly algorithm is introduced for collaborative optimization and drift localization. The initial fruit fly population size is 30, and the maximum number of iterations is set to 50. In the olfactory search stage, the center point of the semantic feature tensor is used as the initial reference. Individuals in the population randomly generate search positions within a hypersphere with a radius of 5 Euclidean distance units. The similarity between the semantic vector at the current position and the reference vector is used as the smell concentration determination function. Before reaching 20 iterations, the population performs olfactory analysis. In the coarse search, each individual flies randomly within the search space, calculates the flavor concentration value at its own position, and records the current best individual. The flavor concentration determination function uses the Euclidean distance between the semantic vector of the current position and the reference vector. The smaller the distance, the higher the concentration value. After 20 iterations, the search enters the visual fine search stage. At this time, the search radius shrinks to 1 Euclidean distance unit, and the step size decay factor is set to 0.95. Through iterative optimization, when the change of the group's best position is less than the preset threshold of 0.01 for 5 consecutive iterations, convergence is determined, and the coordinates of the global best individual are output. The semantic center point corresponding to this coordinate is the final localization result of the latent semantic drift. This result is stored in the form of a floating-point vector of length 512 for subsequent matching and activation of knowledge subgraphs.

[0061] Step 2: Based on the identified symptoms and time features, dynamically activate and match the corresponding localized knowledge subgraphs in the knowledge graph to narrow the semantic scope and complete the initial semantic scope anchoring. This allows the model to initially focus on knowledge regions closely related to user input, effectively converging the reasoning scope from global generalization to precise local focus. The identified symptoms and time features are then transformed into knowledge graph query instructions. By traversing the symptom nodes and time attribute edges in the graph, a set of candidate knowledge subgraphs that meet the feature conditions is selected, achieving a precise mapping from user intent to graph query and significantly narrowing the scope of data to be processed. The system calculates the semantic similarity between the semantic center of the user input and the set of candidate knowledge subgraphs. It then selects the knowledge subgraph with the highest similarity as the activation object through a graph matching algorithm, thus completing the dynamic activation of the localized knowledge subgraph. This ensures that the activated subgraph is highly matched with the user's current needs, avoids the introduction of interference from irrelevant knowledge, loads the activated localized knowledge subgraph into the working memory, cuts off the generalization connection channel between the large model and the global knowledge base, and precisely anchors the model's reasoning range within the semantic region defined by the knowledge subgraph. This physically isolates global noise and fundamentally eliminates semantic drift during the reasoning process.

[0062] It should be noted that the identified symptom feature vectors (512-dimensional floating-point vectors) and time feature markers (ISO8601 format timestamps, accurate to the day) are converted into structured query commands. Specifically, the query interface maps the symptom feature vectors to the symptom node index table of the knowledge graph. This index table contains a total of 3267 predefined symptom nodes, covering three levels of administrative divisions: country (67), province (342), and city (2858). At the same time, the time feature markers are used to limit the time attributes of relation edges in the graph, filtering out triples whose effective time covers the user's query time window ±30 days. During the traversal, a dynamic matching pattern is constructed using the Cypher query language. From a global knowledge graph containing approximately 120 million triples, candidate knowledge subgraphs are selected where symptom nodes match and time attribute edges meet the criteria. The number of nodes in each candidate subgraph is controlled between 500 and 2000. For the selected candidate knowledge subgraphs, the semantic similarity between them and the user-input semantic center (the optimal individual position coordinates, a 512-dimensional floating-point vector) is calculated. Specifically, each candidate subgraph is mapped to a graph-level semantic vector using a graph neural network encoder, which employs a 3-layer GraphSAGE network. The network has a hidden layer dimension of 256 and an output dimension of 512, aligned with the dimension of the user semantic center vector. Then, a cosine similarity algorithm is used to calculate the similarity value between the user semantic center and each candidate subgraph vector, with a similarity threshold set at 0.75. For candidate subgraphs with similarity higher than the threshold, a subgraph isomorphic matching algorithm is further used for fine-tuning, calculating the graph edit distance. The subgraph with the smallest edit distance (less than 5) and the highest similarity is selected as the final activation object. After the dynamic activation of the localized knowledge subgraph is completed, the subgraph is loaded into the working memory area of ​​the large model inference engine. This working memory is configured as a high-speed cache structure. The size is 16GB, and an LRU cache eviction strategy is adopted to ensure that frequently accessed localized subgraphs can be quickly hit. During loading, the attention range is limited to the range of nodes in the currently loaded subgraph by modifying the attention mask matrix during model inference, thereby cutting off the generalization connection channel between the large model and the global knowledge base. At this time, all knowledge retrieval and inference operations of the model are precisely anchored within the semantic region defined by the localized subgraph. The subgraph usually contains 300 to 800 entity nodes and corresponding relation edges. Entity attributes are stored in key-value pairs, with an average of 12 attributes per entity. After anchoring is completed, a subgraph loading confirmation signal is returned.

[0063] Step 3: Based on real-time interactive feedback, a reinforcement learning mechanism is introduced to drive the graph neural network to perform dynamic relationship reasoning on the localized subgraph. This automatically mines and completes missing temporary relationship paths in the graph, repairs reasoning breakpoints, ensures the integrity and timeliness of the knowledge chain, and prevents the reasoning process from being interrupted due to missing relationships. The initially anchored localized knowledge subgraph is input into the graph neural network for initial state encoding. Simultaneously, a real-time interactive feedback receiving channel is constructed, converting the user's click stream and follow-up questions into reward signals for reinforcement learning. This achieves spatial vectorization representation of the subgraph state and quantitative transformation of implicit user feedback, providing... The subsequent reasoning optimization lays the foundation by inputting reward signals into the reinforcement learning model, driving the graph neural network to explore paths between nodes in the knowledge subgraph with the goal of maximizing cumulative rewards. It identifies breakpoints in the reasoning chain and actively discovers weak links in the reasoning path through the reward-driven mechanism to ensure the accuracy of breakpoint location. For the identified breakpoints, a temporary relationship mining thread is launched to extract evidence from an external corpus that is updated in real time. Temporary relationship connection edges are dynamically generated between the breakpoints to achieve real-time completion of missing relationships and effectively repair the connectivity of the reasoning chain to support complete logical deduction.

[0064] It should be noted that a three-layer GraphSAGE network is used as the encoder, with each hidden layer having a dimension of 256 and the output dimension aligned to 512. During the encoding process, node features are fused with entity type embeddings and attribute value vectors, and relation edge features include relation type indices and timeliness markers. After encoding, each entity node is mapped to a 512-dimensional initial state vector. The structural and semantic information of the entire subgraph is transformed into a continuous state space representation that can be processed by the reinforcement learning model. At the same time, a real-time interactive feedback receiving channel is constructed to collect user clickstream data and follow-up text on the front-end interface. The system maps user behavior to three reward signals using a pre-trained classifier: a positive click reward of +1.0, a negative bounce reward of -0.5, and a follow-up clarification reward of +0.8. A deep Q-network is used as the reinforcement learning framework, with the node states encoded by the graph neural network as the input environment state. The agent's action space is defined as jumping from the current node along relation edges to adjacent nodes. During training, the discount factor is set to 0.95, the experience replay pool capacity is set to 10,000 entries, and the batch sampling size is 64. The agent starts with the core entity input by the user and... The subgraph performs path exploration with a maximum of 10 hops. The cumulative reward for each path is calculated by adding the user feedback rewards associated with the nodes along the way. When the exploration path can no longer obtain positive rewards at a certain node and there are no feasible jump edges, that node is marked as a candidate breakpoint in the inference chain. Its location information and contextual features are recorded in the breakpoint cache queue for subsequent processing. The temporary relationship mining thread adopts an independent asynchronous processing mechanism. When the length of the breakpoint cache queue exceeds 5, the mining task is triggered. For each breakpoint, the type labels and key attributes of the entities at both ends are extracted to form the query direction. The system initiates semantic retrieval to a real-time updated local news corpus and government announcement database, with a maximum of 20 retrieval results and a semantic similarity threshold of 0.6. Remote supervised relation extraction is performed on the retrieved text. A pre-trained BERT-BiLSTM-CRF model is used to identify entity pairs and predicate logic in the text. Extraction results with a confidence level higher than 0.8 are selected as candidate temporary relations. Finally, the selected temporary relations are dynamically written into the localized knowledge subgraph in the form of edges with time-sensitive labels, thereby completing the graph structure and repairing inference breakpoints.

[0065] Furthermore, step 3 also includes: extracting features from the node entities at both ends of the breakpoint, generating query vectors containing entity types and attributes, initiating semantic retrieval to the real-time updated local news and announcement corpus, ensuring that the retrieval content closely revolves around the breakpoint entity, improving the targeting of temporary relationship discovery, extracting relationships from the candidate text returned by the retrieval, identifying the entity pairs and predicate logic contained therein, filtering out high-confidence temporary relationships connecting the entities at the breakpoint, ensuring that the mined temporary relationships are real and reliable, providing high-quality candidates for repairing the breakpoint, and dynamically writing the filtered high-confidence temporary relationships into the localized knowledge subgraph in the form of edges with time-sensitive labels, completing the graph structure, forming a complete dynamic knowledge subgraph, realizing real-time self-repair of the graph structure, and ensuring the physical connectivity of the reasoning chain;

[0066] It should be noted that once the reinforcement learning path exploration identifies the inference breakpoint, the feature extraction program is immediately initiated. For the entity nodes at both ends of the breakpoint, the type labels and attribute key-value pairs stored in their knowledge graphs are read. The entity types cover preset categories such as people, organizations, geographical locations, and policies and regulations. The attribute information includes structured fields such as entity name, establishment time, and main responsible person. After concatenating all the information to form a natural language query text, it is input into the pre-trained Sentence-BERT encoder to generate a 384-dimensional query semantic vector. This vector is then compared with document vectors in the real-time updated local news corpus and announcement database. Cosine similarity was calculated with a similarity threshold of 0.6 and a recall limit of 20 entries. The local news corpus was updated daily with approximately 5,000 entries, and the announcement database contained documents and notices issued by various levels of government departments over the past three years, ensuring the timeliness and authority of the retrieved content. The retrieved candidate texts were fed into the relation extraction model pipeline, which adopted a BERT-BiLSTM-CRF architecture. The BERT base was a pre-trained Chinese version with a hidden layer dimension of 768, and the BiLSTM hidden layer dimension was set to 256. The model first identified all entity references in the text and aligned them with the breakpoint entities, and then passed them through a classifier. Predicting predicate logic between entity pairs, each extraction result is accompanied by a confidence score, which is obtained by normalizing the path probability output by the CRF layer using softmax. Only extraction results with a confidence score higher than 0.8 are retained as candidate temporary relations. Simultaneously, the contextual text fragments and time expressions of the relation occurrence are recorded for subsequent time-sensitive tag generation. Validated candidate temporary relations are encapsulated into graph writing units, each containing a source entity ID, target entity ID, relation type index, and time-sensitive tag. The time-sensitive tag is generated based on the time expressions identified in the text, extracted using the HeidelTime time series recognition tool and standardized to ISO860. 1. If the text does not explicitly contain time information, the current system time is used as the default effective time. The write operation is executed atomically through the graph database transaction interface. The newly added relation edges are accompanied by valid_from and valid_until attributes (valid_from is the effective start time and valid_until is the effective end time). The initial valid_until is set to 2099-12-31 to indicate long-term validity. After the write is completed, the version number and last modification timestamp of the subgraph are updated, and the newly added edge information is synchronized to the state representation of the reinforcement learning model to form a dynamic knowledge subgraph that can support the complete inference chain.

[0067] Step 4: Based on the completed dynamic knowledge subgraph, long-path association reasoning is carried out to construct a complete reasoning link with time-sensitive markers, forming a logically coherent and time-sequential complete reasoning path, providing a structured foundation for subsequent verification. Starting from the core entity in the user input, a graph traversal algorithm is launched on the completed dynamic knowledge subgraph to perform a multi-hop path search combining breadth-first and depth-first along the relation edges, ensuring broad coverage of candidate answers and in-depth mining of complex clues. During the path search process, a current timestamp is attached to each traversed relation edge, and the effective period and data version number are recorded for each node along the path, realizing the time-sensitive markers and data traceability of the entire reasoning process. All the searched paths are assembled and spliced ​​according to the number of hops to generate a complete reasoning link from the starting entity to the target answer entity, with each hop having a time-sensitive marker, forming a transparent reasoning chain with complete time-sequential attributes and logical coherence.

[0068] It should be noted that in the completed dynamic knowledge subgraph, starting from the parsed core entity, a hybrid graph traversal algorithm is initiated, combining breadth-first and depth-first strategies. First, breadth-first search is used to explore the adjacency relationships of the current node, identifying all entity nodes reachable in one hop. Then, for each branch, depth-first search is used to extend along specific relationship paths, ensuring coverage of a broad range of candidate answers while also deeply exploring complex long-path reasoning clues. During traversal, the maximum search depth is limited to 10 hops, and the maximum number of nodes explored per layer is set to 500 to prevent state space explosion. When the traversal path reaches an entity node that may be the target answer, the algorithm automatically terminates the depth search of the current branch and records the complete path, ultimately generating several candidate reasoning links. Simultaneously with path searching, an ISO8601 format timestamp of the current system time is attached to each traversed relationship edge in real time, marking the specific time the relationship was activated. For each entity node along the path, its effective period attribute stored in the knowledge graph is read and recorded. The `valid_from` and `valid_until` fields are used to read the data version number of the node. This version number is represented by an incrementing integer, which increments by 1 each time the node attribute is updated. Timeliness information is appended to the path node in the form of key-value pairs to form an intermediate representation with time-series attributes. The granularity of the timeliness information is recorded at the day level to ensure precise alignment with the time context of the user's question. After traversing all branches, the searched path segments are assembled and spliced ​​according to the number of hops. The splicing process takes the starting entity as the root node and the target answer entity as the leaf node. The intermediate entities and relation edges are connected in the order of access to form a complete inference link from the starting point to the answer. Each link consists of several triple sequences. Each triple contains the source node, relation edge, and target node. Each node is accompanied by its effective period and version number. Each relation edge is accompanied by a timestamp mark during traversal. Finally, an inference link is generated and stored serialized in JSON-LD format. The link length is 3 to 8 hops and contains complete time-series metadata.

[0069] Step 5: Use deep learning to construct a time-series logic scoring card, and perform timeliness verification and credibility scoring on each inference node. Filter out invalid logic, retain highly reliable inference paths, and eliminate invalid or untrustworthy logic nodes to ensure that every piece of information used in the final basis is true and valid.

[0070] Step 6: Based on the multi-dimensional logical verification results, optimize and adjust the generated content to form reliable output text with traceable anchor points, improve users' trust and acceptance of the model output, realize the verifiability and traceability of the generated content, and significantly enhance users' trust in the output results.

[0071] Example 2, as Figure 1 , Figure 2As shown, based on Embodiment 1, the present invention provides a technical solution: Step 5 includes the following steps: Constructing a temporal logic scoring card model, taking each node entity in the inference link and its attached time attribute as input parameters, loading the feature evaluation dimension of the scoring card, realizing standardized measurement and quantitative evaluation of the features of the inference nodes, using a deep learning classifier to perform binary classification judgment on each node, identifying whether its timestamp is within the time window of the user's question, outputting the timeliness verification result of each node, accurately determining the effective state of the node information under the time of the user's question, combining the verification result with the semantic relevance weight of the node, calculating the comprehensive credibility score of each node, and selecting nodes with scores higher than the preset threshold to form a highly reliable inference path, ensuring that the nodes that finally participate in the generation all have high timeliness and strong relevance;

[0072] It should be noted that the temporal logic scorecard model uses logistic regression as its basic framework. Feature dimensions include node timeliness, semantic relevance, and structural importance. Timeliness features are extracted from the `valid_from` and `valid_until` fields attached to the node. Semantic relevance features are obtained by calculating the cosine similarity between the node vector and the user query vector. Structural importance features are measured using the betweenness centrality of nodes in the inference chain. Each feature dimension is discretized by binning, and each bin is assigned a WOE code and corresponding score based on the training data. The total score of the scorecard is obtained by linearly weighted summing of the scores of each feature. The weight coefficients are obtained from the annotations through maximum likelihood estimation. The data is learned from the training samples. The deep learning classifier is built on the Transformer architecture. The input is a concatenated sequence of node text description and time attribute. The output is the effective probability of the node in the context of the user's question time. The timeliness window of the user's question is centered on the question time and extended forward and backward by 30 days as the default window range. The effective probability output by the classifier is mapped to a timeliness score of 0 to 100. It is then weighted and fused with the basic credibility score output by the scorecard. The fusion weight is set to timeliness score of 0.6, semantic relevance score of 0.3, and structural importance score of 0.1. Finally, each node obtains a comprehensive credibility score, with a score range of 0 to 100.

[0073] The formula for calculating the overall credibility score of a node is as follows:

[0074] ;

[0075] In the formula: This represents the overall credibility score of the node. The weight for timeliness score is 0.6; The timeliness score is output by a deep learning classifier. This is the weight for the semantic relevance score, with a value of 0.3; Score the semantic relevance; This is the weight for the structural importance score, with a value of 0.1; Score the importance of the structure;

[0076] The formula for calculating the timeliness score is as follows:

[0077] ;

[0078] ;

[0079] In the formula: The effective probability output by the deep learning classifier; This is a function for a classifier model built on the Transformer architecture; The input feature sequence is composed of node text descriptions and time attributes concatenated together.

[0080] The expression for calculating the semantic relevance score is as follows:

[0081] ;

[0082] ;

[0083] In the formula: The semantic vector representation of the node has 512 dimensions; The semantic vector representation of the user query has 512 dimensions; Let L2 be the norm (magnitude) of the vector.

[0084] The formula for calculating the structural importance score is as follows:

[0085] ;

[0086] ;

[0087] ;

[0088] In the formula: For nodes Primitive values ​​of betweenness centrality in the inference chain; It is the set of all nodes in the inference link; For the node To the node The number of all shortest paths; For the node To the node The shortest path passes through the nodes Quantity; The minimum-maximum normalization function maps betweenness centrality to the interval [0, 1]. It is the minimum betweenness centrality value of all nodes in the link; The maximum betweenness centrality value of all nodes in the link;

[0089] The inference path is filtered based on the overall confidence score of the nodes. Only nodes with a score higher than a preset threshold are retained to form a highly reliable inference path. The threshold setting adopts a dynamic adjustment strategy, with a default initial value of 75 points. It can be configured to a range of 60 to 85 points depending on the specific application scenario. The filtering process first traverses the entire path and marks all nodes below the threshold as nodes to be removed. Then, the connectivity of the removed path is checked. If the removal causes the path to break, a local path reconnection mechanism is activated. The dynamic relationship of the completed path is used to try to reconnect the broken part. After successful reconnection, the final highly reliable inference path is output. This path contains 3 to 8 nodes, all of which have an overall confidence score of not less than 75 points and the path connectivity is complete.

[0090] Furthermore, step 5 also includes: extracting the timestamp attribute of each node in the inference chain and the implicit time context in the user's question, inputting the two into a time-series comparison encoder to calculate the relative time-series relationship, achieving accurate quantification of the alignment between the node's timeliness and the user's time requirement, inputting the calculated relative time-series relationship vector into a pre-trained failure logic detection model, outputting the probability value of whether the node has failed in the current question context, completing a reliable probability determination of whether the inference node has logically failed. If the probability value exceeds the preset failure threshold, the node is determined to be logically failed and is filtered out from the inference chain to ensure that failed nodes do not enter the subsequent generation stage and pollute the output results; if it is lower than the preset failure threshold, it is determined to be valid and retained in the candidate path to ensure that the retained nodes meet the time context requirements of the user's question in terms of timeliness.

[0091] It should be noted that a time-series comparison encoder is constructed to calculate the relative relationship between node time and user query time. This encoder receives two types of input data: first, timestamp attributes extracted from inference link nodes, specifically including the node's effective start time (valid_from) and effective end time (valid_until), both stored in ISO8601 format with precision down to the day; second, implicit time context extracted from user queries through a semantic parsing module, forming a time window centered on the query time and extending 30 days before and after. The encoder internally employs a multi-layer Transformer structure to correlate node time intervals with user query times. The window is mapped to a relative temporal relationship vector with a dimension of 128. During the encoding process, a self-attention mechanism is used to capture the time alignment between node timeliness and user needs. The output is used to characterize the time adaptation state of the node in the context of the user's question. After the relative temporal relationship vector is calculated, it is input into a pre-trained failure logic detection model for failure probability determination. This failure logic detection model adopts a multilayer perceptron architecture with an input layer dimension of 128, aligned with the output dimension of the temporal comparison encoder. The hidden layer has two dimensions: the first layer has a dimension of 256, and the second layer has a dimension of 128. The activation function is ReLU, and the output layer uses S... The igmoid activation function outputs the probability value of whether a node is invalid in the current query context. During model training, labeled historical query data is used, with positive samples representing valid nodes and negative samples representing invalid nodes. The binary cross-entropy loss function is used, and the Adam optimizer is selected with a learning rate of 0.001. The training iterations are 50, and the batch size is 32. An early stopping mechanism prevents overfitting, ensuring the model has stable failure detection capabilities in practical applications. After the failure logic detection model outputs the probability value, node validity is determined based on a preset failure threshold, set to 0.65. This threshold is determined by maximizing the F1 score on the validation set. To ensure a balance between recall and precision, the failure probability of each node is compared with 0.65 during execution. If the probability is greater than 0.65, the node is considered logically invalid in the current query context and is immediately removed from the inference chain. If the probability is less than or equal to 0.65, the node is considered valid and retained in the candidate path. After the removal operation is completed, a link connectivity check is automatically triggered. If the removal of a node causes a path break, the dynamic relation edge library is called for local reconnection. The reconnection process only uses temporary relations with a confidence level higher than 0.8 to ensure that the final retained candidate path meets the requirements in terms of both logical integrity and timeliness.

[0092] Step 6 includes the following steps: parsing the output highly reliable inference path into structured logical chain data, extracting key node entities on the chain as fact anchors for content generation, ensuring that the generated content is strictly based on the verified fact path, eliminating the risk of information illusion from the source, inputting the fact anchors and logical relationships into the text generator of the large model, and forcibly constraining the generated content to cover the anchor entities and their associated logic during the decoding process, achieving forced alignment between the generated text and the fact skeleton, ensuring the logical integrity and accuracy of the output content, and in the final generated text, marking the source of each key conclusion with its corresponding inference anchor through superscript marks or hyperlinks, forming a traceable and verifiable reliable output text, giving users the ability to trace the source of the conclusions, and fundamentally improving the trust in the generated results;

[0093] It should be noted that the highly reliable inference path is parsed from JSON-LD format into structured logical chain data. The parser traverses each node and relation edge in the path, extracting key node entities as fact anchors for content generation. Each anchor includes entity ID, name, type, and comprehensive credibility score. The comprehensive credibility score of a node is calculated by weighted fusion of timeliness score, semantic relevance score, and structural importance score, where timeliness score has a weight of 0.6, semantic relevance score has a weight of 0.3, and structural importance score has a weight of 0.1. The fused score must be no less than 75 points to be selected as a fact anchor. Simultaneously, the logical relationships between nodes are extracted, and each relationship includes relation type, confidence score, and... The process iterates through timestamps, and only those with a relationship confidence score higher than 0.8 are included in the logical chain. After extraction, the anchor sequence and relationship sequence are assembled into a linear logical chain according to the path access order. Each chain contains 3 to 8 anchor nodes and corresponding relationships, forming a complete skeleton of factual evidence. The parsed fact anchor sequence and logical relationships are input into the text generator of the large model, and a Transformer-based decoding architecture is used for text generation. During the decoding process, the attention mask mechanism is modified to force the generated content to cover all fact anchors and their associated logic. Specifically, the anchor coverage loss is calculated at each decoding step. If the generated content deviates from the anchor entity or omits key relationships, the word probabilities are adjusted accordingly. The generator's temperature parameter is set to 0.7, the kernel sampling value is set to 0.9, and the maximum generation length is limited to 512 tokens. The position of anchor entities in the generated text is globally optimized using a dynamic programming algorithm to ensure a balanced and logically coherent distribution of key information. During generation, the semantic matching degree between each sentence and the anchor set is checked in real time, with a matching degree threshold set to 0.75. Candidates with a matching degree below the threshold are resampled until the constraints are met. After the generated text is completed, an anchor source annotation program is implemented, adding the corresponding reasoning anchor source information to each key conclusion. The annotation method uses a combination of superscript numbers and hyperlinks, with superscript numbers placed at the bottom or side of the text. The anchor points in each column correspond one-to-one, with hyperlinks pointing to the detailed information pages of the corresponding entities in the knowledge graph. The anchor point descriptions include the entity name, effective period, data version number, and comprehensive credibility score. The effective period is identified by the fields valid_from and valid_until, in the ISO8601 standard format accurate to the day. The data version number is represented by an incrementing integer. The final output text is accompanied by a complete inference chain metadata file, including the timeliness markers of each node in the path, traversal timestamps, and relationship confidence. Users can view the complete inference process of any conclusion through the interactive interface, enabling the verification and traceability of the model's output content, significantly improving users' trust and acceptance of the generated results.

[0094] The following section details the workflow of a method for enhancing the reliability of large model outputs based on heterogeneous dual models.

[0095] First, after receiving user input, the deep learning semantic parsing module fuses symptom and time feature encodings, and combines the fruit fly algorithm to perform swarm intelligence iterative optimization of the semantic feature space. This accurately locates the implicit semantic drift caused by symptom or time differences in the user input, identifies the implicit symptom and time features, and stores the semantic center point in vector form. Second, based on the identified symptom and time features, the system converts them into knowledge graph query instructions. It selects a set of candidate knowledge subgraphs from the global knowledge graph that meet the conditions of symptom node hit and time attribute edge matching. Through graph neural network encoding and semantic similarity calculation, the system selects the localized knowledge subgraph with the highest matching degree with the user's semantic center as the activation object, loads it into the working memory of the large model inference engine, and precisely anchors the model inference range within the semantic region defined by the subgraph by modifying the attention mask matrix.

[0096] Then, the anchored localized knowledge subgraph is input into the graph neural network for state encoding, a real-time interactive feedback receiving channel is constructed, user click streams and follow-up texts are transformed into reward signals for reinforcement learning, driving the deep Q network to explore paths between subgraph nodes, identify breakpoints in the inference chain and record them in the cache queue, start an asynchronous temporary relation mining thread, initiate semantic retrieval to the real-time updated local news corpus and government announcement database, extract high-confidence temporary relations from candidate texts through a relation extraction model, dynamically write them into the subgraph in the form of edges with time-sensitive labels, complete the graph structure to form a dynamic knowledge subgraph, and then start a hybrid graph traversal algorithm on the dynamic knowledge subgraph starting from the core entity, perform a multi-hop path search combining breadth-first and depth-first along the relation edges, attach the current timestamp to each traversed relation edge, record the effective period and data version number for each node along the path, assemble and splice the searched paths according to the number of hops, and generate a complete inference link with time-sensitive labels;

[0097] Next, a temporal logic scoring card model and a deep learning classifier are constructed to verify the timeliness and credibility of each node in the inference chain. The comprehensive credibility score of the nodes is calculated, and nodes with scores higher than a preset threshold are selected to form a highly reliable inference path. For the link breakage caused by the filtered nodes, the dynamic relation edge library is called to perform local reconnection to ensure the integrity and connectivity of the path. Finally, the highly reliable inference path is parsed into structured logical chain data, and key node entities are extracted as fact anchors. The anchor sequence and logical relationship are input into the large model text generator. During the decoding process, the attention mask mechanism is modified to force the generated content to cover all anchors and their associated logic. After generation, each key conclusion is labeled with anchor source information in the form of superscript numbers and hyperlinks, and a complete inference chain metadata file is attached to form a traceable and verifiable reliable output text.

[0098] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. The scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for enhancing the reliability of large model output based on heterogeneous dual models, characterized in that, Includes the following steps: Step 1: Perform deep learning semantic parsing on user input, integrate symptom and time feature encoding, combine fruit fly algorithm for collaborative optimization, accurately locate implicit semantic drift, and identify implicit symptom and time features; Step 2: Based on the identified symptoms and time features, dynamically activate and match the corresponding localized knowledge subgraphs in the knowledge graph to narrow the semantic scope and complete the initial anchoring of the semantic scope; Step 3: Based on real-time interactive feedback, a reinforcement learning mechanism is introduced to drive the graph neural network to perform dynamic relationship reasoning on the localized subgraph, automatically mining and completing the missing temporary relationship paths in the graph, and repairing reasoning breakpoints. Step 4: Based on the completed dynamic knowledge subgraph, conduct long-path association reasoning to construct a complete reasoning link with time-sensitive tags; Step 5: Use deep learning to construct a temporal logic scoring card, and perform timeliness verification and credibility scoring on each inference node to filter out invalid logic and retain highly reliable inference paths. Step 6: Based on the multi-dimensional logical verification results, optimize and adjust the generated content to form reliable output text with traceable anchor points.

2. The method for enhancing the reliability of large model output based on heterogeneous dual models according to claim 1, characterized in that: Step 1 includes the following steps: We use an encoder based on the Transformer architecture to perform deep semantic parsing of user input. By introducing time-aware location encoding and symptom entity vector matrix, we map the implicit time and symptom features in the text to a high-dimensional semantic space. A symptom-time joint feature embedding layer is constructed, and the mapped features are deeply fused and encoded to generate a semantic feature tensor carrying spatiotemporal context information; By introducing the olfactory search mechanism of the fruit fly algorithm into the feature space, and through swarm intelligence iterative optimization, the symptom weights and time thresholds in the semantic feature tensor are calibrated in a coordinated manner to pinpoint the precise location of implicit semantic drift.

3. The method for enhancing the reliability of large model output based on heterogeneous dual models according to claim 2, characterized in that: Step 1 further includes: The position of the fruit fly swarm in the semantic feature space is initialized, and the Euclidean distance of the semantic feature tensor is used as the taste concentration determination value to drive the swarm to perform a large-scale coarse olfactory search to identify potential drift regions. Based on the odor concentration value fed back by the olfactory search, the flight direction and stride of the fruit fly swarm are dynamically adjusted, and a visual fine search mechanism is used to accurately define the boundaries of symptoms and time features in the candidate drift area. Iteratively perform visual search until the population converges, output the globally optimal individual position coordinates, and use the semantic center point corresponding to the coordinates as the final localization result of implicit semantic drift.

4. The method for enhancing the reliability of large model output based on heterogeneous dual models according to claim 1, characterized in that: Step 2 includes the following steps: The identified symptoms and time features are transformed into knowledge graph query commands. The symptoms nodes and time attribute edges in the graph are traversed to filter out a set of candidate knowledge subgraphs that meet the feature conditions. Calculate the semantic similarity between the semantic center of the user input and the set of candidate knowledge subgraphs, and select the knowledge subgraph with the highest similarity as the activation object through a graph matching algorithm to complete the dynamic activation of the localized knowledge subgraph; The activated localized knowledge subgraph is loaded into the working memory, severing the generalized connection between the large model and the global knowledge base, so that the model's reasoning scope is precisely anchored within the semantic region defined by the knowledge subgraph.

5. The method for enhancing the reliability of large model output based on heterogeneous dual models according to claim 1, characterized in that: Step 3 includes the following steps: The initially anchored localized knowledge subgraph is input into the graph neural network for initial state encoding. At the same time, a real-time interactive feedback receiving channel is constructed to transform the user's click stream and follow-up questions into reward signals for reinforcement learning. The reward signal is input into the reinforcement learning model, which drives the graph neural network to explore paths between nodes in the knowledge subgraph with the goal of maximizing the cumulative reward, and identifies the breakpoints in the reasoning chain. For the identified breakpoints, a temporary relationship mining thread is started to extract evidence from an external corpus that is updated in real time, and temporary relationship connection edges are dynamically generated between the breakpoints.

6. The method for enhancing the reliability of large model output based on heterogeneous dual models according to claim 5, characterized in that: Step 3 also includes: Feature extraction is performed on the node entities at both ends of the breakpoint to generate a query vector containing entity type and attributes, and semantic retrieval is initiated to the real-time updated local news and announcement corpus; Relation extraction is performed on the candidate text returned by the retrieval, the entity pairs and predicate logic contained therein are identified, and high-confidence temporary relations connecting the entities at the breakpoint are selected. The selected high-confidence temporary relationships are dynamically written into the localized knowledge subgraph as edges with time-sensitive labels, thus completing the graph structure and forming a complete dynamic knowledge subgraph.

7. The method for enhancing the reliability of large model output based on heterogeneous dual models according to claim 1, characterized in that: Step 4 includes the following steps: Starting with the core entity in the user input, a graph traversal algorithm is launched on the completed dynamic knowledge subgraph, and a multi-hop path search combining breadth-first and depth-first is performed along the relation edges; During the path search process, a current timestamp is attached to each traversed relation edge, and the effective period and data version number of each node along the path are recorded. All the searched paths are assembled and spliced ​​according to the number of hops to generate a complete reasoning link that runs from the starting entity to the target answer entity, with each hop marked with a time limit.

8. The method for enhancing the reliability of large model output based on heterogeneous dual models according to claim 1, characterized in that: Step 5 includes the following steps: Construct a temporal logic scoring card model, taking each node entity in the inference link and its attached temporal attributes as input parameters, and loading the feature evaluation dimensions of the scoring card. A deep learning classifier is used to perform binary classification on each node to identify whether its timestamp is within the time window of the user's question, and the timeliness verification result of each node is output. By combining the verification results with the semantic relevance weights of the nodes, a comprehensive credibility score for each node is calculated, and nodes with scores higher than a preset threshold are selected to form a highly reliable inference path.

9. The method for enhancing the reliability of large model output based on heterogeneous dual models according to claim 8, characterized in that: Step 5 further includes: Extract the timestamp attribute of each node in the inference chain and the implicit time context in the user's question, and input the two into the time series comparison encoder to calculate the relative time series relationship; The calculated relative temporal relationship vector is input into the pre-trained failure logic detection model, and the probability value of whether the node has failed in the current question context is output. If the probability value exceeds the preset failure threshold, the node is determined to be logically faulty and is removed from the inference chain; if it is below the preset failure threshold, it is determined to be valid and retained in the candidate path.

10. The method for enhancing the reliability of large model output based on heterogeneous dual models according to claim 1, characterized in that: Step 6 includes the following steps: The output highly reliable inference path is parsed into structured logical chain data, and key node entities on the chain are extracted as fact anchors for content generation. Input factual anchors and logical relationships into the text generator of the large model, and force the generated content to cover the anchor entities and their associated logic during the decoding process; In the final generated text, each key conclusion is marked with its corresponding reasoning anchor source through superscript markers or hyperlinks, forming a reliable output text that is traceable and verifiable.

Citation Information

Patent Citations

  • Method and system for enhancing reliability of large language model based on knowledge graph

    CN119443230A