Knowledge exploration method and system based on generative thinking chain and feedback mechanism
Through the generative thinking chain and feedback mechanism, the inference path is dynamically constructed and optimized, and the efficiency and accuracy of the existing intelligent query system under complex problems is solved, and efficient and transparent knowledge services are realized, especially suitable for intelligent query and reasoning of long texts such as laws, regulations and academic papers.
Patent Information
- Application Number
- CN202510331228.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-20
AI Technical Summary
When facing complex problems, existing intelligent query systems have problems such as low retrieval efficiency, single results, large computing resources consumption, slow response speed, unstable inference paths, lagging feedback mechanisms, inaccurate results and lack of transparency, especially in high-risk areas such as law and medical care, which are difficult to meet user needs.
Using generative thinking chain and feedback mechanism, through multimodal query analysis, dynamic weight allocation, meta-reinforcement learning, causal reasoning and game theory feedback aggregation technology, hierarchical reasoning paths are dynamically constructed, inferential strategies are optimized, model parameters are adjusted in real time, and structured and visual reasoning processes are provided.
It realizes accurate, real-time and credible knowledge services in complex scenarios, improves the accuracy, efficiency and transparency of reasoning, and is suitable for intelligent query and reasoning of long texts such as laws, regulations, academic papers, etc.
Smart Images

Figure CN120258137A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence (AI), specifically to the cross - technical field of intelligent query and knowledge reasoning technologies, and more specifically to a knowledge exploration method and system based on generative thought chains and feedback mechanisms. Background Art
[0002] Existing intelligent query systems mostly adopt static rules or single - modality inputs, relying on keyword matching and preset inference templates, and have the following limitations:
[0003] 1. Traditional systems cannot dynamically adjust the inference path according to the context. When facing complex problems (such as legal article conflicts, multi - modality input fusion), the retrieval efficiency is low and the results are single.
[0004] Existing AI models often require large amounts of computing resources and time when dealing with long - text, high - complexity query tasks. Due to the large amount of data, the traditional inference process is slow. Especially when facing high - concurrency query requests, the response speed of the system significantly decreases, resulting in a poor user experience.
[0005] In complex scenarios, such as legal regulations or academic papers, existing models are prone to problems such as inaccurate key information extraction, logical reasoning errors, or context understanding deviations. In such cases, the returned results may not be accurate or relevant enough to meet users' needs for high - quality knowledge services. Existing technologies are difficult to capture deep semantic associations in long texts or multi - turn interactions. Especially in professional fields such as law and medicine, they cannot accurately match users' intentions.
[0006] Most existing technologies adopt static rules or fixed inference path generation strategies, resulting in easily deviating from the expected path during the inference process. This instability affects the accuracy and consistency of knowledge derivation. Especially when dealing with changing query requests, the results are difficult to ensure stability.
[0007] 2. Traditional feedback mechanisms only rely on offline user behavior data and cannot optimize the inference path through real - time interaction, resulting in limited system response speed and personalization level.
[0008] Existing systems lack a mechanism to adjust inference parameters or paths in real - time according to feedback after inference is completed. As the application environment changes, the model cannot be updated in time to maintain high performance. Over time, the system performance gradually declines, making it difficult to continuously provide efficient inference services.
[0009] Existing models often exhibit "black box" characteristics, and the output results lack a detailed display of the reasoning process. Especially in high-risk fields such as law and medicine, users expect to understand the specific steps and basis of the reasoning, but this need cannot be effectively met in the existing technology, reducing users' trust in the system results.
[0010] Therefore, how to better overcome the limitations of the above intelligent query system has become an urgent problem for those skilled in the art to solve. Summary of the Invention
[0011] In view of this, the present invention provides a knowledge exploration method and system based on a generative thought chain and a feedback mechanism, which can effectively solve the problems of intelligent query and reasoning in complex scenarios, and realize the precision, real-time and credibility of knowledge services.
[0012] To achieve the above object, the present invention adopts the following technical solutions:
[0013] In a first aspect, an embodiment of the present invention provides a knowledge exploration method based on a generative thought chain and a feedback mechanism, including the following steps:
[0014] S10. Receive a multimodal query input by a user, perform intention analysis to identify the core requirement, writing theme or problem focus;
[0015] S20. Preprocess the multimodal query and extract key information, and decompose complex problems into multiple sub-problems or sub-tasks;
[0016] S30. Based on the multiple sub-problems or sub-tasks, dynamically construct a hierarchical reasoning path based on a generative thought chain, and label the logical basis for each step of the reasoning path;
[0017] S40. Optimize the priority of the reasoning path through dynamic weight allocation, and combine the meta-reinforcement learning framework to dynamically adjust the reasoning strategy online to optimize the reasoning path;
[0018] S50. On the basis of optimizing the reasoning path, use a causal reasoning analysis tool to verify the reasoning path and generate a counterfactual path to correct the wrong nodes;
[0019] S60. Output structured text and a visual reasoning process, and optimize the model parameters in real time based on the multi-person feedback aggregation of game theory.
[0020] In one embodiment, in step S40, optimizing the priority of the reasoning path through dynamic weight allocation includes:
[0021] Dynamically calculate the weights of each node in the reasoning path according to the task context, semantic similarity and node importance, and the weights are calculated by the following formula:
[0022]
[0023] In the formula, w i represents the weight of the inference node n i ; q represents the current query task; n i represents the i-th inference node; j represents the number of all inference nodes in the current path; sim(q, n i ) represents the semantic similarity between the current task q and the inference node n i ; τ represents the temperature coefficient, which is used to control the path exploration intensity;
[0024] According to multiple nodes decomposed by the inference path, each node corresponds to a subtask, and the optimal selection of the path is achieved by maximizing the total reward value through the reward model; among them, the formula for maximizing the total reward value is as follows:
[0025]
[0026] Among them, T represents the query task, and R i (T) represents the reward value of the node, k represents the number of nodes in the current path; p represents the dynamically adjusted normalized weight coefficient, p ∈ [0, 1].
[0027] In one embodiment, in the step S40, optimizing the inference path priority through dynamic weight allocation further includes:
[0028] Fusing multi-modal input information, and realizing semantic alignment of text and image through the alignment loss function, and its loss function is defined as:
[0029]
[0030] Among them, Ⅱ ab ∈ {0, 1} represents the modal association indicator function, when the text T a is related to the image I b it is 1, otherwise it is 0; f T (·) and f I (·) respectively represent the text encoder and the image encoder, and output the normalized embedding vectors; I k represents the k-th sample in the image modality, k ≠ b; cos(·, ·) represents the cosine similarity, τ represents the temperature coefficient; N represents the number of text samples, and M represents the number of image samples.
[0031] In one embodiment, in the step S40, the meta-reinforcement learning framework dynamically adjusts the inference measurement online, including:
[0032] Offline meta-pre-training, using the MAML framework to pre-train the model on the cross-domain task set;
[0033] Online fine-tuning generates adversarial samples through real-time user feedback and dynamically updates model parameters;
[0034] A composite reward mechanism integrating interpretability metrics designs a multi-objective reward function to balance inference accuracy, efficiency, and interpretability.
[0035] In one embodiment, in step S50, using a causal reasoning analysis tool to verify the inference path includes:
[0036] Generating counterfactual queries when the inference result deviates, comparing the actual path with the counterfactual path to locate error nodes and trigger corrections;
[0037] Defining the causal strength between nodes by embedding causal probability edges in knowledge, and calculating logical consistency scores, task relevance scores, case support scores, and interpretability scores to enhance inference credibility and interpretability.
[0038] In one embodiment, in step S50, using a causal reasoning analysis tool to verify the inference path further includes:
[0039] In each step of generating the inference path, a reward model is used to optimize the inference chain, and the reward mechanism is as follows:
[0040] R = w1·Accuracy + w2·Efficiency - w3·Resources
[0041] w1 represents the weight coefficient of accuracy; Accuracy represents the accuracy score of the inference chain; w2 represents the weight coefficient of efficiency; Efficiency represents the efficiency score of the inference chain; w3 represents the weight coefficient of resource consumption; Resources represents the resource consumption score of the inference chain.
[0042] In one embodiment, in step S60, real-time optimization of model parameters based on game theory-based multi-person feedback aggregation includes:
[0043] Based on the Shapley value algorithm of game theory, quantifying the contribution of user feedback, filtering low-quality or malicious feedback, and coordinating multi-user opinions to optimize the inference path;
[0044] Using a meta-reinforcement learning framework to process the aggregated feedback information in real time and dynamically adjust the inference path and model parameters.
[0045] In a second aspect, an embodiment of the present invention further provides a knowledge exploration system based on a generative thought chain and a feedback mechanism, including:
[0046] A data input module for receiving multi-modal queries input by a user, performing intent analysis to identify core requirements, writing topics, or problem focuses;
[0047] A semantic processing module, which is used to preprocess the multimodal query and extract key information, and decompose complex problems into multiple sub-problems or sub-tasks;
[0048] An inference engine module, which is used to dynamically construct a hierarchical inference path based on the generative thought chain according to the multiple sub-problems or sub-tasks, and annotate the logical basis for each step of the inference path;
[0049] A dynamic inference path module, which is used to optimize the priority of the inference path through dynamic weight allocation, and combine the meta-reinforcement learning framework to dynamically adjust the inference strategy online and optimize the inference path;
[0050] A causal inference analysis module, which is used to verify the inference path by using a causal inference analysis tool on the basis of optimizing the inference path, and generate a counterfactual path to correct the wrong nodes;
[0051] A result output and feedback optimization module, which is used to output structured text and a visualized inference process, and optimize the model parameters in real time based on the multi-person feedback aggregation of game theory.
[0052] It can be seen from the above technical solutions that compared with the prior art, the present invention has the following technical advantages:
[0053] Accuracy: Through dynamic weight allocation and causal inference, it can accurately capture the deep logical relationship of the text and avoid missing key information.
[0054] Real-time performance: Based on the online optimization framework of meta-reinforcement learning, it ensures that it can quickly adapt to the dynamic changes of the text.
[0055] Reliability: The results are enhanced by causal inference and interpretability, meeting the strict requirements of high-risk fields.
[0056] Efficiency: The overall architecture design improves the operation efficiency of large-scale inference tasks and is suitable for real-time applications in complex scenarios.
[0057] By integrating a variety of advanced AI technologies, the present invention deeply integrates an inference path generation mechanism with dynamic weight allocation, meta-reinforcement learning, an inference path optimization framework based on meta-reinforcement learning, causal inference, a multi-person feedback aggregation algorithm based on game theory, and interpretability enhancement technology, significantly improving the performance of the intelligent query and knowledge inference system, and providing reliable technical support for knowledge services in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided drawings.
[0059] Figure 1 It is a flowchart of the knowledge exploration method based on the generative thought chain and feedback mechanism provided by the present invention.
[0060] Figure 2 It is a technical roadmap of the knowledge exploration based on the generative thought chain and feedback mechanism provided by the present invention.
[0061] Figure 3 It is an architecture diagram of the knowledge exploration system based on the generative thought chain and feedback mechanism provided by the present invention. Detailed implementation manners
[0062] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0063] The complex scenarios targeted by the present invention involve a type of text with the following characteristics:
[0064] Formal characteristics: Formal and rigorous long texts with complex structures, including multi-level logical relationships; dense terms in the text, high semantic ambiguity, posing high requirements for information extraction and reasoning; frequent dynamic updates of the text content, requiring the system to have efficient real-time update and adaptation capabilities.
[0065] Characteristics in applications: Strong multi-dimensional relevance, requiring reasoning to integrate multiple information sources; the scenarios involve high-risk decisions, demanding extremely high accuracy and transparency of the reasoning results; different users have different interpretation requirements for the same text, and personalized reasoning services need to be supported; some scenarios require multi-disciplinary collaborative reasoning to process multi-source and multi-modal information.
[0066] These characteristics are concentrated in formal and rigorous long texts such as laws and regulations, academic papers, technical specifications / standards documents, government reports, contract documents, medical guidelines, financial reports, user manuals, etc. Their core uses include:
[0067] Laws and regulations: Quickly locate key information in legal provisions to assist professionals in compliance analysis or legal consultation.
[0068] Academic research and technical documents: Support the mining of knowledge points in academic papers and the generation of literature reviews, and accelerate the understanding and citation of technical specifications by technical R & D personnel.
[0069] Business and financial analysis: Analyze financial reports and contract documents to assist in identifying risk terms or conducting financial data analysis.
[0070] Medical and health: Interpret medical guidelines and case materials to provide a reference basis for clinical decision-making.
[0071] Policy interpretation: Process government reports and social policy documents to extract policy highlights and implementation suggestions.
[0072] In modern society, laws and regulations, academic papers, technical specifications / standards documents, government reports, contract documents, medical guidelines, financial reports, user manuals, etc. are all long and formal, rigorous texts. Such texts have the following significant characteristics and laws:
[0073] Large text scale and complex structure: The texts in these scenarios are often long in length and have a complex internal structure, containing various levels of logical relationships, such as progressive, parallel, and conditional associations between clauses.
[0074] Term-intensive and high semantic ambiguity: Formal texts often contain a large number of professional terms, industry abbreviations, and ambiguous expressions, which increase the difficulty of information extraction and logical reasoning.
[0075] Strong multi-dimensional correlation: There is often a high degree of correlation between different paragraphs or chapters in the text, and the derivation of a single conclusion may require reasoning by integrating information from multiple sources.
[0076] Frequent dynamic updates and fast version iterations: Texts such as laws and regulations, technical standards, etc. are often revised frequently due to policy adjustments or technological innovations, which puts high requirements on the real-time update ability of the knowledge base.
[0077] Application scenarios involve high-risk decision-making: In the fields of finance, medical care, law, etc., any reasoning error may lead to major economic losses or legal disputes, so extremely high requirements are placed on the accuracy, reliability, and interpretability of the system.
[0078] Significant differences in diverse user needs: Different users may have different interpretations of the same text. For example, legal advisors may focus on the scope of application of legal terms, while technicians may focus on the specific implementation steps of operating specifications.
[0079] Strong cross-field collaboration: For example, the interpretation of medical guidelines may require combining clinical experience, pharmacological knowledge, and individual patient conditions. This multi-disciplinary cross characteristic makes simple keyword matching unable to meet the needs, and deep reasoning ability must be relied on.
[0080] In summary, the regularity of these scenarios lies in the fact that they pose comprehensive challenges to the knowledge integration ability, reasoning accuracy, dynamic adaptation ability, and multi-dimensional feedback mechanism of the knowledge service system. In view of these characteristics, the present invention adopts a generative thought chain and dynamic feedback optimization technology, which can effectively solve the problems of intelligent query and reasoning in complex scenarios and achieve the precision, real-time, and credibility of knowledge services.
[0081] Refer to Figure 1 As shown, an embodiment of the present invention discloses a knowledge exploration method based on a generative thought chain and a feedback mechanism, including the following steps S10 to S60:
[0082] S10. Receive a multi-modal query input by a user, and perform intention analysis to identify the core requirement, writing theme, or problem focus;
[0083] This method can be implemented through a software program. For example, it is an interactive query platform system that provides a friendly user interface and supports the user to input questions or requirements in natural language. The user first inputs writing requirements through the system interface, and these requirements can be specific writing instructions, natural language descriptions, or queries (QA) in the form of questions. After receiving the input, the system immediately performs intention analysis, parses the user's input, and identifies the core requirement, writing theme, or problem focus, so as to provide a clear direction for subsequent processing.
[0084] S20. Preprocess the multi-modal query and extract key information, and decompose the complex problem into multiple sub-questions or sub-tasks;
[0085] After the intention analysis in step S10 is completed, the system preprocesses the Query input by the user, including natural language processing operations such as word segmentation and part-of-speech tagging to extract key information. Subsequently, the system further extracts keywords and key terms in the Query using natural language processing technology, and these keywords will serve as the basis for generating the reasoning path. For complex problems, the system will also decompose the Query into multiple sub-questions or sub-tasks for subsequent hierarchical reasoning.
[0086] S30. Based on the multiple sub-questions or sub-tasks, dynamically construct a hierarchical reasoning path based on the generative thought chain, and label the logical basis for each step of the reasoning path;
[0087] In this step, based on the generative thought chain technology, the system dynamically constructs a hierarchical reasoning path. The reasoning path starts from the Query input by the user and is gradually decomposed into multiple subtasks, with each subtask corresponding to a node in the reasoning path. The system annotates each step of the reasoning path with clear logical bases, such as causal relationships, rule matches, or semantic associations, which provide a foundation for subsequent weight assignment and optimization. The initial reasoning path is presented in a hierarchical structure to ensure that complex problems can be solved step by step.
[0088] S40. Optimize the priority of the reasoning path through dynamic weight assignment, and combine the meta-reinforcement learning framework to dynamically adjust the reasoning strategy online to optimize the reasoning path;
[0089] In this step, the system calculates the weight of each reasoning node through a dynamic weight assignment mechanism. The weight assignment takes into account the task context, semantic similarity, and the importance of the node to ensure that key nodes and important logical relationships are processed first. According to the dynamic weights, the system re-adjusts the priority of the reasoning path to optimize the efficiency and accuracy of the reasoning path. During the reasoning process, the system also dynamically adjusts the weight assignment according to the intermediate results to further optimize the reasoning path.
[0090] The optimized reasoning path then enters the meta-reinforcement learning framework for online optimization. The meta-reinforcement learning framework uses real-time user feedback and historical data to dynamically adjust the reasoning strategy to adapt to new tasks and changing input scenarios. Through policy gradient updates and adversarial sample generation, the system continuously optimizes the response speed and accuracy of the reasoning path. At the same time, the meta-reinforcement learning framework combines causal reasoning analysis tools to further strengthen the correctness of causal relationships in the reasoning path and avoid false conclusions caused by misjudgment of correlations.
[0091] Among them, the dynamic weight assignment mechanism and the meta-reinforcement learning framework are as follows:
[0092] 1. Reasoning path generation mechanism based on dynamic weight assignment
[0093] Traditional thought chain generation relies on fixed rules or single-modal input, while this system realizes real-time optimization of the reasoning path and multi-modal collaborative reasoning through a dynamic weight assignment algorithm and a cross-modal joint embedding space, specifically including:
[0094] Dynamic reasoning node weight calculation:
[0095] Propose an adaptive weight assignment formula based on the task context to transform traditional static path selection into a dynamic optimization problem:
[0096]
[0097] w i : represents the reasoning node n iWeight.
[0098] q: Represents the current query task.
[0099] n i : Represents the i-th inference node.
[0100] j represents the number of all inference nodes in the current path;
[0101] sim(q, n i ): Represents the semantic similarity between the current task q and the inference node n i .
[0102] τ: Temperature coefficient, used to control the intensity of path exploration.
[0103] Multi-modal joint inference path generation:
[0104] Based on the joint embedding method of cross-modal alignment loss, solve the problem of multi-modal information (text, image, speech) fusion deviation.
[0105] The following takes text and image as examples to illustrate and define the alignment loss function:
[0106]
[0107] Among them, Ⅱ ab ∈ {0, 1} is the modal association indicator function. When the text T a is related to the image I b , it is 1, otherwise it is 0.
[0108] f T (·) and f I (·) are the text encoder and the image encoder respectively, and output the normalized embedding vectors.
[0109] I k represents the k-th sample in the image modality; k ∈ (1, M), but k ≠ b;
[0110] cos(·, ·) represents the cosine similarity, and τ is the temperature coefficient (set to 0.07 in the experiment).
[0111] N is the number of text samples, and M is the number of image samples.
[0112] Objective: Achieve cross-modal semantic alignment by maximizing the similarity of relevant text-image pairs and minimizing the similarity of irrelevant pairs.
[0113] Normalization and temperature coefficient: The embedding vectors are L2-normalized, and the temperature coefficient τ controls the sharpness of the similarity distribution to avoid gradient saturation.
[0114] The system can generate a logical reasoning path based on the query problem input by the user and finally obtain the answer to the problem. The key lies in dynamically constructing the reasoning path, that is, in the context of each problem, the reasoning chain should be automatically generated and adjusted at any time according to the content of the query.
[0115] According to the query type, decompose complex problems into multiple actionable subtasks or sub-problems. For legal domain problems, tasks may include subtasks such as legal provision retrieval, liable party analysis, case retrieval, etc.
[0116] According to the problem type (such as legal liability, medical problems, business decisions, etc.), the system selects appropriate reasoning frameworks. These reasoning frameworks include rule-based reasoning and knowledge reasoning.
[0117] The core idea of constructing the reasoning path is to dynamically select the subtasks of the reasoning path according to the specific needs of the user's query and adjust the path according to the intermediate results. Suppose the system has multiple reasoning nodes n1, n2,..., n k , each node represents a reasoning task; k represents the number of nodes in the current path. For a given input query, the path selection is determined by the following dynamic weight assignment:
[0118]
[0119] where T is the query task, W i is the weight of node n i , and R i (T) is the reward value of the node (calculated according to the effectiveness and efficiency of the current reasoning result). The path selection is completed by maximizing the total reward of the path:
[0120]
[0121] where p represents the dynamically adjusted normalized weight coefficient, p ∈ [0, 1].
[0122] 2. Online Optimization Framework for Reasoning Path Based on Meta-Reinforcement Learning
[0123] Aiming at the lag problem of the traditional feedback mechanism, a hybrid architecture of meta-reinforcement learning and online supervised learning is proposed to realize the real-time correction and long-term stability optimization of the reasoning path:
[0124] Dual-stage meta-learning strategy:
[0125] Offline meta-pretraining: Adopt the MAML (Model-Agnostic Meta-Learning) framework to pre-train the model on a cross-domain task set (law, medicine, finance) to enable it to quickly adapt to new tasks.
[0126] Online fine-tuning: Generate adversarial samples through real-time user feedback, dynamically update model parameters, and solve the domain drift problem.
[0127] Multi-objective reward function design:
[0128] Propose a composite reward mechanism that integrates interpretability metrics:
[0129] R = λ1·Accuracy + λ2·Efficiency + λ3·Interpretability + λ4·UserSatisfaction
[0130] Where λ1, λ2, λ3, λ4 are weight coefficients, Accuracy is accuracy, Efficiency is efficiency, Interpretability is interpretability, and UserSatisfaction represents user satisfaction. Determine the optimal combination through Pareto front analysis, satisfying λ1 + λ2 + λ3 + λ4 = 1. The time decay factor r = 0.95 dynamically adjusts the weight of user satisfaction to ensure the long-term stability of the system performance.
[0131] For each relationship r, the system generates a structured mapping rule by defining its corresponding entity category Ener in the template:
[0132] SchemaNER(r) = {(h, t)|h ∈ HeadEntities(r), t ∈ TailEntities(r)}
[0133] For example, for the relationship r = "responsibility", define:
[0134] SchemaNER("responsibility") = {("company", "employee"), ("legal person", "shareholder")}
[0135] Guide the model to generate SchemaNERS through task examples to ensure the standardization and consistency of entity prediction.
[0136] To help the large language model understand the generation background of named entity patterns, task background prompts can be provided to the model. This prompt contains background information of a specific domain, aiming to guide the large model to better understand and generate correct entity patterns in tasks of legal and responsibility classification.
[0137] Task background prompt: The task background prompt provides the large model with the generation background information, including the specific definition of the potential relationship r, examples of the entity category Ener, and the relevant legal context. By giving the task background prompt, the system can guide the model to be more precise in processing entity annotation, reduce irrelevant predictions, and focus on the target task.
[0138] Template Filling and Data Generation: The process of filling the template guides the model to generate the corresponding SchemaNER through given task examples. For example, given the potential relationship in the task description, such as "the responsibilities of a company towards its employees", the template will automatically fill in the entity categories and generate a schema that includes "company" and "employee" as the head and tail entities, ultimately generating a complete named entity schema SchemaNER.
[0139] The system can ensure that each potential relationship can accurately match the relevant legal provisions and entities when generating the named entity schema. This typical mapping defines the structured mapping between relationships and entities, ensuring the standardization and consistency of annotation.
[0140] Example Mapping: Suppose a certain legal provision involves the responsibility issues of "company" and "employee". Through the mapping relationships "company - responsibility" and "employee - responsibility", the system defines "company" and "employee" as potential entity categories in the named entity template to help the subsequent NER model understand their semantic roles.
[0141] Relationship Constraints: The system provides clear entity prediction boundaries for the model by restricting the potential relationship r in the task. For example, the task background prompt clearly stipulates that the "responsibility" relationship can only involve the two entities of company and employee, thereby reducing the prediction of other irrelevant entities.
[0142] Over - prediction Control: Through the mapping control between relationships and entities, the system can ensure that the output of the large - model is limited to entities related to the task, avoiding the generation of irrelevant or redundant entity categories. This makes the model more focused and accurate when dealing with complex legal relationships.
[0143] S50. On the basis of optimizing the inference path, use the causal reasoning analysis tool to verify the inference path and generate a counterfactual path to correct the wrong nodes;
[0144] On the basis of the optimized inference path, the system uses the causal reasoning analysis tool to further identify and verify the causal relationship between variables in the inference path. Through counterfactual path analysis, the system compares the differences between the actual path and the counterfactual path, locates the wrong inference nodes and triggers corrections. At the same time, causal probability edges are embedded in the knowledge graph to define the causal strength between nodes, further enhancing the interpretability and credibility of the inference path.
[0145] Among them, the causal - driven inference path correction technology is as follows:
[0146] To solve the credibility problem of the black - box model, introduce causal intervention technology and construct a traceable inference chain correction process:
[0147] Counterfactual Path Analysis:
[0148] When the inference result deviates from the expectation, the system automatically generates a counterfactual query:
[0149] "How will the result change if the conclusion of node N_K is ignored?"
[0150] By comparing the differences between the actual path and the counterfactual path, locate the misreasoning node and trigger correction.
[0151] Causal graph enhancement:
[0152] Embed causal probability edges in knowledge and define the causal strength between nodes: For example, this technology enables the system to identify confounding variables and output causal explanations in scenarios such as medical liability determination. By calculating the logical consistency score, task relevance score, case support score, and interpretability score, the inference credibility and interpretability are enhanced.
[0153] Logical consistency score: Calculate the logical coherence between the steps of the inference chain to ensure that there are no contradictory conclusions.
[0154] Task relevance score: Calculate the relevance between the user query and the generated thought chain through vector similarity. Combine the pre-trained language model to calculate the matching degree between the generated content and the target question.
[0155] Case support score: Retrieve relevant legal cases from the case library and calculate the support degree of the current inference chain. If the inference chain highly matches the known legal cases, a higher score is given to enhance the inference credibility.
[0156] Interpretability score: Evaluate whether the inference chain provides enough intermediate inference steps to make the final conclusion verifiable. If the inference path is transparent and convenient for manual review, a higher score is given.
[0157] To improve the accuracy and robustness of the reward model, this embodiment uses the reinforcement learning method for optimization. Use the policy gradient method to optimize the inference path of the large model to make it more in line with the scoring criteria of the reward model. Generate multiple candidate inference paths through the autoregressive method and select the path with the highest score as the final output. Combine the high-quality inference chains with manual annotation to fine-tune the reward model to make it more in line with the judgment criteria of legal experts.
[0158] In each step of generating the inference path, the system optimizes the inference chain through the reward model. The reward mechanism is as follows:
[0159] R = w1·Accuracy + w2·Efficiency - w3·Resources
[0160] w1: The weight coefficient of Accuracy. It represents the importance of the accuracy of the reasoning chain in the reward model. Accuracy generally refers to the degree of consistency between the reasoning result and the actual result or the known correct answer. A higher accuracy weight means that the system is more inclined to select those paths that can provide correct or precise reasoning results.
[0161] Accuracy: It refers to the accuracy score of the reasoning chain, which can be calculated by comparing the matching degree between the reasoning result and the real situation or known cases. In the legal reasoning system, this involves multiple aspects such as logical consistency score, task relevance score, and case support score.
[0162] w2: The weight coefficient of Efficiency. It represents the importance of the efficiency of the reasoning process in the reward model. Efficiency generally refers to the time and resources required to complete the reasoning. A higher efficiency means that the system can generate reasoning results faster.
[0163] Efficiency: This refers to the efficiency score of the reasoning chain, which can be evaluated by calculating the time required for the reasoning process, computing resource consumption, etc. In practical applications, improving efficiency can help the system respond to user queries faster, thereby enhancing the user experience.
[0164] w3: The weight coefficient of Resources. It represents the importance of the negative impact of resource consumption in the reasoning process in the reward model. Resource consumption generally refers to the computing resources, storage space, etc. required to complete the reasoning.
[0165] Resources: This refers to the resource consumption score of the reasoning chain, which can be calculated by evaluating the computing resources used and storage requirements during the reasoning process. In an environment with limited resources, reducing resource consumption can improve the scalability and cost-effectiveness of the system.
[0166] S60. Output the structured text and the visualized reasoning process, and optimize the model parameters in real time based on the multi-person feedback aggregation of game theory.
[0167] After the reasoning path that has undergone causal reasoning verification and reinforcement, it finally enters the text generation step. The system generates text content that meets the user's needs in a structured form according to the optimized reasoning path. The generated text not only meets the user's writing needs but also comes with a detailed description of the reasoning process, intuitively showing the reasoning logic through methods such as timeline diagrams and causal relationship diagrams, enhancing the transparency of the results and the user's trust in the system.
[0168] Users evaluate the generated text content to provide feedback on the quality of the system's generated results. First, the system processes the users' evaluations through a multi-person feedback aggregation algorithm based on game theory, quantifying the feedback contribution of each user and coordinating the evaluation opinions of different users to ensure the fairness and stability of the optimization process. Subsequently, the system uses a meta-reinforcement learning framework to process the aggregated feedback information in real time, dynamically adjusting the inference path and model parameters to meet the diverse needs of users. Based on the users' feedback, the system further continuously updates the model parameters through online supervised learning to continuously improve the stability and long-term performance of the inference path. Through this closed-loop optimization mechanism, the system can accurately meet the users' needs and achieve intelligent knowledge exploration and text generation services.
[0169] Among them, the above-mentioned multi-person feedback aggregation algorithm based on game theory is described as follows:
[0170] Regarding the problem of multi-user feedback conflicts, a Shapley value feedback aggregation mechanism is proposed to ensure the fairness and stability of the optimization process:
[0171] Quantification of feedback contribution degree:
[0172]
[0173] Among them, φ u (v) represents the feedback contribution degree of user u, v(S) represents the feedback value of user subset S, which is used to filter low-quality or malicious feedback. H represents the set of all users participating in the feedback, and! represents the factorial.
[0174] When the query contains multiple modalities (such as text, image, audio, etc.), the system fuses various modality information by introducing a joint embedding space. For example, assuming T represents the text modality and I represents the image modality, the system can use the following joint embedding e c to uniformly represent the two modalities:
[0175] e c = f(T,1) = α·f T (T)+(1-α)·f I (I)
[0176] Among them, f T (T) and f I (I) are the embedding representations of text and image respectively, and α∈[0.1] is the weighting coefficient, which is used to control the degree of information fusion between text and image. This fusion method ensures that different modality information can jointly affect the path selection in the inference path, thus improving the flexibility and accuracy of the inference chain.
[0177] Vector Representation: Word embedding and semantic embedding techniques not only transform entities into vectors but also capture the semantic relationships between entities. For example, the relationship between "company" and "legal person" will have a high similarity in the vector space, while the similarity between "company" and "fine" is relatively low. The system judges the semantic relationships between entities through the similarity measurement of these vectors, providing a basis for subsequent reasoning and the generation of the chain of thought.
[0178] The system adjusts the vector representation according to the input query and context. When processing user queries, the model can dynamically update the vectors based on the user's feedback to optimize semantic understanding. For example, when the user queries "the scope of company liability", the system will understand the specific meanings of "company" and "liability" in different legal fields (such as civil law or criminal law) through the large model and adjust the vector representation to generate a more appropriate chain of thought.
[0179] When facing new tasks, the system quickly adjusts the reasoning path through transfer learning and meta-learning techniques. Through MAML in meta-learning, the reasoning path can quickly adapt to new query requirements from a small amount of task data:
[0180]
[0181] θ * It means optimizing the model parameter θ through meta-learning to minimize the expected loss on the task distribution d(T). d(T) is the task distribution, representing the probability distribution of all possible tasks. ET~d(T) means taking the expected value of the loss L for all tasks T. T Take the expected value.
[0182] Among them, θ is the model parameter, θ' is the model parameter after task adaptation, and L T is the loss function of task T. By training an initial parameter θ, the model can obtain the optimal solution through adjustment with a small amount of data when encountering new tasks.
[0183] Refer to Figure 2 As shown, it describes the entire technical process of this method from multi-modal input to the final output result, as well as how to optimize this process through feedback:
[0184] 1. Multi-modal Input: Receive input data from different modalities, such as text, images, and speech.
[0185] 2. Information Structuring: The input information is structured to facilitate subsequent processing. This step ensures the organization and format of the information, making it suitable for machine processing.
[0186] 3. Task Identification: Identify the specific task requirements of the user, providing a direction for subsequent reasoning and processing.
[0187] 4. Vectorization processing: Information is converted into vector form for semantic processing. This step transforms text, image, and speech data into numerical forms that can be processed by machine learning models.
[0188] 5. Semantic conflict detection: Detect and handle semantic conflicts in the input information. This step ensures the consistency and accuracy of the information.
[0189] 6. Dynamic inference path generation: Generate dynamic inference paths according to task requirements. This step allows the system to dynamically adjust its inference strategy according to specific tasks.
[0190] 7. Dynamic weight assignment: Assign weights to each node in the inference path to optimize the inference process. This step improves the efficiency and accuracy of inference by adjusting the importance of different inference steps.
[0191] 8. Meta-reinforcement learning optimization: Use meta-reinforcement learning techniques to optimize the inference path and improve the adaptability and efficiency of the system. This step continuously improves the inference process by learning the optimal strategy.
[0192] 9. Multi-level inference: Conduct multi-level inferences to obtain the final result. This step improves the depth and breadth of inference by deeply analyzing and synthesizing information at different levels.
[0193] 10. Visualization of the inference process: Visualize the inference process for users to understand. This step improves the interpretability of the result by graphically displaying the inference process.
[0194] 11. Output result: Output the inference result to the user. This step ensures the clarity and ease of understanding of the result.
[0195] 12. Counterfactual path analysis: Analyze possible counterfactual paths to optimize the inference process. This step improves the flexibility and adaptability of the system by exploring different inference paths.
[0196] 13. User interaction feedback: Collect user feedback, including explicit and implicit signals. This step provides a basis for system optimization by collecting users' feedback on the result.
[0197] 14. Multi-party feedback processing: Process user feedback from different aspects to further optimize the system. This step improves the comprehensiveness and accuracy of the system by synthesizing feedback from different sources.
[0198] 15. Structured decision-making advice: Provide structured decision-making advice based on the inference result. This step improves the practicality of the result by transforming the inference result into specific action suggestions.
[0199] 16. Reasoning process backtracking: Allows users and the system to backtrack the reasoning process for better understanding and improvement. This step enhances the transparency and interpretability of the system by providing a detailed record of the reasoning process.
[0200] The present invention deeply integrates a reasoning path generation mechanism with dynamic weight allocation, meta-reinforcement learning, a reasoning path optimization framework based on meta-reinforcement learning, causal reasoning, a multi-person feedback aggregation algorithm based on game theory, and an interpretability enhancement technology, significantly improving the performance of the intelligent query and knowledge reasoning system and providing reliable technical support for knowledge services in complex scenarios.
[0201] 1) Reasoning path generation mechanism with dynamic weight allocation:
[0202] Through dynamic weight allocation, the present invention can intelligently allocate attention to different reasoning paths, prioritize key nodes, and ensure the logic and efficiency of the reasoning process. This solves the problem of missing key points that may be caused by fixed weights in traditional methods.
[0203] 2) Meta-reinforcement learning and reasoning path optimization framework:
[0204] The meta-reinforcement learning framework allows the system to automatically adjust the reasoning strategy during operation to adapt to dynamically changing input scenarios. This mechanism can quickly optimize the reasoning path in a real-time environment, improving accuracy and response speed.
[0205] 3) Causal reasoning and multi-dimensional correlation analysis:
[0206] Based on causal reasoning technology, the system can identify the true causal relationships between variables and avoid incorrect reasoning conclusions due to correlation. This is particularly important for dealing with multiple causal relationships in complex texts.
[0207] 4) Multi-person feedback aggregation algorithm based on game theory:
[0208] In response to the diverse needs of different users, the system uses game theory methods to balance multi-party feedback and find the optimal reasoning result. This method can coordinate the views of different stakeholders and improve the universality of the final conclusion.
[0209] 5) Interpretability enhancement technology:
[0210] Through the interpretability enhancement technology, not only the final conclusion is output, but also the reasoning process and basis are detailedly displayed, enhancing the transparency and credibility of the result.
[0211] Based on the same inventive concept, an embodiment of the present invention also provides a knowledge exploration system based on a generative thought chain and a feedback mechanism. Referring to Figure 3 as shown, it includes:
[0212] A data input module for receiving multimodal queries input by a user, performing intent analysis to identify the core requirements, writing themes, or problem focuses;
[0213] A semantic processing module for preprocessing the multimodal queries and extracting key information, and decomposing complex problems into multiple sub-problems or sub-tasks;
[0214] An inference engine module for dynamically constructing a hierarchical inference path based on the generative thought chain according to the multiple sub-problems or sub-tasks, and annotating the logical basis for each step of the inference path;
[0215] A dynamic inference path module for optimizing the priority of the inference path through dynamic weight allocation, and online dynamically adjusting the inference strategy in combination with the meta-reinforcement learning framework to optimize the inference path;
[0216] A causal inference analysis module for verifying the inference path by using a causal inference analysis tool on the basis of optimizing the inference path, and generating a counterfactual path to correct the wrong nodes;
[0217] A result output and feedback optimization module for outputting structured text and visualizing the inference process, and real-time optimizing the model parameters based on the multi-person feedback aggregation of game theory.
[0218] As Figure 3 shown, the system architecture depicts the complete process from data input to user feedback, as well as the interactions and data flows between the various modules.
[0219] Among them, the data input module is the starting point for the user to interact with the system, responsible for receiving the user's query requests, which can be multimodal data in the form of text, images, or voices. And perform SchemaNER entity recognition and structuring; after the data input, the system uses SchemaNER technology to recognize and structure the entities in the input data, and perform intent analysis to identify the core requirements, writing themes, or problem focuses.
[0220] The semantic processing module can perform dynamic vectorization processing on the input information, and is responsible for converting the user's query into a semantic vector that can be understood by the machine. Through the multi-modal joint embedding space technology, the system can process the fusion of text, images, and voices, and realize the unified representation of information. It can preprocess the multimodal queries and extract key information, and decompose complex problems into multiple sub-problems or sub-tasks.
[0221] The inference engine module is used for thought chain generation; this module is the core of the system and is responsible for generating the inference path. It can dynamically construct a hierarchical inference path based on the generative thought chain according to multiple sub-problems or sub-tasks, and annotate the logical basis for each step of the inference path.
[0222] The dynamic reasoning path module generates reasoning paths dynamically according to task requirements through the dynamic reasoning path tree generation technology, and assigns node weights; and combines with the meta-reinforcement learning framework to dynamically adjust the reasoning strategy online and optimize the reasoning path; in addition, the system also expands the breadth and depth of reasoning through the intention expansion and divergent thinking technology to obtain the reasoning result.
[0223] The causal reasoning analysis module, on the basis of optimizing the reasoning path, uses the causal reasoning analysis tool to verify the reasoning path and generate counterfactual paths to correct the wrong nodes;
[0224] The result output and feedback optimization module converts the reasoning result into a structured explanation and presents it in a user-friendly way, making the result easy to understand and operate. User feedback collection (explicit / implicit signals) is used to collect user feedback information. This feedback can be explicit (such as direct evaluations or suggestions) or implicit (such as click behavior or usage patterns). The feedback optimization part includes a meta-reinforcement learning optimizer
[0225] (Meta-RL policy update), supervised learning corrector (fine-tuning of labeled data), and causal counterfactual analysis (backtracking of wrong paths). These techniques are used to optimize the reasoning path and model parameters in real time, improving the adaptability and accuracy of the system.
[0226] The knowledge exploration system based on the generative thinking chain and feedback mechanism provided by the present invention is designed specifically for complex long text scenarios and is applicable to the processing of formal and rigorous long texts such as laws and regulations, academic papers, technical specifications / standards documents, government reports, contract documents, medical guidelines, financial reports, and user manuals. The following uses 4 application embodiments to illustrate the solution of the present invention:
[0227] Embodiment 1, an intelligent search and retrieval system for laws and regulations:
[0228] Traditional law information retrieval systems usually rely on simple keyword matching, which results in low relevance and accuracy of query results. Especially when facing complex legal issues, it is difficult to screen out the clauses that truly meet the requirements.
[0229] When applying the present invention to search or retrieve laws and regulations, the system can intelligently generate multi-level reasoning paths according to the query requirements input by the user through the reasoning path generation mechanism with dynamic weight allocation, combined with the semantic understanding and context analysis capabilities of laws and regulations. The system can expand the reasoning path from multiple dimensions, such as the scope of application, effective conditions, exceptions, etc. of legal clauses, to help users quickly locate key legal provisions and relevant cases. Through dynamic weight allocation, the system can prioritize important clauses, reduce the interference of irrelevant information, and achieve more accurate retrieval results.
[0230] In addition, the system can also utilize meta-reinforcement learning technology to optimize the inference path based on user feedback. For example, when the user marks certain results as "irrelevant" or "relevant", the system will adjust the weight allocation strategy in real time to gradually optimize the accuracy of search results.
[0231] Example 2, Legal Q&A System:
[0232] In the legal Q&A scenario, the questions raised by users usually involve complex legal backgrounds and may require comprehensive analysis of multiple laws, regulations, and judicial interpretations. Traditional systems are difficult to associate multiple legal provisions in the same query and cannot deeply explore the logical relationships between laws, resulting in incomplete answers.
[0233] When applying the present invention as a legal Q&A system, through the dynamic inference path generation technology, the system can automatically generate a dynamic inference path based on the user's query and form a multi-level legal analysis in combination with legal provisions and power and responsibility classifications. For example, when dealing with issues related to "enterprise compliance review", the system can not only quote relevant provisions of the "Company Law" but also combine relevant provisions of the "Administrative Penalty Law" and the "Civil Procedure Law" to construct a complete legal reasoning chain.
[0234] In addition, the system also supports a multi-round feedback optimization mechanism. Users can rate the answer or put forward supplementary requirements, and the system will dynamically adjust the inference path according to the feedback to gradually improve the quality and personalization of the answer. This method is particularly suitable for dealing with complex legal matters such as intellectual property disputes and contract performance disputes.
[0235] Example 3, Legal Document Retrieval and Analysis System:
[0236] Facing a vast amount of legal documents (such as case laws, legal interpretations, lawyer's opinions, etc.), traditional systems usually lack in-depth reasoning ability and semantic understanding ability and are difficult to quickly locate the most relevant documents.
[0237] When conducting legal document retrieval and analysis based on the system of the present invention, through intelligent query path generation and causal reasoning technology, it can identify the logical relationships between documents, such as the citation relationship of legal provisions, the corresponding relationship between cases and legal clauses, etc., to help users quickly obtain the required information. For example, when retrieving cases applying a certain legal principle, the system can combine causal reasoning to screen out relevant real cases and sort them according to importance.
[0238] Through the dynamic feedback mechanism, the system can also perform intelligent recommendations according to the user's preferences and behavior habits (such as frequently consulted legal fields, commonly used legal terms, etc.) to further improve the retrieval efficiency and reduce the interference of invalid information.
[0239] Example 4, Legal Consultation and Intelligent Customer Service:
[0240] In legal consultation services, users may need to obtain personalized legal advice for specific issues. Traditional legal consultation systems usually can only provide static knowledge base content and cannot be dynamically adjusted according to the specific situation of users.
[0241] When applying the system of the present invention for legal consultation, the system can generate personalized legal advice based on the user's query history and behavior data by combining a feedback mechanism and a dynamic optimization framework. For example, when a user asks "how to handle labor arbitration cases", the system will not only quote the relevant provisions of the Labor Law but also combine local judicial practices and similar cases to provide more specific advice.
[0242] In addition, the system supports a multi-round interaction mode, can gradually refine the user's query requirements in the conversation, and finally generate a complete legal analysis report. The user can also rate the inference results of each step, and the system will adjust the next inference strategy according to the feedback, thereby continuously improving the service quality.
[0243] The present invention introduces a feedback mechanism based on meta-reinforcement learning to achieve real-time inference path optimization. Through user feedback, the system can dynamically adjust the weight distribution of the inference path and optimize the accuracy of the query results. For example, the system can update the weight parameters in real time through the user's scoring of the results or behavior data (such as the number of clicks to view), and gradually improve the relevance of the retrieval results. For the conflict problem of multi-user feedback, the system adopts the Shapley value feedback aggregation mechanism to ensure the fairness and stability of the optimization process. By quantifying the feedback contribution of each user, the system can exclude the interference of low-quality feedback and optimize the selection strategy of the inference path.
[0244] In addition, the system also supports feedback-driven online supervised learning, uses user feedback to continuously update the model parameters, and improves the stability and long-term performance of the inference path. For example, the system can generate adversarial samples after each user feedback, simulate the inference path in extreme situations, and test and correct potential inference biases.
[0245] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0246] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A knowledge exploration method based on generative thought chains and feedback mechanisms, characterized in that, It includes the following steps: S10. Receive a multimodal query input by the user, perform intent analysis to identify the core requirement, writing theme, or problem focus; S20. Preprocess the multimodal query and extract key information, decompose complex problems into multiple sub-problems or sub-tasks; S30. Based on the multiple sub-problems or sub-tasks, dynamically construct a hierarchical reasoning path based on generative thinking chains, and annotate logical bases for each step of the reasoning path; S40. Optimize the priority of the reasoning path through dynamic weight allocation, combine the meta-reinforcement learning framework to dynamically adjust the reasoning strategy online, and optimize the reasoning path; S50. On the basis of optimizing the reasoning path, use a causal reasoning analysis tool to verify the reasoning path and generate a counterfactual path to correct error nodes; S60. Output structured text and a visual reasoning process, and optimize model parameters in real time based on the multi-person feedback aggregation of game theory.
2. The knowledge exploration method based on generative thinking chain and feedback mechanism according to claim 1, characterized in that In step S40, optimizing the priority of the reasoning path through dynamic weight allocation includes: Dynamically calculate the weights of each node in the reasoning path according to task context, semantic similarity, and node importance. The weights are calculated by the following formula: where w i represents the weight of the inference node n i ; q represents the current query task; n i represents the i-th inference node; j represents the number of all inference nodes in the current path; sim(q, n i ) represents the semantic similarity between the current task q and the inference node n i ; τ represents the temperature coefficient, which is used to control the path exploration intensity; According to the multiple nodes decomposed by the reasoning path, each node corresponds to a sub-task, and the optimal selection of the path is achieved by maximizing the total reward value through a reward model; among them, the formula for maximizing the total reward value is as follows: Among them, T represents the query task, and R i (T) represents the reward value of the node, k represents the number of nodes in the current path; p represents the dynamically adjusted normalized weight coefficient, where p ∈ [0, 1].
3. The knowledge exploration method based on generative thinking chains and feedback mechanisms according to claim 2, wherein In step S40, optimizing the priority of the reasoning path through dynamic weight allocation also includes: Fuse multimodal input information, and achieve semantic alignment of text and images through an alignment loss function. The loss function is defined as: Among them, Ⅱ ab ∈ {0, 1} represents the modal correlation indicator function. When the text T a is related to the image I b , it is 1; otherwise, it is 0; f T (·) and f I (·) represent the text encoder and the image encoder respectively, and output the normalized embedding vectors; I k represents the k-th sample in the image modality, k ≠ b; cos(·, ·) represents the cosine similarity, τ represents the temperature coefficient; N represents the number of text samples, and M represents the number of image samples.
4. A knowledge exploration method based on generative thought chains and feedback mechanisms according to claim 2, characterized in that In step S40, the meta-reinforcement learning framework dynamically adjusts the reasoning measurement online, including: Offline meta-pre-training, using the MAML framework to pre-train the model on a cross-domain task set; Online fine-tuning generates adversarial samples through real-time user feedback and dynamically updates model parameters; A composite reward mechanism integrating interpretability metrics, through the design of a multi-objective reward function, balances reasoning accuracy, efficiency, and interpretability.
5. A knowledge exploration method based on generative thought chains and feedback mechanisms according to claim 2, characterized in that In step S50, using a causal reasoning analysis tool to verify the reasoning path includes: Generate a counterfactual query when the reasoning result deviates, compare the actual path with the counterfactual path to locate error nodes and trigger corrections; Define the causal strength between nodes by embedding causal probability edges in knowledge, and calculate the logical consistency score, task relevance score, case support score, and interpretability score to enhance the credibility and interpretability of reasoning.
6. A knowledge exploration method based on generative thought chains and feedback mechanisms according to claim 5, characterized in that In step S50, using a causal reasoning analysis tool to verify the reasoning path also includes: In each step of the reasoning path generation, optimize the reasoning chain through a reward model, and the reward mechanism is as follows: R = w1·Accuracy + w2·Efficiency - w3·Resources w1 represents the weight coefficient of accuracy; Accuracy represents the accuracy score of the reasoning chain; w2 represents the weight coefficient of efficiency; Efficiency represents the efficiency score of the reasoning chain; w3 represents the weight coefficient of resource consumption; Resources represents the resource consumption score of the reasoning chain.
7. A knowledge exploration method based on generative thinking chains and feedback mechanisms according to claim 1, characterized in that, In the step S60, the parameters of the real-time optimization model for multi-person feedback aggregation based on game theory include: Based on the Shapley value algorithm of game theory, quantify the contribution of user feedback, filter out low-quality or malicious feedback, and coordinate the opinions of multiple users to optimize the inference path; Use the meta-reinforcement learning framework to process the aggregated feedback information in real time, and dynamically adjust the inference path and model parameters.
8. A knowledge exploration system based on generative thought chains and a feedback mechanism, characterized in that, It includes: A data input module for receiving multi-modal queries input by users, performing intention analysis to identify the core requirements, writing themes or problem focuses; A semantic processing module for preprocessing the multi-modal queries and extracting key information, and decomposing complex problems into multiple sub-problems or sub-tasks; An inference engine module for dynamically constructing a hierarchical inference path based on the generative thinking chain according to the multiple sub-problems or sub-tasks, and annotating the logical basis for each step of the inference path; A dynamic inference path module for optimizing the priority of the inference path through dynamic weight allocation, combining the meta-reinforcement learning framework to dynamically adjust the inference strategy online, and optimizing the inference path; A causal inference analysis module for verifying the inference path using causal inference analysis tools on the basis of optimizing the inference path, and generating counterfactual paths to correct error nodes; A result output and feedback optimization module for outputting structured text and visualizing the inference process, and real-time optimizing the model parameters based on multi-person feedback aggregation of game theory.
Citation Information
Patent Citations
Small sample knowledge graph completion method based on reinforcement learning
CN117150041A
Task processing method and device based on pre-training language model, equipment and medium
CN117217201A
Intelligent doorbell visitor identification method and system based on deep learning
CN118761034A
Digital factory operation virtual simulation teaching method and system
CN119396096A
Software-defined vehicle
WO2024226722A2
Cited By
Technology development situation awareness system and method
CN120429414A
A technology development situation awareness system and method
CN120429414B
Information extraction method and storage medium
CN120430420A
Task processing model evaluation method, role playing model evaluation method and task processing method
CN120744423A
Generative AI text reasoning and characterization method based on loyalty perception mechanism
CN120996214A