Multi-path legal reasoning engine based on mcp protocol

By introducing the Model Context Protocol (MCP) and combining it with hierarchical training and multi-path reasoning mechanisms, the shortcomings of existing legal intelligent question answering systems in terms of context management and reasoning consistency are addressed. This improves the logical consistency and interpretability of the legal question answering system, enhancing its professionalism and credibility in complex legal scenarios.

CN120764667BActive Publication Date: 2026-03-31UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing legal intelligent question-answering systems have shortcomings in constructing complex reasoning chains, the accuracy of legal retrieval, the consistency of article citation, and the interpretability of answers. In particular, they have weak context management capabilities in multi-jurisdictional and multi-level legal scenarios, and the reasoning process is opaque, the results are inconsistent, and there is a lack of traceability.

Method used

The Model Context Protocol (MCP) is introduced to uniformly manage model state, inference path and knowledge activation throughout the entire process of model training, inference decision and result output. Through hierarchical training, Monte Carlo Tree Search (MCTS) and Retrieval Enhancement Generation (RAG) mechanism, the logical consistency and interpretability of legal question answering are ensured.

Benefits of technology

The system has improved the legal question-and-answer system’s capabilities in terms of professionalism, transparency, and structured output, achieving structural transparency of the reasoning chain, controllability of decision-making behavior, and consistency of the context of the output results, thereby enhancing the system’s credibility and interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120764667B_ABST
    Figure CN120764667B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of large language models, and discloses a multi-path legal reasoning engine based on an MCP protocol, which comprises a legal question input module, a legal question answering module and a question and answer result output module; the legal question answering module utilizes a large language model, combines a Monte Carlo tree search algorithm, a multi-path deduction mechanism, a retrieval enhancement mechanism and a user feedback optimization mechanism, and relies on an MCP unified management model to manage internal states, knowledge activation and reasoning path information, and generates an interpretable legal answer corresponding to a legal query question input by a user. The application integrates key information streams with the MCP protocol, outputs answers, reasoning chains, cited bases and credibility evaluations in a visual manner, and significantly improves the professionalism, interpretability and user trust of the whole legal reasoning process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of interdisciplinary research in artificial intelligence and legal technology, specifically relating to an intelligent legal question-answering system that integrates a Model Context Protocol (MCP) and a Large Language Model (LLM). This system comprehensively utilizes the language understanding and generation capabilities of the large language model, combined with Monte Carlo Tree Search (MCTS), causal consistency judgment, multi-path inference mechanisms, retrieval enhancement, and user feedback optimization. With the MCP protocol at its core, it permeates the entire process of model training, inference execution, and result output, ensuring the contextual consistency, logical traceability, and reliability of legal question-answering results. It is particularly suitable for legal intelligent question-answering scenarios with high logicality, high interpretability, and strong professional requirements. Background Technology

[0002] In recent years, Large Language Models (LLMs) have made groundbreaking progress in the field of Natural Language Processing (NLP). Pre-trained models based on the Transformer architecture (such as the GPT series) have demonstrated strong generalization capabilities in text generation, multi-turn question answering, and task transfer, especially as the number of parameters continues to increase, showing broad adaptability in few-shot and zero-shot scenarios. However, traditional LLMs generally face problems such as opaque inference paths, inconsistent outputs, and lack of controllability in practical applications, especially showing significant shortcomings in high-precision semantic domains such as law.

[0003] To improve the security and preference consistency of large-scale model-generated results, scholars have proposed alignment paradigms such as RLHF (Reinforcement Learning Based on Human Feedback) and DPO (Direct Preference Optimization). Although DPO provides an efficient training path without relying on explicit reward models, these techniques have not fundamentally solved the problems of contextual coherence and inference chain integrity in complex contexts for large language models.

[0004] Meanwhile, the continuous development of legal AI is placing higher demands on model capabilities. Tasks such as legal question answering, application of regulations, and case analysis not only require language comprehension but also rely heavily on the ability to model the legal rule system, the interpretability of the reasoning process, and the credibility of the cited evidence. Current mainstream legal models (such as ChatLaw, Lawformer, Lawyer, and LLaMA) generally suffer from the following technical limitations:

[0005] The reasoning process is opaque, lacking transparency and traceability.

[0006] • Frequent semantic mismatches and version conflicts in legal citations;

[0007] • Weak context management capabilities in multi-jurisdictional and multi-level legal scenarios;

[0008] • Model training relies heavily on general corpora and lacks structured legal knowledge and path supervision.

[0009] Even after introducing the Retrieval Enhancement Generation (RAG) framework, the existing system still faces the following bottlenecks:

[0010] There is a lack of consistent context between the search results and the content generated by the model.

[0011] The lack of a unified mechanism for evaluation and ranking after multi-path recall affects the stability of answers.

[0012] The legal citation chain is complex and lacks structured tracing and path basis.

[0013] Therefore, there is an urgent need for a novel mechanism that can uniformly manage model states, inference paths, and knowledge activation to overcome the problems of context fragmentation, inference jumps, and unreliable results in traditional systems. This invention introduces a Model Context Protocol (MCP) to form a consistent context control structure across the stages of model training, inference decision-making, knowledge retrieval, and result output. This improves the logical consistency, accuracy of citations, and interpretability of output in legal question answering, solving key technical challenges in the professionalism, transparency, and structured output of existing legal intelligent question answering systems. Summary of the Invention

[0014] To address the shortcomings of existing legal intelligent question-answering systems in constructing complex reasoning chains, the accuracy of legal retrieval, the consistency of text citation, and the interpretability of answers, this invention proposes a multi-path legal reasoning engine and method based on the MCP protocol to effectively overcome the aforementioned technical challenges and improve the system's practicality and reliability in professional legal scenarios.

[0015] The core objective of this invention is to construct a legal intelligent question-answering system based on the MCP unified context management mechanism, enhancing the comprehensive capabilities of large language models in handling multi-step legal reasoning, legal citation paths, context preservation, and entity consistency. MCP, as the master control protocol throughout the system's training, reasoning, feedback, and output, is responsible for the unified management of the model's internal state, reasoning path evolution, knowledge activation, and citation records, thereby ensuring the logical consistency and interpretability of the entire legal question-answering chain. To achieve the above objectives, this invention proposes the following technical solution:

[0016] According to one aspect of the present invention, a multi-path legal reasoning engine based on the MCP protocol is provided, comprising:

[0017] Legal question input module;

[0018] Legal Q&A module;

[0019] Question and answer result output module.

[0020] in:

[0021] The legal question input module receives natural language legal queries from users and supports uploading legal documents, case materials, etc., so that the system can identify the user's intent and legal context, and initialize the MCP (Multi-Channel Programming) context.

[0022] Document status;

[0023] The legal question answering module uses a layered, trained large language model to plan reasoning paths using the Monte Carlo Tree Search (MCTS) algorithm, integrates Retrieval Enhancement Generation (RAG) mechanisms to locate legal information, and continuously tracks the state and optimizes the path under the support of the MCP protocol. Ultimately, it generates a legal question answering module that matches the query question and provides relevant information.

[0024] The solution is supported by legal force and has transparent reasoning;

[0025] The question-and-answer result output module is based on the unified structure of the MCP protocol and outputs legal answers, visual reasoning chains, and citations.

[0026] Based on legal provisions and system confidence scores, the verifiability and traceability of the output content are ensured. Furthermore, the legal question-answering module includes:

[0027] Model hierarchical training module: used to build a two-stage training framework that supports the MCP protocol. The first stage completes the basic legal task capability training, and the second stage performs fine-tuning and optimization for multi-jurisdictional preference and legal conflict tasks, and introduces DPO and path consistency supervision.

[0028] Decision reasoning module: Based on user intent and current MCP state, it uses MCTS strategy for path planning, and combines MCP history records to constrain the consistency of reasoning path, thus realizing a closed loop of causal chain;

[0029] Search enhancement module: integrates keyword retrieval, semantic similarity, MCP path context, and supervised re-ranking strategy.

[0030] (Omitted) This allows for personalized and precise positioning of legal paragraphs.

[0031] Legal Answer Generation Module: Utilizing the unified state context, legal basis, and reasoning trajectory managed by MCP, this module drives the large language model to generate professional legal answers.

[0032] Furthermore, the model hierarchical training module includes the following sub-modules:

[0033] in:

[0034] Open-source data acquisition module: Systematically collects legal data such as legal provisions, judicial interpretations, and judgments, and

[0035] Construct a multi-layer training set by combining general instruction data;

[0036] Basic Legal Task Alignment Module: Structures general legal data into task sets, using fused knowledge to maintain and...

[0037] The loss function of the MCP consistency regularization term is used to train the model to master the basic abilities of legal language expression, structural analysis and reasoning.

[0038] Legal Task Alignment Module: Based on the Self-Instruct strategy, SFT and human preference datasets are constructed, focusing on complex tasks such as multi-jurisdictional conflicts, provision adaptation, and reference tracking. This enhances the model's ability to understand and judge legal context, and achieves context consistency and path memory mechanism during training through the MCP protocol.

[0039] In the first level of training, to ensure that the language comprehension and contextual consistency expression abilities for basic legal tasks are effectively mastered, this invention introduces a joint loss function that integrates a knowledge preservation mechanism and MCP contextual protocol consistency constraints, the expression of which is:

[0040]

[0041] in:

[0042] ·L base This represents the total loss value during the basic training phase;

[0043] ·x t This represents the t-th word in the input legal text sequence;

[0044] .θ represents the model parameters;

[0045] ·L KP To maintain regularity for knowledge, the model's ability to memorize core legal knowledge is constrained;

[0046] ●L MCPThis is an MCP protocol consistency item used to measure the degree of consistency between the model-generated content and the context state and path of the MCP record;

[0047] ·λ1,λ2 are hyperparameters that adjust the contribution of each loss term.

[0048] In the second-level training, the system employs a dual strategy of Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) to improve the model's performance and path rationality under specific legal preference scenarios. The SFT stage continues to use the aforementioned joint loss function, while the DPO stage introduces a policy model and a reference model, and strengthens the MCP path consistency constraint. Its loss function expression is as follows:

[0049]

[0050] in:

[0051] ·L DPO+ Indicates the final loss;

[0052] ●(x,y w ,y l () indicates the better and worse responses in the sample input and preference data;

[0053] ·π θ These are the output distributions of the current policy model and the reference model, respectively;

[0054] .D consistency For path logic consistency items;

[0055] .C confidence Answer the confidence constraint terms for the model;

[0056] .R MCP This is a path consistency regularization term for MCP, used to ensure that the optimized answer remains consistent with the reasoning history recorded by MCP in terms of logical structure and state transitions.

[0057] •β1,β2,β3,β4 are the weighting coefficients of the loss term.

[0058] Through this structured hierarchical training mechanism, while mastering basic language and logic capabilities, the model can establish a stable and clear legal reasoning chain under the guidance of the MCP protocol, achieving effective synergy between general capabilities and legal task adaptability.

[0059] Furthermore, the decision reasoning module, supported by the MCP protocol, includes a state space definition module, an action definition module, and a Monte Carlo tree search module. These modules are used to construct, maintain, and optimize the entire multi-step reasoning path, and dynamically update the state, knowledge, and path records in the MCP in a structured manner. Wherein:

[0060] The state space definition module is used to abstract legal issues into an initial state, combine user intent and context, dynamically generate intermediate states through legal retrieval, context completion, fact confirmation, etc., and finally enter the termination state marked by MCP (including completing the solution or reaching the reasoning saturation).

[0061] The action definition module is used to define structured reasoning actions, including semantic retrieval of the legal database, adjustment of regulatory priorities, generation of user follow-up questions, context expansion, and clarification of regulatory citations. The selection and execution of these actions are all controlled by the MCP protocol, with path recording and state maintenance.

[0062] The Monte Carlo tree search module is used to generate dynamic paths by combining state and action execution. It generates multiple candidate paths through a simulated search mechanism and uses information from the MCP feedback to evaluate the consistency and value of each path.

[0063] The execution process of this MCTS module includes the following steps:

[0064] 1. Selection Phase: Starting from the root node, high-value nodes are selected based on improved strategies such as PUCT and historical MCP states.

[0065] 2. Expansion Phase: Invoke LLM to generate new states and add them to the inference tree, and update the MCP records;

[0066] 3. Simulation Phase: Simulate the future reasoning process under the guidance of the strategy to generate a multi-path structure;

[0067] 4. Feedback Phase: The logical information and value signals from the simulation process are fed back to the upper-level node and the statistics are updated;

[0068] 5. Strategy Decision Stage: Based on the consistency constraint between path scores and MCP records, the optimal path is selected as the final reasoning guide.

[0069] By leveraging the MCP protocol throughout the inference structure, the system can not only improve the logical integrity of path selection, but also achieve visualization, traceability, and auditability of inference behavior through structured path control, thereby enhancing the professionalism and credibility of the model output.

[0070] Furthermore, under the unified scheduling of the MCP protocol, the retrieval enhancement module integrates a Retrieval-Augmented Generation (RAG) mechanism to quickly, accurately, and controllably locate legal passages related to the user's legal query, while ensuring the integrity of the citation chain, consistency of context, and traceability of path information. Its execution process includes the following steps:

[0071] Preliminary retrieval: The MCP provides the current reasoning state and context labels, guiding the system to perform preliminary retrieval in the underlying legal database based on the keywords, entity labels and task intent input by the user, and to calculate the similarity between the user query vector and the legal paragraph vector based on the semantic vector model to obtain a list of candidate paragraphs;

[0072] Multi-strategy recall: Combining BM25, Dense Embedding similarity, and rule constraint matching, multiple strategies are used to perform parallel recall operations. The MCP protocol is used to record the context fit and historical hit rate of each strategy result, which serves as an important reference for subsequent re-ranking.

[0073] Unified reordering: By mapping the structure between the MCP path context and candidate regulatory paragraphs, the multi-strategy recall results are embedded in a unified vector space. Then, a reordering model that integrates legal semantics, entity structure and MCP context consistency is used to prioritize the results and optimize the logical fit and reasoning interpretability of the regulatory matching results.

[0074] Citation extraction: The MCP calls the citation extraction tool to identify information such as article number jumps and cross-references in the sorted paragraphs, trace back the source structure of citations, ensure the consistency of the citation chain with the legal context, and update the citation mapping index in the MCP.

[0075] Global Information Refinement: For lengthy legal content, the information refinement module performs noise reduction processing by leveraging the task context and reasoning objectives managed by the MCP protocol, retaining only key clauses or sentences that contribute to reasoning, filtering out invalid background information, and forming a set of legal materials with greater decision-making value.

[0076] Reasoning chain guided filtering: The large language model constructs a causal chain sketch based on the current reasoning path recorded by the MCP, further identifies the subset of laws and regulations most relevant to the target answer, extracts key nodes of the path from candidate laws and regulations, and forms a refined combination of law and regulation segments for the final input model to support the output of explanatory legal question answering.

[0077] In addition, the similarity calculation in the initial search uses the following fusion expression:

[0078] Similarity final (q,d)=α·CosineSimilarity(q,d)+(1-α)·S(q,d;MCP)

[0079] in:

[0080] ·q represents the semantic vector of a user's legal query after processing by the Embedding model;

[0081] ●d represents the embedded representation of candidate paragraphs in the regulatory database;

[0082] · Indicates traditional cosine similarity;

[0083] S(q,d;MCP) represents a structural semantic scoring function that combines the legal context, citation path, and historical reasoning state recorded in the MCP protocol.

[0084] α is the fusion weight hyperparameter, which controls the trade-off between semantic similarity and contextual structure adaptation.

[0085] The aforementioned similarity score is not only used for initial screening of text retrieval, but also achieves personalized reordering and citation chain consistency control through the dynamic participation of MCP, significantly improving the accuracy of legal provision matching, contextual integrity, and the professionalism and credibility of the final model generation.

[0086] According to another aspect of the present invention, a legal question-answering method based on a multi-path legal reasoning engine using the MCP protocol is proposed. This method is characterized by structured legal reasoning and full-chain interpretability, and the MCP protocol uniformly coordinates model state, knowledge activation, and path information, maintaining logical consistency and contextual coherence throughout the training, reasoning, and output processes. The method specifically includes the following steps:

[0087] 1. Question Input Stage: Users input legal query questions in natural language and can choose to upload relevant legal documents, case background, or factual materials; the system uses MCP to initialize the user's context state and uses the intent recognition and tag extraction module to parse the query target and semantic structure, providing contextual anchors for subsequent reasoning.

[0088] 2. Intelligent Solution Phase: The system invokes a large language model embedded and optimized using the MCP protocol. Combining Monte Carlo Tree Search (MCTS) and a multi-path inference mechanism, it progressively expands, filters, and optimizes reasoning paths under MCP control. Simultaneously, it utilizes the RAG retrieval enhancement mechanism to accurately locate and match semantically and structurally appropriate legal paragraphs and citation chains to the query target. MCP dynamically records the reasoning state, path evolution, and knowledge activation, using this information to filter logical contradictions or path redundancy, achieving a structurally sound and causally clear interpretable reasoning process.

[0089] 3. Output Stage: With the support of state and path management provided by the MCP protocol, the system outputs structured legal solutions, including the final answer, corresponding provisions, sources of citation, reasoning chain, citation relationship diagram, and credibility score. All output information is based on the context tracking and reasoning trajectory preserved in MCP, supporting transparent reasoning, visualized paths, and traceable citations, significantly enhancing user understanding and trust.

[0090] This invention has the following technical effects and practical value:

[0091] 1) This invention introduces a unified MCP protocol into the intelligent legal question-answering framework for the first time, comprehensively integrating model training, state control, knowledge retrieval, and path output, effectively improving the system's ability to model legal contexts and its logical consistency performance in complex tasks. The dual-level hierarchical training mechanism, through basic capability alignment and preference reinforcement learning (including Self-Instruct and DPO+ path consistency optimization), simultaneously ensures the stability and flexibility of semantic learning and inference chain construction within the MCP framework.

[0092] 2) By embedding Monte Carlo Tree Search (MCTS) and multi-path reasoning mechanism into the MCP protocol framework, the system has the ability to perform multi-step legal logic reasoning for complex queries. It can adjust the strategy path in real time, complete the context information, select the optimal logic chain, and rely on MCP to perform path consistency evaluation and historical state comparison to achieve structural transparency of the reasoning chain and state controllability of reasoning behavior.

[0093] 3) The RAG module constructed in this invention, based on the historical knowledge activation trajectory and legal citation map provided by MCP, can accurately extract the citation relationships, hierarchical citation mapping structure, and applicable boundaries from legal paragraphs, and dynamically filter high-weight information blocks under the guidance of reasoning. MCP records all knowledge call behaviors and path switching decisions during this process, thereby improving the system's ability to model and interpret highly complex semantic relationships. It demonstrates strong professionalism, interpretability, and reliability in scenarios such as judicial consultation, legal review, and case reasoning. Attached Figure Description

[0094] To more clearly illustrate the technical solution of the present invention, the accompanying drawings provide schematic illustrations of relevant modules and processes in the embodiments. Obviously, the following drawings are merely illustrative of one or more feasible embodiments of the present invention, and those skilled in the art can make equivalent substitutions without creative effort.

[0095] Figure 1 : Overall engine framework diagram of the present invention;

[0096] Figure 2 : Structure and workflow diagram of the retrieval enhancement module integrating the MCP protocol;

[0097] Figure 3 Flowchart of an intelligent legal question-answering method based on MCP;

[0098] Figure 4 Flowchart of MCTS inference path planning and feedback update integrating MCP. Detailed Implementation

[0099] To further illustrate the technical solution of the present invention, several specific embodiments are provided in conjunction with the accompanying drawings to assist in explaining the system architecture and method flow. Those skilled in the art can implement equivalent substitutions and modifications based on the drawings and description without inventive effort.

[0100] Terminology Definition

[0101] • LLM (Large Language Model): Refers to a pre-trained language model based on the Transformer architecture with hundreds of millions of parameters, possessing natural language generation, reasoning, and understanding capabilities. In this invention, LLM refers to a legal language model trained in multiple stages and layers and integrated with the MCP protocol.

[0102] • Monte Carlo Tree Search (MCTS): A search algorithm that follows a "select-expand-simulate-return" process.

[0103] The four-stage state space exploration is the core mechanism for the system to plan legal logic paths.

[0104] • RAG (Retrieval-Augmented Generation): A retrieval-enhanced generation technology that integrates document retrieval and language generation modules to effectively improve the factual consistency and legal citation accuracy of the model's answers.

[0105] .MCP (Model Context Protocol): The Model Context Protocol is the key mechanism proposed in this invention. It runs through the model training, inference and output stages and is used to uniformly manage model state, knowledge activation, inference path and context information. It is the core supporting component for realizing inference interpretability, state consistency and output transparency.

[0106] According to one embodiment of the present invention, such as Figure 1 and Figure 2 As shown, a multi-path legal reasoning engine based on the MCP protocol is provided, including a legal question input module, a legal question answering module, and a question-and-answer result output module.

[0107] in:

[0108] The legal question input module is used to receive legal queries input by users in natural language, and supports uploading legal provisions, contract documents, factual materials and other information. The system initializes the user context state based on MCP for subsequent state management and path scheduling.

[0109] The legal question answering module is the core of the system, calling upon a large language model trained and optimized using the MCP protocol, combined with...

[0110] The MCTS algorithm constructs a multi-step reasoning path, integrates the RAG mechanism to locate legal information, and dynamically adjusts the reasoning strategy under the support of the MCP unified context structure to generate professional and credible legal answers.

[0111] With the support of MCP path and status information, the question-and-answer result output module outputs structured legal answers, corresponding provisions, reasoning chain diagrams, citation paths, and credibility scores, achieving full-process transparency, visualization, and contextual consistency.

[0112] In the legal question answering module, the model hierarchical training module constructs a two-stage training system based on the MCP protocol to achieve decoupling and alignment between general language ability and legal professional ability.

[0113] Its interior includes:

[0114] • Open source data acquisition module: systematically collects legal corpora such as laws and regulations, judgment documents, and judicial interpretations, and combines them with general task data for training preparation;

[0115] • Basic Legal Task Alignment Module (L0): This module trains the model's basic understanding, logical expression, and language generation capabilities using a structured task set. During training, a joint loss function combining knowledge preservation and MCP context consistency terms is introduced to ensure the stability of state transitions and contextual coherence during training. The loss function takes the following form:

[0116]

[0117] Among them, L KP L represents the knowledge retention term. MCP This represents the MCP consistency regularization term.

[0118] • Legal Task Alignment Module (L1): Employs a Self-Instruct strategy to construct the SFT dataset and preference dataset, refining the model's training for legal tasks. It focuses on adapting the model to tasks such as regulation adaptation, multi-jurisdictional conflict, provision citation, and causal chain reasoning. It uses the MCP-enhanced DPO+ loss function.

[0119]

[0120] Among them, R MCP The inference path and state consistency regularization term is used to supervise whether the generated model results are consistent with the historical states and paths of the MCP.

[0121] This training phase incorporates scenarios such as legal instruction tasks, judicial reasoning, regulatory adaptation, and case analysis into a fine-tuning system. After training, the model is evaluated through multiple legal benchmarks such as LawBench, LexGLUE, and LexEval to verify its professionalism and robustness.

[0122] Guided by the MCP framework, the decision-making reasoning module employs the MCTS algorithm to complete multi-step path planning for solving legal problems. The system uses the state information in the MCP as the foundation of the MCTS state space, and dynamically defines the initial, intermediate, and final states by combining user intent, knowledge activation status, and historical reasoning paths. The action set includes semantic retrieval, clause tracing, user supplementation, and context expansion, among other behaviors. The state-sharing mechanism provided by the MCP supports information synchronization between modules.

[0123] In practice, MCTS includes the following four stages:

[0124] 1) Selection phase: Starting from the current state node defined by MCP, the potential optimal child node is selected by combining the access frequency and the MCP path consistency score;

[0125] 2) Expansion Phase: Generate unexplored actions based on LLM, expand intermediate state nodes, and record the new states in MCP;

[0126] 3) Simulation phase: Multiple rounds of simulation are conducted based on the strategy-guided path, leading to the termination state;

[0127] 4) Feedback phase: Based on path performance and result credibility, the value is fed back and the state score and path weight in MCP are updated to improve the quality of subsequent reasoning behavior.

[0128] In summary, the specific implementation of this invention, through the MCP protocol, uniformly manages the context state, reasoning path, and knowledge call throughout the entire process of model training, inference execution, and result output. This achieves structural transparency of the reasoning chain, controllable state of decision-making behavior, and consistent context of result output, providing high-accuracy, high-interpretability, and high-reliability system support for legal intelligent question answering systems.

[0129] Specifically, it includes the following sub-modules and inference strategies. The state management, inference path tracking, and knowledge retrieval of all modules are uniformly managed by the Model Context Protocol (MCP), achieving end-to-end state traceability and path transparency. State Space Definition Module:

[0130] • Initial state: The initial state consists of the natural language questions raised by the user and the background information such as the legal texts and case facts uploaded by the user. The system initializes and records this state through MCP.

[0131] • Intermediate state: The intermediate state represents the state update node formed after the model performs actions such as legal retrieval, user-guided questioning, or context expansion during the reasoning process. The MCP protocol records the state evolution path and the corresponding context switch.

[0132] Termination State: The termination state indicates that the system has collected sufficient legal evidence or reached the preset reasoning depth threshold in the current path. MCP locks the path structure here and triggers the state freeze mechanism for result feedback and interpretability output.

[0133] Action definition module:

[0134] With the support of MCP, the system defines the following scheduled inference actions for MCTS to dynamically select and combine for execution:

[0135] 1. Legal database semantic retrieval: Based on the context tags and knowledge activation status stored in MCP, the semantic matching mechanism and legal structure parser are invoked to accurately extract the articles, applicable requirements and interpretation opinions most relevant to the current status;

[0136] 2. Online-assisted retrieval: Combining user context and MCP preference paths, publicly available legal resources are retrieved to supplement external knowledge fragments such as cases, industry standards, and guidelines;

[0137] 3. User Question Guidance: When MCP determines that there is semantic ambiguity or missing key elements in the current path, it proactively triggers a follow-up question strategy to help users clarify the boundaries of the problem and the scope of the legal jurisdiction;

[0138] 4. Contextual Expansion: LLM is used to generate associative extended content based on the MCP path, covering relevant legal background, historical cases or similar reasoning structures, in order to improve contextual coherence and reasoning coverage.

[0139] Monte Carlo Tree Search Module:

[0140] This module constructs a dynamic asymmetric search tree with the support of the MCP protocol and optimizes the inference path through multi-round interactive simulation. It includes the following steps:

[0141] 1. Selection: Starting from the root node, based on the improved PUCT strategy, the node's MCP context weight, number of visits, and path consistency score are comprehensively considered to select the most promising node in the tree;

[0142] 2. Expansion: If the current node does not meet the termination condition, the LLM executes the action that has not yet been tried in the MCP state space to generate a new state node. This state and path information are written to the MCP in real time to update the inference trajectory.

[0143] 3. Simulation: Perform simulated reasoning along the new extended path, using the policy network and historical behavior trajectories recorded in the MCP to generate multiple steps until the termination state is reached;

[0144] 4. Backpropagation: The result score and inference path transparency index of the termination node are propagated upward from the leaf node, the statistical values ​​of all relevant nodes in the search tree are updated, and the data is synchronously written to the MCP path database.

[0145] 5. Action Decision: Based on the cumulative MCP path score and state credibility of all candidate child nodes,

[0146] Selecting the optimal node as the next execution path improves the rationality of the decision and the causal consistency of the final answer. It is worth noting that the "Legal Database Semantic Retrieval" action in this module is the core part of the RAG (Retrieval Enhanced Generation) strategy under the constraints of MCP path integration. Under the scheduling and management of MCP, this action not only accurately matches the provisions but also adaptively adjusts the breadth and depth of the retrieval based on contextual information. This scheduling mechanism ensures that the execution timing and direction of the RAG module in the path branches always remain consistent with historical states and preferred paths, effectively reducing semantic drift and provision citation deviations.

[0147] The retrieval enhancement module, guided by MCP, undertakes the tasks of recalling regulatory content, restoring the structure of provisions, and refining context. This module achieves content focusing and information integration through the following steps:

[0148] 1. In the initial recall phase, the system performs multi-strategy retrieval under the knowledge tags and jurisdictional limitations provided by MCP.

[0149] (Including BM25, semantic vector similarity, rule triggers, etc.);

[0150] 2. The unified reordering stage maps candidate paragraphs to the same semantic space and integrates information such as user preferences and historical call frequency recorded in MCP for personalized reordering;

[0151] 3. The citation relationship extraction stage identifies the hierarchical citation relationships and cross-article jump structures between legal provisions, and traces the integrity of the citation chain through the MCP path recorder;

[0152] 4. Context restoration and information refinement stage: Based on the MCP path construction and causal chain logic, key information nodes are filtered out from long texts and noise is eliminated;

[0153] 5. Finally, the RAG module inputs the reordered and refined legal paragraphs along with the user's questions into the LLM. Under the management of the MCP protocol, it generates legal answers that conform to legal logic, are consistent in citation, and are clearly expressed, ensuring the interpretability and professionalism of the system in handling complex legal issues.

[0154] In legal question-and-answer tasks, many scenarios involve efficient retrieval and content refinement of extremely long legal texts with complex contexts. To address this, this embodiment introduces a legal retrieval enhancement module based on the RAG (Retrieval Enhancement Generation) mechanism, coupled with a structured-trained LLM and a unified MCP throughout the training and reasoning process, to achieve accurate legal provision location and authoritative answer generation.

[0155] The system's underlying retrieval capabilities are based on a regulatory index database built on Elasticsearch (ES), supporting full-text search, structured field matching, and vectorized semantic queries. The system uses MCP to uniformly manage user intent, context state, and search history. When constructing queries, it integrates vector embedding and rule-based retrieval logic. In the initial retrieval stage, it integrates keyword Boolean search and semantic vector matching. The similarity calculation method between the user query vector q and the regulatory text vector d is as follows:

[0156] In the initial retrieval phase, the system performs Boolean logic retrieval based on keywords in the user's query and merges the text query results with vector retrieval results based on embedding representation. The basic similarity between the user query vector q and the regulatory text vector d is calculated using the following formula:

[0157] Similarity final (q,d)=α·CosineSimilarity(q,d)+(1-α)·S(q,d;MCP)

[0158] To mitigate the discrepancies in scoring metrics among different recall strategies, this invention integrates historical preferences, contextual relevance, and path scoring information through MCP (Multi-Channel Propagation), uniformly guiding the ranking model to aggregate and rank multi-source recall results, ensuring the diversity and relevance of legal texts. Addressing the "context truncation" and "information drift" problems frequently encountered by LLM (Limited Language Management) when processing extremely long legal documents, this invention proposes a three-stage structured retrieval enhancement strategy driven by the MCP protocol:

[0159] 1. Citation Relationship Extraction Stage: With the help of the citation context and paragraph jump logic recorded by MCP, the regulatory citation chain is extracted from the reordered text and mapped back to the original document context to achieve a complete restoration of the regulatory network.

[0160] 2. Global Information Filtering Stage: The system combines MCP path context and reference chain information to denoise content with high semantic density but high logical risk, retains traceable and verifiable legal provisions, and optimizes the quality of the context used for generation through the semantic compression module.

[0161] 3. Reasoning Chain Guidance Phase: Based on the user's question and the path state diagram maintained by the MCP protocol, the LLM's chain-of-thought mechanism is triggered to dynamically select clauses that directly support the answer to the question.

[0162] Eliminate redundant content and enhance the focus and causal consistency of the output.

[0163] The processed content, along with the original user query, is input into the large language model. The MCP protocol coordinates the loading of historical states, knowledge activation records, and reference path information, guiding the model to perform structured fusion reasoning and generate professional answers with clear textual support, legal logic chains, and authoritative citations.

[0164] The legal answer generation module, based on the input control and status labeling functions provided by MCP, can automatically load the cited evidence, the position of the reasoning chain, and the user intent markers during the answer generation process, and finally output answer content with a clear causal structure.

[0165] The question-and-answer result output module uses the knowledge flow and path tracking information managed by MCP to display the final answer in a graphical form, including legal answer summary, links to cited provisions, reasoning chain diagram, typical case annotations and system confidence score, effectively improving the interpretability of the results and user trust.

[0166] According to another embodiment of the present invention, such as Figure 3 and Figure 4 As shown, a multi-path legal reasoning engine-based legal question-answering method based on the MCP protocol is provided, including the following steps:

[0167] ●S1: Question Input and Regulation Upload. Users input natural language questions and upload relevant regulatory documents. The system initializes the context state and path management structure based on MCP.

[0168] .S2: System Reasoning and Question Answering Generation. The system uses MCP to construct a reasoning tree and combines it with MCTS for dynamic selection and retrieval, knowledge activation, context expansion, and feedback collection to gradually build multi-path answer candidates. All path status information and selection processes are included in MCP tracking.

[0169] • S3: Result Generation and Visualization Output. The system selects the optimal scoring path according to the MCP protocol, schedules a structured generation process, outputs professional answer text, and displays the core conclusions, applicable provisions, causal paths, and confidence scores in a graphic and textual structure through the output module.

[0170] In summary, the intelligent legal question-answering system and method proposed in this invention innovatively introduces the Model Context Protocol (MCP) to achieve contextual consistency management, path transparency, and information traceability throughout the entire process from training to inference to output, thereby comprehensively improving the system's accuracy, interpretability, and credibility in professional legal task scenarios.

Claims

1. A multi-path legal reasoning engine based on MCP protocol, characterized in that, The engine comprises: a legal problem input module for receiving a user-input legal query and uploading relevant legal norm files; a legal problem solving module based on an LLM, combining Monte Carlo tree search, multi-path reasoning, causal consistency judgment, retrieval enhancement mechanism and user feedback mechanism for legal logic reasoning, relying on the MCP unified management model to manage internal state, knowledge activation and reasoning path information, to ensure the context consistency, information persistence and path transparency of the reasoning process, thereby generating an interpretable legal answer corresponding to the user-input problem; a question and answer result output module that outputs legal answers, reasoning chains, cited basis and credibility assessment in a visual manner by integrating the reasoning path, state information and knowledge source of the MCP, to enhance user understanding and trust; The legal problem solving module comprises: a model training framework that embeds the MCP protocol in the training whole process to form a unified context management mechanism, including label awareness and hierarchical training, structured instruction templates, knowledge preservation mechanisms, legal conflict simulators and direct preference optimization strategies, for improving the generalization and refinement capabilities of the model under different legal tasks; a reasoning engine based on the reasoning process management of the MCP protocol, covering large language models, core reasoning mechanisms, multi-path reasoning modules, causal consistency judgment modules, retrieval enhancement modules and user feedback optimization modules; MCP protocol components that run through the stages of model training and reasoning execution, for unified management of state transfer, knowledge activation, path tracking and context sharing; The model training framework comprises: a label awareness and hierarchical pre-training module; introducing structured instruction templates and knowledge preservation mechanisms in the basic legal task training stage, combined with the MCP to manage the consistency of the training context; the professional legal task training stage combines the legal conflict simulator and preference optimization strategy under the MCP protocol to achieve deep fine-tuning for complex legal contexts; The basic legal task training stage uses a joint loss function that integrates knowledge preservation items and MCP context consistency constraints, and its expression is: ; wherein: : the tth word in the original legal text sequence; : model parameters; : knowledge preserving regularizer : Context agreement item, measures the difference in agreement of the prediction result with the MCP historical state; : a hyperparameter that regulates the loss term weight; The professional legal task training stage uses an improved DPO loss function that integrates the MCP path management mechanism, and its expression is: ; : input question; : better and worse answers, respectively; : current policy model; : reference model; : sigmoid function; : weight hyperparameters; : Inference Path Consistency Item; : generating an answer confidence term; : Inference path consistency regular term based on MCP protocol, for constraining the consistency of answer logic and history path record.

2. The engine of claim 1, wherein, The MCP component in the reasoning engine supports the following functions: Reasoning state tracking: recording the state of each step of the model decision in time sequence; Knowledge activation mechanism: dynamically invoking structured legal knowledge according to the current context; Path recording and scoring: scoring each candidate path for structural consistency and interpretability; State sharing mechanism: to realize information sharing between different modules in the reasoning process.

3. The engine of claim 2, wherein, The retrieval enhancement mechanism comprises: a legal entity semantic correction module; a legal conflict early warning module; a multi-strategy recall and large language model supervised reordering module that combines the user context and regulation preferences stored in the MCP to achieve personalized sorting.

4. The engine of claim 3, wherein, The multi-strategy recall and large language model supervised reordering module performs similarity score calculation in the preliminary retrieval step, which is calculated as follows: ; : vector representation of user query question; : vector representation of a regulation paragraph; : cosine similarity; : fusion proportion hyperparameter, controls the relative weight of the two parts; : Weighted score function that takes into account legal semantics, regulation context and MCP path weights.

5. The engine of any one of claims 1 to 4, wherein, The question and answer result output module is based on the MCP protocol to manage the structured reasoning track, and the complete decision logic and reliability are visualized in the form of a causal chain diagram, reference basis, knowledge source annotation and state evolution diagram.

6. A multi-path legal reasoning legal Q&A method based on the MCP protocol, implemented based on the multi-path legal reasoning engine based on the MCP protocol in claim 1, characterized in that, It comprises the following steps:

1. Receiving a legal query and extracting key tags and entities, initializing the MCP context state; 2. Building an input that fuses the MCP historical state and the user Prompt, calling a large language model for reasoning; 3. The MCP protocol dynamically maintains the reasoning state, knowledge calling record and path structure; 4. Based on the MCP context, low confidence paths are filtered by combining multi-path simulation and causal consistency evaluation mechanism; 5. With the help of the retrieval enhancement module, call the regulations, and fuse the user context for semantic reordering; 6. The question and answer result output module provides structured information flow based on MCP for visual display, including reasoning chain, reference basis, legal answer and reliability score.

Citation Information

Patent Citations

  • All-weather RAG intelligent agent automatic newspaper design method and device and electronic equipment

    CN119476247A

  • Multi-round knowledge-guided question and answer method and system fusing large language model and knowledge graph

    CN120069068A