A Case Retrieval Method and System Based on Multi-Agent Collaboration
By employing a multi-agent collaborative case retrieval method, and utilizing multi-turn interactions between case library receptionist and administrator agents and large language model technology, the lack of connection between the user's natural language expression and the structured needs of case retrieval in the existing system is resolved, thereby achieving dynamic clarification of user intent and improving the accuracy of case recommendations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI JIAOTONG UNIV
- Filing Date
- 2026-03-05
- Publication Date
- 2026-06-30
AI Technical Summary
Existing legal case retrieval systems lack the ability to perform in-depth semantic analysis and multi-round interactive clarification when users ask questions directly in natural language. They are unable to dynamically adapt to users' ambiguous intentions, resulting in insufficient accuracy and coverage of search results, and they ignore user preferences and result diversity.
A multi-agent collaborative case retrieval method is adopted. Through multiple rounds of interaction between the case database receptionist and administrator agents, the method can clarify users' legal questions, understand semantics, analyze case facts, extract case attributes, and predict the cause of action. Combined with Large Language Model (LLM) technology, the retrieval process is dynamically adjusted to improve accuracy.
It improves the accuracy of legal case recommendations and user experience, lowers the professional threshold for users, and enhances the applicability of the system in complex legal issues.
Smart Images

Figure CN122309653A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of legal information retrieval technology, specifically relating to a case retrieval method and system based on multi-agent collaboration. Background Technology
[0002] With the rapid development of legal informatization and judicial data disclosure, the scale of judicial case databases continues to expand. Traditional legal case retrieval often relies on keyword matching or Boolean logic queries, requiring users to have strong legal expertise to construct effective search queries, which poses a high barrier to entry for non-professional users or grassroots legal workers.
[0003] To enhance the intelligence of search capabilities, existing technologies have proposed various solutions. One approach focuses on deep semantic understanding, employing pre-trained language models (such as BERT and Lawformer) to encode the full text of cases and making recommendations based on the similarity of semantic vectors. Another approach focuses on the identification and matching of legal elements, using natural language processing techniques to extract key elements such as legal entities, points of contention, and causes of action from case texts, and then performing searches based on the matching degree of these elements. Furthermore, some solutions attempt to construct legal knowledge graphs or employ multi-task, multimodal learning to integrate various information such as text and event sequences to improve the accuracy and interpretability of recommendations.
[0004] However, these existing methods typically assume that the input is relatively standardized, structured case text (such as a complete case description). In practical applications, when users (especially non-experts) ask questions directly in natural language, existing systems generally lack the ability to perform in-depth legal semantic analysis and multi-round interactive clarification of the questions, making it difficult to dynamically adapt to the user's ambiguous intentions. This results in search results that fail to meet the requirements in terms of accuracy and coverage. Furthermore, existing systems often neglect the balance between user preferences and result diversity, limiting their effectiveness in practical applications.
[0005] Therefore, there is an urgent need for a legal case retrieval system that is user-centric, enables multi-round interaction to clarify issues, analyze cases, extract case attributes and predict causes of action, and combines intelligent retrieval and recommendation. Summary of the Invention
[0006] This invention addresses the lack of effective connection between existing legal case retrieval technologies and the structured needs of users' natural language expressions, especially when users' legal questions are incomplete or unclear. Existing technologies struggle to continuously clarify and dynamically model users' true search intentions, leading to insufficient matching, singular results, or results deviating from users' actual needs. This invention proposes a case retrieval method and system based on multi-agent collaboration. This system utilizes natural language processing and Large Language Model (LLM) technology, combined with multi-round interactions between case database receptionist and administrator agents, to achieve clarification of users' legal questions, semantic understanding, case analysis, case attribute extraction and cause-of-fact prediction, and case retrieval and recommendation, thereby improving the accuracy of legal case recommendations and user experience.
[0007] To achieve the above objectives, the technical solution of the present invention includes the following:
[0008] A case retrieval method based on multi-agent collaboration, the method comprising: Standardized questions that meet the evaluation criteria are obtained based on the user's original question; wherein, the user's original question includes the natural language question content entered by the user at the current interaction time step and the historical interaction information corresponding to the natural language question content; By combining the case database, we can deduce the selected cases corresponding to this standardized problem; Calculate the matching score between the user's original question and the featured case, and output the featured case when the matching score reaches a set score threshold.
[0009] Furthermore, standardized questions that meet the evaluation criteria are obtained based on the user's original question, including: Obtain the original question input by the user; The original questions and clarification texts were evaluated according to evaluation criteria, which included: information completeness, coverage of legal elements, and logical consistency. For original questions and clarification texts that do not meet the evaluation criteria, a large language model is used to analyze them and generate rhetorical questions to obtain user feedback and clarification texts for evaluation. This process continues until the original questions and clarification texts meet the evaluation criteria, resulting in the final question description. The final problem description is semantically parsed using a large language model to generate a standardized problem.
[0010] Furthermore, by combining the case database, we can deduce the selected cases corresponding to this standardized question, including: By combining the case database, we can deduce the candidate cases corresponding to this standardized problem; The quantitative score between each candidate case and the original question is calculated based on a large language model. The quantitative scoring factors include: factual context, legal relationship and focus of dispute, legal subjects and liability structure, and applicable legal provisions and legal logic. Based on the quantitative scoring results, selected cases were chosen from the candidate cases.
[0011] Furthermore, the case database includes: a case text database; The process of reasoning about candidate cases corresponding to the standardized problem using a case database includes: Based on the standardized question, a search is performed in the case text database to obtain the candidate cases corresponding to the standardized question.
[0012] Furthermore, the case database includes: a case text database and a case vector database; The process of reasoning about candidate cases corresponding to the standardized problem using a case database includes: Based on the large language model, the case analysis of the standardization problem is carried out, and Boolean logic retrieval expression is generated based on the elements obtained from the case analysis. Use large language models to extract case attributes from standardized problems; Based on a pre-built database of case types and causes of action, and combined with a large language model, cause of action identification is performed on standardized issues; Based on standardized questions, case attributes, and causes of action, a search is conducted in the case text database to obtain the first set of candidate cases; Based on Boolean logic search terms, case attributes, and causes of action, a search is performed in the case vector database to obtain a second set of candidate cases. Merge the first candidate case set and the second candidate case set to obtain the candidate cases corresponding to the standardized problem.
[0013] Furthermore, after calculating the matching score between the user's original question and the selected case, the method further includes: If the matching score is less than a set score threshold, obtain the problem understanding result generated during the reasoning process of the standardized problem, and evaluate the problem understanding result; Based on the evaluation results, the process of obtaining standardized questions that meet the evaluation criteria based on the user's original question is re-executed, or the process of reasoning about the selected cases corresponding to the standardized question based on the case database is re-executed, until the matching score of the selected cases reaches the set score threshold.
[0014] Furthermore, after outputting the selected case, the method also includes: Build a curated collection of case studies; Legal scenario analysis is performed on the selected case set. When the analysis results indicate that the selected case set involves multiple legal scenario preferences, a scenario clarification question is issued to the user based on a large language model. The large language model analyzes the user's response to the clarification question in the context in order to obtain specific selected cases from the selected case set; wherein, the specific selected cases are selected cases involving the legal context of the user's original question.
[0015] A case retrieval system based on multi-agent collaboration, the system comprising: The Case Library Receptionist Intelligent Agent module is used to obtain standardized questions that meet the evaluation criteria based on the user's original questions. The Case Library Administrator intelligent agent module is used to infer the selected cases corresponding to the standardized question by combining the case database; calculate the matching score between the user's original question and the selected cases; and output the selected cases when the matching score reaches a set score threshold.
[0016] A computer device, characterized in that the computer device comprises: a processor and a memory storing computer program instructions; the processor, when executing the computer program instructions, implements the case retrieval method based on multi-agent cooperation as described above.
[0017] A computer-readable storage medium, characterized in that the computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the case retrieval method based on multi-agent cooperation as described above.
[0018] Compared with the prior art, the present invention has at least the following beneficial effects.
[0019] (1) By constructing a multi-agent collaborative mechanism with clear division of responsibilities, different technical functions such as user question clarification, retrieval decision management and case search are decoupled and executed collaboratively, so that the system can gradually converge the semantics of the question during the user interaction process, thereby reducing the retrieval deviation caused by incomplete or inaccurate initial user input and improving the overall effectiveness of legal case retrieval.
[0020] (2) The case library administrator intelligent agent integrates functions such as case analysis, case prediction, case retrieval, retrieval result evaluation and re-interaction triggering within the same processing framework, enabling the system to dynamically adjust the retrieval process according to the coverage and matching degree of the retrieval results, avoiding the problem of single or mismatched results caused by a fixed retrieval process, and improving the intelligence and adaptability of the retrieval process.
[0021] (3) By using the question-and-clarification mechanism triggered by the analysis of search results, the system can proactively obtain user preference information and use it for subsequent search decisions when the search results cannot fully meet the user's goals or the results are distributed in a single way. This allows the system to continuously optimize the search results without increasing the user's burden, thereby improving the matching degree and stability of the case recommendations. Attached Figure Description
[0022] Figure 1 This is an overall flowchart of the multi-agent collaborative case retrieval method described in this invention. Detailed Implementation
[0023] The present invention will be further described below with reference to the accompanying drawings and specific implementation steps. It should be noted that the technical solution disclosed in the present invention is not limited to the following embodiments. Without departing from the overall technical concept and technical effects of the present invention, those skilled in the art can make equivalent substitutions or adjustments to the relevant implementation methods, all of which should be considered to fall within the protection scope of the present invention.
[0024] This invention proposes a case retrieval method based on multi-agent collaboration. It introduces a case database receptionist agent and a case database administrator agent with different decision-making permissions and functional boundaries, forming a clearly defined and collaborative intelligent processing mechanism. Through the collaborative cooperation of these two agents, the system can complete multiple rounds of questioning and clarification of legal issues, progressive semantic convergence, and case analysis during user interaction. Based on this, it extracts case attributes, predicts case causes, and performs case retrieval and refined recommendation based on the analysis results. This achieves dynamic modeling of user legal needs and output of highly relevant cases at the overall process level. The following will combine... Figure 1 The overall process shown provides a detailed description of the specific implementation methods and systems of the present invention.
[0025] Step S1: Case Library Receptionist Agent Module.
[0026] This module serves as a hub for user interaction, clarifying user intent through multiple rounds of dialogue. The specific implementation steps are as follows.
[0027] Step S11: User input and clarification of questions.
[0028] The system receives descriptions of legal issues input by users via a web interface or API. The input text is unstructured natural language, for example: "An outsourced employee was involved in a traffic accident during his / her employment. A third party sued the actual employing unit. The employing unit refused to assume responsibility on the grounds that the employee had no employment contract with the unit, and argued that the relationship between the two units was a service outsourcing relationship, and that the employing unit with which the employee had an employment contract should bear the responsibility. Would this be supported? Please provide a successful case."
[0029] The receptionist agent performs real-time analysis based on LLM technology. The system generates rhetorical questions to clarify ambiguous intentions. For example, given the above input, it might ask: "If an outsourced employee has an accident while performing work tasks, even without an employment contract, the actual employing unit may still be considered the responsible party. The key is whether the employee was directly managed or directed by the employing unit, and whether the accident occurred during the performance of assigned tasks. Specifically, who usually arranges this employee's work content, attendance, and vehicle usage?" This fills in missing details. Clarification through rhetorical questions is achieved through dynamic instruction prompts, focusing on inquiring about case details, legal relationships, or key claims, thus lowering the professional threshold for users.
[0030] Step S12: End clarification judgment.
[0031] Based on the context of multi-turn dialogue, the agent evaluates the completeness and clarity of the user's response. If the question description is sufficiently clear (e.g., the user adds key facts), the clarification process ends; otherwise, it continues to generate follow-up questions. Evaluation criteria include information completeness, coverage of legal elements, and logical consistency, ensuring that the input question can be accurately processed by subsequent modules.
[0032] Step S13: Problem standardization.
[0033] After the clarification is completed, obtain the final problem description. Standardized problem descriptions conforming to the legal context are generated through LLM semantic parsing. For example, the above example could be restructured as follows: "In a service outsourcing relationship, an outsourced employee is involved in a traffic accident during their work, causing damage to a third party. The actual employing unit refuses to assume responsibility on the grounds that there is no labor contract with the employee, and claims that the outsourcing unit should bear the responsibility. Would the court support this defense? Please provide a successful case." This ensures that the issue is focused, the logic is smooth, and the legal details and claims are highlighted.
[0034] Step S2: Case Library Administrator Intelligent Agent Module.
[0035] This module, as the core processing unit, has dynamic decision-making capabilities—it skips in-depth analysis and directly retrieves results for simple problems (such as "traffic accidents"); it executes the entire process for complex problems; and when the results are unsatisfactory, it first conducts a self-evaluation before deciding to re-analyze or trigger a follow-up question for clarification.
[0036] Step S21: Case analysis and Boolean search expression generation.
[0037] Description of standardization issues In-depth case analysis is conducted using instruction-driven LLM (such as the Qwen series or deepseek model) to understand the issue from dimensions such as legal subject relationships, time span, points of contention, and legal basis, thus clarifying the search target. Specifically, the system uses case analysis prompts as shown in Table 1 to combine elements (such as "employer" and "exemption") using AND, OR, and NOT operators to generate Boolean logic search expressions. For example, the generated search query is: "(Outsourced employee OR Labor dispatch employee OR Dispatch employee) AND (Traffic accident OR Car crash) AND (During employment OR Working period OR Performing duties) AND (Employing unit OR Actual employing unit) AND (Service outsourcing relationship OR Outsourcing relationship) AND (Defense OR Refusal to assume liability OR Exemption from liability) AND (Support OR Win the case) AND (Employer's liability OR Vicarious liability OR Civil Code OR Tort liability)". This search query can be directly used for structured database queries. Table 1: Key Words for Case Analysis Step S22: Case attribute extraction and case cause prediction.
[0038] This function is based on standardized questions after semantic understanding of the questions. The specific implementation method for extracting structured case attributes and predicting case causes is as follows.
[0039] Case attribute extraction: The system uses LLM instruction technology (as shown in Table 2 for case attribute extraction prompts) to extract attributes from standardized issues. Extract case attributes This includes case date, court of trial, court level, trial procedure, document type, and case region. To ensure data consistency, the extracted attribute values are standardized using metadata from the local case database. For example, in the question "What is the conviction, sentencing, and legal basis for a case involving embezzlement of 25 million yuan in the first instance at the Intermediate People's Court?", the extracted court level is standardized to "Intermediate People's Court," and the trial procedure is "first instance." This process only extracts attributes explicitly present in the user description and does not perform inference generation. Table 2: Case Attribute Extraction Hints Case Cause Prediction: This function, based on a pre-built database of case types and causes of action, combined with LLM technology, predicts standardized issues. Perform case identification. Use the prompts shown in Table 3 to drive the LLM output of case types. and specific causes of action Specifically, the cause of action and case type are linked through a mapping function. Ensure that each cause of action One case type corresponds to one unique case. This enables unified modeling for case cause prediction and type identification. The prediction results are also standardized based on the case cause database. For example, for the question "A case heard by a Hainan court where an elderly person injured another person while protecting their kidnapped grandson, and the court considered it legitimate self-defense," the prediction can identify the case types related to "civil" ("disputes over the right to life, health, and bodily integrity") and "disputes over liability for excessive self-defense"), and "criminal" ("intentional injury"), ensuring consistency between the output and the case cause fields in the case database. Table 3: Case Prediction Hints Step S23: Case retrieval and ranking.
[0040] This step involves an integrated process of dual-path retrieval and fine-grained ranking of cases, and the specific implementation method is as follows.
[0041] 1) Dual-path case retrieval.
[0042] First search path (standardized question + case attributes + cause of action): Standardized question description generated by the question semantic understanding module. Attribute values output by the case attribute extraction module and the predicted results of the case. As input, standardized problem descriptions are used in the ES case text database. As a full-text search field, combined with the case attributes of this question. and cause of action As a filter condition, a matching search is performed. Simultaneously, the standardization problem is... Converted into vectors using the bge-m3 semantic embedding model Vector retrieval is performed in the Milvus vector database, calculating the cosine similarity with the case text vector, and then applying the case attributes. and cause of action Filter and return the cases with the highest similarity.
[0043] The second search path (Boolean logic search expression + case attributes + cause of action): based on the Boolean logic search expression generated by the problem case analysis module. Case attributes and the predicted results of the case. As input. In the ES database, the Boolean logic search expression... As a query condition, it is also combined with case attributes. and cause of action Perform filtering and retrieval. Simultaneously, convert the Boolean logic search expression into a semantic vector. When performing similarity retrieval in a vector database, certain conditions must be set before returning the results. and Filtering.
[0044] The two paths execute in parallel, searching both text fields from the Elasticsearch library and semantic vectors from the vector library. The system records the matching text (e.g., semantically highlighted text or highlighted logical terms) for each candidate case. Based on a unique case identifier (e.g., case number), the system deduplicates and merges identical cases from both paths. The final output is a set of deduplicated candidate cases. ,in The number of candidate cases after deduplication is used as input for case ranking.
[0045] 2) Case study layout.
[0046] Using LLM and sorting hints, each candidate case Quantitative scoring ,in The questions are legal issues originally entered by the user. The system uses prompts as shown in Table 4 to drive LLM case-by-case analysis and matching scores. The scoring factors include: (1) Factual context: the similarity between the facts of the case and the user's description of the situation. (2) Legal relationship and focus of dispute: the degree of alignment between the legal relationship and the focus of dispute. (3) Legal subject and liability structure: the consistency between the subject's identity and the division of liability. (4) Applicable legal provisions and legal logic: the relevance between the legal basis and the reasoning logic. Table 4: Case Study Layout Tips The system sorts candidate cases from highest to lowest based on the matching score output by the LLM, and each candidate case... It consists of the case number, cause of action, highlighted text, and case document content (including the findings of this court, the opinion of this court, and the judgment).
[0047] Step S24: Trigger a counter-question for clarification.
[0048] To achieve more accurate case recommendations and user interaction, after the case ranking results are obtained, the administrator AI agent uses the matching score... and score threshold Initiate different processes: When there are no high-match cases, i.e., case match score All below the threshold The agent first assesses the accuracy of its understanding of the user's current problem. For example, if the user's problem is "the actual employer is not liable," and the system initially misinterprets it as "the employer is exempt from liability," then this understanding is inconsistent with the user's core demand. After identifying this misunderstanding, the agent re-executes the problem understanding process and performs case retrieval again. If the assessment result indicates that the current understanding is accurate, the agent triggers a clarification module to guide the user to supplement information related to their retrieval needs.
[0049] When highly matched cases exist, the agent further analyzes whether these cases correspond to different legal situation preferences, meaning that under the current context, the system cannot accurately infer the user's true legal situation. For example, regarding "issues concerning the refund of deposits in a home purchase contract," after recommending case results, if it finds that the highly matched cases simultaneously include case types with opposing responsible parties, the agent triggers a clarification mechanism, asking the user, "Case results with different scenarios have been detected. Are you more concerned about the deposit handling rules under the 'seller's breach' or 'buyer's breach' scenarios?" This feature integrates user feedback on search results and adjusts subsequent search logic or result filtering rules to form a closed-loop processing mechanism based on user interaction: "search—evaluation—interaction—optimization," thereby improving the accuracy and relevance of case recommendations.
[0050] In another embodiment of the present invention, a case retrieval system based on multi-agent collaboration is also disclosed. The system includes: a user interaction interface, a case database receptionist agent module, a case database administrator agent module, a case database, and a recommendation output interface. Each functional module communicates through a predefined data protocol and application programming interface. Data is transferred between modules via HTTP API, ensuring a highly cohesive and loosely coupled system architecture. This forms a loosely coupled, highly cohesive system architecture. The system can be deployed in a cloud computing environment or a local server environment to support concurrent user access and large-scale case data processing needs.
[0051] 1. User Interaction Interface: Used to receive user input of legal question descriptions and interactive feedback, and to pass the interactive content to the case library receptionist intelligent agent.
[0052] 2. Case Library Receptionist Intelligent Agent Module: This module is used to perform multiple rounds of questioning and clarification on legal questions entered by users, as well as semantic normalization. Through continuous interaction, it gradually clarifies the user's search target and sends the clarified question information to the Case Library Administrator Intelligent Agent.
[0053] Specifically, the receptionist agent module in this case library serves as the hub for system-user interaction, responsible for engaging in multi-round dialogues with users to clarify and define their legal issues. Specific functions include: Clarification through counter-question: Based on LLM technology, the description of legal issues input by the user. The system performs real-time analysis and generates follow-up questions to clarify user intent, such as inquiring about missing case details, legal relationships, or key points of the claim, thus making the user's problem description clearer and more complete. This feature helps lower the professional barrier for users and improves the accuracy of subsequent processing.
[0054] Ending Clarification: Based on the context of the multi-turn dialogue, the AI assesses the completeness and clarity of the user's response to determine whether the user's question is sufficiently clear. If the question description is sufficiently clear, the clarification ends; otherwise, it continues with follow-up questions for clarification.
[0055] Question standardization: After the clarification process, obtain a description of the legal issues encountered in the user interaction. By using LLM to perform semantic parsing of the text, standardized problem descriptions that conform to the legal context are generated. The aim is to focus on the issue, ensure logical flow, and highlight the plot and the demands.
[0056] 3. Case Library Administrator Intelligent Agent Module: As the decision-making and management unit of the case retrieval process, it is used to perform case analysis, case prediction, case retrieval and fine sorting based on the problem information obtained from multiple rounds of interaction; it is also used to analyze and judge the retrieval results, and trigger the questioning and clarification process when the retrieval results cannot meet the user's goals or involve different legal situation preferences, and dynamically adjust the retrieval strategy according to the newly added interaction information.
[0057] Specifically, this intelligent agent, as the core processing and decision-making unit of the system, possesses multiple functional tools and can dynamically invoke them based on the question context and search results, forming an intelligent closed loop integrating self-assessment, attribution analysis, and interactive optimization. Its core decision-making logic is as follows: for simple and clear questions like "traffic accident," it can skip in-depth case analysis and directly invoke the case search function; for complex questions, it executes the complete process. Crucially, when the search results are unsatisfactory, the intelligent agent does not simply directly question the user, but first reflects on its own understanding of the user's question: if the understanding is incorrect, it automatically re-analyzes to correct itself; if the understanding is correct but the result is not accurate enough, it triggers a clarification function to interact with the user and obtain preference information to optimize the search. Its specific functions include: Case analysis and Boolean query generation: Standardized problem descriptions that conform to the legal context. In-depth case analysis is conducted using instruction-driven LLM to understand the issues from aspects such as the type of legal subject relationship, time span, focus of the problem, and legal basis. This clarifies the search objectives and analyzes relevant case types (civil, criminal, administrative, enforcement, and state compensation) and circumstances, ensuring both breadth and depth of understanding and analysis. Boolean logic relationships (AND, OR, or NOT) are then established based on the analyzed subject identities, relationship types, case circumstances, legal rights, and legal basis to generate Boolean logic search expressions that conform to the case database retrieval standards. .
[0058] Case attribute extraction and cause prediction: from standardized problem description Extract case attribute information This includes case date, court of trial, court level, trial procedure, document type, and case region. All of the above attribute information are user-defined constraints used to retrieve information from the case database. LLM (Limited Ledger Modeling) is used only to extract attribute information, not for inference or generation. To standardize the extracted attribute information, it is processed based on case attribute data stored in the local case database. The system, based on an established database of case types and causes of action, combines LLM and prompt word technology to standardize the problem descriptions. Perform multi-category prediction of case type and cause of action, and output the corresponding case type. and specific causes of action Similarly, the predicted cause of action names are standardized based on information from the cause of action database.
[0059] Case retrieval and ranking: This function performs a dual-path retrieval; the first path uses a standardized problem description. Predicted Case and extracted case attributes The core is the second path, which uses Boolean logic retrieval. Predicted Case and case attributes The core functionality involves two search paths that simultaneously execute search tasks in two different case databases. First, matching searches are performed in the case text database built using Elasticsearch (ES) to quickly filter structured text results. Second, based on semantic embedding vectors, semantic similarity searches are performed in the Milvus vector database to obtain the most semantically relevant cases. This feature uses standardized question descriptions... Boolean logic retrieval expression and its corresponding semantic embedding vector , As core query criteria, searches are conducted in Elasticsearch and a vector database to match relevant cases. Subsequently, the results are filtered based on case attributes and cause of action, precisely identifying cases that meet the user's needs. The cases returned by both filtering methods are then deduplicated and merged to obtain candidate cases. ,in To determine the number of duplicate candidate cases, we then use sorting suggestions and LLM to quantify the relevance of each case to the user's problem description. Matching score , Indicates the first One case, This represents the original question description that involves multiple rounds of user interaction. Cases are assigned matching scores. The cases are ranked from highest to lowest, and the scoring criteria include factors such as the factual context, legal relationships and points of contention, legal subjects and liability structure, and applicable legal provisions and legal logic. This invention does not limit the specific entity responsible for building the case database, only requiring that it possess the capability to retrieve text content and metadata.
[0060] Triggering a clarification question: If the result after fine-ranking the case does not have a matching score... Higher than the preset threshold In cases where the AI agent encounters a problem, it will initiate a self-reflection mechanism to assess the accuracy of its understanding of the user's problem description. If it determines that there is a misunderstanding, it will re-execute the understanding process and search for relevant cases again. If the understanding is correct, the AI agent will invoke the question-and-clarification module to provide a friendly prompt to the user, facilitating further information interaction. If there are matching results higher than a preset threshold in the ranked cases, after recommending case search results, the AI agent will further analyze whether the cases involve legal situations with different preferences. If such situational disagreements are detected, the AI agent will also trigger the question-and-clarification function to obtain the user's clear preferences and assist in subsequent decision-making.
[0061] 4. Case Database: This includes an inverted index database for storing case text and structured metadata, and a vector database for storing case semantic vectors.
[0062] 5. Recommended Output Interface: Used to receive the final case ranking results processed by the case library administrator agent and output them to the user terminal.
[0063] By constructing a multi-round clarification mechanism oriented towards the user interaction process and a dynamic case analysis and retrieval strategy driven by clarification results, this invention can continuously converge the user's retrieval intent during the legal consultation process, improve the accuracy of case matching and the rationality of result coverage, thereby enhancing the applicability of the legal case retrieval system in complex legal problem handling scenarios, and is suitable for application scenarios such as legal consultation and judicial assistance.
[0064] The above embodiments specifically illustrate the implementation method and operation flow of the multi-agent collaborative processing mechanism proposed in this invention in the legal case retrieval scenario, verifying the feasibility and stability of this technical solution in the linked processing of user question clarification, case understanding, and case retrieval. Without changing the division of responsibilities and collaborative logic of the multi-agents established in this invention, the language model type, case database storage format, and data interaction method between agents can all be configured or adapted according to the specific application environment. These changes do not affect the overall concept of the technical solution of this invention or the technical effects it achieves.
Claims
1. A case retrieval method based on multi-agent collaboration, characterized in that, The method includes: Obtain standardized questions that meet the evaluation criteria based on the user's original questions; By combining the case database, we can deduce the selected cases corresponding to this standardized problem; Calculate the matching score between the user's original question and the featured case, and output the featured case when the matching score reaches a set score threshold.
2. The method according to claim 1, characterized in that, Based on the user's original question, standardized questions that meet the evaluation criteria are obtained, including: Obtain the original question input by the user; The original questions and clarification texts were evaluated according to evaluation criteria, which included: information completeness, coverage of legal elements, and logical consistency. For original questions and clarification texts that do not meet the evaluation criteria, a large language model is used to analyze them and generate rhetorical questions to obtain user feedback and clarification texts for evaluation. This process continues until the original questions and clarification texts meet the evaluation criteria, resulting in the final question description. The final problem description is semantically parsed using a large language model to generate a standardized problem.
3. The method according to claim 2, characterized in that, Based on the case database, the selected cases corresponding to this standardized question are deduced, including: By combining the case database, we can deduce the candidate cases corresponding to this standardized problem; The quantitative score between each candidate case and the original question is calculated based on a large language model. The quantitative scoring factors include: factual context, legal relationship and focus of dispute, legal subjects and liability structure, and applicable legal provisions and legal logic. Based on the quantitative scoring results, selected cases were chosen from the candidate cases.
4. The method according to claim 3, characterized in that, The case database includes: a case text database; The process of reasoning about candidate cases corresponding to the standardized problem using a case database includes: Based on the standardized question, a search is performed in the case text database to obtain the candidate cases corresponding to the standardized question.
5. The method according to claim 3, characterized in that, The case database includes: a case text database and a case vector database; The process of reasoning about candidate cases corresponding to the standardized problem using a case database includes: Based on the large language model, the case analysis of the standardization problem is carried out, and Boolean logic retrieval expression is generated based on the elements obtained from the case analysis. Use large language models to extract case attributes from standardized problems; Based on a pre-built database of case types and causes of action, and combined with a large language model, cause of action identification is performed on standardized issues; Based on standardized questions, case attributes, and causes of action, a search is conducted in the case text database to obtain the first set of candidate cases; Based on Boolean logic search terms, case attributes, and causes of action, a search is performed in the case vector database to obtain a second set of candidate cases. Merge the first candidate case set and the second candidate case set to obtain the candidate cases corresponding to the standardized problem.
6. The method according to any one of claims 1 to 5, characterized in that, After calculating the matching score between the user's original question and the featured case, the method further includes: If the matching score is less than a set score threshold, obtain the problem understanding result generated during the reasoning process of the standardized problem, and evaluate the problem understanding result; Based on the evaluation results, the process of obtaining standardized questions that meet the evaluation criteria based on the user's original question is re-executed, or the process of reasoning about the selected cases corresponding to the standardized question based on the case database is re-executed, until the matching score of the selected cases reaches the set score threshold.
7. The method according to any one of claims 1 to 5, characterized in that, After outputting this selected case, the method further includes: Build a curated collection of case studies; Legal scenario analysis is performed on the selected case set. When the analysis results indicate that the selected case set involves multiple legal scenario preferences, a scenario clarification question is issued to the user based on a large language model. The large language model analyzes the user's response to the clarification question in the context in order to obtain specific selected cases from the selected case set; wherein, the specific selected cases are selected cases involving the legal context of the user's original question.
8. A case retrieval system based on multi-agent collaboration, characterized in that, The system includes: The Case Library Receptionist Intelligent Agent module is used to obtain standardized questions that meet the evaluation criteria based on the user's original questions. The Case Library Administrator intelligent agent module is used to infer the selected cases corresponding to the standardized question by combining the case database; calculate the matching score between the user's original question and the selected cases; and output the selected cases when the matching score reaches a set score threshold.
9. A computer device, characterized in that, The computer device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements the case retrieval method based on multi-agent cooperation as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the case retrieval method based on multi-agent collaboration as described in any one of claims 1-7.