Case retrieval method, device and equipment based on large language model and vector retrieval
Patent Information
- Application Number
- CN202610719450.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2046-05-25
AI Technical Summary
[0005]因此,为了克服上述现有技术的缺点,本发明提供一种基于大语言模型和向量检索的类案检索方法、装置和设备,通过大语言模型基于不同维度对案情案由进行多角度拆解从而解决现有技术中类案检索单维度信息易丢失、召回噪音大以及缺乏意图区分的问题
[0016]与现有技术相比,本发明的优点在于:通过大语言模型基于不同维度对案情案由进行多角度拆解,多维度的向量拆解确保了案例搜索不会遗漏事实或争议焦点等任一重要维度,极大地提高了类案检索的准确率和全面性。另外,在进行多维度拆解时,还根据不同维度在本次分析中的重要性实时调整各维度的检索权重,使得该次类案检索更具有针对性。而且本申请还通过大语言模型对案情事情进行拆解实现了降维提取检索词(案例检索元素),而后基于案例检索元素进行向量检索,再通过大语言模型对召回的初步召回案例进行多次过滤,将不能支持用户法律研究任务的无关案例剔除,克服了传统RAG(检索增强生成)技术在长文本、高专业门槛的法律领域中容易产生的幻觉现象,保证了引用的准确性;而后基于多次过滤后的目标案例进行大语言模型总结,不仅提高分析效率,还提高了分析对比的精度。而且本申请还通过大语言模型精准判断用户意图,保证系统能有效处理泛化的法律研究,提高了用户体验。
Smart Images

Figure CN122262205B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data retrieval, and in particular to a case retrieval method, apparatus, and device based on large language models and vector retrieval. Background Technology
[0002] Case retrieval is a process by which judges, prosecutors, or lawyers search for previously effective judgments that share similar basic facts, points of contention, and applicable laws based on given factual descriptions. This provides judges, prosecutors, or lawyers with references or helps to standardize judgments. With the continuous development of information technology, the legal industry is also gradually transforming towards digitalization and intelligence. Traditional legal document retrieval methods rely on manual searching and archiving, which is not only time-consuming and labor-intensive, but also faces increasing demands for efficiency and accuracy as the volume of legal documents continues to grow. To address these challenges, in recent years, more and more legal practitioners and researchers have begun to focus on using artificial intelligence technology to improve the efficiency and accuracy of legal document retrieval. AI-based legal retrieval systems can quickly and accurately analyze and retrieve massive amounts of legal document data through technologies such as natural language processing and machine learning, thereby providing users with more convenient legal information query services.
[0003] However, existing AI-based legal document retrieval systems still have some shortcomings in practical applications. First, existing systems have poor single-dimensional retrieval capabilities. Complex cases often involve multiple claims, intricate factual processes, and multiple points of contention, but traditional retrieval typically transforms the entire case into a single query vector, leading to significant information loss and the easy retrieval of incomplete cases. Second, retrieval is mostly based on semantic similarity calculations, often resulting in the retrieval of irrelevant cases that are superficially similar but have completely different legal logic / cause of action, resulting in high retrieval noise and false positives. In addition, the retrieval system cannot adaptively distinguish the intent of user commands, leading to a mismatch in retrieval strategies.
[0004] Therefore, how to accurately improve the retrieval of similar cases for multi-dimensional information is an urgent problem to be solved. Summary of the Invention
[0005] Therefore, in order to overcome the shortcomings of the prior art, the present invention provides a case retrieval method, apparatus and device based on large language model and vector retrieval. By using large language model to decompose the case details and causes from multiple perspectives based on different dimensions, the present invention solves the problems of easy loss of single-dimensional information, large recall noise and lack of intent differentiation in the prior art.
[0006] To achieve the above objectives, this invention provides a case retrieval method based on a large language model and vector retrieval, comprising: receiving a case retrieval query instruction sent by a terminal; invoking a large language model to perform semantic feature analysis on the case retrieval query instruction, determining the corresponding back-end processing branch, wherein the back-end processing branch includes at least a regular case retrieval; when determined to be a regular case retrieval, using a large language model to decompose the case information and causes carried by the case retrieval query instruction according to different case dimensions, generating a predetermined number of case retrieval elements with a text length not less than a preset threshold, and assigning a retrieval weight to each case retrieval element; and based on the multiple case retrieval elements... The initial vector retrieval request is input into the case vector database, asynchronously and concurrently invoked, and a preliminary recall case set corresponding to each case dimension is obtained based on the similarity value. The preliminary recall cases in the preliminary recall case set are then filtered based on the similarity value and the retrieval weight to generate a candidate case set containing a preset number of candidate cases. The candidate cases and the case details are then batch-processed and input into the large language model for relevance Boolean value determination, filtering out weakly related or irrelevant candidate cases to obtain the target case set. Finally, the large language model is used to analyze and generate a case study analysis report based on the case details and the target cases in the target case set.
[0007] In one embodiment, the step of using a large language model to decompose the case information and causes of action carried by the case retrieval query instruction according to different case dimensions, generating a predetermined number of case retrieval elements with a text length not less than a preset threshold, and assigning retrieval weights to each case retrieval element includes: analyzing the case information and causes of action using a large language model to determine the text content corresponding to the case dimensions, wherein the case dimensions are at least one of the following: case claims, factual circumstances, and points of contention; summarizing the case information and causes of action based on the text content of different case dimensions using a large language model to obtain multiple summary texts; when the text length of the summary text corresponding to the case dimension is less than a preset threshold, adding strategy constraints for that case dimension and regenerating the summary text; setting the summary texts with a text length not less than the preset threshold as case retrieval elements, and assigning retrieval weights to each case retrieval element based on the case dimensions.
[0008] In one embodiment, setting the summary text with a text length not less than a preset threshold as a case retrieval element includes: extracting at least one of the time field, administrative division field, legal field, and court level field from the case details and causes of action, and generating structured parameters based on the extracted fields; setting the summary text with a text length not less than the preset threshold and the structured parameters as case retrieval elements.
[0009] In one embodiment, the step of extracting at least one of a time field, an administrative division field, a legal field, and a court level field from the case facts and causes of action, and generating structured parameters based on the extracted fields, includes: when extracting an administrative division field and / or a court level field from the case facts and causes of action, generating corresponding court level parameters and geographical scope parameters as structured parameters; when extracting a time field from the case facts and causes of action, determining whether the time range of the time field is clear; if the time range is determined to be unclear, determining whether the case facts and causes of action contain a legal field, and obtaining the implementation date of the specific legal field as a start time parameter and setting it as a structured parameter based on the legal field determined to be a specific legal field; otherwise, generating specific start time parameters and end time parameters based on the current system time of the server and setting them as structured parameters.
[0010] In one embodiment, the step of filtering the preliminary recalled cases in the preliminary recalled case set based on the similarity value and the retrieval weight to generate a candidate case set containing a preset number of candidate cases includes: deserializing the preliminary recalled case set and performing hash storage using the case number of the preliminary recalled case as the primary key; if a primary key conflict is detected, storing the higher similarity value between the preliminary recalled case and the case retrieval element; sorting the preliminary recalled cases in each preliminary recalled case set separately according to the similarity value, and selecting a specified number of preliminary recalled cases with large similarity values as candidate cases and storing them in the candidate case set; mixing the remaining preliminary recalled cases in all preliminary recalled case sets that have not yet entered the candidate case set, performing global sorting according to the similarity value and the retrieval weight, and sequentially adding preliminary recalled cases as candidate cases to the candidate case set until the number of cases in the candidate case set reaches a preset number.
[0011] In one embodiment, the candidate cases and the case details are input into the large language model in batch processing for relevance Boolean value determination, filtering out weakly related or irrelevant candidate cases to obtain a target case set. This includes: obtaining case details corresponding to each candidate case in the candidate case set; encapsulating the case details of all candidate cases and the case details into a batch processing queue task and sending it to the large language model for relevance Boolean value determination; deleting candidate cases corresponding to case details with negative relevance Boolean values, and retaining candidate cases corresponding to case details with positive relevance Boolean values as target cases, thus obtaining a target case set.
[0012] In one embodiment, the step of using the large language model to analyze and generate a case study analysis report based on the case facts and the target cases in the target case set includes: extracting case details from the case facts and the target cases in the target case set, and combining the case details from the case facts and the target cases into the input context of the large language model according to the context length limit of the large language model; the large language model performs multi-step reasoning based on its internal knowledge base and the input context, performs cluster analysis on the case details of different target cases, identifies different judicial opinions; and extracts key legal facts and judgment logic from each target case to form coherent and professional natural language text, and outputs a structured case study analysis report.
[0013] In one embodiment, the background processing branch further includes specific case search, which obtains the party's name and / or precise case number through regular expression or lightweight extraction mechanism, and performs precise matching retrieval in the case vector database based on the obtained party's name and / or precise case number to obtain the corresponding specific case.
[0014] A case retrieval device based on a large language model and vector retrieval, the device comprising: an instruction receiving module for receiving a case retrieval query instruction sent by a terminal; a branch analysis module for calling a large language model to perform semantic feature analysis on the case retrieval query instruction, determining the corresponding back-end processing branch, wherein the back-end processing branch includes at least a regular case retrieval; an element decomposition module for, when determined to be a regular case retrieval, using a large language model to decompose the case information and causes carried by the case retrieval query instruction according to different case dimensions, generating a predetermined number of case retrieval elements with a text length not less than a preset threshold, and assigning a retrieval weight to each case retrieval element; and a preliminary case recall module for, based on multiple case retrieval elements... The system initiates a vector retrieval request to input case vector databases, asynchronously and concurrently calling and obtaining a preliminary recall case set corresponding to each case dimension based on similarity values. A primary filtering module is used to filter the preliminary recall cases in the preliminary recall case set based on the similarity values and retrieval weights, generating a candidate case set containing a preset number of candidate cases. A secondary filtering module is used to input the candidate cases and the case details into the large language model in batch processing for relevance Boolean value determination, filtering out weakly related or irrelevant candidate cases to obtain a target case set. An analysis module is used to use the large language model to analyze and generate a case study analysis report based on the case details and the target cases in the target case set.
[0015] A computer device includes a memory and a processor, the memory storing a computer program, characterized in that the processor executes the computer program to implement the steps of the above-described method.
[0016] Compared with existing technologies, the advantages of this invention are as follows: By using a large language model to decompose case facts and causes of action from multiple perspectives based on different dimensions, the multi-dimensional vector decomposition ensures that case searches do not miss any important dimensions such as facts or points of contention, greatly improving the accuracy and comprehensiveness of case retrieval. Furthermore, during multi-dimensional decomposition, the retrieval weights of each dimension are adjusted in real time according to their importance in the analysis, making the case retrieval more targeted. Moreover, this application also achieves dimensionality reduction and extraction of search terms (case retrieval elements) by decomposing case facts using a large language model, followed by vector retrieval based on these elements. The initial recalled cases are then filtered multiple times using the large language model, eliminating irrelevant cases that do not support the user's legal research tasks. This overcomes the illusion phenomenon that traditional RAG (Retrieval Augmentation) technology is prone to in long texts and highly specialized legal fields, ensuring the accuracy of citations. Finally, the large language model summarizes the target cases after multiple filtering, improving not only analytical efficiency but also the accuracy of comparative analysis. Furthermore, this application uses a large language model to accurately determine user intent, ensuring that the system can effectively handle generalized legal research and improving user experience. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating the case retrieval method based on a large language model and vector retrieval in an embodiment of the present invention; Figure 2 This is a structural block diagram of a case retrieval device based on a large language model and vector retrieval in an embodiment of the present invention; Figure 3 This is an internal structural diagram of a computer device in an embodiment of the present invention. Detailed Implementation
[0019] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0020] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. This application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] It should be noted that the following description covers various aspects of embodiments within the scope of protection of this invention. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this application, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number and aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using other structures and / or functionalities besides one or more of the aspects set forth herein.
[0022] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. The drawings only show the components related to this application and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0023] Furthermore, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0024] This application provides a case retrieval method based on a large language model and vector retrieval, applied on a server or terminal. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable smart devices. The server can be a standalone server or a server cluster composed of multiple servers. For example... Figure 1 As shown, taking the application of this case retrieval method based on a large language model and vector retrieval on a server as an example, the method includes the following steps: Step 101: Receive the case search query instruction sent by the terminal.
[0025] The server receives case retrieval query instructions sent by the terminal. The terminal can be a smart device held by a user, who can be a judge, prosecutor, lawyer, or citizen. The case retrieval query instruction carries the search content entered by the user, which can be legal research on a specific issue or a search for a specific case. When the search content is legal research on a specific issue, the case retrieval query instruction can carry the relevant case details and causes of action. Case details refer to the process, circumstances, and background of the case, including details of the facts, events, actions, time, place, and relationships involved. Cause of action refers to the nature and category of the case, reflecting the legal relationships and points of contention involved, and at least includes the case type (specifying the legal category to which the case belongs, such as contract dispute, tort, criminal offense, etc.). When the case retrieval query instruction is a search for a specific case, it can carry the case number and / or the names of the parties involved.
[0026] Step 102: Call the large language model to perform semantic feature analysis on the case retrieval query command, and determine the corresponding back-end processing branch. The back-end processing branch includes at least the regular case retrieval.
[0027] The server invokes a large language model to perform semantic feature analysis on the case retrieval query command, determining the corresponding backend processing branch. This backend processing branch includes at least a regular case retrieval. The large language model analyzes the case details and causes of action in the case retrieval query command, determining that the query is a legal study addressing a specific issue. The large language model then identifies the backend processing branch as a regular case retrieval. The server performs the case retrieval based on the determination result of this backend processing branch.
[0028] The back-end processing branch may also include a fallback interception branch. When the search content in the case retrieval query is unrelated to the law, the large language model determines that the back-end processing branch is a fallback interception branch and outputs a command for the server to enter the fallback interception branch. The server stops subsequent searches and sends an adjustment command to the user's terminal to guide the user to modify the search content.
[0029] Step 103: When the case is determined to be a regular case search, the large language model is used to decompose the case information and cause of action carried by the case search query instruction according to different case dimensions, generate a predetermined number of case search elements with a text length not less than a preset threshold, and assign search weights to each case search element.
[0030] When a case is determined to be a routine case search, the server again uses a large language model to break down the case information and causes carried by the case search query command according to different case dimensions, generating a predetermined number of case search elements (cvector) with a text length not less than a preset threshold. The server then adjusts the weights in real time based on the importance of different dimensions in this analysis, assigning search weights to each case search element. The sum of the search weights of all case search elements is 1.
[0031] Case dimensions can be any of the following: claims, factual circumstances, and points of contention. Claims can be specific rights asserted by the parties to the court. Factual circumstances refer to the objective course and state of the case, generally arranged chronologically or causally, outlining the chain of events involving time, place, tasks, actions, and results. Points of contention are the core legally significant disagreements between the parties. In some embodiments, case dimensions can be any of the following: claims, factual circumstances, applicable law, judgment, and points of contention. The large language model can decompose the case details based on different case dimensions, forcibly splitting a single instruction into multiple rich-feature retrieval elements with a text length not less than a preset threshold. The preset threshold can be determined based on the analytical performance of the large language model, for example, setting it to 200 characters. The large language model can decompose the case details into 4-6 rich-feature retrieval elements exceeding 200 characters each based on case dimensions, thereby reducing feature interpretation issues in cases with multiple points of contention.
[0032] The large language model breaks down case details and causes based on different case dimensions. For example, in a case retrieval analysis, the factual dimension is more important, so case retrieval elements decomposed based on the factual dimension are assigned high weights (e.g., 0.5), while case retrieval elements decomposed based on other dimensions are assigned low weights (e.g., the remaining 5 dimensions are all 0.1). The server verifies through logic code that the sum of the floating-point numbers of this group of retrieval weights must equal 1.
[0033] Step 104: Initiate a vector retrieval request based on multiple case retrieval elements and input it into the case vector database. Asynchronously and concurrently call the database and obtain a preliminary recall case set corresponding to each case dimension based on the similarity value.
[0034] The server initiates vector retrieval requests based on multiple case retrieval elements, inputting them into the case vector database. It asynchronously and concurrently calls the database and obtains a preliminary recall case set corresponding to each case dimension based on similarity values. Cases in the case vector database can be publicly available cases downloaded from the People's Court Case Database or the China Judgments Online website. Based on the case retrieval elements with different weights output in step 103, the server establishes a concurrent thread pool of size 10. The background program initiates multiple parallel asynchronous RPC calls to access the case vector database. The server can set a minimum similarity threshold (minScore=0.8) to recall multiple sets of preliminary recall cases corresponding to different dimensions.
[0035] Step 105: Based on similarity value and retrieval weight, the preliminary recalled cases in the preliminary recalled case set are screened to generate a candidate case set containing a preset number of candidate cases.
[0036] The server performs a preliminary screening of the initial recalled cases based on similarity values and retrieval weights, generating a candidate case set containing a preset number of candidate cases. The server can multiply the similarity values and retrieval weights, then sort all the initial recalled cases from largest to smallest based on the product, selecting the preset number of initial recalled cases at the top as candidate cases. Alternatively, the server can first determine the number of candidate cases to be selected for each dimension based on the preset number and the retrieval weights of the case dimensions, then sort the initial recalled cases for each dimension by similarity values, selecting the cases with the highest similarity values corresponding to that selection number as candidate cases.
[0037] Step 106: Input the candidate cases and case details into the large language model in batch processing to determine the relevance Boolean value, filter out weakly related or irrelevant candidate cases, and obtain the target case set.
[0038] The server inputs candidate cases and their case details into the large language model in batches for relevance Boolean value determination, filtering out weakly or irrelevant candidate cases to obtain the target case set. The server can extract all candidate case records and reassemble them into batch processing queues (Batch Size). Each task combines the case details of one candidate case with the original case details and asynchronously calls the large model's filtering interface. The server can configure the interface's responseFormat, restricting it to only return Boolean values (True / False) containing the is_relevant field. The server can iterate through the return values; the memory for candidate cases with a Boolean value of False is directly released, retaining only the memory for candidate cases with a Boolean value of True (as target cases) to form the target case set.
[0039] Step 107: Using a large language model, based on the case details and the target cases in the target case set, analyze and generate a case study report.
[0040] The server uses a large language model to analyze and generate a case study report based on the case details, causes of action, and target cases in the target case set. The server can execute loop detection logic, converting the filtered target case set into a string and accumulating the character length (total_length) item by item. When the accumulated value exceeds the system's large model context processing limit (e.g., setting a maximum safety threshold max_length = 70000 characters), it immediately executes a break to exit the loop, discarding redundant cases at the end, thus achieving safe truncation to prevent overflow. The server merges the truncated target case set with the original instruction and inputs it into the final text analysis large model interface. Based on the layout constraints and placeholder mapping rules set by the backend system, it automatically renders and generates multi-level headings and bolded conclusions, and generates a case study report by replacing identifiers, which is then distributed to the user's terminal for display.
[0041] The aforementioned method, through a large language model, decomposes case facts and causes of action from multiple perspectives based on different dimensions. This multi-dimensional vector decomposition ensures that case searches do not overlook any important dimensions, such as facts or points of contention, significantly improving the accuracy and comprehensiveness of case retrieval. Furthermore, during multi-dimensional decomposition, the retrieval weights of each dimension are adjusted in real time according to their importance in the analysis, making the case retrieval more targeted. Moreover, this application uses a large language model to decompose case facts and achieve dimensionality reduction in extracting search terms (case retrieval elements). Vector retrieval is then performed based on these elements, and the large language model further filters the initially recalled cases multiple times, eliminating irrelevant cases that do not support the user's legal research task. This overcomes the illusion phenomenon that traditional RAG (Retrieval Augmentation) technology is prone to in long texts and highly specialized legal fields, ensuring the accuracy of citations. Finally, a large language model summary is performed based on the filtered target cases, improving both analytical efficiency and the accuracy of comparative analysis. Furthermore, this application uses a large language model to accurately determine user intent, ensuring the system can effectively handle generalized legal research and improving the user experience.
[0042] In one embodiment, a large language model is used to decompose the case facts and causes of action carried by the case retrieval query instruction according to different case dimensions, generating a predetermined number of case retrieval elements with a text length not less than a preset threshold, and assigning retrieval weights to each case retrieval element. This includes: using a large language model to analyze the case facts and causes of action, determining the text content corresponding to the case dimensions, where the case dimensions are at least one of the following: litigation claims, factual circumstances, and points of contention; using a large language model to summarize the case facts and causes of action based on the text content of different case dimensions, obtaining multiple summarized texts; when the text length of the summarized text corresponding to a case dimension is less than a preset threshold, adding strategy constraints for that case dimension and regenerating the summarized text; setting the summarized texts with a text length not less than the preset threshold as case retrieval elements, and assigning retrieval weights to each case retrieval element based on the case dimensions.
[0043] The server uses a large language model to analyze the case details and causes of action, determining the text content corresponding to each case dimension. Each case dimension includes at least one of the following: the claim, the factual circumstances, and the points of contention. Before applying the large language model, it can learn from a large number of real cases in a case vector database to understand the text content corresponding to each case dimension. The large language model can be a GLM or ChatGLM series model based on the Transformer architecture, etc.
[0044] The server then uses a large language model to summarize the case details and causes of action based on the text content from different case dimensions, resulting in multiple summarized texts. For example, to guide the model in generating search elements that meet the requirements, the server constructs the following prompt: "Based on the above patent infringement dispute case text, please summarize from the three dimensions of 'claims,' 'facts,' and 'points of contention.' Requirements: 1. The generated content must be logically clear and semantically coherent; 2. It must clearly cover the plaintiff's claims, the defendant's defenses, the technical facts ascertained by the court, and the points of contention regarding the application of law."
[0045] Based on the aforementioned prompts, the large language model can output multiple preliminary summary texts.
[0046] The server performs a word count check on the generated summary text. For example, a word count threshold of no less than 200 words is defined. When the summary text has fewer than 200 words, the server determines that the generated content has insufficient information density or is too brief, and automatically triggers a regeneration process. The server can extract a portion of the text corresponding to the case dimension of the summary text as a strategy constraint, and combine this constraint with previous prompts to send it to the large language model to regenerate the summary text. When the summary text has more than 200 words, the server can directly set it as a case retrieval element. In one embodiment, the server also uses the large language model to analyze whether the summary text simultaneously contains content corresponding to the case's claims, factual circumstances, and points of contention; if it is determined that the summary text simultaneously contains content corresponding to these three aspects, the server sets the summary text as a case retrieval element.
[0047] The above method, by introducing a multi-dimensional prompting engineering and word count verification mechanism, effectively avoids the problem of missing key information caused by overly brief case summaries in traditional case retrieval. It ensures that the generated case retrieval elements simultaneously possess the three core elements of litigation claims, factual circumstances, and points of contention, significantly improving the accuracy of subsequent vectorized retrieval and similarity matching.
[0048] In one embodiment, setting a summary text with a text length not less than a preset threshold as a case retrieval element includes: extracting at least one of the following fields from the case details and causes of action: time field, administrative division field, legal field, and court level field; and generating structured parameters based on the extracted fields; and setting the summary text with a text length not less than the preset threshold and the structured parameters as case retrieval elements.
[0049] The server extracts at least one of the following fields from the case details: time, administrative division, regulations, and court level. Based on these extracted fields, it generates structured parameters. The administrative division and court level fields clearly identify the level of the court handling the case, which can be the Supreme People's Court, a provincial high court, an intermediate court, or a basic-level court. Under the same similarity level, the higher the level, the greater the reference value of the extracted case. The time field helps clarify the time of the case ruling; legal interpretations or judgment standards may change over time. Under the same similarity level, the more recent the time, the greater the reference value of the extracted case. The regulations field can quickly filter out preliminary recall cases similar to the case details.
[0050] The server sets summary text with a text length not less than a preset threshold and structured parameters as case retrieval elements.
[0051] The above methods can further increase the accuracy of case retrieval elements and ensure that the initial recalled cases are more relevant to the case details and causes.
[0052] In one embodiment, at least one of the following fields—time field, administrative division field, legal field, and court level field—is extracted from the case details and causes of action, and structured parameters are generated based on the extracted fields. This includes: when the administrative division field and / or court level field are extracted from the case details and causes of action, corresponding court level parameters and geographical scope parameters are generated as structured parameters; when the time field is extracted from the case details and causes of action, it is determined whether the time range of the time field is clear; if the time range is determined to be unclear, it is determined whether the case details and causes of action contain a legal field, and based on the legal field determined to be a specific legal field, the effective date of the specific legal field is obtained as a start time parameter and set as a structured parameter; otherwise, specific start time parameters and end time parameters are generated based on the current system time of the server and set as structured parameters.
[0053] When the administrative division field and / or court level field are extracted from the case details and cause of action, the server generates the corresponding court level parameter and geographical scope parameter as structured parameters.
[0054] When extracting the time field from the case details, the server determines whether the time range of the time field is clear. For example, if the time field is a specific time (the specific time can be information about the actual year, month, and day, or information that can be deduced from the actual year, month, and day by combining other specific year information with lunar calendar information), then the server determines that the time field is clear; otherwise, the server determines that the time field is unclear.
[0055] When the time frame for determination is unclear, the server checks whether a legal field exists in the case details. Based on the legal field indicating a specific legal regulation, it obtains the effective date of that regulation as the start time parameter and sets it as a structured parameter. Otherwise, it generates specific start and end time parameters based on the server's current system time and sets them as structured parameters. For example, if the specific regulation is the "New Company Law," the server will determine that the actual law is the "Company Law of the People's Republic of China," which came into effect on July 1, 2024. The server will then use the effective date of that specific regulation, "2024-07-01," as the start time parameter and set it as a structured parameter.
[0056] The above method further refines the search terms, rapidly improving the accuracy of similar case searches.
[0057] In one embodiment, the preliminary recalled cases in the preliminary recalled case set are screened based on similarity value and retrieval weight to generate a candidate case set containing a preset number of candidate cases. This includes: deserializing the preliminary recalled case set and performing hash storage using the case number of the preliminary recalled case as the primary key; if a primary key conflict is detected, storing the higher similarity value between the preliminary recalled case and the case retrieval element; sorting the preliminary recalled cases in each preliminary recalled case set separately according to the similarity value, and selecting a specified number of preliminary recalled cases with high similarity values as candidate cases and storing them in the candidate case set; mixing the remaining preliminary recalled cases in all preliminary recalled case sets that have not yet entered the candidate case set, sorting them globally according to similarity value and retrieval weight, and sequentially adding the preliminary recalled cases as candidate cases to the candidate case set until the number of cases in the candidate case set reaches a preset number.
[0058] The server deserializes the initial recall case set and performs hash storage using the case number of the initial recall case as the primary key. If a primary key conflict is detected, the server associates the initial recall case with the case retrieval element with the higher similarity value. The large language model can return the initial recall case set to the server in JSON format. The server can then convert the JSON-formatted initial recall case set into a set of objects that can be manipulated at runtime, specifically parsing the JSON string into in-memory object instances (such as a CaseResult object in Java or a dict structure in Python) for subsequent field-level access and comparison. The server performs hash storage using the case number of the initial recall case (as the unique identifier of the case) as the primary key. Then, the server checks whether there are duplicate initial recall cases in all the initial recall case sets. If a duplicate is found, the server obtains the different similarity values between the duplicate initial recall case and the case retrieval element, associates the initial recall case with the case retrieval element with the higher similarity value, and then deletes the initial recall case with the low similarity value from the initial recall case set.
[0059] The server sorts the initial recalled cases in each initial recalled case set separately based on similarity values, and selects a specified number of initial recalled cases with high similarity values as candidate cases to be stored in the candidate case set. The specified number of cases truncated after sorting each initial recalled case set is a pre-set fixed value, independent of the retrieval weight of each case dimension, to ensure coverage of each business dimension. For each individual cvector recall array (initial recalled case set), the system processor performs a local sort (Sorted by score), and rigidly truncates the top 10 cases to be stored in an independent target set, ensuring that no business dimension is missing.
[0060] The server mixes all remaining preliminary recalled cases from the initial recalled case set that have not yet entered the candidate case set, performs a global sort based on similarity value and retrieval weight, and sequentially adds the initial recalled cases as candidate cases to the candidate case set until the number of cases in the candidate case set reaches a preset number. Specifically, the server mixes the remaining case records in memory that have not yet entered the target set, performs a second global sorting based on the product of similarity value and retrieval weight, and sequentially adds the records to the target set until the number of cases in the candidate case set reaches 100. The global sorting is based on the product of similarity value and retrieval weight, ensuring the relevance of the cases to the current search.
[0061] The method described above first filters candidate cases based on business dimensions, and then performs a second global ranking by multiplying the similarity value and the retrieval weight. This ensures that all business dimensions are included while also guaranteeing that the similarity of the selected candidate cases better matches the needs of this case retrieval. Furthermore, this method resolves conflicts at the memory level, significantly reducing I / O overhead and ensuring that the final set of cases participating in the ranking and display has higher relevance and authority.
[0062] In one embodiment, candidate cases and their case details are input into a large language model in batch processing for relevance Boolean value determination. Weakly or irrelevant candidate cases are filtered out to obtain a target case set. This includes: obtaining case details corresponding to each candidate case in the candidate case set; encapsulating the case details and case details of all candidate cases into a batch processing queue task and sending it to the large language model for relevance Boolean value determination; deleting candidate cases corresponding to case details with negative relevance Boolean values, and retaining candidate cases corresponding to case details with positive relevance Boolean values as target cases, thus obtaining the target case set.
[0063] The server retrieves the case details corresponding to each candidate case in the candidate case set. For example, the server extracts the case details of 100 candidate cases from the candidate case set.
[0064] The server encapsulates the case details and cause of action of all candidate cases into a batch processing queue task and sends it to the large language model for relevance boolean value determination. The server assembles the case details and cause of action of all candidate cases into a batch processing queue task (Batch Size=100). Each task combines one case detail with the original query and asynchronously calls the large model's filtering interface. The server sets the responseFormat of this interface, constraining it to only return a boolean value (True / False) containing the is_relevant field.
[0065] The server deletes candidate cases corresponding to case details with negative correlation Boolean values, and retains candidate cases corresponding to case details with positive correlation Boolean values as target cases, thus obtaining the target case set.
[0066] The method described above achieves automated and high-precision filtering of a large number of candidate cases by introducing batch processing queue tasks, response format constraints, and a dynamic memory filtering mechanism based on Boolean values. This method effectively eliminates interfering cases, significantly improves the semantic relevance and legal applicability consistency of the final output target case set, and optimizes system resource utilization by releasing irrelevant records in memory in a timely manner.
[0067] In one embodiment, a large language model is used to analyze and generate a case study analysis report based on the case facts and target cases in the target case set. This includes: extracting case details from the case facts and target cases in the target case set, and combining the case details from the case facts and target cases into the input context of the large language model according to the context length limit of the large language model; the large language model performs multi-step reasoning based on its internal knowledge base and input context, performs cluster analysis on the case details of different target cases, identifies different judicial opinions, and extracts key legal facts and judgment logic from each target case to form a coherent and professional natural language text, and outputs a structured case study analysis report.
[0068] The server extracts the case details and cause of action from the target case set, and combines these details into the input context of the large language model according to the context length limit. The server converts the filtered target case set into a string and accumulates the character length `total_length` for each case. When the accumulated value exceeds the system's large model context processing limit (e.g., setting a maximum safety threshold `max_length = 70000 characters`), it immediately executes `break` to exit the loop, discarding any redundant cases at the end, thus achieving safe truncation to prevent overflow. The server sends the truncated input context to the large language model. The server can construct prompt words with the following structure: You are a senior intellectual property law expert. Based on the following pending cases and several target similar cases, please write a professional case study and analysis report.
[0069] The case details and cause of action for the pending case are XXXXX; List of target cases (including key points of judgment): Case details of Case 1, Case details of Case 2, ...; Report requirements: 1. The report should include the following sections: a) Case-based judgment trend statistics: Quantitative analysis of the proportion of cases in which a certain claim is supported / rejected in the target case; b) Summary of core points of contention: Extract the legal and technical issues that repeatedly arise in the above-mentioned cases; c) Analysis of the Applicable Law Path: Analyzing the main legal provisions applied by the courts in similar cases and their interpretation logic; d) Handling of case conflicts: If there are discrepancies between cases, the following order should be followed: higher court level > lower court level, newer judgment time > older judgment time, and decision made by the adjudication committee > general case.
[0070] 2. The report should be well-organized and the arguments should be based on the provided target case studies.
[0071] Based on its internal knowledge base and input context, the large language model performs multi-step reasoning, clusters case details of different target cases to identify different judicial opinions, extracts key legal facts and judgment logic from each target case, forms coherent and professional natural language text, and outputs structured case study analysis reports.
[0072] The above method automates the transformation from unstructured case texts into professional, highly readable research reports, enabling the output of comprehensive reports containing data statistics, key point summaries, and strategic recommendations within seconds. This greatly improves users' legal search efficiency and case analysis quality.
[0073] In one embodiment, the backend processing branch also includes specific case search, which obtains the party's name and / or precise case number through regular expression or lightweight extraction mechanism, and performs precise matching and retrieval in the case vector database based on the obtained party's name and / or precise case number to obtain the corresponding specific case.
[0074] The server analyzes the case retrieval query commands using a large language model. When the large language model determines that no case details or causes of action exist, it determines that the backend processing branch is for specific case search. The server obtains the party's name and / or precise case number from the case retrieval query commands using regular expressions or lightweight extraction mechanisms. Based on the obtained party's name and / or precise case number, it performs an exact match search in the case vector database to obtain the corresponding specific case.
[0075] The above method directly triggers the precise retrieval and matching process of the scalar database, greatly reducing system response latency.
[0076] In one embodiment, such as Figure 2As shown, a case retrieval device based on a large language model and vector retrieval is provided. The device includes an instruction receiving module 201, a branch analysis module 202, an element decomposition module 203, a preliminary case recall module 204, a primary screening module 205, a secondary screening module 206, and an analysis module 207.
[0077] The instruction receiving module 201 is used to receive case search query instructions sent by the terminal.
[0078] The branch analysis module 202 is used to call the large language model to perform semantic feature analysis on the case retrieval query command and determine the corresponding back-end processing branch. The back-end processing branch includes at least the regular case retrieval.
[0079] The element decomposition module 203 is used to decompose the case information and cause of action carried by the case retrieval query instruction according to different case dimensions using a large language model when the case is determined to be a regular case retrieval, generate a predetermined number of case retrieval elements with a text length not less than a preset threshold, and assign retrieval weights to each case retrieval element.
[0080] The case preliminary recall module 204 is used to initiate vector retrieval requests based on multiple case retrieval elements, input the case vector database, asynchronously and concurrently call and obtain a preliminary recall case set corresponding to each case dimension based on the similarity value.
[0081] The first-stage filtering module 205 is used to perform a first-stage filtering on the preliminary recalled cases in the preliminary recalled case set based on similarity value and retrieval weight, and generate a candidate case set containing a preset number of candidate cases.
[0082] The secondary screening module 206 is used to input candidate cases and case details into the large language model in batch processing to determine the relevance of the Boolean values, filter out weakly related or irrelevant candidate cases, and obtain the target case set.
[0083] Analysis module 207 is used to generate a case study analysis report by using a large language model based on the case details and the target cases in the target case set.
[0084] Specific limitations regarding the case retrieval device based on large language models and vector retrieval can be found in the limitations of the case retrieval method based on large language models and vector retrieval mentioned above, and will not be repeated here. Each module in the aforementioned case retrieval device based on large language models and vector retrieval can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0085] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 3 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and the database. The internal memory provides the environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores data such as preliminary recall cases, candidate cases, and target cases. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a case retrieval method based on a large language model and vector retrieval.
[0086] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0087] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to perform the following steps: receiving a case retrieval query instruction sent by a terminal; invoking a large language model to perform semantic feature analysis on the case retrieval query instruction, determining the corresponding back-end processing branch, the back-end processing branch including at least regular case retrieval; when determined to be a regular case retrieval, using a large language model to decompose the case information and causes of action carried by the case retrieval query instruction according to different case dimensions, generating a predetermined number of case retrieval elements with a text length not less than a preset threshold, and assigning a method to each case retrieval element. The process involves: retrieving weights; initiating vector retrieval requests based on multiple case retrieval elements and inputting them into a case vector database; asynchronously and concurrently calling and obtaining a preliminary recall case set corresponding to each case dimension based on similarity values; filtering the preliminary recall cases in the preliminary recall case set based on similarity values and retrieval weights to generate a candidate case set containing a preset number of candidate cases; inputting the candidate cases and case details into a large language model in batch processing for relevance Boolean value determination, filtering out weakly related or irrelevant candidate cases to obtain a target case set; and using the large language model to analyze and generate a case study analysis report based on the case details and target cases in the target case set.
[0088] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it performs the following steps: receiving a case retrieval query instruction sent by a terminal; invoking a large language model to perform semantic feature analysis on the case retrieval query instruction, determining the corresponding back-end processing branch, wherein the back-end processing branch includes at least a regular case retrieval; when it is determined to be a regular case retrieval, using a large language model to decompose the case information and causes of action carried by the case retrieval query instruction according to different case dimensions, generating a predetermined number of case retrieval elements with a text length not less than a preset threshold, and assigning a retrieval weight to each case retrieval element. The system initiates vector retrieval requests based on multiple case retrieval elements, inputs them into the case vector database, asynchronously and concurrently calls and obtains a preliminary recall case set corresponding to each case dimension based on similarity values; it then filters the preliminary recall cases in the preliminary recall case set based on similarity values and retrieval weights, generating a candidate case set containing a preset number of candidate cases; it then inputs the candidate cases and case details into a large language model in batch processing for relevance Boolean value determination, filtering out weakly related or irrelevant candidate cases to obtain the target case set; finally, it uses the large language model to analyze and generate a case study analysis report based on the case details and the target cases in the target case set.
[0089] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A case retrieval method based on a large language model and vector retrieval, characterized in that, include: Receive case search query instructions sent by the terminal; The large language model is invoked to perform semantic feature analysis on the case retrieval query command to determine the corresponding back-end processing branch, which includes at least regular case retrieval. When the case is determined to be a regular case search, a large language model is used to decompose the case information and cause of action carried by the case search query instruction according to different case dimensions, generate a predetermined number of case search elements with a text length not less than a preset threshold, and assign search weights to each case search element. Based on multiple case retrieval elements, a vector retrieval request is initiated, and a concurrent thread pool is established by inputting the case vector database. The thread pool is then called in parallel and asynchronously, and a preliminary recall case set corresponding to each case dimension is obtained based on the similarity value. Candidate cases are selected from the preliminary recall case set corresponding to each of the case dimensions. Then, the remaining preliminary recall cases in the preliminary recall case set are filtered once based on the similarity value and the retrieval weight to generate a candidate case set containing a preset number of candidate cases. The candidate cases and the case details are input into the large language model in batch processing to determine the relevance Boolean value, and weakly or irrelevant candidate cases are filtered out to obtain the target case set. Using the aforementioned large language model, based on the aforementioned case details and cause of action and target cases from the target case set, a case study analysis report is generated. The process of selecting candidate cases from the preliminary recall case set corresponding to each of the aforementioned case dimensions, and then filtering the remaining preliminary recall cases in the preliminary recall case set based on the similarity value and the retrieval weight to generate a candidate case set containing a preset number of candidate cases, includes: The initial recall case set is deserialized and hashed using the case number of the initial recall case as the primary key. If a primary key conflict is detected, the case is associated with the case retrieval element with a higher similarity value. Based on the similarity value, the preliminary recalled cases in each of the preliminary recalled case sets are sorted separately, and a specified number of preliminary recalled cases with large similarity values are selected as candidate cases and stored in the candidate case set. All remaining preliminary recall cases in the preliminary recall case set that have not yet entered the candidate case set are mixed together, globally sorted according to the similarity value and the retrieval weight, and the preliminary recall cases are added to the candidate case set as candidate cases in sequence until the number of cases in the candidate case set reaches a preset number.
2. The case retrieval method according to claim 1, characterized in that, The method employs a large language model to decompose the case details and causes of action carried by the case retrieval query command according to different case dimensions, generating a predetermined number of case retrieval elements with a text length not less than a preset threshold, and assigning retrieval weights to each case retrieval element, including: The case facts and causes of action are analyzed using a large language model to determine the text content corresponding to the case dimensions, wherein the case dimensions are at least one of the following: litigation claims, factual circumstances, and points of contention. The case details and causes were summarized using a large language model based on the text content of different case dimensions, resulting in multiple summary texts; When the length of the summary text corresponding to the case dimension is lower than a preset threshold, a portion of the text is extracted from the case details and causes of action corresponding to the case dimension as a strategy constraint, and the strategy constraint is combined with prompt words and resent to the large language model to regenerate the summary text until the text length reaches the preset threshold. The summary text with a length not less than a preset threshold is set as the case retrieval element, and a retrieval weight is assigned to each case retrieval element based on the case dimension. The sum of the retrieval weights of all case retrieval elements is 1.
3. The case retrieval method according to claim 2, characterized in that, The step of setting summary texts with a length not less than a preset threshold as case retrieval elements includes: Extract at least one of the following fields from the case details and causes of action: time field, administrative division field, legal field, and court level field, and generate structured parameters based on the extracted fields; The summary text with a length not less than a preset threshold and the structured parameters are set as case retrieval elements.
4. The case retrieval method according to claim 3, characterized in that, The step of extracting at least one of the following fields from the case details and causes of action: time field, administrative division field, legal field, and court level field, and generating structured parameters based on the extracted fields, includes: When the administrative division field and / or court level field are extracted from the case details and causes of action, the corresponding court level parameter and geographical scope parameter are generated as structured parameters. When extracting the time field from the case details and causes of action, determine whether the time range of the time field is clear; When the time range is unclear, determine whether the case details or cause of action have a legal field, and based on the legal field that is determined to be a specific legal field, obtain the implementation date of the specific legal field as the start time parameter and set it as a structured parameter; otherwise, generate specific start time parameters and end time parameters based on the current system time of the server and set them as structured parameters.
5. The case retrieval method according to claim 1, characterized in that, The candidate cases and the case details are input into the large language model in batch processing for relevance Boolean value determination. Weakly related or irrelevant candidate cases are filtered out to obtain the target case set, including: Obtain the case details corresponding to each candidate case in the candidate case set; The case details and case facts of all the candidate cases are encapsulated into a batch processing queue task and sent to the large language model for relevance Boolean value determination; Remove the candidate cases corresponding to the case details with negative correlation Boolean values, and retain the candidate cases corresponding to the case details with positive correlation Boolean values as target cases, thus obtaining the target case set.
6. The case retrieval method according to claim 1, characterized in that, The process of using the large language model to analyze and generate a case study report based on the case details and the target cases in the target case set includes: Extract the case details and cause of action from the case and the case details of the target cases in the target case set, and combine the case details and cause of action and the case details of the target cases into the input context of the large language model according to the context length limit of the large language model; The large language model performs multi-step reasoning based on its internal knowledge base and the input context, performs cluster analysis on the case details of different target cases, identifies different judicial opinions, extracts key legal facts and judgment logic from each target case, forms coherent and professional natural language text, and outputs a structured case study analysis report.
7. The case retrieval method according to claim 1, characterized in that, The background processing branch also includes specific case lookup. The names of the parties involved and / or the precise case numbers are obtained through regular expressions or lightweight extraction mechanisms. Based on the obtained names of the parties involved and / or the precise case numbers, a precise matching search is performed in the case vector database to obtain the corresponding specific cases.
8. A case retrieval device based on a large language model and vector retrieval, characterized in that, The device includes: The instruction receiving module is used to receive case search query instructions sent by the terminal; The branch analysis module is used to call the large language model to perform semantic feature analysis on the case retrieval query command and determine the corresponding back-end processing branch. The back-end processing branch includes at least the regular case retrieval. The element decomposition module is used to decompose the case information and cause of action carried by the case retrieval query instruction according to different case dimensions using a large language model when the case is determined to be a regular case retrieval, generate a predetermined number of case retrieval elements with a text length not less than a preset threshold, and assign retrieval weights to each case retrieval element. The case preliminary recall module is used to initiate vector retrieval requests based on multiple case retrieval elements, input case vector database to establish a concurrent thread pool, and concurrently and asynchronously call and obtain a preliminary recall case set corresponding to each case dimension according to the similarity value. A primary filtering module is used to select candidate cases from the preliminary recall case set corresponding to each of the case dimensions, and then to perform a primary filtering on the remaining preliminary recall cases in the preliminary recall case set based on the similarity value and the retrieval weight, thereby generating a candidate case set containing a preset number of candidate cases. The secondary screening module is used to input the candidate cases and the case details into the large language model in batch processing to determine the relevance Boolean value, filter out weakly related or irrelevant candidate cases, and obtain the target case set. The analysis module is used to generate a case study analysis report by employing the large language model based on the case details and the target cases in the target case set. The primary screening module includes the following steps: The initial recall case set is deserialized and hashed using the case number of the initial recall case as the primary key. If a primary key conflict is detected, the case is associated with the case retrieval element with a higher similarity value. Based on the similarity value, the preliminary recalled cases in each of the preliminary recalled case sets are sorted separately, and a specified number of preliminary recalled cases with large similarity values are selected as candidate cases and stored in the candidate case set. All remaining preliminary recall cases in the preliminary recall case set that have not yet entered the candidate case set are mixed together, globally sorted according to the similarity value and the retrieval weight, and the preliminary recall cases are added to the candidate case set as candidate cases in sequence until the number of cases in the candidate case set reaches a preset number.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Judicial field case similar matching retrieval system and method
CN121092696A