Law question and answer method, device and equipment based on large model and storage medium

By rewriting the model and using multi-channel recall technology, combined with a large legal response model, we have solved the problem of quickly generating professional legal responses in the legal field, and achieved efficient and accurate legal consulting services for non-professional users.

CN120821799APending Publication Date: 2025-10-21BAIDU COM TIMES TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510900687.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

In the legal field, existing technologies have difficulty in providing professional and authoritative legal consulting services quickly and accurately. Especially for non-professional users, existing technologies have difficulty in handling complex legal issues and generating high-quality legal responses.

Method used

A large-model-based legal question-answering method is adopted. Natural language expressions are converted into information suitable for retrieval through a rewriting model. Combined with multi-way recall and screening technology, candidate corpora are obtained from authoritative and timely corpora, and a large-scale legal response model is used to generate professional legal response results.

Benefits of technology

It provides fast and accurate legal consulting services to non-professional users, improves the professionalism and practicality of legal responses, reduces users' time and energy investment, and enhances user satisfaction and trust.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120821799A_ABST
    Figure CN120821799A_ABST
Patent Text Reader

Abstract

The invention provides a legal question and answer method, device and equipment based on a large model and a storage medium, and relates to the technical field of artificial intelligence, in particular to the technical fields of data processing, large models and the like. According to the specific implementation scheme, natural language expression about the legal problem provided by a target object is obtained; under the condition that the natural language expression meets a rewriting condition, inputting the natural language expression into a rewriting model to obtain at least one piece of to-be-retrieved information output by the rewriting model; adopting a multi-path recall mode to recall a plurality of candidate corpora matched with the at least one piece of to-be-retrieved information; screening out a plurality of target corpora meeting a preset condition from the plurality of candidate corpora; the preset condition comprises a correlation index and an optimization index, so that the target corpus meets corpus of authority requirements and / or timeliness requirements and meets semantic correlation; and inputting the plurality of target corpora and the natural language expression into the large legal reply model to obtain a legal reply result for the natural language expression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to technical fields such as data processing, large models, and deep learning. Background Art

[0002] Large models are deep learning models with enormous parameter sizes (often reaching billions or even tens of billions of parameters), trained using vast amounts of data and powerful computing resources. These models, such as large language models and multimodal models, are examples. With the continuous advancement and development of AI technology in various fields, question-answering systems based on large models have emerged, providing a diverse channel for consultation. This not only improves the efficiency of answering questions but also reduces consultation costs. Summary of the Invention

[0003] The present disclosure provides a legal question-answering method, apparatus, device, and storage medium based on a large model.

[0004] According to one aspect of the present disclosure, a large model-based legal question answering method is provided, comprising:

[0005] Obtaining natural language expressions about legal issues provided by the target object;

[0006] When the natural language expression satisfies the rewriting condition, the natural language expression is input into the rewriting model to obtain at least one information to be retrieved output by the rewriting model;

[0007] Using a multi-way recall method, multiple candidate corpora that match at least one piece of information to be retrieved are recalled;

[0008] Filtering multiple target corpora that meet preset conditions from multiple candidate corpora; the preset conditions include relevance indicators and optimization indicators; the optimization indicators are used to filter corpora that meet authority requirements and / or timeliness requirements; the relevance indicators are used to filter corpora that meet the semantic relevance requirements with the corresponding information to be retrieved;

[0009] Multiple target corpora and natural language expressions are input into the legal response model to obtain legal response results based on the natural language expressions.

[0010] According to another aspect of the present disclosure, a large model-based legal question-answering device is provided, comprising:

[0011] An acquisition module is used to obtain natural language expressions about legal issues provided by the target object;

[0012] a rewriting module, configured to input the natural language expression into a rewriting model when the natural language expression satisfies the rewriting condition, and obtain at least one information to be retrieved output by the rewriting model;

[0013] A recall module, configured to recall multiple candidate corpora that match at least one piece of information to be retrieved using a multi-way recall method;

[0014] A screening module is used to screen out multiple target corpora that meet preset conditions from multiple candidate corpora; the preset conditions include relevance indicators and optimization indicators; the optimization indicators are used to screen out corpora that meet authority requirements and / or timeliness requirements; the relevance indicators are used to screen out corpora that meet the semantic relevance requirements with the corresponding information to be retrieved;

[0015] The response module is used to input multiple target corpora and natural language expressions into the legal response model to obtain legal response results for the natural language expressions.

[0016] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0017] at least one processor; and

[0018] a memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.

[0020] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.

[0021] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any one of the methods according to the embodiments of the present disclosure.

[0022] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0024] Figure 1 This is a scenario architecture diagram of a large-model-based legal question-answering method according to an embodiment of the present disclosure;

[0025] Figure 2 1 is a flowchart of a large-model-based legal question-answering method according to an embodiment of the present disclosure;

[0026] Figure 3 is a schematic diagram of a process for screening target corpus according to an embodiment of the present disclosure;

[0027] Figure 4 is a flowchart of generating a legal response result according to an embodiment of the present disclosure;

[0028] Figure 5 This is an overall architecture diagram of a large-model-based legal question-answering method according to an embodiment of the present disclosure;

[0029] Figure 6 2 is a schematic structural diagram of a large-model-based legal question-answering device according to an embodiment of the present disclosure;

[0030] Figure 7 It is a block diagram of an electronic device used to implement the large model-based legal question-answering method of an embodiment of the present disclosure. DETAILED DESCRIPTION

[0031] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0032] The terms "first," "second," and the like in this disclosure are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. Furthermore, the terms "including," "comprising," and "having," and any variations thereof, are intended to cover non-exclusive inclusions, such as, for example, inclusion of a series of steps or elements. A method, system, product, or apparatus is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such process, method, product, or apparatus.

[0033] It should be noted that, unless it is explicitly stated that there is a sequence of execution between different operations shown in the flowchart in the embodiments of the present disclosure, or there is a sequence of execution between different operations in technical implementation, otherwise, the execution order between multiple operations may not be prioritized, and multiple operations may also be executed simultaneously.

[0034] The legal field is a rigorous and specialized body of knowledge, characterized by a vast and complex multi-layered structure, with varying legal effects at different levels. Furthermore, legal language is precise and rigorous; even the slightest nuance in a single word, such as "shall," "may," or "shall not," can lead to completely different legal consequences. Legal logic is also rigorous, requiring rigorous logic at every step, from determining case facts to citing legal provisions to deriving conclusions. Otherwise, the conclusions could be erroneous. Furthermore, it requires a high level of professional knowledge and practical experience. These complex factors mean that the use of large models in the legal field requires not only speed and convenience but also a certain level of professional expertise.

[0035] In view of this, an embodiment of the present disclosure provides a legal question-answering method based on a large model.

[0036] The embodiments of the present disclosure provide Figure 1 An application scenario architecture 100 is shown to which the big model-based legal question answering method of the present disclosure can be applied.

[0037] like Figure 1 As shown, scenario architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 is used to provide a medium for communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables. Server 105 is provided with an agent 1051 for providing legal question-and-answer functionality. Server 105 or agent 1051 is used to execute the large model-based legal question-and-answer method provided in the embodiments of the present disclosure.

[0038] Users can use terminal devices 101, 102, and 103 to interact with an agent 1051 in a server 105 via a network 104 to receive or send messages, etc. Various applications for enabling information communication between the terminal devices 101, 102, and 103 and the server 105 can be installed, such as complex task processing applications, browser applications, and instant messaging applications.

[0039] Terminal devices 101, 102, 103 and server 105 can be either hardware or software. When terminal devices 101, 102, 103 are hardware, they can be various electronic devices with display screens, including but not limited to smartphones, tablet computers, laptop computers, and desktop computers. When terminal devices 101, 102, 103 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software or software modules, or as a single software or software module, and are not specifically limited here. When server 105 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules, or as a single software or software module, and are not specifically limited here.

[0040] The server 105 can provide various services through various built-in applications. Taking the complex task processing application as an example, the server 105 can achieve the following effects when running the complex task processing application: the server 105 can store and maintain various legal knowledge resources; after receiving the natural language expression of the legal issue sent by the target object through the terminal device, the server 105 can rewrite the natural language expression to obtain at least one piece of information to be retrieved suitable for retrieving relevant legal knowledge resources, and then use a multi-way recall method to retrieve the legal knowledge resources to obtain multiple candidate corpora; select the target corpora that meet the requirements from the multiple candidate corpora, and then provide them to a large number of large response models to refer to the target corpora to give legal response results.

[0041] Furthermore, the server 105 may also transmit the legal response result back to the terminal devices 101, 102, 103 via the network 104, so that the terminal devices 101, 102, 103 display the received legal response result to the user to complete the legal question and answer.

[0042] It should be noted that, in addition to being obtained from the terminal devices 101, 102, and 103 via the network 104, any form of natural language input can also be pre-stored locally on the server 105 in various ways. Therefore, when the server 105 detects that such data is already stored locally (for example, when starting to process a previously reserved task), it can choose to directly obtain such data locally. In this case, the exemplary scenario architecture 100 may also not include the terminal devices 101, 102, and 103 and the network 104.

[0043] Since processing complex tasks requires more computing resources and stronger computing power, the legal question-and-answer method based on a big model provided in the subsequent embodiments of this disclosure is generally executed by a server 105 with stronger computing power and more computing resources. Correspondingly, the legal question-and-answer device based on a big model is generally also set in the server 105.

[0044] It should be understood that Figure 1 The numbers of terminal devices, networks, and servers are only for reference. The corresponding numbers may be set based on implementation needs.

[0045] It should be noted that the acquisition, storage and application of the target object's personal information involved in the technical solution of this disclosure are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0046] like Figure 2 The figure is a flow chart of the legal question-answering method based on a large model provided by an embodiment of the present disclosure. The method can be Figure 1 The server 105 shown executes, including the following:

[0047] S201, obtaining a natural language expression of a legal issue provided by a target object.

[0048] This means obtaining legal questions for response provided by the target subject in any form of natural language. Examples include inquiries about regulatory cases, popularization of required legal knowledge, acquisition of legal information, and consultation on legal action recommendations. This natural language expression can be voice or text, and the disclosed embodiments do not limit the form of natural language expression.

[0049] S202: When the natural language expression satisfies the rewriting condition, the natural language expression is input into the rewriting model to obtain at least one information to be retrieved output by the rewriting model.

[0050] During implementation, the acquired natural language expressions are analyzed to determine whether they meet the rewriting criteria for legal questions and answers. If they meet the rewriting criteria, they are input into the rewriting model, which converts them into a format more suitable for legal knowledge resource retrieval, outputting at least one piece of information to be retrieved. For example, complex natural language expressions can be broken down according to different legal fields, simplified, key information extracted from natural language expressions can be expanded, and ambiguous terms in natural language expressions can be converted into more specific legal terms.

[0051] S203: Recall multiple candidate corpora that match at least one piece of information to be retrieved by using a multi-way recall method.

[0052] Candidate corpus refers to various possible legal knowledge resources related to the information to be retrieved, such as legal provisions, cases, interpretations of legal provisions, etc.

[0053] During implementation, multiple candidate corpora matching at least one piece of information to be retrieved are retrieved from a large corpus knowledge base through various recall methods. The corpus knowledge base encompasses a wide range of legal text materials, including but not limited to legal and regulatory texts, judicial case databases, legal academic literature, articles from professional legal information platforms, and lawyers' practical experience summaries.

[0054] S204, screening out multiple target corpora that meet preset conditions from multiple candidate corpora; the preset conditions include relevance indicators and optimization indicators; the optimization indicators are used to screen out corpora that meet authority requirements and / or timeliness requirements; the relevance indicators are used to screen out corpora that meet the requirements for semantic relevance to the corresponding information to be retrieved.

[0055] That is, from the recalled multiple candidate corpora, multiple target corpora are obtained by screening according to the relevance index and optimization index in the preset conditions.

[0056] Among them, the relevance index is used to measure the semantic relevance between the candidate corpus and the corresponding information to be retrieved. During implementation, the candidate corpus that is semantically closely related to the information to be retrieved can be screened out by calculating semantic similarity, keyword matching, and other methods. The optimization index is used to screen out candidate corpora that meet the authority requirements and / or timeliness requirements. The authority requirement can ensure that the selected target corpus comes from authoritative sources, such as officially issued laws and regulations, authoritative judicial interpretations, and works and interpretations of well-known legal experts; the timeliness requirement ensures that the selected target corpus is the latest and in line with current legal practices and legal changes, and avoids the use of outdated or repealed legal provisions or cases.

[0057] S205: Input multiple target corpora and natural language expressions into the legal response model to obtain legal response results for the natural language expressions.

[0058] The selected target corpora are input into the legal response model together with the original natural language expression, and the legal response model generates legal response results for the natural language expression.

[0059] In the disclosed embodiment, the large legal response model can be a large language model that has been trained with a large amount of legal text data. The Large Language Model (LLM) is an artificial intelligence model based on deep learning technology, with large-scale parameters and powerful language understanding and generation capabilities. It can understand the natural language expression of legal issues and, combined with relevant information in the target corpus, generate a response result for the natural language expression of the legal issue. The legal response results obtained for the natural language expression include but are not limited to answering legal questions, providing relevant legal basis, analyzing possible legal consequences, and giving treatment suggestions.

[0060] The legal response model can also be a multimodal model. A multimodal model can include an encoder for text and an encoder for media resources, such as images and videos. After being encoded by the encoders in different modalities, the decoder processes the encoded data to produce the legal response.

[0061] In the disclosed embodiment, the natural language expression provided by the target object is converted into more professional and accurate information to be retrieved through a rewriting model. The rewriting model can adapt to different user groups, and the user may be a user who does not understand legal profession and legal terminology. Through the rewriting model, the natural language expression can be made more suitable for the legal field, so as to accurately recall relevant content from a large amount of corpus, avoid retrieval bias caused by the ambiguity or diversity of natural language, and thus improve the accuracy of the subsequently generated legal response results. Furthermore, a multi-way recall method is used to recall multiple candidate corpora that match at least one piece of information to be retrieved, and candidate corpora can be obtained from multiple angles and sources, ensuring the comprehensiveness of the information. At the same time, the use of relevance indicators and optimization indicators to screen the target corpus can ensure that the corpus is closely related to the problem and that it is authoritative and / or timely, so that the final legal response result can screen out target corpora with sufficient and reliable basis from a large amount of candidate corpora as much as possible for the legal response large model to perform reasoning and analysis. In summary, the methods provided by the disclosed embodiments facilitate the raising of legal questions by target clients, eliminating the need for specialized legal knowledge or specialized expression methods. This saves time and effort for target clients, provides a positive service experience, and helps improve their satisfaction with and trust in legal consulting services. By optimizing the entire process of model rewriting, multi-channel recall, and corpus screening, the generation of legal response information is enhanced, improving the professionalism and practicality of responses.

[0062] During implementation, once a natural language expression regarding a legal issue is obtained from a target subject and satisfies a rewriting condition, the natural language expression can be input into the rewriting model to obtain at least one piece of information to be retrieved as output by the rewriting model. The rewriting condition may include at least one of the following:

[0063] (1) The length of the natural language expression exceeds the length threshold;

[0064] The length threshold is a pre-set criterion for measuring the length of a natural language expression. If the number of characters or words in the natural language expression regarding a legal issue entered by the target exceeds the pre-set length threshold, the natural language expression can be input into the rewriting model, which then outputs the information to be retrieved.

[0065] For example, if the length threshold is set to n characters, and the target user inputs a natural language expression of m characters, and m exceeds the length threshold of n characters, then the natural language expression needs to be rewritten. Where n and m are both positive integers greater than 0.

[0066] For example, the natural language expression provided by the target object regarding a legal issue is a detailed description of legal field A, but it contains a lot of background information and irrelevant details. Therefore, the length of the language expression exceeds the set length threshold. The rewriting model can shield the redundant language expressions and convert the language expression into information to be retrieved that can be recognized by the legal response model.

[0067] (2) Natural language expressions encompass multiple legal fields.

[0068] Since the law involves different sub-fields, when the target object provides a natural language expression involving multiple different legal fields, the natural language expression can be input into the rewriting model, and the rewriting model outputs the information to be retrieved.

[0069] For example, a target subject might provide a natural language expression about a legal issue that involves both legal fields A and B. Since this type of natural language expression encompassing multiple legal fields is complex, the rewriting model can be used to break it down and rewrite it into more targeted information to be retrieved, specifically for each legal field. This allows for more accurate recall of candidate corpora related to the corresponding fields from the corpus knowledge base.

[0070] In the disclosed embodiments, longer natural language expressions may contain a large amount of redundant information. Rewriting them by setting a length threshold can remove unnecessary details and highlight key information, making it easier to search for core corpora of concern, thereby improving the efficiency of the legal response model and the professionalism of the responses. Rewriting natural language expressions that contain multiple legal fields can classify and organize the content according to different legal fields, and rewrite it into more targeted information to be retrieved for different legal fields, thereby improving the accuracy and pertinence of recalling candidate corpora through retrieval information and generating legal responses through the legal response model.

[0071] In the embodiment of the present disclosure, when any of the above conditions are met for rewriting, the rewriting model can rewrite the natural language expression in the corresponding form according to the situation, including but not limited to semantic rewriting, problem decomposition and keyword expansion. Among them, semantic rewriting can be understood as rewriting part of the semantics of the natural language expression so that the semantics meet the description requirements of the legal profession. Problem decomposition can be understood as decomposing the nested description sentences in the natural language expression into independent questions. When implemented, it can be decomposed according to the legal field to obtain independent questions in different legal fields. Keyword expansion, that is, keyword expansion of questions in different fields, so that in the retrieval stage, different keywords can be used to recall relevant corpus as much as possible to improve the retrieval quality. When implemented, the problem decomposition can be performed first, and then the semantic rewriting and keyword expansion can be performed on each question obtained by decomposition.

[0072] After obtaining at least one information to be retrieved output by the rewriting model, a multi-way recall method can be used to recall multiple candidate corpora that match the at least one information to be retrieved. The multi-way recall method includes at least two of the following:

[0073] (1) Full-text search based on keywords;

[0074] That is, keywords are extracted from the information to be retrieved, and then a full-text search is performed in the entire corpus knowledge base to obtain multiple candidate corpora that match at least one piece of information to be retrieved.

[0075] (2) Retrieval method based on semantic vectors;

[0076] This involves converting the information to be retrieved into a semantic vector and then recalling multiple candidate texts from the corpus that are similar to the semantic vector. For example, a Bidirectional Generalized Encoder (BGE) model, a text-to-vector model, is used to convert the information to a semantic vector. The similarity between the semantic vector of the information to be retrieved and the semantic vectors of text in the corpus is calculated, and texts semantically related to the information to be retrieved are retrieved based on the similarity.

[0077] (3) Real-time recall method based on real-time web page library.

[0078] That is, obtaining content related to the information to be retrieved from a web page library that is updated in real time.

[0079] During implementation, a search can be conducted in a real-time updated webpage library based on the information to be retrieved, and multiple real-time candidate corpora that match the information to be retrieved can be retrieved.

[0080] AIAPI (Artificial Intelligence Application Programming Interface) can be used to retrieve relevant candidate corpora from relevant web pages in real time.

[0081] It should be noted that regardless of which of the above recall methods is used, the candidate corpora recalled will include their sources. These sources include, but are not limited to, authoritative legal databases, official legal platforms, professional academic resources, and rigorously reviewed legal forums. When the information to be retrieved encompasses multiple legal fields, the corpora are divided according to the different legal fields. Furthermore, the corpora are divided according to the quality level of the legal sources, for example, ranking them from high to low by authoritative legal databases, official legal platforms, professional academic resources, and rigorously reviewed legal forums.

[0082] In the disclosed embodiment, the full-text search method based on keywords can directly perform accurate matching in the entire corpus according to the keywords in the information to be retrieved, and quickly find the candidate corpus containing the relevant keywords. The search method based on semantic vectors can understand the semantics of the information to be retrieved, and improve the comprehensiveness of the recalled candidate corpus. The real-time recall method based on the real-time web page library can timely obtain the latest relevant information on the Internet in real time, making the recalled candidate corpus more timely and avoiding the answering and handling of legal issues due to untimely information updates. The combination of multiple recall methods can fully utilize the advantages of various retrieval methods, reduce the performance bottlenecks that may be brought about by a single retrieval method, and thus improve the efficiency, accuracy and real-time performance of recalling candidate corpora.

[0083] Since the number of candidate corpora obtained by the multi-way recall method is usually large, some of them may be not highly relevant to the information to be retrieved or inaccurate. In order to further make the legal response results generated by the legal response model more accurate and better meet the needs of the target audience.

[0084] During implementation, multiple target corpora that meet the preset conditions can be screened out from multiple candidate corpora. The specific implementation method is as follows: Figure 3 As shown, including the following:

[0085] S301: Deduplication is performed on multiple candidate corpora to obtain remaining corpora after deduplication.

[0086] Deduplication refers to identifying and removing duplicate content from multiple candidate corpora, and retaining only the content of the corpora with higher relevance.

[0087] During implementation, multiple candidate corpora can be deduplicated based on the following steps to obtain the remaining corpora after deduplication:

[0088] Step A1, determining the semantic similarity between multiple candidate corpora;

[0089] Semantic similarity refers to the degree of similarity between different candidate corpora at the semantic level.

[0090] During implementation, the semantic similarity between multiple candidate corpora can be determined based on the corresponding natural language processing technology. For example, each candidate corpus can be converted into a semantic vector, and then the similarity between the semantic vectors of the candidate corpora can be calculated as the semantic similarity between the candidate corpora.

[0091] In step A2, based on semantic similarity, multiple candidate corpora are deduplicated to obtain the remaining corpora.

[0092] During implementation, after determining the semantic similarity between multiple candidate corpora, a specific similarity threshold can be set according to the actual business situation. When the semantic similarity between two candidate corpora is higher than the specific threshold, the two candidate corpora are considered to be semantically duplicated or nearly duplicated. For candidate corpora with excessively high semantic similarity, only one of them is retained, and the other similar corpora are removed. In this way, multiple candidate corpora are deduplicated to obtain the remaining corpora. Among them, the corpora to be retained can be screened out based on optimization indicators and relevance quality. For example, corpora with high authority, high relevance to the corresponding information to be retrieved, and high timeliness are given priority.

[0093] In the disclosed embodiment, deduplication is performed by determining the semantic similarity between the multiple recalled candidate corpora, which can remove candidate corpora with substantially the same or highly similar content, thereby avoiding a large amount of repeated information in each candidate corpus, making the remaining corpus obtained after deduplication more refined, and more accurately reflecting the information differences contained in different corpora, avoiding interference caused by repeated corpora in the subsequent processing process, thereby improving the accuracy and reliability of the legal response results.

[0094] S302: Obtain ranking scores of the plurality of first intermediate corpora based on preset conditions.

[0095] Based on the content described above, the pre-set conditions include relevance indicators and optimization indicators. By analyzing and evaluating the candidate corpora according to the pre-set conditions, a ranking score is calculated for each candidate corpus. This ranking score reflects the comprehensive degree of semantic relevance, authority, and timeliness of the corpus to the information being retrieved.

[0096] During implementation, obtaining the ranking scores of the plurality of first intermediate corpora based on the preset conditions can be achieved by following the steps below:

[0097] Step B1: for each candidate corpus, determine the rough ranking score of the candidate corpus in the multi-way recall results.

[0098] The coarse ranking score is positively correlated with the relevance score, negatively correlated with the candidate's ranking position in each recall, and positively correlated with the optimization index. The relevance score is determined based on the candidate and the corresponding information to be retrieved. In other words, the relevance score expresses the degree of relevance between the candidate and the corresponding information to be retrieved.

[0099] The relevance score is determined based on the candidate corpus and the corresponding information to be retrieved. The coarse ranking score is positively correlated with the relevance score. That is, the higher the match between the candidate corpus and the corresponding information to be retrieved, the higher its relevance score. When all other conditions are equal, the coarse ranking score will also be higher.

[0100] The coarse ranking score is negatively correlated with the candidate's ranking position in each recall path. That is, during a multi-path recall, each candidate will have a ranking position under different recall paths. The higher the candidate's ranking position, the more it meets the requirements under the corresponding recall path, and the higher its coarse ranking score. For example, if a candidate ranks relatively high in both keyword-based recall and semantic-based recall, and all other conditions are equal, then its coarse ranking score in the multi-path recall results will be higher.

[0101] The optimization index is used to select corpora that meet the authority and / or timeliness requirements. It is positively correlated with the optimization index. That is, the higher the authority and timeliness of a candidate corpus, the higher its rough ranking score will be, assuming all other conditions are equal.

[0102] During implementation, the rough ranking scores of candidate corpora in multi-channel recall results can be determined based on the RRF (Reciprocal Rank Fusion) method combined with corpus weights, so as to facilitate preliminary screening of recalled corpora based on the rough ranking scores, giving priority to corpora with authoritative sources and strong timeliness.

[0103] For example, for each candidate corpus, a correlation score between the candidate corpus and the corresponding search information can be determined. Specifically, when the candidate corpus is retrieved by a full-text search method based on keywords or a real-time recall method based on a real-time web page library, the correlation score between the keyword and the candidate corpus can be determined based on an algorithm such as TF-IDF (Term Frequency-Inverse Document Frequency). When the candidate corpus is retrieved by a search method based on a semantic vector, the similarity between the candidate corpus and the semantic vector used to retrieve the candidate corpus is determined to obtain a correlation score between the candidate corpus and the corresponding information to be retrieved.

[0104] In addition to calculating the relevance score of the candidate corpus, it is also necessary to calculate the optimization score of the optimization index of the candidate corpus, which includes the timeliness score and the source authority score.

[0105] For timeliness scores, we can use the current time as a benchmark to divide the data into multiple time windows. These windows are then normalized to obtain timeliness scores for each window. For each candidate corpus, we determine the time window that corresponds to the candidate corpus and then use the timeliness score for that window as the candidate's timeliness score.

[0106] For source authority scores, authority levels for different sources can be pre-established, with different authority levels corresponding to their respective source authority scores. For each candidate corpus, the source corresponding to the candidate corpus is obtained, and based on the authority level corresponding to the source, the source authority score corresponding to the candidate corpus is obtained. The authority levels of the same source can be divided into the same level, and different sub-sections within the same source can also be divided into their own authority levels. Taking a website as an example, the authority level of the same website can be divided into the same level, and different columns within the same website can also be divided into their own authority levels.

[0107] Then, the corpus weight of each candidate corpus is calculated based on its timeliness score, relevance score, and source authority score. The specific calculation formula is shown in formula (1):

[0108] W i =α*F1+β*F2+γ*F3 (1)

[0109] In formula (1), W iis the corpus weight of the i-th candidate corpus; F1 is the timeliness score of the candidate corpus, α is the timeliness score weight of the candidate corpus; F2 is the relevance score of the candidate corpus, β is the relevance score weight of the candidate corpus; F3 is the source authority score of the candidate corpus, and γ is the source weight of the candidate corpus. The setting ratio of each weight can be set according to the actual situation.

[0110] Based on the RRF algorithm and the corpus weight of the candidate corpus, multiple ranking results are fused and the relevance of the ranking results is adjusted to obtain the rough ranking score of the candidate corpus in the multi-way recall results, which can be described by formula (2):

[0111]

[0112] In formula (2), RRF i represents the rough ranking score of the i-th candidate corpus; W i is the corpus weight of the i-th candidate corpus; j is the ranking position of the i-th candidate corpus in the j-th recall path; n represents the total number of recall paths, that is, if there are n-way recall methods, there are corresponding n-way recall paths.

[0113] In step B2, the natural language expression and the plurality of candidate corpora are input into a semantic re-ranking model to obtain the semantic scores of the plurality of candidate corpora.

[0114] The semantic reranking model can be selected using BGE-Reranker (Bidirectional Generator Encoder Reranker), a model trained to understand the semantic relationship between natural language expressions and candidate corpora. After the natural language expression and each candidate corpus are input into the model, the model calculates a semantic score for each candidate corpus based on factors such as the degree of semantic match between the two. This semantic score reflects the semantic fit between the candidate corpus and the natural language expression. For example, the more accurate and comprehensive the candidate corpus's explanation of the legal issues involved in the natural language expression, the higher its semantic score.

[0115] Step B3: Fusing the rough ranking scores and semantic scores of the multiple candidate corpora to obtain ranking scores of the multiple candidate corpora.

[0116] To more accurately evaluate candidate corpora by comprehensively considering both the coarse ranking score and the semantic score, these two scores need to be fused. This fusion can be achieved through methods such as weighted summation, where the coarse ranking score and the semantic score are weighted to produce a composite ranking score. This ranking score more comprehensively reflects the quality of the candidate corpus and its degree of match with the information being retrieved, providing a more reliable basis for subsequently selecting target corpora from multiple candidate corpora.

[0117] In the disclosed embodiment, by comprehensively considering multiple factors of the candidate corpus in the multi-path recall results, it is possible to preliminarily measure the importance and relevance of the candidate corpus from multiple dimensions based on the rough ranking score. The correlation index and optimization index between each information to be retrieved and the recalled candidate corpus are taken into account in the rough ranking score, and different candidate corpora can be comprehensively considered from different information to be retrieved and different recall paths. Furthermore, since the information to be retrieved is obtained by rewriting the natural language, the association between the natural language expression and the candidate corpus input is further considered through the semantic reranking model, and the degree of matching between the candidate corpus and the natural language expression can be deeply understood from the semantic level of the natural language expression. The fusion processing of the rough ranking score and the semantic score can fully utilize the advantages of both. The rough ranking score provides branching information and macro-ranking information based on the multi-path recall results, while the semantic score is refined and corrected from the micro-semantic level. The fusion of these two scoring methods can make the final ranking score more comprehensively reflect the true relevance of the candidate corpus to the information to be retrieved, avoiding the one-sidedness that may be caused by a single scoring method, thereby improving the accuracy of corpus sorting, and helping to more efficiently screen out target corpora that are highly matched with the needs of the target objects, providing better basic data for subsequent tasks such as legal response generation, thereby improving the accuracy and professionalism of legal response results.

[0118] It should be noted that, in order to retain as many candidate corpora with different contents as possible and to screen out candidate corpora with high scores, step S301 and step S302 can be performed in parallel.

[0119] S303: Filter out at least one intersection corpus from the intersection of the first intermediate corpus and the remaining corpus.

[0120] The first intermediate corpus is the set of corpora with the highest ranking scores among multiple candidate corpora. The remaining corpus is the set of corpora remaining after deduplication of multiple candidate corpora. The intersection corpus is the portion that belongs to both the first intermediate corpus and the remaining corpus. By screening the intersection corpus, we can find corpora that meet the initial screening criteria and have no duplicates. These corpora are more reliable in subsequent processing.

[0121] S304 : When at least one intersection corpus does not meet the preset number of items, supplementary corpora are screened out from the first intermediate corpus according to the ranking scores to obtain a plurality of target corpora including the supplementary corpora and the intersection corpus.

[0122] The preset number of entries is a standard set based on actual needs, used to determine the required number of target corpora for the final screening. If the number of intersection corpora does not meet the preset number of entries, it means that the intersection corpus alone cannot reach the required number of corpora. In this case, supplementary corpora are selected from the first intermediate corpus in descending order of ranking score to ensure that the quality of the resulting multiple target corpora meets the preset requirements and provides suitable corpus support for subsequent input into the large legal response model.

[0123] In the disclosed embodiment, by deduplicating multiple candidate corpora, duplicate candidate corpora can be removed, data redundancy can be reduced, and multiple analyses and screenings of the content in the same candidate corpora can be avoided in subsequent processing. While improving processing efficiency, the remaining corpora can also be made more refined and more accurately reflect different key information points. By assigning ranking scores to multiple first intermediate corpora according to preset conditions, the corpora can be quantitatively evaluated and ranked according to the degree of matching with the preset conditions, which facilitates the subsequent screening of the corpora that best meet the requirements. Deduplication and sorting are performed in parallel, and the differences between the candidate corpora can be retained while screening out candidate corpora with high scores. By screening out the intersection corpora from the first intermediate corpora and the remaining corpora, corpora that meet the preset conditions and have been deduplicated can be found. These corpora have high credibility in terms of quality and relevance, are important components of the target corpora, and can provide high-quality data sets for the subsequent legal response model to generate legal response results. When the intersection corpus does not meet the preset number of items, supplementary corpus is selected from the first intermediate corpus according to the ranking points, which can ensure that the final number of target corpora meets the preset requirements. While ensuring the integrity of the corpus, it avoids information loss due to insufficient intersection corpus, and provides sufficient data support for the subsequent legal response model to generate legal response results, so as to improve the accuracy and professionalism of the response.

[0124] In the embodiment of the present disclosure, multiple target corpora and natural language expressions are input into the legal response model to obtain legal response results for the natural language expressions. The specific implementation method is as follows: Figure 4 As shown, including the following:

[0125] S401, generating prompt words based on multiple target corpora and natural language expressions; the prompt words include logical self-consistent constraints.

[0126] The target corpus is high-quality, relevant to the information being retrieved, obtained through the aforementioned series of screening operations. The natural language expression is the target's initial description of the legal issue. Based on the target corpus and the natural language expression, prompt words are generated and used to guide the legal response model in responding.

[0127] Among them, the prompt words include logical self-consistency constraints, which are to ensure that the generated legal response is logically coherent, reasonable and non-contradictory.

[0128] For example, it can be set up to follow a specific chain of thought patterns. For example, first, the legal response model is required to understand the core and specific needs of the target object regarding legal issues. Next, the corpus is screened from the target corpus in combination with the corpus knowledge base and background knowledge. Then, the selected corpus and questions are analyzed in depth. After that, answers are given from multiple perspectives based on the analysis results. Finally, the answers are reviewed and optimized, and the accuracy, completeness, and logic of the answers are checked, and the answers are improved to ensure that high-quality responses are provided to the target object. In this way, when there are different legal viewpoints or treatment methods regarding the natural language expressions of legal issues provided by the target object in the target corpus, the logical self-consistency constraints will prompt the large model to reasonably integrate and interpret when generating responses, avoiding inconsistencies.

[0129] S402: Input the prompt words into the legal response macromodel, so that the legal response macromodel can select reference corpora from multiple target corpora based on logically consistent constraints to generate logically consistent legal response results.

[0130] When a prompt containing logically consistent constraints is input into the legal response model, the model will filter out at least one of the most relevant and reliable pieces of target text from the multiple target texts based on the logically consistent constraints as a reference text. Then, based on this selected reference text, it generates a logically coherent, reasonable, and accurate legal response.

[0131] In the disclosed embodiments, by generating prompts containing logically consistent constraints from multiple target corpora and natural language expressions, the large legal response model can be prevented from being confused or erroneously identified by excessive information. This allows the large legal response model to more accurately select, from the multiple target corpora, reference corpora relevant to the natural language expressions of legal issues provided by the target subject. Furthermore, the logically consistent constraints ensure that the generated legal response conforms to legal logic and reasoning rules, and is logically coherent, reasonable, and consistent, thereby enhancing the accuracy and professionalism of the generated legal response.

[0132] During implementation, in order to further enhance the practicality and accuracy of the generated legal response results, the generated legal response results also include generated legal interpretations and similar cases; and, the old legal provisions in the legal response results are corrected to the latest version to obtain the final answer used to respond to the target object.

[0133] That is, in the generated legal response results, legal interpretations and similar cases related to the natural language expression of legal issues provided by the target object are generated, and old legal provisions that have become invalid are verified and corrected to ensure that the latest versions of laws and regulations are used, and to ensure the effectiveness and accuracy of the cited legal knowledge content.

[0134] During implementation, in the legal response, the target audience must be clearly explained the source, specific meaning, scope of application, legislative purpose, etc. of the relevant legal provisions involved. This will help the target audience better understand the legal provisions and understand the specific basis for judging their legal issues within the legal framework.

[0135] As for references to old legal provisions, an update library can be pre-established during implementation, which can maintain changes to legal provisions. The legal provisions involved in the response results can be reviewed based on the pre-maintained legal provision update library, and by comparing them with the latest legal database, officially released legal documents, etc., old legal provisions that have been modified, repealed, or have new judicial interpretations issued can be identified. If it is determined to be an old legal provision, it will be replaced and revised in accordance with the latest legal provisions, so that the references and interpretations involving the legal provision are consistent with the latest version, so that the final response content accurately reflects the currently effective legal provisions, avoiding the target object receiving incorrect or inaccurate legal information due to citing old provisions.

[0136] The purpose of generating similar cases is to provide the target audience with a more intuitive understanding of how similar legal issues were handled and the outcomes they achieved in actual judicial practice. This process can be implemented by searching past judicial precedents, guiding cases, or other relevant cases to identify cases that share similarities with the current legal issue in terms of factual circumstances and legal relationships. By presenting the basic facts, decision reasons, and outcomes of these cases to the target audience, the target audience can further understand the application of legal provisions in specific contexts through real-world examples, thereby enhancing the persuasiveness and practicality of legal responses.

[0137] In the disclosed embodiment, providing legal interpretations in the legal response can make the response more professional and authoritative, so that the target object can understand the specific basis for judging its legal issues within the legal framework. Correcting old legal provisions to the latest version can avoid the inclusion of erroneous legal advice in the generated legal response results due to outdated laws. The generated similar cases can apply abstract legal provisions to specific practical scenarios, so that the target object can more intuitively understand how the law is applied in similar situations and the possible legal consequences; at the same time, different cases may present a variety of handling methods and response strategies, which can also provide the target object with ideas and methods for solving problems. The combination of legal interpretations and similar cases in the generated legal response results can be argued from both the legal provisions and practical applications levels, making the generated legal response results more convincing.

[0138] In summary, the overall framework of the legal question-answering method provided in the embodiments of the present disclosure is as follows: Figure 5 As shown in Figure 1, it includes four parts: pre-retrieval layer, retrieval layer, enhancement layer, and generation layer. The following is an explanation of each part:

[0139] 1. Pre-retrieval layer

[0140] That is, the natural language expressions about legal issues provided by the target object are preprocessed, including semantic rewriting, problem decomposition, keyword expansion, etc., and the natural language expressions are rewritten into information to be retrieved that can be recognized by the legal response model.

[0141] 2. Retrieval Layer

[0142] That is, a multi-way recall method is used to recall multiple candidate corpora that match at least one piece of information to be retrieved, which may include:

[0143] (1) A keyword-based full-text search method extracts keywords from the information to be retrieved, and then performs a full-text search in the entire corpus to obtain multiple candidate corpora that match at least one piece of information to be retrieved;

[0144] (2) The retrieval method based on semantic vectors uses BGE vector retrieval technology to retrieve multiple candidate corpora similar to the semantic vector of the information to be retrieved from the corpus.

[0145] (3) A real-time recall method based on a real-time web page library obtains content related to the information to be retrieved from the real-time updated web page library.

[0146] 3. Enhancement Layer

[0147] This involves screening and optimizing the recalled candidate corpora to ensure the final legal response is authoritative and accurate, including:

[0148] (1) Through the rough ranking model, the candidate corpus to be recalled is preliminarily screened according to the relevance index and optimization index of the corpus using the reciprocal ranking fusion (RRF) method.

[0149] (2) Through the refined ranking model, the candidate corpus is further screened through the semantic reranking model (BGE-Reranker) to ensure that the candidate corpus is highly relevant to the semantics of the legal issues provided by the target object.

[0150] The target corpus is selected through a deduplication operator and a fusion operator. Deduplication removes candidate corpora with high semantic similarity. Sorting and deduplication are performed in parallel, selecting high-scoring candidate corpora while preserving the differences between them. Fusion combines the deduplicated candidate corpora with the final ranked candidate corpora. Merging involves sorting the deduplicated candidate corpora to obtain a certain amount of corpus content, resulting in the target corpus.

[0151] 4. Generate Layer

[0152] The generation layer is responsible for calling the legal response model based on the filtered target corpus to generate the final legal response results and perform post-processing, which may include:

[0153] (1) Generate prompt words based on multiple target corpora and natural language expressions; the prompt words include logically consistent constraints.

[0154] (2) Inputting the prompt words into the legal response model, so that the legal response model can select reference corpora from multiple target corpora based on logically consistent constraints to generate legal response results, and generate logically consistent legal response results.

[0155] (3) Post-regulatory matching, that is, generating legal interpretations in legal response results.

[0156] (4) Post-case matching, which is to generate similar cases in the legal response results.

[0157] (5) Regulatory error correction zipper, which corrects the old legal provisions in the legal response results to the latest version to ensure the effectiveness and accuracy of the cited legal knowledge content.

[0158] In summary, the legal question-and-answer method provided by the disclosed embodiments ensures the authority, timeliness, and accuracy of generated legal responses through multi-channel recall, corpus screening, and generation optimization. By combining keyword matching, semantic matching, and real-time recall, it comprehensively covers the natural language corpus sources provided by the target subject regarding legal issues. Through coarse and fine sorting models, it ensures the quality of the screened corpus. Through post-matching of regulations and cases, it enhances the practicality of generated legal responses. Through a regulatory error correction zipper, it ensures the accuracy of generated legal responses. This method effectively supports intelligent question-and-answer and knowledge services in the legal field.

[0159] Based on the same technical concept, the embodiment of the present disclosure also provides a legal question-answering device 600 based on a large model, such as Figure 6 Shown, including:

[0160] Acquisition module 601, for acquiring a natural language expression of a legal issue provided by a target object;

[0161] A rewriting module 602 is configured to input the natural language expression into a rewriting model when the natural language expression satisfies the rewriting condition, and obtain at least one information to be retrieved output by the rewriting model;

[0162] A recall module 603 is configured to recall multiple candidate corpora that match at least one piece of information to be retrieved using a multi-way recall method;

[0163] The screening module 604 is used to screen out multiple target corpora that meet preset conditions from multiple candidate corpora. The preset conditions include a relevance index and an optimization index. The optimization index is used to screen out corpora that meet authority requirements and / or timeliness requirements. The relevance index is used to screen out corpora that meet the required semantic relevance to the corresponding information to be retrieved.

[0164] The response module 605 is used to input multiple target corpora and natural language expressions into the legal response model to obtain legal response results for the natural language expressions.

[0165] In some embodiments, the screening module comprises:

[0166] The deduplication unit is used to perform deduplication operations on multiple candidate corpora to obtain the remaining corpora after deduplication;

[0167] a ranking score determining unit, configured to obtain a ranking score for each of the plurality of first intermediate corpora based on a preset condition;

[0168] A first screening unit is configured to screen out at least one intersection corpus from the intersection of the first intermediate corpus and the remaining corpus;

[0169] The second screening unit is configured to screen out supplementary corpora from the first intermediate corpora according to the ranking scores when at least one intersection corpus does not meet a preset number of items, so as to obtain a plurality of target corpora including the supplementary corpora and the intersection corpora.

[0170] In some embodiments, the ranking score determination unit is specifically configured to:

[0171] For each candidate corpus, determine its rough ranking score in the multi-channel recall results; the rough ranking score is positively correlated with the relevance score, negatively correlated with the candidate corpus's ranking position in each channel of recall, and positively correlated with the optimization index; the relevance score is determined based on the candidate corpus and the corresponding information to be retrieved;

[0172] Input the natural language expression and multiple candidate corpora into the semantic re-ranking model to obtain the semantic scores of the multiple candidate corpora;

[0173] The rough ranking scores and semantic scores of multiple candidate corpora are fused to obtain the ranking scores of multiple candidate corpora.

[0174] In some embodiments, the deduplication unit is specifically configured to:

[0175] Determine the semantic similarity between multiple candidate corpora;

[0176] Based on semantic similarity, multiple candidate corpora are deduplicated to obtain the remaining corpora.

[0177] In some embodiments, the multi-way recall method includes at least two of the following:

[0178] Full-text search method based on keywords;

[0179] Retrieval method based on semantic vectors;

[0180] Real-time recall method based on real-time web page library.

[0181] In some embodiments, the rewriting condition includes at least one of the following:

[0182] The length of the natural language expression exceeds the length threshold;

[0183] Natural language expressions encompass multiple legal domains.

[0184] In some embodiments, the reply module includes:

[0185] A generation unit is used to generate prompt words based on multiple target corpora and natural language expressions; the prompt words include logical self-consistent constraints;

[0186] The reply unit is used to input the prompt words into the legal reply model so that the legal reply model can filter out the reference corpus from multiple target corpora based on the logical self-consistency constraint conditions to generate a legal reply result, and generate a logically consistent legal reply result.

[0187] In some embodiments, the reply module is further configured to:

[0188] Generate legal interpretations and similar cases in legal response results; and

[0189] Correct the old legal provisions in the legal response results to the latest version, and obtain the final answer that is ultimately used to respond to the target object.

[0190] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0191] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0192] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0193] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0194] like Figure 7 As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0195] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0196] The computing unit 701 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as the legal question-answering method. For example, in some embodiments, the legal question-answering method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the legal question-answering method described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform the legal question-answering method by any other suitable means (e.g., by means of firmware).

[0197] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0198] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0199] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0200] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0201] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0202] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0203] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0204] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A legal question-answering method based on a large model, comprising: Obtaining natural language expressions about legal issues provided by the target object; If the natural language expression satisfies the rewriting condition, inputting the natural language expression into a rewriting model to obtain at least one information to be retrieved output by the rewriting model; Recalling multiple candidate corpora matching the at least one information to be retrieved by using a multi-way recall method; Filtering a plurality of target corpora that meet preset conditions from the plurality of candidate corpora; the preset conditions include a relevance index and an optimization index; the optimization index is used to filter corpora that meet authority requirements and / or timeliness requirements; the relevance index is used to filter corpora that meet the required semantic relevance to the corresponding information to be retrieved; The plurality of target corpora and the natural language expressions are input into a large legal response model to obtain a legal response result for the natural language expressions.

2. The method according to claim 1, wherein The step of selecting a plurality of target corpora that meet preset conditions from the plurality of candidate corpora includes: Performing a deduplication operation on the plurality of candidate corpora to obtain the remaining corpora after deduplication; and Obtaining ranking scores for each of the plurality of first intermediate corpora based on the preset conditions; Selecting at least one intersection corpus from the intersection of the first intermediate corpus and the remaining corpus; When the at least one intersection corpus does not meet the preset number of items, supplementary corpora are screened out from the first intermediate corpus according to the ranking points to obtain the plurality of target corpora including the supplementary corpora and the intersection corpus.

3. The method according to claim 2, wherein: The obtaining of ranking scores of the plurality of first intermediate corpora based on the preset condition includes: For each candidate corpus, determining a rough ranking score of the candidate corpus in the multi-way recall results; the rough ranking score is positively correlated with the relevance score, negatively correlated with the ranking position of the candidate corpus in each recall, and positively correlated with the optimization index; the relevance score is determined based on the candidate corpus and the corresponding information to be retrieved; Inputting the natural language expression and the plurality of candidate corpora into a semantic re-ranking model to obtain a semantic score for each of the plurality of candidate corpora; The rough ranking scores and semantic scores of the multiple candidate corpora are fused to obtain ranking scores of the multiple candidate corpora.

4. The method according to claim 2, wherein: The performing a deduplication operation on the plurality of candidate corpora to obtain the remaining corpora after deduplication includes: Determining semantic similarity between the plurality of candidate corpora; Based on the semantic similarity, a deduplication operation is performed on the multiple candidate corpora to obtain the remaining corpora.

5. The method according to claim 1, wherein the multi-way recall method includes at least two of the following: Full-text search method based on keywords; Retrieval method based on semantic vectors; Real-time recall method based on real-time web page library.

6. The method according to any one of claims 1 to 5, wherein The rewriting condition includes at least one of the following: The length of the natural language expression exceeds a length threshold; The natural language expression includes multiple legal fields.

7. The method according to any one of claims 1 to 6, wherein Inputting the plurality of target corpora and the natural language expressions into the legal response macro model to obtain a legal response result for the natural language expressions includes: generating prompt words based on the plurality of target corpora and the natural language expressions; wherein the prompt words include logical self-consistent constraints; The prompt words are input into the legal response model so that the legal response model selects reference corpora from the multiple target corpora based on the logical self-consistency constraint conditions to generate the logically consistent legal response result.

8. The method according to any one of claims 1 to 7, further comprising: Generate legal interpretations and similar cases in the legal response results; as well as, Correct the old legal provisions in the legal response result to the latest version, and obtain the final answer used to respond to the target object.

9. A legal question-answering device based on a large model, comprising: An acquisition module is used to obtain natural language expressions about legal issues provided by the target object; a rewriting module, configured to input the natural language expression into a rewriting model when the natural language expression satisfies a rewriting condition, and obtain at least one information to be retrieved output by the rewriting model; A recall module, configured to recall multiple candidate corpora matching the at least one information to be retrieved by adopting a multi-way recall method; a screening module for screening out a plurality of target corpora that meet preset conditions from the plurality of candidate corpora; the preset conditions include a relevance index and an optimization index; the optimization index is used to screen out corpora that meet authority requirements and / or timeliness requirements; the relevance index is used to screen out corpora that meet the required semantic relevance to the corresponding information to be retrieved; The reply module is used to input the multiple target corpora and the natural language expressions into the legal reply model to obtain a legal reply result for the natural language expressions.

10. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.

12. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.