Legal system class case retrieval method and system
By evaluating the ambiguity and semantic vector analysis of legal constituent elements, a class of cases with subtle differences was selected, which solved the problem of difficulty in capturing differences between cases in the existing technology, and improved the depth and accuracy of class of cases search.
Patent Information
- Application Number
- CN202510829676.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The existing case search system is difficult to capture the subtle differences between cases, resulting in differentiation of referee results and lack of differentiated search capabilities.
By determining the ambiguity of the legal components of the cases to be retrieved according to the law, based on semantic vector analysis, cases with the total similarity greater than the first threshold and the highest ambiguity less than the second threshold are screened to achieve similar case searches focusing on differences.
The identification and presentation of subtle differences affecting the judgment conclusions in generally similar cases has been achieved, supporting the judge to understand the application of the law more comprehensively and the defense lawyer to find more ways to defend.
Smart Images

Figure CN120353922A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information retrieval, and in particular to a method and system for retrieving similar cases in a legal affairs system. Background Art
[0002] In legal practice, retrieving similar cases refers to finding highly similar historical cases from a huge case database based on a certain legal case, so as to compare and analyze its legal application, fact determination, judgment logic, and even the final judgment result, which has irreplaceable important value in legal practice, research and study, etc.
[0003] Currently, the similar case retrieval systems in the prior art mainly provide "coarse-grained" similarity search, often centered around keyword matching, co-occurrence of legal provisions, case labels, identities of the parties, categories of basic case facts, etc. The similarity criteria relied on by this retrieval mechanism are relatively superficial, focusing on the proximity of cases in terms of structured attributes (such as basic case facts, charges) or text semantics, and lacking in-depth exploration and presentation of the "key differences" between cases. However, in judicial judgments, the differences in many cases are precisely hidden behind the "seemingly similar" cases. Such cases - which we can call cases where subtle differences lead to different judgment results - are the parts that existing similar case retrieval systems are difficult to capture.
[0004] Therefore, it is extremely necessary to propose a differential similar case retrieval scheme. The goal of this type of retrieval is no longer just to find the "most similar" cases, but to consciously identify and compare those subtle differences that may affect the judgment conclusion on the premise of general similarity (such as consistent basic case facts and legal application frameworks), and present these differences in an interpretable manner. This not only helps judges understand the application of each similar clause more comprehensively and deeply, but also enables defenders to find more ways of defense. Summary of the Invention
[0005] The present invention provides a method and system for retrieving similar cases in a legal affairs system to solve the defect that it is difficult to perform differential retrieval in the prior art's similar case retrieval methods.
[0006] The present invention provides a method for retrieving similar cases in a legal affairs system, including: Receiving a case query text corresponding to the case to be retrieved, which includes basic case facts and judgment results; Determining the ambiguity of each legal element of the law on which the case to be retrieved is based; the higher the ambiguity of any legal element, the higher the possibility that the any legal element includes multiple legal interpretations; Retrieving a case database based on the law on which the case to be retrieved is based to obtain an intermediate query result; Determine the similarity of each legal element corresponding to the any case and the case to be retrieved, and the total similarity of the any case and the case to be retrieved, based on the semantic vectors of the first fact text segments corresponding to each legal element in the basic case, the semantic vectors of the second fact text segments corresponding to each legal element in any case in the intermediate query result, and the ambiguity of each legal element; Filter out the cases from the intermediate query result whose total similarity with the case to be retrieved is greater than the first threshold and the similarity of the legal element corresponding to the highest ambiguity is less than the second threshold, to obtain the query result.
[0007] According to a method for retrieving similar cases in a legal system provided by the present invention, the ambiguity of any legal element of the case to be retrieved based on a legal provision is determined based on the following steps: Retrieve the set of cases in the case library whose legal provisions are the same as those of the case to be retrieved; For any legal element, determine the initial semantic vectors of the fact text segments corresponding to each case in the case set for the any legal element; Based on the differences between the initial semantic vectors of the fact text segments corresponding to each case for the any legal element, determine the ambiguity of the any legal element.
[0008] According to a method for retrieving similar cases in a legal system provided by the present invention, determining the ambiguity of the any legal element based on the differences between the initial semantic vectors of the fact text segments corresponding to each case for the any legal element includes: Cluster the initial semantic vectors of the fact text segments corresponding to each case for the any legal element to obtain one or more clusters; If there is only one cluster, determine the ambiguity of the any legal element as the first preset value; Otherwise, based on the number of clusters of which the distance between the cluster centers is greater than the preset threshold, determine the ambiguity of the any legal element.
[0009] According to a method for retrieving similar cases in a legal system provided by the present invention, the semantic vector of the first fact text segment or the second fact text segment corresponding to any legal element is determined based on the following steps: Based on a language model, obtain the initial semantic vector of the target fact text segment corresponding to any legal element; the target fact text segment is the first fact text segment or the second fact text segment; Input the initial semantic vector of the target fact text segment into a non-linear layer to obtain the directional vector of the target fact text segment; the directional vector has the same dimension as the initial semantic vector; For any dimension, take the vector value corresponding to this dimension in the initial semantic vector of the target factual text segment as the real part, and the vector value corresponding to this dimension in the directional vector as the imaginary part to obtain a complex value corresponding to this dimension; Combine the complex values corresponding to each dimension into the semantic vector of the target factual text segment.
[0010] According to a method for retrieving similar cases in a legal affairs system provided by the present invention, the similarity between any case and the case to be retrieved corresponding to any legal element is calculated based on the following steps: Calculate the similarity between the initial semantic vector of the second factual text segment corresponding to any legal element of any case and the initial semantic vector of the first factual text segment corresponding to any legal element of the case to be retrieved as the first similarity; Calculate the dot product of the conjugate vector of the semantic vector of the second factual text segment corresponding to any legal element of any case and the semantic vector of the first factual text segment corresponding to any legal element of the case to be retrieved; Determine the degree of difference based on the phase angle of the dot product; Based on the first similarity and the degree of difference, determine the similarity between any case and the case to be retrieved corresponding to any legal element.
[0011] According to a method for retrieving similar cases in a legal affairs system provided by the present invention, the total similarity between any case and the case to be retrieved is calculated based on the following steps: Based on the fuzziness of each legal element, perform a weighted sum of the similarities between any case and the case to be retrieved corresponding to each legal element to obtain the total similarity between any case and the case to be retrieved.
[0012] The present invention also provides a system for retrieving similar cases in a legal affairs system, including: A query text receiving unit for receiving a case query text corresponding to the case to be retrieved, including the basic case situation and the judgment result; A fuzziness evaluation unit for determining the fuzziness of each legal element of the case to be retrieved according to the legal provisions; the higher the fuzziness of any legal element, the higher the possibility that any legal element contains multiple legal interpretations; A first retrieval unit for retrieving a case library based on the legal provisions on which the case to be retrieved is based to obtain an intermediate query result; A similarity evaluation unit, configured to determine the similarity of each legal element corresponding to any case and the case to be retrieved, and the total similarity between any case and the case to be retrieved, based on the semantic vectors of the first factual text segments corresponding to each legal element in the basic case, the semantic vectors of the second factual text segments corresponding to each legal element in any case in the intermediate query result, and the ambiguity of each legal element; A second retrieval unit, configured to screen, from the intermediate query result, cases whose total similarity with the case to be retrieved is greater than a first threshold and the similarity of the legal element with the highest corresponding ambiguity is less than a second threshold, to obtain a query result.
[0013] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method for retrieving similar cases in the legal affairs system as described in any one of the above is implemented.
[0014] The present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for retrieving similar cases in the legal affairs system as described in any one of the above is implemented.
[0015] The present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for retrieving similar cases in the legal affairs system as described in any one of the above is implemented.
[0016] A method and system for retrieving similar cases in a legal affairs system provided by the present invention determine the ambiguity of each legal element of a case to be retrieved according to a legal provision, and retrieve a case library based on the legal provision on which the case to be retrieved is based, to obtain an intermediate query result. Then, based on the semantic vectors of the first factual text segments corresponding to each legal element in the basic case, the semantic vectors of the second factual text segments corresponding to each legal element in any case in the intermediate query result, and the ambiguity of each legal element, the similarity of each legal element corresponding to the case and the case to be retrieved, and the total similarity between the case and the case to be retrieved are determined. Furthermore, cases whose total similarity with the case to be retrieved is greater than a first threshold and the similarity of the legal element with the highest corresponding ambiguity is less than a second threshold are screened from the intermediate query result as query results and returned to the user. By introducing legal element ambiguity evaluation, legal element-level comparison based on semantic vectors, and a screening strategy that combines similarity and difference conditions, case retrieval focusing on differences is realized. Description of the Drawings
[0017] Figure 1 is a schematic flowchart of the method for retrieving similar cases in the legal affairs system provided by the present invention; Figure 2It is a schematic flowchart of the method for determining the fuzziness of the constituent elements provided by the present invention; Figure 3 It is a schematic flowchart of the method for extracting semantic vectors provided by the present invention; Figure 4 It is a schematic structural diagram of the case retrieval system for similar cases in the legal affairs system provided by the present invention; Figure 5 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0018] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0019] Figure 1 It is a schematic flowchart of the method for retrieving similar cases in the legal affairs system provided by the present invention. As Figure 1 shown, the method includes: Step 110: Receive a case query text corresponding to the case to be retrieved, which includes the basic case situation and the judgment result; Step 120: Determine the fuzziness of each legal constituent element of the law on which the case to be retrieved is based; the higher the fuzziness of any legal constituent element, the higher the possibility that the any legal constituent element includes multiple legal interpretations; Step 130: Retrieve the case library based on the law on which the case to be retrieved is based to obtain an intermediate query result; Step 140: Determine the similarity of each legal constituent element between any case in the intermediate query result and the case to be retrieved, and the total similarity between any case and the case to be retrieved, based on the semantic vectors of the first factual text fragments corresponding to each legal constituent element in the basic case situation, the semantic vectors of the second factual text fragments corresponding to each legal constituent element in any case in the intermediate query result, and the fuzziness of each legal constituent element; Step 150: Screen out cases from the intermediate query result whose total similarity with the case to be retrieved is greater than a first threshold and the similarity of the legal constituent element with the highest fuzziness is less than a second threshold to obtain a query result.
[0020] Specifically, the user can input the case query text corresponding to the case to be retrieved through the human-computer interaction interface. Among them, the case query text includes the case description, the judgment result, and the basic facts of the case (i.e., the description content in the fact-finding part). Using natural language processing technology, based on the existing language model, the semantic structural analysis of the case query text can be carried out to extract the legal provisions on which the judgment result is based, and the factual text fragments related to each legal element corresponding to the legal provisions. Among them, the legal elements corresponding to any legal provision can be obtained from the pre-constructed legal knowledge base, and the factual text fragments related to any legal element may or may not meet the conditions of the legal element in terms of fact-finding.
[0021] After extracting the above text structure information from the case query text, it enters the stage of identifying the fuzziness of legal elements. Each legal provision usually contains multiple legal elements in judicial application, such as "subjective intention", "damage fact", etc. However, due to the fact that some legal elements themselves have a certain degree of fuzziness, these fuzzy legal concepts may lead to different legal interpretations. For example, "negligence" means that the actor did not foresee the possible harmful consequences, but if under the same conditions, the actor should have foreseen the consequences, it can be determined as negligence. However, in actual cases, the word "negligence" has quite a lot of fuzziness. How to judge whether the actor "should" have foreseen the result often leads to completely different legal judgments. It can be seen that the legal interpretations of fuzzy legal elements may vary in different cases, and these differences may lead to completely different judgment results. Therefore, in order to retrieve cases that are generally similar (such as the basic facts and the legal application framework are the same) but the subtle differences affect the judgment conclusion, the fuzzy legal elements can be locked.
[0022] For this reason, the fuzziness of each legal element of the legal provisions relied on by the case to be retrieved can be quantitatively evaluated to obtain the fuzziness of each legal element. Among them, the higher the fuzziness of any legal element, the higher the possibility that the legal element contains multiple legal interpretations in practice, the easier it is to cause disputes, and the more likely it is to be the fundamental reason for the difference in judgment results.
[0023] In some embodiments, as Figure 2 shown, the fuzziness of any legal element of the legal provisions relied on by the case to be retrieved is determined based on the following steps: Step 210, retrieve the case set in the case library that is based on the same legal provisions as the case to be retrieved; Step 220, for any legal element, determine the initial semantic vectors of the factual text fragments corresponding to each case in the case set for the any legal element; Step 230: Determine the ambiguity of any legal element based on the differences between the initial semantic vectors of the factual text segments corresponding to each case for any legal element.
[0024] Legal provisions are often structured and contain several specific legal elements. In traditional retrieval, these elements are often processed integrally, making it difficult to identify potential ambiguity factors. In this embodiment, legal elements are taken as the core semantic units and processed and evaluated separately. Here, in this embodiment, in order to quantitatively identify the ambiguity degree of each legal element in the case to be retrieved, so as to support the subsequent case screening logic of "focusing on differences", a method for calculating the ambiguity of legal elements based on the difference in semantic vector distribution is designed. This method analyzes the language expression differences in the description of the same legal element in different historical cases under the same legal provision in the case library, so as to quantify the "ambiguity" or "degree of divergence" of this legal element in practice.
[0025] Among them, in order to evaluate the ambiguity of each legal element, first perform a clause-level screening in the case library, and select from all historical cases a set of cases with the same legal provision as the case to be retrieved. This set constitutes the sample pool for ambiguity evaluation. Next, for each specific legal element under this legal provision, the system will extract from each case in the case set the factual text segment corresponding to this legal element. After obtaining the factual text segments corresponding to this legal element level in each case, based on a language model (such as LegalBERT), each factual text segment is transformed into an initial semantic vector. Subsequently, measure the differences between the initial semantic vectors of the factual text segments corresponding to the same legal element in different cases, and thus determine the ambiguity of this legal element based on this difference.
[0026] In some embodiments, a statistical distribution method can be adopted to perform a difference analysis on the initial semantic vectors corresponding to the same legal element in different cases. For example, the sum of variances of each initial semantic vector in each semantic dimension can be calculated as the ambiguity of the legal element. The larger the variance, the greater the expression difference of the same legal element in different cases, indicating that there are more possible interpretation paths. In some other embodiments, unsupervised clustering can be performed on the initial semantic vectors of the factual text segments corresponding to the legal element in each case to obtain one or more clusters. If there is only one cluster, the ambiguity of the legal element is determined to be a first preset value (which can be a relatively small value). Otherwise, based on the number of pairs of clusters whose distances between cluster centers are greater than a preset threshold (by combining clusters pairwise to obtain several pairs of clusters, and then calculating the distances between the two cluster centers in each pair of clusters to filter out the pairs of clusters whose distances between cluster centers are greater than the preset threshold), the ambiguity of the legal element is determined. Among them, the larger the number of pairs of clusters whose distances between cluster centers are greater than the preset threshold, the more obvious separated semantic clusters are formed by the vectors of the legal element, and the greater the possibility of multiple interpretations of the legal element in practice. Therefore, the ambiguity of the legal element is determined to be greater.
[0027] It can be seen that the above embodiments realize the accurate modeling of the ambiguity of legal elements by introducing an ambiguity recognition mechanism based on semantic vector distribution, making the entire similar case retrieval system more sensitive when dealing with complex cases, marginal cases or case differences. During the entire similar case retrieval process, this ambiguity information will be used as the core parameter for subsequent case screening and weighting.
[0028] After the above data preparation is completed, based on the legal provisions relied on by the case to be retrieved, cases with the same legal provisions but different judgment results are retrieved from the case library to form an intermediate query result. This important link ensures that the subsequent screening steps focus not on all similar cases, but on those cases that are factually similar but have differences in legal consequences, thus providing materials for analyzing the legal application logic behind the subtle differences.
[0029] Next, a fine-grained comparison at the semantic level is performed for each case in the intermediate query result and the case to be retrieved. Specifically, for any case in the intermediate query result, a factual text fragment related to the legal elements isomorphic to the case to be retrieved is extracted from the case, and semantic vector encoding is performed on them respectively. For the sake of description, the factual text fragment corresponding to each legal element in the basic facts of the case to be retrieved is referred to as the first factual text fragment, and the factual text fragment corresponding to each legal element in the basic facts of the case in the intermediate query result is referred to as the second factual text fragment. The semantic vectors of the first factual text fragment and the second factual text fragment corresponding to the same legal element are obtained. Based on the semantic vectors of the first factual text fragment and the second factual text fragment corresponding to each legal element, combined with the ambiguity of each legal element, the similarity of the corresponding case in the intermediate query result and the case to be retrieved corresponding to each legal element and the total similarity of the case and the case to be retrieved can be determined. In some embodiments, based on the similarity of the case and the case to be retrieved corresponding to each legal element, the ambiguity of each legal element can be introduced for weighting, so as to obtain the total similarity of the case and the case to be retrieved. By weighting the similarity and ambiguity corresponding to each legal element, for a legal element with a smaller ambiguity, its similarity needs to be higher to ensure a higher total similarity, while for a legal element with a higher ambiguity, the requirement for its similarity does not need to be too high. The significance of this is that it can focus on cases where the facts corresponding to the deterministic legal elements (i.e., legal elements with lower ambiguity) are similar but the facts corresponding to the ambiguous legal elements (i.e., legal elements with higher ambiguity) are different.
[0030] Considering that the particularity of the differential retrieval task compared to the conventional retrieval task lies in the need to simultaneously focus on "similarity" and "difference", it is necessary to specifically design a semantic vector extraction scheme applicable to the differential retrieval task and a similarity measurement scheme based on this. For the semantic vector extraction scheme, in some embodiments, as Figure 3 shown, the semantic vector of the first factual text fragment or the second factual text fragment corresponding to any legal element can be determined based on the following steps: Step 310, obtaining an initial semantic vector of the target factual text fragment corresponding to any legal element based on a language model; the target factual text fragment is the first factual text fragment or the second factual text fragment; Step 320, inputting the initial semantic vector of the target factual text fragment into a non-linear layer to obtain a directional vector of the target factual text fragment; the dimensionality of the directional vector is the same as that of the initial semantic vector; Step 330: For any dimension, use the vector value corresponding to this dimension in the initial semantic vector of the target factual text segment as the real part, and the vector value corresponding to this dimension in the directional vector as the imaginary part to obtain a complex value corresponding to this dimension. Step 340: Combine the complex values corresponding to each dimension into the semantic vector of the target factual text segment.
[0031] In this embodiment, to further improve the semantic expression ability of the factual expressions at the legal element level in legal case law, a complex vector structure is introduced to enhance the directionality and ambiguity expression ability of the semantic vector. Especially when facing the detailed descriptions of case details of legal elements with large legal interpretation differences, this mechanism can provide a richer and more refined representation of the "directionality" of factual text segments in specific legal semantic dimensions, thereby supporting more accurate legal element alignment and semantic difference recognition.
[0032] Specifically, for any legal element involved in a case to be retrieved and any case in the intermediate query results, the semantic vectors of the first factual text segment and the second factual text segment corresponding to this element can be constructed. The semantic vector is expressed in complex form, and each dimension not only contains the basic semantics (real part), but also integrates the directional characteristics (imaginary part) shown by this semantics in the context. Among them, first, the target factual text segment (the target factual text segment is the first factual text segment or the second factual text segment) is encoded based on a pre-trained language model to obtain the initial semantic vector of this target factual text segment. This vector is a real number vector, that is, each vector value is a real number. This vector mainly reflects the basic semantic components carried in the corresponding factual text segment, such as the involved subject behavior, described situation, etc., but lacks the ability to model the "direction" of semantics, that is, it cannot depict the differences in some semantic interpretation paths. To make up for the above defects, after obtaining the initial semantic vector, a non-linear layer is designed, and the initial semantic vector obtained in the previous step is input into this non-linear layer (for example, a fully connected layer with ReLU or tanh activation), and a directional vector with the same dimension as the initial semantic vector is obtained. This directional vector can be understood as the semantic deviation trend possessed by each dimension of semantics. If the interpretations of two factual text segments corresponding to the same legal element for this legal element are different, then this difference can be reflected in the corresponding directional vectors, that is, the directional vectors of these two factual text segments will be quite different.
[0033] To enable the non-linear layer to have the above capabilities, the non-linear layer can be pre-trained based on the following steps: Input the initial semantic vectors of each sample factual text corresponding to the above legal element into the non-linear layer to obtain the directional vectors of each sample factual text output by the non-linear layer; Based on the legal interpretation annotations of any two sample factual texts (if the interpretations of the sample legal elements by the two sample factual texts are different, their corresponding legal interpretation annotations are different), and the distances between the directional vectors of any two of the above-mentioned sample factual texts pairwise, determine the model loss corresponding to any two of the above-mentioned sample factual texts; wherein, if the legal interpretation annotations of any two of the above-mentioned sample factual texts are different, the greater the distance between the directional vectors of any two of the above-mentioned sample factual texts, the smaller the model loss corresponding to any two of the above-mentioned sample factual texts; if the legal interpretation annotations of any two of the above-mentioned sample factual texts are the same, the smaller the distance between the directional vectors of any two of the above-mentioned sample factual texts, the smaller the model loss corresponding to any two of the above-mentioned sample factual texts; Based on the sum of the model losses corresponding to any two sample factual texts, determine the total model loss, and adjust the parameters of the non-linear layer based on the total model loss.
[0034] By repeating the above process, a trained non-linear layer can be obtained, and thus the directional vector of the target factual text segment can be obtained based on the non-linear layer.
[0035] After obtaining the initial semantic vector (real part) and the directional vector (imaginary part) of the target factual text segment, these two sets of vectors are combined dimension by dimension. That is, for any dimension, the vector value of the initial semantic vector of the target factual text segment corresponding to this dimension is used as the real part, and the vector value of the directional vector corresponding to this dimension is used as the imaginary part, and a complex value corresponding to this dimension is combined. The complex values corresponding to each dimension are combined into the semantic vector of the target factual text segment, thereby obtaining the complex vector representation of the target factual text segment in the semantic space.
[0036] For the similarity measurement scheme, in some embodiments, to better measure whether the factual expressions of any case in the intermediate query results and the case to be retrieved are similar in each legal element, and at the same time reflect the semantic differences, a fusion matching method combining real number similarity and complex number phase difference degree is proposed, further enhancing the recognition ability of the similar case retrieval system for cases that are generally similar but the subtle differences affect the judgment conclusion. This scheme is mainly based on the operation of the semantic vectors of the first factual text and the second factual text corresponding to the same legal element, including the vector similarity of the real part (the first similarity) and the phase angle difference (difference degree) between the semantic vectors, to construct a legal element-level similarity that comprehensively reflects semantic commonality and direction difference. Specifically, the similarity between any case in the intermediate query results and the case to be retrieved corresponding to any legal element can be calculated based on the following steps: Calculate the similarity (which can be calculated based on the cosine similarity calculation method) between the initial semantic vector of the second factual text segment corresponding to this case for the legal element and the initial semantic vector of the first factual text segment corresponding to the case to be retrieved for the legal element, as the first similarity; the first similarity measures whether the two factual text segments involve similar legal element descriptions and behavioral structures; Calculate the dot product of the conjugate vector of the semantic vector of the second factual text segment corresponding to this case for the legal element (the conjugate vector is the vector obtained by reversing the signs of the imaginary parts of each dimension in the semantic vector) and the semantic vector of the first factual text segment corresponding to the case to be retrieved for the legal element; this dot product is a complex number; Based on the phase angle of the above dot product, determine the degree of difference; where the phase angle represents the direction offset angle of the two sets of semantic vectors in the complex space, that is, it depicts the direction difference of the semantic vectors of the two factual text segments. Therefore, the larger the phase angle of this dot product, the greater the degree of difference; Based on the above first similarity and the above degree of difference, determine the similarity between this case and the case to be retrieved for the legal element. For example, the ratio of the above first similarity to the above degree of difference can be determined as the similarity between this case and the case to be retrieved for the legal element.
[0037] After obtaining the total similarity between all cases in the intermediate query result and the case to be retrieved, a screening strategy based on the similarity-difference combination criterion will be executed, that is, screen out the cases in the intermediate query result whose total similarity with the case to be retrieved is greater than the first threshold and the similarity of the legal element with the highest ambiguity is less than the second threshold to obtain the query result. This dual-threshold screening mechanism effectively takes into account the dual requirements of "overall similarity" and "key differences". The finally selected cases are consistent with the case to be retrieved in most legal elements, but show different semantic characteristics in specific and highly ambiguous elements, resulting in different judgment results. These different cases have high reference value for case handlers, researchers, etc., and can be used to deeply study issues such as the applicable boundaries of legal elements, differences in judgment scales, and judicial unity.
[0038] In summary, the case retrieval method provided by the embodiments of the present invention determines the fuzziness of each legal element of the legal provisions based on the case to be retrieved, and retrieves the case library based on the legal provisions relied on by the case to be retrieved to obtain an intermediate query result. Then, based on the semantic vectors of the first factual text fragments corresponding to each legal element in the basic case, the semantic vectors of the second factual text fragments corresponding to each legal element in any case in the intermediate query result, and the fuzziness of each legal element, it determines the similarity of each legal element of the case to the case to be retrieved and the total similarity of the case to the case to be retrieved. Furthermore, it screens out the cases from the intermediate query result whose total similarity to the case to be retrieved is greater than the first threshold and the similarity of the legal element with the highest corresponding fuzziness is less than the second threshold as the query result to be returned to the user. By introducing the evaluation of the fuzziness of legal elements, the comparison of legal elements at the level of semantic vectors, and the screening strategy combining similarity and difference conditions, the case retrieval focusing on differences is realized.
[0039] The case retrieval system of the legal affairs system provided by the present invention will be described below. The case retrieval system of the legal affairs system described below can be mutually corresponding and referred to the case retrieval method of the legal affairs system described above.
[0040] Based on any of the above embodiments, Figure 4 is a schematic structural diagram of the case retrieval system of the legal affairs system provided by the present invention. As Figure 4 shown, the system includes: A query text receiving unit 410, configured to receive a case query text corresponding to the case to be retrieved, which includes the basic case and the judgment result; A fuzziness evaluation unit 420, configured to determine the fuzziness of each legal element of the legal provisions based on the case to be retrieved; the higher the fuzziness of any legal element, the higher the possibility that the any legal element includes multiple legal interpretations; A first retrieval unit 430, configured to retrieve a case library based on the legal provisions relied on by the case to be retrieved to obtain an intermediate query result; A similarity evaluation unit 440, configured to determine the similarity of each legal element of the any case to the case to be retrieved and the total similarity of the any case to the case to be retrieved based on the semantic vectors of the first factual text fragments corresponding to each legal element in the basic case, the semantic vectors of the second factual text fragments corresponding to each legal element in any case in the intermediate query result, and the fuzziness of each legal element; A second retrieval unit 450, configured to screen out the cases from the intermediate query result whose total similarity to the case to be retrieved is greater than the first threshold and the similarity of the legal element with the highest corresponding fuzziness is less than the second threshold to obtain a query result.
[0041] The system provided by the embodiment of the present invention determines the fuzziness of each legal element of the legal provision on which the case to be retrieved is based, retrieves the case library based on the legal provision on which the case to be retrieved is based, and obtains an intermediate query result. Then, based on the semantic vectors of the first factual text segments corresponding to each legal element in the basic case, the semantic vectors of the second factual text segments corresponding to each legal element in any case in the intermediate query result, and the fuzziness of each legal element, it determines the similarity of each legal element of the case to the case to be retrieved and the total similarity of the case to the case to be retrieved. Furthermore, it screens out the cases in the intermediate query result whose total similarity to the case to be retrieved is greater than the first threshold and the similarity of the legal element with the highest corresponding fuzziness is less than the second threshold as the query result to be returned to the user. By introducing the evaluation of the fuzziness of legal elements, the comparison of legal elements at the level of semantic vectors, and the screening strategy that combines similarity and difference conditions, the retrieval of similar cases focusing on differences is realized.
[0042] Based on any of the above embodiments, the fuzziness of any legal element of the legal provision on which the case to be retrieved is based is determined according to the following steps: Retrieve the set of cases in the case library that are based on the same legal provision as the case to be retrieved; For any legal element, determine the initial semantic vectors of the factual text segments corresponding to the legal element in each case in the case set; Based on the differences between the initial semantic vectors of the factual text segments corresponding to any legal element in each case, determine the fuzziness of the legal element.
[0043] Based on any of the above embodiments, determining the fuzziness of any legal element based on the differences between the initial semantic vectors of the factual text segments corresponding to the legal element in each case includes: Cluster the initial semantic vectors of the factual text segments corresponding to any legal element in each case to obtain one or more clusters; If there is only one cluster, determine the fuzziness of the legal element as the first preset value; Otherwise, based on the number of clusters whose distances between cluster centers are greater than the preset threshold, determine the fuzziness of the legal element.
[0044] Based on any of the above embodiments, the semantic vector of the first factual text segment or the second factual text segment corresponding to any legal element is determined according to the following steps: Obtain the initial semantic vector of the target factual text segment corresponding to any legal element based on the language model; the target factual text segment is the first factual text segment or the second factual text segment; Input the initial semantic vector of the target factual text segment into the non-linear layer to obtain the directional vector of the target factual text segment; the dimensionality of the directional vector is the same as that of the initial semantic vector; For any dimension, use the vector value corresponding to this dimension in the initial semantic vector of the target factual text segment as the real part and the vector value corresponding to this dimension in the directional vector as the imaginary part to obtain the complex value corresponding to this dimension; Combine the complex values corresponding to each dimension into the semantic vector of the target factual text segment.
[0045] Based on any of the above embodiments, the similarity between any case and the case to be retrieved corresponding to any legal element is calculated based on the following steps: Calculate the similarity between the initial semantic vector of the second factual text segment of any case corresponding to any legal element and the initial semantic vector of the first factual text segment of the case to be retrieved corresponding to any legal element as the first similarity; Calculate the dot product of the conjugate vector of the semantic vector of the second factual text segment of any case corresponding to any legal element and the semantic vector of the first factual text segment of the case to be retrieved corresponding to any legal element; Determine the degree of difference based on the phase angle of the dot product; Based on the first similarity and the degree of difference, determine the similarity between any case and the case to be retrieved corresponding to any legal element.
[0046] Based on any of the above embodiments, the total similarity between any case and the case to be retrieved is calculated based on the following steps: Based on the fuzziness of each legal element, perform weighted summation on the similarities between any case and the case to be retrieved corresponding to each legal element to obtain the total similarity between any case and the case to be retrieved.
[0047] Figure 5 It is a schematic structural diagram of the electronic device provided by the present invention, as Figure 5As shown in the figure, the electronic device may include: a processor 510, a memory 520, a communications interface 530, and a communication bus 540. Among them, the processor 510, the memory 520, and the communication interface 530 communicate with each other through the communication bus 540. The processor 510 may call the logical instructions in the memory 520 to execute the case retrieval method for the legal system. The method includes: receiving a case query text corresponding to the case to be retrieved, which includes the basic case situation and the judgment result; determining the ambiguity of each legal element of the case to be retrieved based on the legal provisions; the higher the ambiguity of any legal element, the higher the possibility that the any legal element includes multiple legal interpretations; retrieving a case library based on the legal provisions on which the case to be retrieved is based to obtain an intermediate query result; determining the similarity of any case to each legal element of the case to be retrieved and the total similarity of any case to the case to be retrieved based on the semantic vectors of the first factual text fragments corresponding to each legal element in the basic case situation, the semantic vectors of the second factual text fragments corresponding to each legal element of any case in the intermediate query result, and the ambiguity of each legal element; screening out cases from the intermediate query result whose total similarity to the case to be retrieved is greater than a first threshold and the similarity of the legal element with the highest ambiguity is less than a second threshold to obtain a query result.
[0048] In addition, when the logical instructions in the above-mentioned memory 520 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0049] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the case retrieval method for the legal system provided by the above-mentioned various methods. The method includes: receiving a case query text corresponding to a case to be retrieved, which includes the basic case situation and the judgment result; determining the ambiguity of each legal element of the case to be retrieved based on the legal provisions; the higher the ambiguity of any legal element, the higher the possibility that the any legal element includes multiple legal interpretations; retrieving a case library based on the legal provisions on which the case to be retrieved is based to obtain an intermediate query result; based on the semantic vectors of the first factual text segments corresponding to each legal element in the basic case situation, the semantic vectors of the second factual text segments corresponding to each legal element in any case in the intermediate query result, and the ambiguity of each legal element, determining the similarity of any case to each legal element of the case to be retrieved and the overall similarity of the any case to the case to be retrieved; screening out cases from the intermediate query result whose overall similarity to the case to be retrieved is greater than a first threshold and the similarity of the legal element with the highest corresponding ambiguity is less than a second threshold to obtain a query result.
[0050] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the case retrieval method for the legal system provided by the above-mentioned various methods. The method includes: receiving a case query text corresponding to a case to be retrieved, which includes the basic case situation and the judgment result; determining the ambiguity of each legal element of the case to be retrieved based on the legal provisions; the higher the ambiguity of any legal element, the higher the possibility that the any legal element includes multiple legal interpretations; retrieving a case library based on the legal provisions on which the case to be retrieved is based to obtain an intermediate query result; based on the semantic vectors of the first factual text segments corresponding to each legal element in the basic case situation, the semantic vectors of the second factual text segments corresponding to each legal element in any case in the intermediate query result, and the ambiguity of each legal element, determining the similarity of any case to each legal element of the case to be retrieved and the overall similarity of the any case to the case to be retrieved; screening out cases from the intermediate query result whose overall similarity to the case to be retrieved is greater than a first threshold and the similarity of the legal element with the highest corresponding ambiguity is less than a second threshold to obtain a query result.
[0051] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0052] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0053] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for retrieving similar cases in a legal affairs system, characterized in that, Including: Receiving a case query text corresponding to the case to be retrieved, which includes the basic case situation and the judgment result; Determining the ambiguity of each legal element of the legal provisions on which the case to be retrieved is based; the higher the ambiguity of any legal element, the higher the possibility that the any legal element contains multiple legal interpretations; Retrieving a case library based on the legal provisions on which the case to be retrieved is based to obtain an intermediate query result; Based on the semantic vectors of the first factual text fragments corresponding to each legal element in the basic case situation, the semantic vectors of the second factual text fragments corresponding to each legal element in any case in the intermediate query result, and the ambiguity of each legal element, determining the similarity of any case to each legal element of the case to be retrieved and the overall similarity of any case to the case to be retrieved; Screening from the intermediate query result cases whose overall similarity to the case to be retrieved is greater than a first threshold and the similarity of the legal element with the highest ambiguity is less than a second threshold to obtain a query result.
2. The case retrieval method for the legal affairs system according to claim 1, wherein The ambiguity of any legal element of the legal provisions on which the case to be retrieved is based is determined based on the following steps: Retrieving a set of cases in the case library whose legal provisions are the same as those of the case to be retrieved; For any legal element, determining the initial semantic vector of the factual text fragment corresponding to each case in the case set for the any legal element; Based on the differences between the initial semantic vectors of the factual text fragments corresponding to each case for the any legal element, determining the ambiguity of the any legal element.
3. The method for retrieving similar cases in the legal affairs system according to claim 2, wherein The determining the ambiguity of the any legal element based on the differences between the initial semantic vectors of the factual text fragments corresponding to each case for the any legal element includes: Clustering the initial semantic vectors of the factual text fragments corresponding to each case for the any legal element to obtain one or more clusters; If there is only one cluster, determining the ambiguity of the any legal element as a first preset value; Otherwise, based on the number of pairs of clusters whose distances between cluster centers are greater than a preset threshold, determining the ambiguity of the any legal element.
4. The method for retrieving similar cases in a legal affairs system according to any one of claims 1 to 3, characterized in that, The semantic vector of the first factual text fragment or the second factual text fragment corresponding to any legal element is determined based on the following steps: Obtaining the initial semantic vector of the target factual text fragment corresponding to any legal element based on a language model; the target factual text fragment is the first factual text fragment or the second factual text fragment; Inputting the initial semantic vector of the target factual text fragment into a non-linear layer to obtain the directional vector of the target factual text fragment; the directional vector has the same dimension as the initial semantic vector; For any dimension, taking the vector value corresponding to the dimension in the initial semantic vector of the target factual text fragment as the real part and the vector value corresponding to the dimension in the directional vector as the imaginary part to obtain a complex number corresponding to the dimension; Combining the complex numbers corresponding to each dimension into the semantic vector of the target factual text fragment.
5. The method for retrieving similar cases in the legal affairs system according to claim 4, wherein The similarity between any one of the cases and any legal element of the case to be retrieved is calculated based on the following steps: Calculate the similarity between the initial semantic vector of the second factual text segment corresponding to any legal element of any one of the cases and the initial semantic vector of the first factual text segment corresponding to any legal element of the case to be retrieved as the first similarity; Calculate the dot product of the conjugate vector of the semantic vector of the second factual text segment corresponding to any legal element of any one of the cases and the semantic vector of the first factual text segment corresponding to any legal element of the case to be retrieved; Determine the degree of difference based on the phase angle of the dot product; Determine the similarity between any one of the cases and any legal element of the case to be retrieved based on the first similarity and the degree of difference.
6. The method for retrieving similar cases in the legal affairs system according to claim 4, characterized in that, The total similarity between any one of the cases and the case to be retrieved is calculated based on the following steps: Based on the fuzziness of each legal element, perform a weighted sum of the similarities between any one of the cases and each legal element of the case to be retrieved to obtain the total similarity between any one of the cases and the case to be retrieved.
7. A case retrieval system for a legal affairs system, characterized in that, Including: A query text receiving unit for receiving a case query text corresponding to the case to be retrieved, which includes the basic case situation and the judgment result; A fuzziness evaluation unit for determining the fuzziness of each legal element of the case to be retrieved according to the legal provisions; the higher the fuzziness of any legal element, the higher the possibility that any legal element includes multiple legal interpretations; A first retrieval unit for retrieving a case library based on the legal provisions on which the case to be retrieved is based to obtain an intermediate query result; A similarity evaluation unit for determining the similarity between any one of the cases and each legal element of the case to be retrieved and the total similarity between any one of the cases and the case to be retrieved based on the semantic vectors of the first factual text segments corresponding to each legal element in the basic case situation, the semantic vectors of the second factual text segments corresponding to each legal element in any one of the cases in the intermediate query result, and the fuzziness of each legal element; A second retrieval unit for screening, from the intermediate query result, cases whose total similarity with the case to be retrieved is greater than a first threshold and the similarity of the legal element corresponding to the highest fuzziness is less than a second threshold to obtain a query result.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for retrieving similar cases in the legal system according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for retrieving similar cases in the legal system according to any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for retrieving similar cases in the legal system according to any one of claims 1 to 6.
Citation Information
Patent Citations
Similar case retrieval method, similar case retrieval device and electronic equipment
CN110928994A
Historical legal case similarity recommendation method based on big data analysis
CN118210915A
Class case retrieval method fusing time sequence behavior chain and event type
CN119046410A
Similar law judgment document matching method based on prompt engineering and judgment prediction
CN119783655A
Knowledge graph-based case retrieval method, device and equipment, and storage medium
US20220121695A1