Big data-based legal document processing method and system

By weighting the case description in legal documents, the matching validity and TF-IDF value of each descriptive word are calculated, the problem of inaccurate search results in the prior art is solved, and more accurate search of similar cases is achieved.

CN120067237AActive Publication Date: 2025-05-30GUANGDONG BOWEI CHUANGYUAN TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510533948.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-30
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

When searching similar legal cases, the prior art ignores the impact of keywords in legal documents on case similarity, resulting in inaccurate search results.

Method used

By segmenting the case description, the matching validity of each descriptor is calculated, and the product of the normalized matching validity and the TF-IDF value is used as the weighting coefficient, the semantic vectors of each descriptor are weighted and summed to obtain the case characteristics, thereby achieving accurate retrieval of similar cases.

Benefits of technology

The accuracy of search results of similar cases is improved, ensuring the accurate extraction of case characteristics and the accurate judgment of similar cases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067237A_ABST
    Figure CN120067237A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of text processing, in particular to a legal document processing method and system based on big data, and the method comprises the steps: carrying out the word segmentation of a case description, and obtaining a plurality of descriptors; calculating the matching validity of each description word in the case library; and taking the product of the normalized matching validity and the TF-IDF value of the descriptor as a weighting coefficient to perform weighted summation on the semantic vector of each descriptor to obtain a case feature, and obtaining a similar case described by the case according to the similarity between the case feature and the case feature of the historical case. According to the technical scheme, the accuracy of similar case retrieval results can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of text processing, and in particular, to a method and system for processing legal documents based on big data. Background Art

[0002] With the progress of popularizing the law, more and more people know and understand the law, and more and more people will search for legal solutions online when encountering problems. This has led to a growing demand for retrieving legal cases. As the number of cases gradually increases, how to accurately obtain the text features of the legal documents corresponding to the cases and then achieve accurate retrieval of similar cases is an urgent problem to be solved.

[0003] Currently, the patent application document with the publication number CN110928994A discloses a method for retrieving similar cases, a device for retrieving similar cases, and an electronic device. The method includes: receiving a case to be retrieved, where the case to be retrieved includes at least one of a text description related to the case and a multimedia file; performing content parsing on the text description to perform paragraph recognition and performing dispute focus parsing, legal element parsing, keyword extraction, multi-model semantic processing, and multi-granularity semantic processing on the recognized paragraphs to generate a document parsing result of the case to be retrieved, where multi-model semantic processing is performed on the recognized paragraphs; performing semantic processing on the multimedia file to generate a semantic parsing result of the case to be retrieved; and performing multi-model semantic multi-granularity multi-modal semantic matching based on the document parsing result and the semantic parsing result of the case to be retrieved and the document parsing result and the semantic parsing result of the cases in the case library to obtain a retrieval result.

[0004] The above method realizes the retrieval of similar cases by extracting the document parsing result and the semantic parsing result of the case to be retrieved. When processing the text description, multiple models are used to extract features from the paragraphs in the text description. However, when performing document parsing in units of paragraphs, the influence of keywords such as "civil" and "criminal" in legal documents on the similarity of cases is ignored, resulting in inaccurate retrieval results of similar cases. Summary of the Invention

[0005] In order to solve the technical problem of inaccurate retrieval results of similar cases, this application provides a method and system for processing legal documents based on big data, which can improve the accuracy of retrieval results of similar cases.

[0006] In the first aspect of the present application, a method for processing legal documents based on big data is provided. The search method includes: segmenting the case description to obtain a plurality of descriptive words; calculating the matching effectiveness of each descriptive word in the case library, including: the case library includes a plurality of case pairs with matching labels, and the matching labels include matching and non-matching; counting the first co-occurrence probability of any descriptive word in the case pairs with the matching label being "matching"; counting the second co-occurrence probability of the descriptive word in the case pairs with the matching label being "non-matching", and taking the ratio of the first co-occurrence probability to the second co-occurrence probability as the matching effectiveness of the descriptive word; taking the product of the normalized matching effectiveness and the TF-IDF value of the descriptive word as a weighting coefficient to weighted sum the semantic vectors of each descriptive word to obtain the case characteristics, and obtaining similar cases of the case description based on the similarity between the case characteristics and the case characteristics of historical cases.

[0007] Segment the case description to obtain a plurality of descriptive words of the case description; the case library includes a plurality of case pairs, and the case pairs can be divided into matching case pairs and non-matching case pairs; count the first co-occurrence probability of any descriptive word in the matching case pairs; count the second co-occurrence probability of the descriptive word in the non-matching case pairs, and take the ratio of the first co-occurrence probability to the second co-occurrence probability as the matching effectiveness of the descriptive word, and the matching effectiveness can measure the ability of the descriptive word to judge whether legal documents are similar; further, take the product of the normalized matching effectiveness and the TF-IDF value as a weighting coefficient to weighted sum the semantic vectors of each descriptive word to obtain the case characteristics, determine the weighting coefficient of each descriptive word by comprehensively considering the matching effectiveness and the TF-IDF value, while accurately extracting the case characteristics, ensuring that the extracted case characteristics can accurately judge whether the case characteristics are similar to the historical cases; finally, obtain similar cases of the case description based on the similarity between the case characteristics and the case characteristics of historical cases, ensuring the accuracy of the similar case retrieval results.

[0008] Preferably, jieba segmentation is used to segment the case description.

[0009] Preferably, the matching effectiveness of the descriptive word is: where is the first co-occurrence probability of the descriptive word and is the second co-occurrence probability of the descriptive word.

[0010] Preferably, the method for obtaining the semantic vector of the descriptive word includes: in response to the matching effectiveness of the descriptive word being greater than a preset threshold, taking the word vector of the descriptive word as the semantic vector, otherwise, taking the sum of the word vectors of the descriptive word and multiple descriptive words in the context information as the semantic vector of the descriptive word.

[0011] For the descriptors that play a positive role in the process of judging whether case pairs are similar (i.e., descriptors with a matching effectiveness greater than 1), whether the case pairs are similar can already be distinguished based on their own word vectors, and the word vectors are directly used as semantic vectors. For the descriptors that play no role or a negative role in the process of judging whether case pairs are similar (i.e., descriptors with a matching effectiveness less than or equal to 1), their own word vectors cannot distinguish whether the case pairs are similar. At this time, it is necessary to combine the context information of the case description to accurately obtain the semantic vector of the descriptor.

[0012] Preferably, the word vectors are obtained using the Legal-BERT or Word2Vec model.

[0013] Preferably, taking the sum of the word vectors of the descriptor and multiple descriptors in the context information as the semantic vector of the descriptor includes: setting an initial window, where the initial window includes the context information of the descriptor; in the case pairs that simultaneously contain the descriptor, taking the sum of the word vectors within the initial window as the semantic vector of the descriptor, and calculating the minimum Euclidean distance of the semantic vectors of the descriptor between different historical cases in the case pairs as the semantic deviation of the descriptor; adjusting the size of the initial window, and taking the initial window corresponding to the maximum value of the objective function as the target window, where the objective function is positively correlated with the semantic deviation of the descriptor in the non-matching case pairs and negatively correlated with the semantic deviation of the descriptor in the matching case pairs; extracting the semantic vector of the descriptor in the case description according to the target window.

[0014] The initial window includes the context information of the descriptor. By adjusting the size of the initial window, more context information of the descriptor is continuously introduced until the semantic information of the descriptor can distinguish whether the case pairs are similar, and the target window is obtained, realizing the accurate extraction of the semantic information of the descriptor in the legal document, and the extracted semantic information can effectively distinguish whether the case pairs are similar.

[0015] Preferably, the objective function satisfies the relationship: ; is the semantic deviation of the descriptor in the non-matching case pairs is the semantic deviation of the descriptor in the matching case pairs and are the sets of non-matching case pairs and matching case pairs respectively.

[0016] The objective function can accurately reflect the ability of the semantic vector of the descriptor within the initial window to judge whether case pairs are similar.

[0017] Preferably, after obtaining similar cases of the case description, the processing method further includes: determining the scores of the similar cases according to the user feedback information, and in response to the score of any similar case being greater than the score threshold, regarding the case description and the similar case as a group of matching case pairs with the matching label, and vice versa, regarding the case description and the similar case as a group of matching case pairs with the non-matching label, so as to update the case library.

[0018] Realize continuous update of the case library, gradually optimize the calculation of matching effectiveness, so as to ensure the accuracy of similar cases.

[0019] Preferably, the user feedback information includes adoption, collection and praise.

[0020] In the second aspect of the present application, a legal document processing system based on big data is further provided, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, a legal document processing method based on big data according to the first aspect of the present application is implemented.

[0021] The technical solution of the present application has the following beneficial technical effects: Segment the case description to obtain multiple description words of the case description; the case library includes multiple case pairs, and the case pairs can be divided into matching case pairs and non-matching case pairs; count the first co-occurrence probability of any description word in the matching case pairs; count the second co-occurrence probability of the description word in the non-matching case pairs, and use the ratio of the first co-occurrence probability to the second co-occurrence probability as the matching effectiveness of the description word. The matching effectiveness can measure the ability of the description word to judge whether legal documents are similar. When the description word plays a positive role in judging whether the case pairs are similar, the matching effectiveness of the description word is greater than 1. When the description word does not play a role in judging whether the case pairs are similar, the matching effectiveness of the description word is equal to 1. When the description word plays a negative role in judging whether the case pairs are similar, the matching effectiveness of the description word is less than 1; further, use the product of the normalized matching effectiveness and the TF-IDF value as the weighting coefficient to weighted sum the semantic vectors of each description word to obtain the case features. By comprehensively determining the matching effectiveness and the TF-IDF value, the weighting coefficient of each description word is determined, and while accurately extracting the case features, it is ensured that the extracted case features can accurately judge whether the case features and historical cases are similar; finally, based on the similarity between the case features of the case description and the historical cases, similar cases of the case description are obtained, ensuring the accuracy of the similar case retrieval results. Description of the Drawings

[0022] Figure 1 is a flowchart of a legal document processing method based on big data according to an embodiment of the present application.

[0023] Figure 2 It is a structural block diagram of a legal document processing system based on big data according to an embodiment of the present application. Detailed implementation manners

[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.

[0025] According to the first aspect of the present application, the present application provides a legal document processing method based on big data for realizing the retrieval of similar cases according to the case description. Figure 1 It is a flowchart of a legal document processing method based on big data according to an embodiment of the present application. As Figure 1 shown, the legal document processing method based on big data includes steps S101 to S103, which are described in detail below.

[0026] S101, perform word segmentation on the case description to obtain multiple description words.

[0027] In one embodiment, obtain the case description input by the user, perform word segmentation processing on the case description using jieba word segmentation, and remove stop words after the word segmentation processing to obtain multiple description words; wherein, the stop words are function words or modal particles such as "of", "its", "ah", etc.

[0028] Among them, jieba word segmentation is an existing Chinese word segmentation tool that can split Chinese text into individual words, and one word is a description word.

[0029] S102, calculate the matching effectiveness of each description word in the case library.

[0030] In one embodiment, the matching effectiveness of a description word can measure the ability of the description word to judge whether legal documents are similar. The greater the matching effectiveness, the stronger the ability of the description word to judge whether legal documents are similar.

[0031] Specifically, calculating the matching effectiveness of each description word in the case library includes: the case library includes multiple case pairs with matching labels, and the matching labels include matching and non-matching; count the first co-occurrence probability of any description word in the case pairs with the matching label of matching; count the second co-occurrence probability of the description word in the case pairs with the matching label of non-matching, and take the ratio of the first co-occurrence probability to the second co-occurrence probability as the matching effectiveness of the description word.

[0032] Among them, the case library includes multiple case pairs. A case pair includes two historical cases, and one historical case corresponds to one or more legal documents. One case pair corresponds to a matching label, which is manually labeled and includes "matching" and "not matching". If the matching label is "matching", it is considered that the similarity of the legal documents of the two historical cases in this case pair is 1. If the matching label is "not matching", it is considered that the similarity of the legal documents of the two historical cases in this case pair is 0.

[0033] For any descriptor, there are the following three situations: The first situation: The first co-occurrence probability is relatively large and the second co-occurrence probability is relatively small, indicating that the occurrence probability of this descriptor in the matching case pairs is greater than that in the non-matching case pairs. This descriptor can effectively distinguish between matching case pairs and non-matching case pairs and plays a positive role in the process of judging whether case pairs are similar. The matching effectiveness of this descriptor should be a relatively large value.

[0034] The second situation: The first co-occurrence probability is equal to the second co-occurrence probability, indicating that the number of occurrences of this descriptor in the matching case pairs and non-matching case pairs is the same. This descriptor cannot effectively distinguish between matching case pairs and non-matching case pairs and does not play a role in the process of judging whether case pairs are similar.

[0035] The third situation: The first co-occurrence probability is relatively small and the second co-occurrence probability is relatively large, indicating that the occurrence probability of this descriptor in the non-matching case pairs is greater than that in the matching case pairs. That is to say, in the non-matching case pairs, this descriptor also has a relatively large co-occurrence probability. Although this descriptor can distinguish between matching case pairs and non-matching case pairs, it will misjudge the non-matching case pairs as matching case pairs. Therefore, this descriptor plays a negative role in the process of judging whether case pairs are similar. The matching effectiveness of this descriptor should be a relatively small value.

[0036] Specifically, the matching effectiveness of the descriptor satisfies the following formula: , is the first co-occurrence probability of the descriptor , is the second co-occurrence probability of the descriptor .

[0037] In summary, for any descriptor obtained in step S101, the matching effectiveness of the descriptor can be obtained. When the descriptor plays a positive role in the process of determining whether case pairs are similar, the matching effectiveness of the descriptor is greater than 1. When the descriptor does not play a role in the process of determining whether case pairs are similar, the matching effectiveness of the descriptor is equal to 1. When the descriptor plays a negative role in the process of determining whether case pairs are similar, the matching effectiveness of the descriptor is less than 1, achieving precise quantification of the matching effectiveness of each descriptor.

[0038] S103. Use the product of the normalized matching effectiveness and the TF-IDF value as a weighting coefficient to perform weighted summation on the semantic vectors of each descriptor to obtain the case features. Based on the similarity between the case features of the case description and the historical cases, obtain the similar cases of the case description.

[0039] In one embodiment, directly use the word vectors of each descriptor as the semantic vectors of the corresponding descriptors.

[0040] It should be noted that for descriptors that play a positive role in the process of determining whether case pairs are similar (i.e., descriptors with a matching effectiveness greater than 1), it is already possible to distinguish whether case pairs are similar based on their own word vectors. For descriptors that do not play a role or play a negative role in the process of determining whether case pairs are similar (i.e., descriptors with a matching effectiveness less than or equal to 1), their own word vectors cannot distinguish whether case pairs are similar. At this time, it is necessary to combine the context information of the case description to determine the semantic vector of the descriptor.

[0041] Specifically, the method for obtaining the semantic vector of the descriptor includes: in response to the matching effectiveness of the descriptor being greater than a preset threshold, using the word vector of the descriptor as the semantic vector; otherwise, using the sum of the word vectors of the descriptor and multiple descriptors in the context information as the semantic vector of the descriptor.

[0042] Among them, the value of the preset threshold is 1.

[0043] In one embodiment, using the sum of the word vectors of the descriptor and multiple descriptors in the context information as the semantic vector of the descriptor includes: obtaining the word vector of the descriptor; setting an initial window, where the initial window includes the context information of the descriptor; in case pairs that simultaneously contain the descriptor, using the sum of the word vectors within the initial window as the semantic vector of the descriptor, and calculating the minimum Euclidean distance of the semantic vectors of the descriptor between different historical cases in the case pair as the semantic deviation of the descriptor; adjusting the size of the initial window, and using the initial window corresponding to the maximum value of the objective function as the target window, where the objective function is positively correlated with the semantic deviation of non-matching case pairs and negatively correlated with the semantic deviation of matching case pairs; extracting the semantic vector of the descriptor in the case description according to the target window.

[0044] Among them, the word vectors can be obtained by using the Legal-BERT or Word2Vec model. Legal-BERT is a pre-trained model in the legal field, which is used to obtain the semantic vectors of each descriptive word in legal documents. This is well-known technology for those skilled in the art and will not be elaborated here.

[0045] Among them, the size of the initial window is 3, including one word of the descriptive word and one word of the context of the descriptive word respectively; when adjusting the size of the initial window, each time the initial window is expanded by the distance of one word in the context direction, that is, after one size adjustment, the size of the initial window is 5, including two words of the descriptive word and two words of the context of the descriptive word respectively.

[0046] A case pair includes two historical cases, denoted as historical case 1 and historical case 2 respectively. If the legal documents of the two historical cases both include the descriptive word, then this case pair is used as a case pair that simultaneously includes the descriptive word; respectively obtain the semantic vectors of the descriptive word in the legal documents of historical case 1 and historical case 2. Since the number of occurrences of the descriptive word in historical case 1 and historical case 2 is at least once, take the minimum value of the Euclidean distance between the semantic vectors of historical case 1 and historical case 2 as the semantic deviation in this case pair.

[0047] By adjusting the size of the initial window, more context information of the descriptive word is continuously introduced to more accurately obtain the semantic information of the descriptive word in the legal document; after each adjustment, an initial window and the objective function value of the initial window will be obtained. The objective function is positively correlated with the semantic deviation of the unmatched case pair and negatively correlated with the semantic deviation of the matched case pair, and can reflect the ability of the semantic vector of the descriptive word to judge whether the case pair is similar. Specifically, the objective function satisfies the relational expression: ; is the semantic deviation of the descriptive word in the unmatched case pair ; is the semantic deviation of the descriptive word in the matched case pair ; and are the sets of unmatched case pairs and matched case pairs respectively.

[0048] It can be understood that when the objective function reaches the maximum value, it means that the initial window at this time can maximize the ability of the descriptive word to judge whether the case pair is similar. Take the initial window at this time as the target window of the descriptive word, which is used to extract the semantic vector of the descriptive word in the case description. An optimization algorithm can be used to determine the target window, and the optimization algorithm is a simulated annealing algorithm or a hill climbing algorithm.

[0049] In this way, the semantic vector of each descriptive word in the case description is obtained, the product of the normalized matching validity and the TF-IDF value is used as the weighting coefficient, and the semantic vector of each descriptive word is weighted and summed according to the weighting coefficient to obtain the case feature. The case feature focuses on the descriptive words that play a positive role in judging whether the case pairs are similar, and weakens the descriptive words that play a negative role.

[0050] The ratio of the matching validity of any descriptive word to the sum of the matching validity of all descriptive words in the case description is taken as the normalized matching validity of the descriptive word.

[0051] Among them, the TF-IDF algorithm is a commonly used technical means in text processing. The TF-IDF value of each descriptive word is obtained in the legal documents of all historical cases. The TF-IDF value tends to filter out common descriptive words and retain important descriptive words. The weighting coefficient of each descriptive word is determined by comprehensively considering the matching effectiveness and TF-IDF value. While accurately extracting case features, it ensures that the extracted case features can accurately judge whether the case features are similar to historical cases.

[0052] In one embodiment, the case features of each historical case are obtained in the same way, the similarity between the case features and the case features of the historical case is calculated, and the similarities are sorted in descending order according to the similarity, and the top ranked historical cases are used as similar cases of the case description. The similarity is calculated using a similarity calculation method based on Euclidean distance. If the Euclidean distance between the case features and the case features of any historical case is large, the case description is less similar to the historical case.

[0053] In one embodiment, after obtaining similar cases of the case description, the processing method also includes: determining the scores of similar cases based on user feedback information, and in response to the score of any similar case being greater than a score threshold, treating the case description and the similar cases as a group of case pairs with matching labels; otherwise, treating the case description and the similar cases as a group of case pairs with mismatching labels, thereby updating the case library.

[0054] Among them, user feedback information can be user behaviors such as adoption, collection and like, and each user behavior can be assigned a score. For example, the score for adoption is 4, the score for collection is 3, the score for like is 3, and the total score is 10. When the user performs the corresponding user behavior, the score corresponding to the user behavior is obtained. The score threshold is 5. When the score of any similar case is greater than the score threshold, the case description and similar cases are regarded as a set of case pairs with matching labels. When the score of any similar case is not greater than the score threshold, the case description and similar cases are regarded as a set of case pairs with mismatching labels. The case library is continuously updated, and the matching validity calculation is gradually optimized to ensure the accuracy of similar cases.

[0055] According to the second aspect of the present application, the present application further provides a legal document processing system based on big data. Figure 2 It is a structural block diagram of a legal document processing system based on big data according to an embodiment of the present application. As Figure 2 shown, the system 50 includes a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, it implements a legal document processing method based on big data according to the first aspect of the present application. The system also includes other components well known to those skilled in the art such as a communication bus and a communication interface, and their settings and functions are known in the art, so they will not be elaborated here.

[0056] It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application.

Claims

1. A legal document processing method based on big data, characterized in that: The processing method comprises: Segment the case description to obtain multiple description words; Calculating the matching validity of each descriptive word in a case library, including: the case library includes a plurality of case pairs with matching labels, the matching labels include matching and non-matching; counting the first co-occurrence probability of any descriptive word in the case pairs with matching labels; counting the second co-occurrence probability of the descriptive word in the case pairs with non-matching labels, and taking the ratio of the first co-occurrence probability to the second co-occurrence probability as the matching validity of the descriptive word; The product of the normalized matching validity and the TF-IDF value of the descriptive word is used as a weighting coefficient to weight the sum of the semantic vectors of each descriptive word to obtain case features. Based on the similarity between the case features and the case features of historical cases, similar cases described in the case are obtained.

2. The method for processing legal documents based on big data according to claim 1, characterized in that: Use jieba participle to segment the case description.

3. The method for processing legal documents based on big data according to claim 1, characterized in that: Descriptive words The matching effectiveness for: , Descriptive words The first co-occurrence probability of Descriptive words The second co-occurrence probability of .

4. The method for processing legal documents based on big data according to claim 1, characterized in that: The method for obtaining the semantic vector of the description word includes: In response to the matching validity of the description word being greater than a preset threshold, the word vector of the description word is used as a semantic vector. Otherwise, the sum of the word vectors of the description word and multiple description words in the context information is used as the semantic vector of the description word.

5. The method for processing legal documents based on big data according to claim 4 is characterized in that: The word vector is obtained using the Legal-BERT or Word2Vec model.

6. The method for processing legal documents based on big data according to claim 4, characterized in that: Taking the sum of the word vectors of the description word and multiple description words in the context information as the semantic vector of the description word includes: Setting an initial window, wherein the initial window includes context information of the description word; In case pairs that both contain the descriptive word, the sum of the word vectors in the initial window is used as the semantic vector of the descriptive word, and the minimum Euclidean distance of the semantic vectors of the descriptive word between different historical cases in the case pair is calculated as the semantic deviation of the descriptive word; The size of the initial window is adjusted, and the initial window corresponding to the maximum value of the objective function is used as the target window. The objective function is positively correlated with the semantic deviation of the descriptive words in the unmatched case pairs and negatively correlated with the semantic deviation of the descriptive words in the matched case pairs. The semantic vector of the descriptive words in the case description is extracted based on the target window.

7. The method for processing legal documents based on big data according to claim 6, characterized in that: The objective function Satisfies the relationship: ; For unmatched case pairs The semantic deviation of the descriptor described in For matching case pairs The semantic deviation of the descriptor described in and They are the set of unmatched case pairs and the set of matched case pairs, respectively.

8. The method for processing legal documents based on big data according to claim 1, characterized in that: After obtaining similar cases described in the case, the processing method further includes: The scores of similar cases are determined based on user feedback information. In response to the score of any similar case being greater than a score threshold, the case description and the similar case are regarded as a group of case pairs with matching labels. Otherwise, the case description and the similar case are regarded as a group of case pairs with mismatching labels, thereby updating the case library.

9. The method for processing legal documents based on big data according to claim 8, characterized in that: User feedback information includes adoption, collection and likes.

10. A legal document processing system based on big data, characterized in that: The invention comprises a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the big data-based legal document processing method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Similar case retrieval method, similar case retrieval device and electronic equipment

    CN110928994A

  • Method for computing semantic similarities among short texts

    CN104102626A

  • Court similar case recommendation model based on word vectors and word frequencies

    CN110597949A

  • Data processing method and device, electronic equipment and storage medium

    CN115204154A

  • Text clustering method, text clustering device and text clustering system

    CN116561319A