Intelligent legal lawyer assistant customer service system
Through the document identification and classification, intelligent search and recommendation, relationship network visualization and case analysis and strategy formulation modules of the intelligent legal lawyer assistant customer service system, the problem of insufficient structured processing of legal document information in the existing technology is solved, and accurate retrieval and efficient decision-making are achieved.
Patent Information
- Application Number
- CN202510262233.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art has shortcomings in structuring the legal document information, which leads to inaccurate search results and is difficult to meet the high requirements for information accuracy and structure of complex legal affairs.
An intelligent legal lawyer assistant customer service system is proposed, including document recognition and classification module, intelligent search and recommendation module, relational network visualization module, case analysis and strategy formulation module, and structured processing and accurate retrieval of legal documents through technical means such as OCR scanning, text extraction, metadata indexing, graph theory algorithms, etc.
By realizing the structured processing and accurate retrieval of legal documents, the accuracy and efficiency of document management are improved, the relevance of legal information and the accuracy of user retrieval is enhanced, the difficulty of user retrieval is reduced and decision-making efficiency is improved.
Smart Images

Figure CN120196594A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information retrieval, and in particular to an intelligent legal affairs lawyer assistant customer service system. Background Art
[0002] Information retrieval refers to the process of collecting, storing, analyzing, indexing, retrieving, and sorting massive amounts of data, information, and document resources through computer technology. Information retrieval technologies generally include methods such as natural language processing, semantic analysis, keyword matching, data mining, and machine learning, aiming to help users efficiently and accurately find the required information, and are widely used in fields such as law, medicine, finance, and education.
[0003] The prior art usually only performs keyword matching, simple semantic analysis, and preliminary sorting on massive legal information, and has deficiencies in the structured processing of legal document information. The document management lacks effective classification and indexing means, resulting in inaccurate retrieval results in actual use and being difficult to effectively meet the high requirements for information accuracy and structuring in complex legal affairs. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of the present invention is to solve the deficiencies existing in the prior art and propose an intelligent legal affairs lawyer assistant customer service system.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: An intelligent legal affairs lawyer assistant customer service system includes:
[0006] A document recognition and classification module that performs OCR scanning and text extraction on legal documents, extracts case numbers, related persons, and dates therefrom to generate basic case information; according to the basic case information, classifies the documents into matching cases and legal fields, and establishes a document index library;
[0007] An intelligent retrieval and recommendation module that, based on the document index library, enables users to retrieve by case number, keyword, or date, generates a retrieval result set through metadata indexing; based on the retrieval result set, recommends related documents and cases by analyzing the user's retrieval habits and legal issues of related cases to form a recommendation result;
[0008] A relationship network visualization module that, according to the person, location, and event information in the retrieval result set, uses graph theory algorithms to create a relationship network diagram between cases and generates a simplified relationship network diagram; based on the simplified relationship network diagram, analyzes the key entities and connection strengths in the network to obtain a relationship network diagram;
[0009] A case analysis and strategy formulation module that extracts precedents for case handling and related legal provisions from the recommendation results, collects suggestions for case handling, and establishes a strategy draft.
[0010] Preferably, the steps for obtaining the basic case information are as follows: Collect digital images of legal documents, analyze the image data obtained through text recognition, extract the original text data, and form a preliminary text output;
[0011] Based on the preliminary text output, parse the text content, identify and extract the case number, associated persons, and date information, and generate detailed basic case information;
[0012] Utilize the detailed basic case information to classify and label the information, and obtain structured and queryable basic case information.
[0013] Preferably, the steps for obtaining the document index library are as follows: Apply machine learning to identify the case number, associated persons, and date of the basic case information, file each document into the corresponding case category and legal field, and form a preliminary classification result;
[0014] Based on the formed preliminary classification result, generate document classification information by cross-validating and comparing the classification result with the known case index;
[0015] Utilize the document classification information to construct a document index library. The document index library includes metadata and index tags, and each document is marked according to the case category and legal field to form a document index library.
[0016] Preferably, the steps for obtaining the retrieval result set are as follows: Based on the document index library, receive the retrieval conditions input by the user, including the case number, keywords, and date, and parse the input content to identify different types of retrieval parameters, screen the preliminary document set that meets the retrieval conditions, and obtain a preliminarily matched document data set;
[0017] According to the preliminarily matched document data set, calculate the document relevance score. The calculation formula is:
[0018]
[0019] where R is the document relevance score, m is the total number of preliminarily matched documents, K i is the number of matches of the retrieval keyword in the i-th document, F i is the global occurrence frequency of the retrieval keyword in all documents in the document index library, L i is the position offset of the retrieval keyword in the text content of the i-th document, D i is the index depth of the i-th document, C i is the number of times the i-th document is cited in the document index library, S i is the number of clicks of the i-th document in the retrieval record;
[0020] Based on the document relevance score, sort according to the document relevance score to generate a retrieval result set.
[0021] Preferably, the step of obtaining the recommended result is as follows: Based on the retrieval result set, analyze the user's retrieval records, including retrieval keywords, retrieval time, click behavior, and access duration, to generate a user retrieval behavior data set;
[0022] According to the user retrieval behavior data set, calculate the associated recommendation value of the document. The expression is:
[0023]
[0024] where T is the associated recommendation value of the document, p is the total number of associated retrievals in the user retrieval behavior data set, U j is the number of documents clicked by the user in the jth retrieval, H j is the average residence time of the user accessing each document in the jth retrieval, O j is the offset in the jth retrieval, M j is the number of different retrieval keywords associated in the jth retrieval, E j is the satisfaction degree of the jth retrieval;
[0025] Based on the associated recommendation value of the document, screen the documents that meet the recommendation conditions, combine the associated cases in the retrieval result set, and sort according to the recommendation value to generate the recommended result.
[0026] Preferably, the step of obtaining the simplified relationship network diagram is as follows: Based on the retrieval result set, extract the association relationships between cases. The association relationships include the common characters associated in the cases and the events occurring at the same location, to generate a case association data set;
[0027] According to the case association data set, calculate the connection strength between cases. The calculation formula is:
[0028]
[0029] where G is the connection strength between cases, P k is the number of common characters in the kth pair of cases, CE k is the number of shared events in the kth pair of cases, CD k is the number of times the kth pair of cases occur at different locations, T k is the interval time between the occurrences of the kth pair of cases;
[0030] Based on the connection strength between cases, construct a relationship network diagram between cases, and connect the case nodes in a directed or undirected manner to generate a simplified relationship network diagram.
[0031] Preferably, the steps for obtaining the relationship network diagram are as follows: Based on the simplified relationship network diagram, calculate the structural importance of each node, and the calculation formula is:
[0032]
[0033] where I i is the structural importance of the i-th node, ck i is the degree of the i-th node, d j is the degree of the j-th neighbor node connected to the i-th node, and B i is the betweenness centrality of the i-th node;
[0034] Based on the structural importance, determine the key entities and strong connections in the network to obtain the relationship network diagram.
[0035] Preferably, the steps for obtaining the strategy draft are as follows: Screen associated cases from the recommendation results, extract precedents and associated legal provisions related to the handling of the current case, identify legal documents and clauses, and form a set of legal documents related to the case;
[0036] Based on the set of legal documents related to the case, conduct content comparison, give suggestions for case handling, and obtain the strategy draft.
[0037] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0038] In the present invention, through text extraction of legal documents, after extracting information such as case numbers, associated persons, and dates, automatic document classification is realized, an index library is constructed, the structuring and precise retrieval of legal information are realized, and the accuracy and efficiency of document management are improved; through the metadata index retrieval method and in-depth analysis based on the user's retrieval habits and associated legal issues, targeted document and case recommendation results are formed, enhancing the relevance of legal information and the accuracy of user retrieval, reducing the user's retrieval difficulty and improving the decision-making efficiency; with the help of graph theory algorithms, based on the person, location, and event information in the retrieval results, a relationship network between cases is automatically constructed, revealing the complex correlation between case entities and legal issues, and enhancing the intuitive understanding and analysis ability of legal personnel regarding case associations; based on the recommendation results, precedents and associated legal provisions for handling cases are intelligently extracted, and case handling suggestions are automatically integrated to form a guiding strategy draft, improving the scientific nature of case handling and the efficiency of legal strategy formulation, and realizing an overall improvement in the intelligent level of legal services. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is the system flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0041] Please refer to Figure 1 , the present invention provides a technical solution: an intelligent legal assistant customer service system includes:
[0042] A document recognition and classification module performs OCR scanning and text extraction on legal documents, extracts case numbers, associated persons, and dates therefrom, generates basic case information; classifies the documents into matching cases and legal fields based on the basic case information, and establishes a document index library;
[0043] An intelligent retrieval and recommendation module, based on the document index library, the user retrieves by case number, keyword or date, and generates a retrieval result set through metadata indexing; based on the retrieval result set, by analyzing the user's retrieval habits and legal issues of associated cases, recommends associated documents and cases to form a recommendation result;
[0044] A relationship network visualization module, according to the person, location and event information in the retrieval result set, uses graph theory algorithms to create a relationship network diagram between cases, and generates a simplified relationship network diagram; based on the simplified relationship network diagram, analyzes the key entities and connection strengths in the network to obtain the relationship network diagram;
[0045] A case analysis and strategy formulation module extracts precedents for case handling and associated legal provisions from the recommendation results, collects suggestions for case handling, and establishes a draft strategy.
[0046] The steps for obtaining basic case information are as follows: collect digital images of legal documents, analyze the image data obtained through text recognition, extract the original text data, and form a preliminary text output;
[0047] Based on the preliminary text output, parse the text content, identify and extract case number, associated person, and date information, and generate detailed basic case information;
[0048] Using the detailed basic case information, classify and mark the information to obtain structured and queryable basic case information.
[0049] Specifically, digital images of legal documents are collected. The original images containing text and graphics are obtained from actual paper documents or electronic scans, and the resolution is unified according to the clarity of each image. Then, a pre-established convolutional neural network recognition model is used for training and inference. During training, 3000 to 5000 sample images are first collected, and specific text regions are labeled for each image. Supervised training is performed for 100 rounds with a learning rate of 0.001 and a batch size of 32. The recognition accuracy of each round during the training process is compared with the validation set, and training is stopped when the accuracy exceeds 95%. Subsequently, the image to be recognized is converted into grayscale and input into the trained model to obtain text candidate results. A character recognition confidence threshold of 0.85 is set to filter out low-confidence characters. This threshold is determined based on multiple comparison experiments and selected after gradual testing between a confidence level of 0.80 and 0.90. In each case where the confidence level is lower than 0.85, the text region is binarized and recognized again. All the recognized characters are concatenated and summarized in the order of their positions in the image to form a preliminary text output.
[0050] Based on the preliminary text output, the text content is parsed, including reading the recognized text information line by line and using a regular expression-based retrieval method to obtain three types of elements: case numbers, associated persons, and dates. First, a regular expression for the case number format, such as "[A-Z]{1,5}\d{1,8}", is defined in the text, and a matching confidence threshold of 0.90 is set. This threshold is obtained through actual comparison of different format combinations and multiple statistics in combination with the training text. Strings in the text that match this regular expression and have a matching confidence higher than 0.90 are marked as case number candidate values. At the same time, a dictionary list for personal names is defined, and strings similar to the name structures in the list and containing two to three Chinese characters are marked as associated person candidates through pattern matching. Then, for date information, the expression "\d{4}-\d{1,2}-\d{1,2}" is used for matching, and the recognized content is compared with a pre-agreed date range. For example, the year is compared between 1900 and 2100, the month is compared between 1 and 12, and the date is compared between 1 and 31. If it is outside the range, it is discarded and the next line of parsing is executed. All the parsed case numbers, associated persons, and date information are classified and indexed according to the text order to generate detailed basic case information.
[0051] Using the detailed case basic information, classify and label the information. First, create an index data structure in memory and assign a unique entry to each case number. Map the case number, associated persons, and date information obtained in the previous step to the corresponding entries. Further set multiple filtering conditions in this structure according to actual application needs. For example, when the number of associated persons is more than five, it is regarded as a large case and labeled as "Case scale = large". If the number of associated persons does not exceed five, it is labeled as "Case scale = small". At the same time, when the date information shows that the time span exceeds two years, it can be labeled as "Duration = long" in the field. In the case of not exceeding two years, it is labeled as "Duration = short". The thresholds or intervals referred to by these labels are derived from industry statistical data and can be set through empirical estimation. After completing the annotation of the key information, regard each entry as a retrievable unit and associate query keywords and index pointers with it to obtain structured and queryable case basic information.
[0052] The steps to obtain the document index library are as follows: Apply machine learning to identify the case number, associated persons, and date of the case basic information, and file each document into the corresponding case category and legal field to form a preliminary classification result.
[0053] Based on the formed preliminary classification result, generate document classification information by cross - validating and comparing the classification result with the known case index.
[0054] Use the document classification information to construct a document index library. The document index library includes metadata and index labels, and each document is labeled according to the case category and legal field to form a document index library.
[0055] Specifically, machine learning is applied to identify the case number, related persons, and date of the basic case information. A training set is established based on 5,000 collected legal documents, and each document is accurately labeled with case number, related persons, and date information. On this basis, a convolutional neural network model is selected, the learning rate is set to 0.001, and the batch size is set to 32. Through the supervised training method, 120 rounds of iteration are carried out, and the precision and recall rate are calculated at the end of each round to observe the recognition situation of the model. To ensure the stability of the data distribution during the training process, the data order is randomly shuffled before each round of training, and 80% of the documents are used for training while 20% of the documents are used for validation. When the comprehensive recognition rate on the validation set is lower than 0.92, an additional 10 rounds of training are carried out to refine the parameter adjustment and observe whether the comprehensive recognition rate can reach or exceed this threshold. This threshold is set through multiple experiments and obtained by combining actual evaluation statistics. To obtain the case number information, a specific format rule is established in advance, and the recognized text is matched with this rule item by item. If a certain text paragraph conforms to the rule and the judgment probability of its deep neural network is higher than 0.85, it is marked as a case number candidate value. For the detection of related persons, it is based on the common name thesaurus in this field and judges whether it matches the syllable or glyph features defined in the thesaurus through character embedding. The extraction of the date also uses regular expressions, and after the extraction is completed, the year, month, and date values are respectively compared with the ranges of 1900 to 2100, 1 to 12, and 1 to 31. If it exceeds the range, the corresponding text is discarded. After obtaining the case number, related persons, and date, they are matched according to the internally maintained case category list and common legal field list, and the most likely category and legal field are determined for each document. When the matching score is higher than 0.80 and this score is obtained by comparing the actual classification conclusions in the historical samples, the final filing operation is performed to form a preliminary classification result.
[0056] Based on the formed preliminary classification result, the classification result is compared with the known case index through cross-validation. First, several document entries with clear case category and legal field annotations are extracted from the existing classification results and stratified sampling is performed on them to ensure that the document quantity distributions of different categories are uniform. Then, a ten-fold cross-validation operation is performed on these documents, and in each fold, 90% of the documents are used for training and the remaining 10% of the documents are used for testing. To quantitatively describe the classification accuracy, the accuracy formula is used Comprehensive analysis is carried out by combining precision and recall. If the accuracy of a certain fold is lower than 0.88, defects in the classification rules or recognition features are searched according to the distribution of FP (misjudged as positive samples) and FN (misjudged as negative samples) in that fold and corrected by adjusting the model parameters. When adjusting the model parameters, the learning rate will be temporarily increased to 0.002 and retrained for 5 to 10 rounds to observe whether the accuracy can be stabilized between 0.88 and 0.95. These interval ranges are determined by the average error rate previously statistically obtained on the training set. In order to compare with the known case index, the case number of each document is corresponded one by one with the registered number, and whether the associated person and date fields match are checked and compared. Once it is determined that there are a large number of in-corresponding case numbers, it indicates that there is an obvious deviation in classification and needs to be corrected again. After multiple cross-validations reach the expected range, the hit rate data of the obtained classification results are summarized to generate document classification information.
[0057] When constructing the document index library using the document classification information, the above information is integrated with the document metadata and index tags. In implementation, a unique number is first created according to the case category and legal field corresponding to each document, and a hash mapping structure is established in memory. For the convenience of subsequent searches, multi-level index arrangement is carried out according to the three main dimensions of "case number", "associated person", and "date". Each document is attached with a list containing keywords and semantic concepts, and a correlation threshold between the keywords and the document text content is defined. When the similarity score of a certain keyword with the document exceeds 0.70, it is added to the search tag range of the document. This value of 0.70 is obtained based on the text vectorization method and referring to the similarity distribution law in the training set. In order to further refine the search requirements in different legal fields during the retrieval stage, a domain-specific tag set is introduced, and after finding the corresponding domain in the document classification information, the corresponding tags are docked. The repeated tags are merged to avoid generating redundant entries. When all the classification information and metadata are collected, the index library sample can be quickly generated and its structured preview can be formed to form the document index library.
[0058] The steps to obtain the retrieval result set are as follows: Based on the document index library, receive the retrieval conditions input by the user, including case number, keyword, and date, and parse the input content to identify different types of retrieval parameters, and filter the preliminary document set that meets the retrieval conditions to obtain the preliminarily matched document data set;
[0059] According to the preliminarily matched document data set, calculate the document relevance score. The calculation formula is:
[0060]
[0061] where R is the document relevance score, m is the total number of preliminarily matched documents, K iTo retrieve the number of matches of a keyword in the i-th document, F i To retrieve the global frequency of occurrence of keywords in all documents in the document index library, L i To retrieve the position offset of the keyword in the text content of the i-th document, D i is the index depth of the i-th document, C i is the number of citations of the i-th document in the document index library, S i is the number of clicks on the i-th document in the search record;
[0062] Based on the document relevance score, the documents are sorted according to the document relevance score to generate a retrieval result set.
[0063] Specifically, the search conditions are read based on the document index library and in combination with the case number, keyword and date entered by the user. First, the case number filled in by the user is received in the interface and matched one by one with the case number range registered in the index library. For example, the input case number is compared with the number of each item in the current index. If they are completely consistent, it is classified as an exact match result of the case number. If there is a partial similarity, it is classified as a fuzzy match result and marked as a possible entry, so that it can be further screened according to keywords and dates. The keywords are compared with the common vocabulary list stored in the index library and a correlation threshold is set, such as 0.75. This value is obtained by counting the frequency of occurrence of common words in archived documents and matching and analyzing multiple batches of documents. Each time the input keyword is compared with a specific word in the list, if the similarity calculation result exceeds 0.75, it is considered to have a linkage relationship. In the date processing part, the year, month, and day are referred to in the range of 1900 to 2100, 1 to 12, and 2100, respectively. The text content is located according to the corresponding entry of the case number recorded previously after all the analysis of the search conditions input by the user is completed, and it is judged whether it meets the screening requirements in conjunction with the keyword and date conditions. If the case number appears in the exact match or fuzzy match candidate set and the keyword similarity exceeds the relevance threshold and the date is also within the statistical range, the document is packaged and stored in the preliminary matching document list. Otherwise, the document is excluded from the next round of comparison. If the number of matches exceeds 100, the data will be grouped internally in the order of the case number and the first 10 in each group will be retained as candidate texts. The remaining entries will be temporarily classified as spare entries and called again when the user needs them. Through this series of judgments and successive comparisons, a preliminary document set that meets the search conditions is obtained, and a preliminary matching document data set is obtained.
[0064] The advantage of the formula is that by comprehensively considering information such as the number of times a keyword matches in a document, its global occurrence frequency, the text position offset, and the citation and click of the document, the relevance of each document can be more objectively quantified during the retrieval phase.
[0065] The steps to obtain the m parameter are as follows: This parameter represents the total number of initially matched documents. First, extract the set of documents that match the input conditions from the index library, and then count the number of these documents to obtain m.
[0066] K i The steps to obtain the parameter are as follows: This parameter represents the number of times the retrieval keyword matches in the i-th document. For a specific document, word-by-word comparison is performed on all paragraphs in the text, and the actual number of occurrences of this keyword is counted as K. i , in order to make the statistics more accurate, the word form of the keyword needs to be standardized in advance. For example, convert the input keyword to lowercase and remove the leading and trailing spaces, and then match the text body sentence by sentence. If the keyword appears 8 times in a certain document, then K i = 8.
[0067] F i The steps to obtain the parameter are as follows: This parameter represents the global occurrence frequency of the retrieval keyword in all documents in the document index library. The specific method is to first summarize the word information of all documents in the system and accumulate the number of occurrences of the keyword, and then divide the accumulated number by the count value of the total number of words or total number of words in the document library to obtain the frequency value. If there are a total of 5000 documents in the system and the total number of words is 1,800,000 words, and assuming the keyword "contract" appears 4500 times, then
[0068] L i The steps to obtain the parameter are as follows: This parameter represents the position offset of the retrieval keyword in the text content of the i-th document. It is necessary to first read the character sequence number or line sequence number where the keyword first appears, and calculate L by taking the ratio of it to the length or number of lines of the entire document. i , if the keyword is at the beginning of the text, then L i The value is relatively small. If the keyword is at the end of the text, then L i The value is relatively large. Assuming that the 7th document has 2000 characters and the keyword first appears at the 500th character, then
[0069] D i The steps to obtain the parameter are as follows: This parameter represents the index depth of the i-th document. Usually, it is quantified by combining the document's category, sub-category, and additional label hierarchy. If the document is in the third layer after nesting by case number and legal field, then D i = 3. To obtain D iIt is necessary to record the hierarchical index path of the document in the document index library, and use the number of layers or tag depth of the path as D i As a numerical value, if a document is retrieved after 4 levels of classification, then D i = 4. In some systems, more refined secondary or tertiary nodes are also weighted. For the purpose of giving an example, the index depth of the second document can be set to 2, D2 = 2.
[0070] C i The steps to obtain the parameter are as follows: This parameter represents the number of times the i-th document is cited in the document index library. It is necessary to first track the citation relationship between documents in the system. Whenever a document is marked in the reference part of another document, its citation count is incremented by 1. Assuming that a judgment document on debt disputes is repeatedly cited by 5 other relevant cases, then its C i = 5. If a new document is added to cite it, then C i will become 6. Such citation relationships can be counted when entering documents or accumulated when establishing a mapping table later. The corresponding value can also change with the update of the database. If the 10th document in the example is cited 8 times in total, then C 10 = 8.
[0071] S i The steps to obtain the parameter are as follows: This parameter represents the number of times the i-th document is clicked in the retrieval record. It is necessary to listen for the behavior of the user clicking on the document title or abstract in the retrieval result list each time in the system and accumulate the quantity. If a document is clicked 12 times, then S i = 12. To obtain this parameter, an event capture mechanism can be set between the front end and the back end of the system. Whenever the user completes a click, the mapping relationship between the document ID and the click count is updated to the database. Assuming that a document has been clicked 20 times in the past week, then S i = 20.
[0072] Calculation process:
[0073] Assume that m = 3 preliminary matching documents are obtained in this retrieval, which are called Document 1, Document 2, and Document 3 respectively. Let the K i 、F i 、L i 、D i 、C i 、S i values of each document are as follows:
[0074] Document 1: K1 = 2, F1 = 0.0015, L1 = 0.12, D1 = 2, C1 = 3, S1 = 5;
[0075] Document 2: K2 = 5, F2 = 0.0015, L2 = 0.60, D2 = 3, C2 = 1, S2 = 2;
[0076] Document 3: K3 = 1, F3 = 0.0015, L3 = 0.40, D3 = 1, C3 = 7, S3 = 10;
[0077] First, calculate the numerator part (K i ·F i ·L i ). For Document 1, it is 2 × 0.0015 × 0.12 = 0.00036. For Document 2, it is 5 × 0.0015 × 0.60 = 0.0045. For Document 3, it is 1 × 0.0015 × 0.40 = 0.0006. Then, calculate the denominator (1 + ln(1 + D i +C i +S i ). For Document 1, the denominator is 1 + ln(1 + 2 + 3 + 5) = 1 + ln(11) = 1 + 2.397895 = 3.397895. For Document 2, the denominator is 1 + ln(1 + 3 + 1 + 2) = 1 + ln(7) = 1 + 1.945910 = 2.945910. For Document 3, the denominator is 1 + ln(1 + 1 + 7 + 10) = 1 + ln(19) = 1 + 2.944439 = 3.944439. Divide the numerator by the denominator to get:
[0078]
[0079] Finally, add the three values to get R = 0.000106 + 0.001527 + 0.000152 ≈ 0.001785.
[0080] This result indicates that when the value of R is relatively small (such as 0.001785 in this example), the overall document relevance score is not high. If more documents retrieved later show a larger K i or a higher S i under the same keywords, it will cause the value of R to increase significantly, thus ranking higher in the system.
[0081] Based on document relevance scoring and sorting according to the scores, it is necessary to organize the scoring results obtained in the previous step from largest to smallest in sequence. First, store the scores of all documents as a list and record additional information such as the corresponding document identifiers, score values, and keyword matching counts in the list. Then, when sorting this list, quicksort can be used or the sorting function in the existing data structure can be directly called to generate an order from high scores to low scores. If documents with similar scores need to be further stratified, a stratification threshold can be added on the basis of the aforementioned scores. For example, a 0.05 interval can be regarded as the same stratification. This threshold is set by referring to the score differences of similar documents in the existing search scenarios. For example, if the score differences between some documents are only 0.03, they are regarded as being sorted in the same layer and marked as a candidate group with smaller differences. If the score difference between two documents reaches 0.2, they are regarded as being in different layers. At this time, when displaying, the high-score documents will be placed in the front row and the low-score documents will be placed in the back row. If the score of the highest-score document in a single sorting result exceeds 1.0, this document can be separately marked as the most relevant, and documents with scores lower than 0.001 can be separately grouped at the end and an additional "low matching degree" mark will be added to them. When the user completes the retrieval and views several documents, the system will cumulatively update the newly generated click counts to S i and use the latest click counts to participate in the scoring when retrieving again later. Through such a sorting process, the final retrieval result set can be generated.
[0082] The steps to obtain the recommended results are as follows: Based on the retrieval result set, analyze the user's retrieval records, including retrieval keywords, retrieval time, click behavior, and access duration, to generate a user retrieval behavior data set;
[0083] According to the user retrieval behavior data set, calculate the associated recommendation value of the document. The expression is:
[0084]
[0085] where T is the associated recommendation value of the document, p is the total number of associated retrievals in the user retrieval behavior data set, U j is the number of documents clicked by the user in the jth retrieval, H j is the average stay time of the user accessing each document in the jth retrieval, O j is the offset in the jth retrieval, M j is the number of different retrieval keywords associated in the jth retrieval, E j is the satisfaction degree in the jth retrieval;
[0086] Based on the associated recommendation value of the document, screen the documents that meet the recommendation conditions, combine the associated cases in the retrieval result set, and sort them according to the recommendation value to generate the recommended results.
[0087] Specifically, based on the retrieval result set and combined with the retrieval time, retrieval keywords, click behavior, and access duration submitted by the user, the retrieval records are analyzed. First, all retrieval logs associated with the user account are read from the database, and the retrieval keyword entries therein are disassembled and extracted. For example, there may be multiple keywords entered sequentially in the same query, and all keywords need to be integrated and summarized under the same record number. Subsequently, the retrieval time information is parsed, and the detailed items such as year, month, day, hour, and minute recorded in the log are sorted into a unified timestamp format. Then, combined with the click behavior, an access count operation is performed on the document numbers listed in each log to count the number of times each document is clicked and the click order in this retrieval. If the same document is clicked more than 20 times cumulatively within the past 28 days, it is marked as a highly concerned document in the behavior dataset. Here, both the 28 days and the 20 - click upper limit are obtained through statistics of historical retrieval behaviors and are empirical values in the current regular application stage of the system. By sampling the retrieval habits of all users within 3 months, the daily click frequency is evaluated and a high - frequency interval is summarized. Therefore, the demarcation value of 20 times is determined. The access duration is calculated by recording the timestamp after the user clicks on the document and starts loading and the timestamp before leaving the document page. If an access duration is not between 0 seconds and 600 seconds, it is regarded as abnormal and archived for further review because the normal retrieval and reading time range in the system is concentrated between 1 second and 600 seconds and has been extended to this interval after being confirmed by the administrator. After all retrieval behavior data is parsed, it is aggregated based on their respective account IDs. If an account conducts multiple retrievals on the same day, the retrieval times and click frequencies will be recorded separately in the final dataset. Finally, those documents with a residence duration longer than 300 seconds are separately summarized as in - depth reading results and mapped to the corresponding document numbers. Through this process, a user retrieval behavior dataset is obtained.
[0088] The benefit of the formula is that it comprehensively considers information such as the number of documents clicked by the user, the average residence time, and the actual access offset, and also incorporates the diversity of associated keywords in the retrieval and the user satisfaction into the operation to evaluate the degree of matching between the document and the user's needs.
[0089] The steps to obtain the p parameter are as follows: p is the total number of retrievals associated in the user retrieval behavior dataset. First, lock a certain document in the system as the target object, and then search for all retrieval records associated with this document in all logs and count them to obtain p.
[0090] U j The steps to obtain the parameter are as follows: U jIndicates the number of documents clicked by the user in the j-th search. This parameter is obtained by the search system through click statistics for each query operation. To specifically quantify it, the number of times the user clicks on the document title or abstract can be recorded in the request communication between the front end and the back end. Click counting is performed for each search event number j. If the user clicks on 3 documents in this search, then U j = 3.
[0091] H j The steps to obtain the parameter are: H j Represents the average residence time of the user accessing each document in the j-th search. It is necessary to extract the time stamps of entering and exiting the document from each click behavior and calculate the difference. Then, the total duration of all clicks is aggregated and divided by the number of clicks to form an average value. For example, if the user clicks on 4 documents in the current search and the cumulative residence duration is 720 seconds, then
[0092] O j The steps to obtain the parameter are: O j Is the offset in the j-th search, specifically used to measure the difference in the scrolling or jumping position of the user on the search result page. When collecting, a position monitoring function can be set on the front end to read the distance of the user's scroll bar from the initial position to the position where the current document link is located. If the user directly clicks on a document with a higher ranking in the search results, then O j Is smaller. If the user flips to the last few pages and then clicks, then O j Is larger. To quantify this parameter, it is necessary to count according to the height of the search result page or the number of pagination. For example, it is stipulated that each time the page is scrolled down one page is recorded as an offset of 2. If the user flips 3 pages and selects the first document, then O j = 6.
[0093] M j The steps to obtain the parameter are: M j Is the number of different search keywords associated with the j-th search, used to represent how many non-repeated search terms or phrases the user uses in a single search. The acquisition method is to extract all keywords from the query statement corresponding to the search event number j and perform deduplication statistics. If the two phrases "liability for breach of contract" and "contract term" are entered in this search, then M j = 2.
[0094] E j The steps to obtain the parameter are: E j Is the satisfaction of the j-th search. It is necessary to collect this search experience through a simple scoring mechanism popped up by the system after the user browses the document. Either a full score of 5 or 10 can be used, but it must be converted into a unified numerical record for subsequent calculation. If the 5-point system is selected and a 3-point evaluation is given, then E j = 3.
[0095] Calculation process:
[0096] Taking p = 3 as an example where there are 3 retrievals of the corresponding document within a week, for each retrieval denoted as j = 1, 2, 3, substitute the respective U j , H j , O j , M j , E j into the following operations. First, for the first retrieval: U1 = 3, H1 = 120, O1 = 2, M1 = 2, E1 = 4; for the second retrieval: U2 = 2, H2 = 60, O2 = 1, M2 = 1, E2 = 5; for the third retrieval: U3 = 4, H3 = 180, O3 = 4, M3 = 3, E3 = 3;
[0097] Calculate separately for each retrieval first
[0098] In the first retrieval Its numerator is 1.732 + 118 = 119.732. Divided by 1.414, it is approximately 84.74. Then multiplied by ln(E1 + 1) = ln(4 + 1) = ln(5) ≈ 1.609, the first partial value is 84.74×1.609 ≈ 136.47;
[0099] In the second retrieval
[0100] Its numerator is 1.414 + 59 = 60.414. Then multiplied by ln(E2 + 1) = ln(5 + 1) = ln(6) ≈ 1.791, the second partial value is 60.414×1.791 ≈ 108.23;
[0101] In the third retrieval Its numerator is 2 + 176 = 178. Divided by 1.732, it is approximately 102.72. Then multiplied by ln(E3 + 1) = ln(3 + 1) = ln(4) ≈ 1.386, the third partial value is 102.72×1.386 ≈ 142.97;
[0102] Add the results of the three retrievals to get T = 136.47 + 108.23 + 142.97 = 387.67;
[0103] This result indicates that by organizing and summarizing the retrieval behavior data of users, a T value with a relatively high correlation with the document can be obtained. If T exceeds 300, it means that the user's attention or matching degree to the target document is relatively high. For subsequent automated recommendations, documents with high T values can be presented preferentially. If T is lower than 30, it means that the contribution of the document to this user group is relatively small.
[0104] Based on the document-based associated recommendation values and screening all candidate documents, it is necessary to first collect the associated recommendation values calculated in the previous step and perform numerical sorting. If the recommendation values of some documents are greater than 200, they are regarded as highly meeting the user's needs and classified into the high-level recommendation sequence. Then, compare the high-level recommendation sequence with the associated cases in the retrieval result set. If the case numbers they belong to have a coincidence degree with the topics concerned by the retrieval user exceeding a matching threshold, such as 0.8, and this threshold is obtained from the analysis of 1000 historical retrieval behaviors, these documents are aggregated into the priority position marking group and arranged in descending order according to the recommendation values. If the recommendation value is less than 10, it is ranked at the end of the final list and marked as a candidate document with low matching degree. Subdivided data such as click times and stay time will also be attached to the corresponding document entries during this process for visual viewing. Finally, the recommendation results are generated on the interface and fed back to the user. If the user clicks on some documents with high recommendation values again, the recommendation values and retrieval records will be further updated to form a new round of relevance screening.
[0105] The steps to obtain the simplified relationship network diagram are as follows: Based on the retrieval result set, extract the association relationships between cases. The association relationships include the common characters associated in the cases and the events occurring at the same location, and generate a case association data set;
[0106] According to the case association data set, calculate the connection strength between cases. The calculation formula is:
[0107]
[0108] Among them, G is the connection strength between cases, and P k is the number of common characters in the kth pair of cases, CE k is the number of shared events in the kth pair of cases, CD k is the number of times the kth pair of cases occur at different locations, and T k is the time interval between the occurrences of the kth pair of cases;
[0109] Based on the connection strength between cases, construct the relationship network diagram between cases, connect the case nodes in a directed or undirected manner, and generate a simplified relationship network diagram.
[0110] Specifically, based on the retrieval result set, the association relationships between cases are extracted. First, the text content of all cases in the retrieval result set is parsed to find records of common characters or events occurring at the same location. If the retrieval records show that the name of a person appearing in one case also appears in another case, or the core locations of these two cases are consistent after coordinate or place name comparison, then these two cases are marked as potentially associated. To ensure clear judgment criteria, a preset range for spelling comparison of person names needs to be set within the system, and a location comparison accuracy threshold, such as 20 meters, is set. This threshold is a reasonable value obtained by collecting a large amount of place name information and calculating in combination with the geographical coordinate difference. When it is found that any case has a match with other cases in terms of characters or locations, it is first regarded as potentially associated, and the corresponding case IDs are recorded to establish a matching comparison table. Then, the matching comparison table is summarized and screened based on relevant information. For example, it is checked whether the number of common characters is greater than or equal to 1 or whether the same location is within the same street range. If these conditions are met, the current case pair is listed as effectively associated, and the verification results are recorded one by one and sorted in the order of case ID and event occurrence time. If it is found that a case has multiple common characters with another case or occurs at multiple exactly the same locations, it is itemized and counted to obtain a more accurate association attribute. For the common characters or locations that may repeatedly appear between the same cases, they need to be compared again to prevent duplicate data entry. If there are multiple aliases for person information in the text, the possible aliases or abbreviations need to be unified in advance to avoid statistical errors. All the association relationships confirmed in the above way are finally summarized into the core data table and a case association data set is generated.
[0111] The benefit of the formula is that it simultaneously multiplies the number of common characters by the number of shared events and jointly considers the number of occurrences at different locations and the time interval between cases, so as to make a refined quantitative assessment of the degree of association between two cases under the same index.
[0112] P k The steps to obtain the parameter are as follows: This parameter represents the number of common characters in the k-th pair of cases. It is necessary to first retrieve all the involved persons in this pair of cases from the case association data set and construct a complete list, and then compare the personnel lists of the two cases one by one and count the number of items with the same name or ID number. To ensure the accuracy of person name recognition, identity documents or other fixed identification information will be checked in the early stage. If the number of common characters in two cases is 3, then P k = 3.
[0113] CE kThe steps to obtain the parameter are as follows: This parameter represents the number of events shared in the k-th pair of cases. It is necessary to compare the case descriptions of the two cases item by item and retrieve whether there are identical or extremely similar plot descriptions. For example, a dispute or transaction activity that occurred at a certain location on the same day is recorded in both cases. In order to quantify CE k A list of the main events in the case will be listed and strictly matched against indicators such as the case occurrence date, location, and nature of the event. For example, events with the same litigation date and similar dispute reasons are considered to overlap. If it is confirmed after verification that both cases record 3 identical events, then CE k = 3.
[0114] CD k The steps to obtain the parameter are as follows: This parameter represents the number of times the k-th pair of cases occurred at different locations. It is necessary to consult the actual occurrence location or geographical area in the case documents and compare the locations of the two cases. If it is found during the comparison that one case has branch events occurring in multiple cities or regions while the other case has no overlap in the same city or region, it can be recorded as one occurrence at a different location. To quantify this value, key locations will be searched in the case body and numbered for statistics. During this process, the city level will be gradually matched with more refined county or street levels to confirm the cumulative number of different locations between the two cases. If it is finally verified that there are 5 obvious different locations and there are case records for the corresponding time period, then CD k = 5.
[0115] T k The steps to obtain the parameter are as follows: This parameter represents the time interval between the occurrences of the k-th pair of cases. It is necessary to compare the case-filing times and key event time periods of the two cases and calculate the time difference, which can be directly recorded as the number of days or months. For example, by querying the case-filing dates registered in the system for the two cases, if one case was filed on June 10, 2021, and the other case was filed on June 10, 2022, then the time interval is 365 days, which can be converted to 12 months to obtain T k = 12.
[0116] Calculation process:
[0117] For a certain scenario with 3 pairs of cases respectively recorded as pair 1, pair 2, and pair 3, through statistics, it can be obtained that:
[0118] Pair 1: P1 = 2, CE1 = 1, CD1 = 3, T1 = 6;
[0119] Pair 2: P2 = 4, CE2 = 2, CD2 = 1, T2 = 2;
[0120] Pair 3: P3 = 1, CE3 = 3, CD3 = 2, T3 = 12;
[0121] First, calculate the numerator ∑(Pk ·CE k ):
[0122] P1·CE1 = 2×1 = 2, P2·CE2 = 4×2 = 8, P3·CE3 = 1×3 = 3. Add these three
[0123] together: 2 + 8 + 3 = 13;
[0124] Then calculate the denominator and make a sub - item understanding before accumulation. Here, calculate the denominators for 1, 2, and 3 respectively, then add up the ratios of the numerators and denominators for each pair and then perform a unified summation and division or weighted calculation. Different implementation methods can be processed according to actual needs, but usually this formula is directly regarded as the G - value for each pair of cases and then averaged or summarized. Now, demonstrate it in an integrated way:
[0125]
[0126] If the numerator 13 is equally divided among each pair of cases, divided by the corresponding denominator, and then summed, or the numerator is distributed in other reasonable ways, calculate with one of the examples:
[0127]
[0128] This result shows that the correlation strength formed by the three pairs of cases is approximately 4.124. If the G - value exceeds 3, it indicates that there are obvious common characters and cumulative shared events among the cases, so these pairs of cases can be classified into the closely - related group in the case management process. If the G - value is less than 1, it means that the shared elements are very limited.
[0129] To construct a relationship network diagram based on the connection strength between cases, it is necessary to first summarize all cases and their corresponding connection strength values into the same structure, and record the number, name of each case, and the G value between it and other cases. Each time the value of G exceeds a certain threshold, it is determined that a connection should be formed between these two cases in the diagram. For example, this threshold can be set to 1.5. This value is obtained by statistically analyzing the historical connection strength distribution of more than 100 pairs of cases, sorting each G value, and selecting the top 30% percentile as the discrimination basis. If G is greater than 1.5, it indicates a high degree of closeness between the two cases and a connection needs to be established; otherwise, the connection can be temporarily not made or the line type can be set to a dashed line. Then, according to the chronological order of occurrence between cases, it is decided whether to use a directed connection. For example, when the main event of a certain case precedes another case by more than six months, a directed arc can be drawn in the diagram from the case that occurred earlier to the case that occurred later. If the time difference between events is less than a statistical range, such as 30 days, and the character relationship is relatively equal, an undirected connection is drawn in the diagram. To facilitate reading in the diagram, the coordinates of the nodes are planned, and the nodes are distinguished by size or color according to indicators such as the number of common characters or the number of shared events. For example, P k Cases with a cumulative value exceeding 10 can be set as large-sized nodes to highlight the relatively large number of people involved. When drawing the connection lines, if CD k is too large, additional explanatory text can be added to the connection line to indicate the events that occurred at different locations. At this time, the corresponding text annotation logic needs to be set up in the system, and the number of repeated locations is directly attached to the connection line to distinguish it from those with fewer locations. After the layout of all case nodes and connection lines is completed, a case relationship network diagram can be presented on the interface, and the node and connection line information is integrated into the graphic database for subsequent retrieval and viewing, and finally a simplified relationship network diagram is generated.
[0130] The steps to obtain the relationship network diagram are as follows: According to the simplified relationship network diagram, calculate the structural importance of each node. The calculation formula is:
[0131]
[0132] where, I i is the structural importance of the i-th node, ck i is the degree of the i-th node, d j is the degree of the j-th neighbor node connected to the i-th node, and B i is the betweenness centrality of the i-th node;
[0133] Based on the structural importance, determine the key entities and strong connections in the network to obtain the relationship network diagram.
[0134] Specifically, based on structural importance, first retrieve all nodes from the parsed network data and list the attribute information marked for each node in the previous steps, including the cases or entities associated with the node and the connection relationships between the node and its neighbor nodes. Then, check the role types of these nodes one by one in the network, such as whether they belong to natural persons or legal entities, and record whether there is a one-way or two-way connection between the nodes. Centralize all the mapping information between nodes in a node relationship list. In this list, distinguish the quantity and occurrence time points of known associations and perform extended queries in combination with the communication paths between nodes. If there is a situation where the same entity appears multiple times in different nodes with similar attributes within the system, merge them under the same central node and classify the remaining duplicate entries into the auxiliary reference scope. To distinguish the influence of node pairs on the relationship strength, additional markings can be made for key persons or key institutions that appear more than 5 times. Here, the threshold value of 5 is the result obtained from the analysis of nearly 100 cases, specifically used to indicate that when the frequency of a node's appearance in the network reaches a relatively concentrated level, it is regarded as an important node. Subsequently, based on the connections between all the entries marked as important nodes and other nodes, judge the main connection paths in the network. For example, in a sub-network containing 20 nodes, if it is found that 3 nodes are directly connected to more than half of the other nodes respectively, these 3 nodes can be listed as potential key entities. Then, focus on sorting out these potential key entities and the connection paths between them, and count whether the differences in dimensions such as date or location of these connections are within the same range. If there are connections spanning more than two years, continue to retrieve whether they correspond to different case stages. Here, the two-year threshold value is obtained by referring to the regional judicial statistical data and determined through experience accumulation in case comparisons. Finally, summarize a set of strong connection candidate records based on node attributes and connections and assign unified numbers, sort these nodes and their most frequent connections, and mark them in the graphic data structure to obtain the relationship network diagram.
[0135] The benefit of the formula is that it integrates the betweenness centrality through information such as the sum of the node degree and the neighbor node degrees, taking into account both the status of the node in the network structure and its neighbor characteristics, so that the evaluation result takes into account both the number of connections of the node itself and the degree of path bridging it undertakes in the network.
[0136] ck i The steps to obtain the parameter are as follows: This parameter represents the degree of the i-th node, that is, the number of neighbor nodes directly connected to this node. To obtain ck i It is necessary to first scan all the edge information in the network and perform matching counting on the nodes connected by each edge. Whenever it is found that one end of an edge is node i, ck iIncrement the count by 1. To ensure accuracy, an adjacency list or adjacency matrix is usually established in the data structure first to record the corresponding relationships between each node and its connected nodes. Then, directly summarize the rows or columns of the corresponding nodes from this data structure to obtain ck i The exact value. For example, if a certain node is connected to 8 other nodes in a graph, then ck i = 8.
[0137] d j The steps to obtain the d parameter are as follows: This parameter represents the degree of the j-th neighbor node connected to the i-th node. First, all neighbor nodes of node i need to be locked, and then for the j-th node among them, count its own degree value, that is, how many other nodes it is connected to. If the j-th neighbor node is located to have 5 connection relationships through edge information in the system, then d j = 5.
[0138] B i The steps to obtain the B parameter are as follows: This parameter represents the betweenness centrality of the i-th node. Usually, it is necessary to count the shortest paths between all node pairs in the graph and calculate how many shortest paths the node i appears in to quantify its control level over the paths. If node i appears in many shortest paths, then B i has a larger value. To obtain B i it is necessary to first list the shortest paths between any two nodes in the graph, and then traverse each shortest path. When it is found that node i is included in it, increment the count of B i After that, standardize according to the total number of paths to obtain a value between 0 and 1. For example, in a graph with 20 nodes, there are a total of node pairs, and each pair of nodes may have 1 or more shortest paths. If node i = 10 appears in 45 of these shortest paths, then
[0139] Calculation process:
[0140] In a small network with 6 nodes, node 1 is connected to nodes 2, 3, 4, node 2 is connected to nodes 1, 3, node 3 is connected to nodes 1, 2, 4, 5, node 4 is connected to nodes 1, 3, node 5 is connected to node 3, and node 6 has no direct connection to other nodes. Then, the calculation of I3 for node 3 can be carried out first:
[0141] ck3 = 4 (neighbors are nodes 1, 2, 4, 5). Respectively count the degrees of each neighbor: the degree of node 1 is 3, the degree of node 2 is 2, the degree of node 4 is 2, and the degree of node 5 is 1. Then:
[0142]
[0143] When calculating the betweenness centrality, first enumerate the shortest paths between each pair of nodes. If node 3 appears on 10 of these shortest paths and the total number of node pairs in the entire graph is pieces, then we can substitute into the formula:
[0144]
[0145] This result indicates that the structural importance value of node 3 is approximately 0.6805. When I i is greater than 0.6, it can be judged that this node has a high degree of connection and centrality in this network. If the I x of a certain node is lower than 0.1, it means its importance in the network is extremely low. By synthesizing the I i of each node, the key nodes and their strong connections with surrounding nodes can be summarized.
[0146] Based on the structural importance and combined with the node interconnection information obtained previously in the network, it is necessary to first extract the I i value corresponding to each node in the system and sort them in descending order. Then compare the sorting result with the connection data between nodes. If the I i of a certain node is significantly higher than the preset threshold, such as 0.5, and this threshold is obtained by statistically analyzing the node importance distribution extracted from the historical network analysis, then list this node as a candidate for key entity. Then check each connection between this node and other nodes one by one. If the strength index of the connection has been recorded as a high value in the previous step and the number of connections is also more than 3 times, it is marked as a strong connection, and when visualizing, it is displayed as an obviously thickened connection line. For nodes with a degree greater than 10, different symbols or shapes are used for identification in the graph. The value of 10 is selected based on the node degree statistics of dozens of network cases of different scales, which can reflect that the node already has extremely frequent external connections. Subsequently, for other nodes with low I i maintain the default connection form and present the connection lines between them and the key nodes in normal thickness. If the I i of a certain node is lower than 0.05, it is moved to the edge position in the graph or displayed in a light color to reflect its weak association property. During the process, if it is necessary to check the degree statistics or betweenness centrality of certain nodes, the previously saved calculation records can be retrieved immediately. After all the identifications are completed, map the node coordinates to the plane space and fine-tune them according to the connection relationship between nodes. During the process, the minimum distance control will be performed on the phenomenon of nodes overlapping or being too close to each other, and the preset minimum distance value is provided by the system graphics rendering function, generally set between 50 and 100 pixels. Through this distance interval, nodes can be prevented from overlapping visually with each other and the recognition degree can be ensured. After the node layout and connection distribution are both completed, the final relationship network graph is generated by summarization.
[0147] The steps to obtain the strategy draft are as follows: screen associated cases from the recommendation results, extract precedents and associated legal provisions relevant to the handling of the current case, identify legal documents and clauses, and form a set of legal documents related to the case;
[0148] Based on the set of legal documents related to the case, conduct content comparison, give suggestions for case handling, and obtain the strategy draft.
[0149] Specifically, to screen associated cases from the recommendation results, first disassemble and extract cases according to the case numbers and legal field tags already included in the recommendation results, read information such as the cause of action, filing time, key participants, and key points of dispute of each case and compare it with the main dispute items of the current case. If it is found that some cases have the same or similar words in the dispute items as the current case, they are considered potentially associated, and the intersection words between the two are recorded. Subsequently, list the judgment results or mediation methods of these potentially associated cases and collect the precedent content related to the handling process of the current case. If the number of associated cases exceeds 10, conduct a refined screening according to key parameters. For example, set that cases with the same nature of dispute should account for at least 50%, and this 50% is confirmed based on the statistics of past similar disputes. When screening, compare the matching degree of the nature of dispute fields item by item. If the matching value exceeds the set threshold, retain the case in the associated queue. Then, retrieve the fields corresponding to the laws and regulations in the extracted case information, such as the specific legal clause numbers appearing in the litigation requests or judgment bases, record the numbers in the legal clause list, and reconfirm the names of the laws and regulations to which these clauses belong. For example, when it is retrieved that a certain text belongs to Article 8 of the "Contract Law", further compare its context to determine the degree of association between the content of this clause and the current case. When the confirmed degree of association is above 0.7, include this clause in the summary table. Here, the value of 0.7 is a discrimination standard obtained through the measurement results of multiple batches of cases and by calculating the word frequency similarity. Summarize the eligible precedents, laws and regulations, and their corresponding clauses to form a set of legal documents related to the case.
[0150] Based on the legal document set related to the case, first arrange the contents of each law and regulation in the order of article numbers and mark their key clauses. Then, compare the disputed matters of the current case one by one with these marked clauses. If a targeted regulation appears in a certain law article for the same disputed matter, record the article number and mark the key points involved in the dispute beside it. Next, screen the precedent judgment results extracted in the previous step for supplementary explanations. For example, retrieve the disputed contents of the liability for breach of contract and the scope of compensation in the precedent and check the corresponding legal article citations. When there are significant differences in the identities of the parties or the amount of economic consideration between the precedent and this case, list this difference as a note but still retain its judgment conclusion for multi-dimensional comparison. If multiple precedents have clear handling methods for the same disputed matter, mark this handling method as a high-frequency conclusion during subsequent summarization and accumulate it according to the number of times the high-frequency conclusion appears. If the number exceeds 3 times, further refer to other supporting clauses in this case for verification. The standard of 3 times here is obtained from the induction of nearly a hundred similar cases. After recording the comparison results, select the clause with the highest dispute rate in the same category from the internal list. Finally, combine the handling methods of all precedents and the order of clause citations to form a coherent case handling idea and write it into a brief description to obtain a draft strategy.
Claims
1. An intelligent legal lawyer assistant customer service system, characterized in that: The system comprises: The document recognition and classification module performs OCR scanning and text extraction on legal documents, extracts case numbers, related persons and dates, and generates basic case information; based on the basic case information, the document is classified into matching cases and legal fields, and a document index library is established; The intelligent retrieval and recommendation module, based on the document index library, allows users to search by case number, keyword or date, and generates a retrieval result set through metadata indexing; based on the retrieval result set, by analyzing the user's retrieval habits and the legal issues of related cases, recommends related documents and cases to form recommendation results; A relationship network visualization module, based on the information of people, places and events in the search result set, uses a graph theory algorithm to create a relationship network diagram between cases and generate a simplified relationship network diagram; based on the simplified relationship network diagram, analyzes key entities and connection strengths in the network to obtain a relationship network diagram; The case analysis and strategy formulation module extracts case handling precedents and related legal provisions from the recommendation results, collects case handling suggestions, and establishes a draft strategy.
2. The intelligent legal lawyer assistant customer service system according to claim 1 is characterized in that: The steps of obtaining the basic case information are: collecting digital images of legal documents, obtaining image data through text recognition analysis, extracting original text data, and forming preliminary text output; Based on the preliminary text output, the text content is parsed to identify and extract the case number, related persons and date information, and generate detailed basic case information; The detailed case basic information is used to classify and mark the information to obtain structured and searchable case basic information.
3. The intelligent legal lawyer assistant customer service system according to claim 1 is characterized in that: The steps of obtaining the document index library are: applying machine learning to identify the case number, related persons and date of the basic information of the case, filing each document into the corresponding case category and legal field, and forming a preliminary classification result; Based on the preliminary classification results formed, generate document classification information by cross-validating and comparing the classification results with known case indexes; The document classification information is used to construct a document index library, which includes metadata and index tags. Each document is marked according to case category and legal field to form a document index library.
4. The intelligent legal lawyer assistant customer service system according to claim 1 is characterized in that: The steps of obtaining the search result set are: based on the document index library, receiving the search conditions input by the user, including case number, keyword and date, parsing the input content, identifying different types of search parameters, screening the preliminary document set that meets the search conditions, and obtaining the preliminary matching document data set; According to the preliminary matched document data set, the document relevance score is calculated using the following formula: Among them, R is the document relevance score, m is the total number of preliminary matching documents, and K i To retrieve the number of matches of a keyword in the i-th document, F i To retrieve the global frequency of occurrence of keywords in all documents in the document index library, L i To retrieve the position offset of the keyword in the text content of the i-th document, D i is the index depth of the i-th document, C i is the number of citations of the i-th document in the document index library, S i is the number of clicks on the i-th document in the search record; Based on the document relevance scores, the documents are sorted according to the document relevance scores to generate a search result set.
5. The intelligent legal lawyer assistant customer service system according to claim 1 is characterized in that: The step of obtaining the recommendation results is: based on the search result set, analyzing the user's search records, including search keywords, search time, click behavior and access duration, to generate a user search behavior data set; According to the user retrieval behavior dataset, the associated recommendation value of the document is calculated, and the expression is: Among them, T is the associated recommendation value of the document, p is the total number of associated retrievals in the user retrieval behavior dataset, and U j is the number of documents clicked by the user in the jth retrieval, H j is the average stay time of users accessing each document in the jth retrieval, O j is the offset in the jth retrieval, M j is the number of different search keywords associated with the jth search, E j is the satisfaction of the jth retrieval; Based on the associated recommendation value of the document, the documents meeting the recommendation condition are screened, and the associated cases in the search result set are sorted by the recommendation value to generate a recommendation result.
6. The intelligent legal lawyer assistant customer service system according to claim 1 is characterized in that: The step of obtaining the simplified relationship network diagram is as follows: based on the search result set, extracting the association relationship between cases, the association relationship including the common persons associated with the cases and the events occurring at the same place, and generating a case association data set; According to the case association data set, the connection strength between cases is calculated using the following formula: Among them, G is the connection strength between cases, P k is the number of common characters in the k-th pair of cases, CE k is the number of events shared in the k-th pair of cases, CD k is the number of times the k-th case occurs in different locations, T k is the interval between the occurrences of the k-th pair of cases; Based on the connection strength between cases, a relationship network diagram between cases is constructed, and the case nodes are connected in a directed or undirected manner to generate a simplified relationship network diagram.
7. The intelligent legal lawyer assistant customer service system according to claim 1 is characterized in that: The step of obtaining the relationship network diagram is: according to the simplified relationship network diagram, the structural importance of each node is calculated, and the calculation formula is: Among them, I i is the structural importance of the i-th node, ck i is the degree of the ith node, d j is the degree of the jth neighbor node connected to the i-th node, B i is the betweenness centrality of the ith node; Based on the structural importance, key entities and strong connections in the network are determined to obtain a relationship network diagram.
8. The intelligent legal affairs lawyer assistant customer service system according to claim 1 is characterized in that: The steps of obtaining the draft strategy are: screening related cases from the recommendation results, extracting precedents and related legal provisions related to the current case handling, identifying legal documents and clauses, and forming a set of case-related legal documents; Based on the set of legal documents related to the case, content comparison is performed, case handling suggestions are given, and a draft strategy is obtained.