Judicial suggestion report generation system and method based on large legal language model
Through the judicial recommendation report generation system based on the legal language model, the problems of low efficiency in writing judicial recommendation reports, strong subjectivity and difficulty in responding to urgent needs are solved, efficient and objective report generation is achieved, and the value of big data analysis is fully utilized.
Patent Information
- Application Number
- CN202510112075.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, the compilation of judicial recommendation reports is inefficient, highly subjective, difficult to respond quickly to urgent needs, and difficult to fully cover information, and fail to make full use of the value of big data analysis.
A judicial recommendation report generation system based on a legal language model is adopted. This system includes an information acquisition module, an element analysis module, a template filling module, a rewrite generation module and a report generation module, and judicial recommendation reports are generated through automated and intelligent means.
It significantly improves the efficiency of generating judicial recommendation reports, reduces the subjectivity of manual writing, improves the objectivity and consistency of reports, can quickly respond to emergency cases needs, fully cover case information, and fully explore the value of historical data and big data analysis.
Smart Images

Figure CN120045718A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of judicial informatization, and particularly to a judicial advice report generation system and method based on a legal large language model. Background Art
[0002] In modern society, the public's demand for legal professional knowledge and services is increasing day by day. However, the current supply far fails to meet this huge demand, and the problem of shortage of legal service resources in reality has become increasingly prominent. The intelligent legal Q&A system has become the key way to solve this problem. In the legal Q&A system, providing users with clear, easy-to-understand legal consultation reports presented in natural language is an important link for accurately conveying legal conclusions.
[0003] With the rapid development of big data and artificial intelligence technologies, various industries have embarked on the path of automation and intelligent transformation, and the judicial field has also actively integrated into this trend. As a governance optimization advice given by the judicial department in the process of case handling, after comprehensively counting and deeply summarizing cases within a specific time period and aiming at different influencing factors and regional differences, the judicial advice report is of crucial significance for early warning, properly arranging work, and effectively preventing contradictions and disputes.
[0004] However, at present, the compilation of traditional judicial advice reports mainly relies on manual work, which exposes many drawbacks that cannot be ignored. On the one hand, the efficiency of manual compilation is extremely low. Especially when faced with a mountain of cases, the compilation personnel need to invest a large amount of time and energy, seriously affecting the work progress. On the other hand, due to the fact that manual compilation is extremely vulnerable to the influence of the compiler's personal experience and subjective views, the objectivity and consistency of judicial advice reports are greatly reduced. In addition, in the face of emergencies or urgent key cases, the speed of manual compilation simply cannot meet the urgent need for rapid response. Moreover, it is very difficult for manual compilation to comprehensively consider all potential factors, often omitting key information and important suggestions, and failing to fully explore the great value contained in historical data and big data analysis, and it is difficult to accurately extract valuable information from a large number of judicial cases. For this reason, we propose a judicial advice report generation system and method based on a legal large language model. Summary of the Invention
[0005] To solve the above technical problems, a judicial advice report generation system and method based on a legal large language model are provided. This technical solution solves the problems of difficult supply of legal resources to meet demand, low efficiency of manual compilation of judicial reports, strong subjectivity and inconsistency of manual compilation, difficulty of manual compilation in quickly responding to demands, and difficulty of manual compilation in comprehensively covering information.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0007] A judicial recommendation report generation system based on a legal large language model, comprising: an information acquisition module, an element analysis module, a template filling module, a rewriting and generation module, and a report generation module;
[0008] The information acquisition module is used to extract the basic case data, and the basic data includes: case party information, case quantity, case type, case filing time, case content, and case handling result, and establish an expert recommendation library;
[0009] The element analysis module is electrically connected to the information acquisition module, and the element analysis module is used to analyze and process the acquired information, extract and analyze case information, obtain relevant judicial recommendations, establish the key point association relationship of the case, and provide content support for report generation;
[0010] The template filling module is electrically connected to the element analysis module, and the template filling module is used to fill the content into a preset report template according to the statistical information and case overview, and adjust the report format and content;
[0011] The rewriting and generation module is electrically connected to the element analysis module, and the rewriting and generation module is used to summarize the case content to generate a case summary, optimize the content of the judicial recommendation, make it conform to the case type and the language is standard and easy to understand;
[0012] The report generation module is electrically connected to the template filling module and the rewriting and generation module, and the report generation module is used to summarize the output content of the previous modules, automatically generate a complete judicial recommendation report, output a report in a common format, and provide a user interface and an artificial review interface, and the common format includes but is not limited to: PDF and Word.
[0013] Preferably, the method for the information acquisition module to extract the basic case data and establish the expert recommendation library is:
[0014] The method for extracting the basic case data is:
[0015] Based on the judicial system database, use the SQL query statement "SELECT party information, case quantity, case type, filing time, case content, handling result FROM case table WHERE relevant conditions" to obtain the basic data related to the case;
[0016] The method for establishing the expert recommendation library is:
[0017] Based on the judicial system database, obtain a historical case data set. Based on the obtained historical case data set, where the historical case set is D = {d 1 , d 2 , …, d n};
[0018] Analyze each case in the historical case data set using the constructed legal large language model to extract a set of key information, where the set of key information is K i ={k i1 , k i2 , …, k im};
[0019] Based on the extracted set of key information, generate a preliminary set of judicial suggestions through the legal large language model, where the set of judicial suggestions is S i ={s i1 , s i2 , …, s ip};
[0020] Manually review the generated preliminary judicial suggestions to judge their rationality and feasibility, and mark the suggestions that pass the review using a function. The expression of the review function is:
[0021] f(S i )
[0022] If f(S i ) = 1, it means that the suggestion passes the review;
[0023] Gather all the suggestions with the review passed mark into a database to complete the establishment of the expert suggestion library.
[0024] Preferably, the construction method of the legal large language model is:
[0025] Collect multi-source legal data, and the legal data includes but is not limited to: court judgment documents, laws and regulations, legal academic papers, legal interpretations, legal case analyses, legal contract texts;
[0026] Based on the collected multi-source legal data, preprocess the data. The preprocessing includes data cleaning, text normalization, and word segmentation, and divide the preprocessed data into a training set and a validation set;
[0027] Construct a legal large language model through the preprocessed multi-source legal data, and conduct supervised training on the constructed model. The supervised training includes: classifying legal texts, classifying judgment documents by case type, and classifying laws and regulations by legal department, using the training set to train the model, evaluating the model performance through the validation set, and measuring the difference between the model prediction result and the true result using the cross-entropy loss function. The calculation method of the cross-entropy loss function is:
[0028]
[0029] In the formula, L represents the loss value, y i represents the true label, p iRepresents the probability predicted by the model;
[0030] Among them, the expression of the legal large language model:
[0031] Assume the input sequence is x = (x 1 , x 2 , …, x n ). After passing through the word embedding layer, the embedding vector E = (e 1 , e 2 , …, e n ) is obtained. Through the layer Transformer encoder, the calculation of each layer is as follows:
[0032] h l = TransformerBlock(h l-1 )
[0033] In the formula, h l represents the output of the l-th layer Transformer encoder, and h 0 = E, that is, the initial input is the embedding vector;
[0034] TransformerBlock contains the multi-head self-attention mechanism MultiHeadAttention(h l-1 ) and the feed-forward neural network FEN(h l-1 ), and its expression is:
[0035] MultiHeadAttention(h l-1 ) = Concat(head 1 , …, head h )W O
[0036]
[0037] FEN(h l-1 ) = ReLU(h l-1 W 1 + b 1 )W 2 + b 2
[0038] Among them, head i = Attention(h l-1 W i Q , h l-1 W i K , h l-1 W i V ) represents the output of the i-th attention head in the multi-head self-attention mechanism, Wi Q , W i K , W i V respectively represent the weight matrices used to project the input into Query, Key, and Value respectively when the i-th attention head calculates, and W O represents the weight matrix used for linear transformation after concatenating the outputs of each attention head in the multi-head self-attention mechanism, and W 1 , W 2 represent the weight matrices of different layers in the feed-forward neural network, and b 1 , b 2 represent the bias terms;
[0039] The final output is:
[0040] y = softmax(h L W y + b y )
[0041] In the formula, y represents the final output of the model, and W y , b y represent the weight matrix and bias term used to map the output h L of the last layer of the Transformer encoder to the final output space.
[0042] Preferably, the factor analysis module includes: a statistical information unit, a case overview unit, and a judicial advice unit;
[0043] The statistical information unit is used to perform data analysis and calculate the relationship between case types and geographical distributions, as well as the changing trend of case processing cycles;
[0044] The case overview unit is used to extract the key elements and core content of the case;
[0045] The judicial advice unit is used to obtain judicial advice similar to the case by using rule matching and similarity algorithms.
[0046] Preferably, in the statistical information unit, calculating the relationship between case types and geographical distributions and calculating the changing trend of case processing cycles specifically include:
[0047] Among them, the method for calculating the relationship between case types and geographical distributions is:
[0048] Suppose there are n case types and m regions. Let A ij represent the number of cases of the i-th case type in the j-th region. By calculation, the proportion of various cases in different regions is obtained. The formula for the proportion of various cases in different regions is:
[0049]
[0050] Among them, the method for calculating the changing trend of the case processing cycle is as follows:
[0051] Obtain the processing start time and end time of each case, and calculate the processing cycle of each case:
[0052] T k = t end,k - t start,k
[0053] In the formula, T k represents the processing cycle of the case, t srart,k represents the processing start time of each case, and t end,k represents the processing end time of each case;
[0054] Sort the cases in chronological order. For two adjacent cases k and k + 1, calculate the change rate of the processing cycle:
[0055]
[0056] In the formula, r k represents the change rate of the processing cycle;
[0057] By analyzing the change rate, draw a changing trend graph of the case processing cycle to visually display the change of the processing cycle and provide data support for subsequent reports.
[0058] Preferably, the method for the case overview module to extract key elements and core content from the case text is as follows:
[0059] Use natural language processing technology to perform word segmentation on the case text and split the text into a sequence of words: W 语 = {w 1 , w 2 , …, w n};
[0060] Through part-of-speech tagging and named entity recognition technology, identify the key entity information. The entity information includes: nouns, verbs, personal names, and place names. Let the identified entity set be E 实 = {e 1 , e 2 , …, e n};
[0061] Use semantic analysis technology to analyze the semantic relationships between words, and determine the subject-predicate-object, attributive-adverbial-complementary relationships between words through dependency syntax analysis. Let the semantic relationship set be R 语 = {r 1 , r 2 , …, r n};
[0062] Construct the key elements and core content of the case based on entity information and semantic relationships, remove duplicate and ambiguous information, and organize it into a unified data format.
[0063] Preferably, the method for obtaining judicial suggestions similar to the case by using rule matching and similarity algorithms is as follows:
[0064] Based on the basic data of the new case and the basic data of historical cases extracted by the information acquisition module, and regard the basic data as key point vectors;
[0065] By calculating the similarity between the new case and historical cases, find the most similar historical case and obtain its judicial suggestion. The calculation method of the similarity between the new case and historical cases is as follows:
[0066]
[0067] In the formula, cos similarity represents the similarity between the new case and historical cases, V new =(v 1 , v 2 ,…, v n ) represents the key point vector of the new case, and V old =(u 1 , u 2 ,…, u n ) represents the key point vector of the historical case.
[0068] Preferably, in the template filling module of the report generation stage, the method for selecting a report template is as follows:
[0069] Establish a template library. The templates in the template library are classified according to different requirements of legal fields and case types, and each template is associated with a specific legal field and case type;
[0070] According to the case type and legal field information in the statistical information, select a template from the template library. The selection function is:
[0071] selectTemplate(T 案 , F 法 )
[0072] In the formula, T 案 represents the case type, and F 法 represents the legal field. This function returns the template that matches the case type and legal field, and fills the content in the statistical information and case overview into the template according to the format requirements of the template.
[0073] Preferably, in the rewriting and generating module of the report generation stage, the method for generating a preliminary summary of the case situation is as follows:
[0074] Preprocess each case text to remove irrelevant information, including stop words and punctuation marks, to obtain the preprocessed text;
[0075] Through the word frequency statistics algorithm, calculate the occurrence frequency of each word in all case texts, and select the words with the highest frequency to form a preliminary summary of the case situation. Among them, the expression of the word frequency statistics algorithm is:
[0076] TF-IDF(w) = TF(w) × IDF(w)
[0077] In the formula, TF-IDF(w) represents the term frequency-inverse document frequency, TF(w) represents the term frequency, and IDF(w) represents the inverse document frequency. The calculation formulas for the term frequency and the inverse document frequency are:
[0078]
[0079] In the formula, count(w, d i ) represents the number of times the word w appears in the text d i |d i | represents the total number of words in the text d i n represents the total number of texts, and |{i: w ∈ d i , i = 1, 2,..., n}| represents the number of case texts containing the word w.
[0080] A method for generating a judicial advice report based on a legal large language model, used to implement the system for generating a judicial advice report based on a legal large language model, including:
[0081] Information acquisition stage: Extract basic data related to the case from the system, including case party information, case quantity, case type, case filing time, case content, and case handling results, establish an expert advice library, use the legal large language model to summarize a large number of historical cases, refine valuable judicial experience, and after manual review, convert it into legal advice for a certain type of case, and establish the key points in the case and their association relationships to obtain judicial advice information related to this type of case;
[0082] Element analysis stage: Conduct data analysis through statistical information units, calculate the relationship between case type and geographical distribution, extract the key elements and core content of the case through the case overview unit, extract and screen information from the case text, and through the judicial advice unit, use rule matching and similarity algorithms to identify the association relationship between the key nodes in the case and the judicial advice, and obtain the judicial advice most similar to this type of case;
[0083] Report generation stage: The template filling module fills the content into a preset report template according to the statistical information and case overview. The rewriting and generation module summarizes all case contents in a certain type of cases within a time period to generate a preliminary case summary, and optimizes the language and professional expression of the judicial advice content according to the current case type. The report generation module aggregates the output contents of the template filling module and the rewriting and generation module, automatically generates a final and complete judicial advice report, adds relevant charts and data analysis results at the same time, and provides an interface for manual review.
[0084] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0085] The judicial advice report generation system and method proposed by the present invention significantly improve the generation efficiency of judicial advice reports in an automated and intelligent manner, reduce the workload of the writing staff, enabling them to handle more cases in a shorter time, improving the overall work efficiency. By using a large legal language model to deeply analyze historical cases, valuable judicial experience is extracted, and combined with the specific circumstances of new cases, more objective and consistent judicial advice is generated, enhancing the accuracy and credibility of the report. It can quickly respond to emergency cases or emergencies, generate relevant judicial advice reports in a timely manner, provide strong support for judicial decision-making, help with early warning and proper arrangement of work, and effectively prevent contradictions and disputes. Through automated report generation, the system can also comprehensively cover case information, avoid omitting key contents and important suggestions. At the same time, by using big data analysis and similarity algorithms, valuable information is accurately extracted from a large number of judicial cases, providing more comprehensive and in-depth data support for judicial work. Description of the Drawings
[0086] Figure 1 It is a flowchart for generating a judicial advice report based on a large legal language model. Detailed Embodiments
[0087] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments in the following description are only examples, and those skilled in the art can think of other obvious variations.
[0088] Refer to Figure 1 As shown, a judicial advice report generation system based on a large legal language model includes: an information acquisition module, an element analysis module, a template filling module, a rewriting and generation module, and a report generation module;
[0089] The information acquisition module plays a vital and basic role in the entire system. It focuses on accurately extracting basic case data. These basic data cover a wide range. The information of the parties involved in the case records in detail the identity, background and other key information of the people involved; the number of cases directly reflects the relevant business volume; the case type clearly distinguishes cases of different natures, such as civil, criminal, administrative, etc.; the case filing time clearly defines the node when the case enters the judicial process; the case content fully presents the ins and outs of the incident; and the case handling results show the final conclusion of the judicial decision. At the same time, this module is also responsible for establishing a comprehensive and practical expert advice library to provide professional reference for subsequent work.
[0090] The element analysis module and the information acquisition module work closely together through electrical connection. It mainly conducts in-depth and detailed analysis and processing of the information extracted by the information acquisition module. In this process, it accurately extracts and thoroughly analyzes case information, extracts key points from massive information, and then obtains targeted and relevant judicial suggestions. In addition, this module also focuses on establishing the correlation between the key points of the case, connecting seemingly scattered information into an organic whole, and providing solid content support for subsequent report generation.
[0091] The template filling module is electrically connected to the factor analysis module. Based on the statistical information and case overview obtained by the factor analysis module, it cleverly fills the corresponding content into the preset report template. Not only that, it will also carefully adjust the report format and content to ensure that the report presents a standardized and clear structure.
[0092] The rewrite generation module is also electrically connected to the element analysis module. Its main task is to summarize the case content and generate a concise and clear summary of the case so that readers can quickly grasp the core of the case. At the same time, the content of the judicial advice is optimized to make it highly consistent with the case type, and the language is standardized and easy to understand, so that all parties can understand and implement it.
[0093] The report generation module is electrically connected to the template filling module and the rewriting generation module. It summarizes the high-quality content output by the previous modules and automatically generates a complete judicial recommendation report. The report can be output in common formats, including but not limited to PDF and Word, which is convenient for users to use and save. In addition, the report generation module also provides a user interface and a manual review interface, which not only ensures the efficiency of report generation, but also takes into account the rigor of manual review.
[0094] The method of extracting case basic data and establishing an expert advice database by the information acquisition module is as follows:
[0095] The method for extracting case basic data is:
[0096] Based on the judicial system database, the basic data related to the cases is obtained through the SQL query statement "SELECT Party Information, Number of Cases, Case Type, Filing Time, Case Content, Disposal Result FROM Case Table WHERE Relevant Conditions".
[0097] The method for establishing the expert advice library is as follows:
[0098] Based on the judicial system database, a historical case data set is obtained. Based on the obtained historical case data set, where the historical case set is D = {d 1 , d 2 , …, d n};
[0099] Using the constructed legal large language model to analyze each case in the historical case data set, and extracting a key information set, where the key information set is K i = {k i1 , k i2 , …, k im};
[0100] Based on the extracted key information set, a preliminary judicial advice set is generated through the legal large language model, where the judicial advice set is S i = {s i1 , s i2 , …, s ip};
[0101] Through manual review of the generated preliminary judicial advice, judge its rationality and feasibility, and use a function to mark the advice for those that pass the review. The review function expression is:
[0102] f(S i )
[0103] If f(S i ) = 1, it means that the advice passes the review;
[0104] Gather all the advice with the review passed mark into a database to complete the establishment of the expert advice library.
[0105] Extracting basic data with an SQL query statement can quickly locate key information in the judicial system database, covering various case elements, with comprehensive and accurate data, enabling subsequent analysis to proceed smoothly. When establishing the expert advice library, the legal large language model gives full play to its powerful analysis ability, efficiently mines the value of a large number of historical cases, and produces professional preliminary advice. Coupled with manual review, strictly control the quality of the advice, screen out unreasonable content, and finally integrate high-quality advice into the database, providing practical and authoritative reference for judicial workers.
[0106] The construction method of the legal large language model is as follows:
[0107] Collect multi-source legal data, which includes but is not limited to: court judgment documents, laws and regulations, legal academic papers, legal interpretations, legal case analyses, legal contract texts;
[0108] Based on the collected multi-source legal data, preprocess the data, and the preprocessing includes data cleaning, text normalization, and word segmentation. Divide the preprocessed data into a training set and a validation set;
[0109] Construct a legal large language model with the preprocessed multi-source legal data, and conduct supervised training on the constructed model. The supervised training includes: classifying legal texts, classifying judgment documents by case type, and classifying legal provisions by legal department. Use the training set to train the model, evaluate the model performance through the validation set, and measure the difference between the model prediction result and the true result through the cross-entropy loss function. The calculation method of the cross-entropy loss function is:
[0110]
[0111] In the formula, L represents the loss value, y i represents the true label, and p i represents the probability predicted by the model;
[0112] Among them, the expression of the legal large language model:
[0113] Assume the input sequence is x=(x 1 ,x 2 ,…,x n ). After passing through the word embedding layer, the embedding vector E=(e 1 ,e 2 ,…,e n ) is obtained. Through the layer Transformer encoder, the calculation of each layer is as follows:
[0114] h l =TransformerBlock(h l-1 )
[0115] In the formula, h l represents the output of the l-th layer Transformer encoder, and h 0 =E, that is, the initial input is the embedding vector;
[0116] TransformerBlock contains the multi-head self-attention mechanism MultiHeadAttention(h l-1 ) and the feed-forward neural network FEN(h l-1 ), and its expression is:
[0117] MultiHeadAttention(h l-1 ) = Concat(head 1 , …, head h )W O
[0118]
[0119] FEN(h l-1 ) = ReLU(h l-1 W 1 + b 1 )W 2 + b 2
[0120] Among them, head i = Attention(h l-1 W i Q , h l-1 W i K , h l-1 W i V ) represents the output of the i-th attention head in the multi-head self-attention mechanism. W i Q , W i K , W i V respectively represent the weight matrices used to project the input into Query, Key, and Value when calculating the i-th attention head. W O represents the weight matrix used for linear transformation after concatenating the outputs of each attention head in the multi-head self-attention mechanism. W 1 , W 2 represent the weight matrices of different layers in the feed-forward neural network. b 1 , b 2 represent the bias terms;
[0121] The final output is:
[0122] y = softmax(h L W y + b y )
[0123] In the formula, y represents the final output of the model. W y , b y represent the weight matrix and bias term used to map the output h L of the last layer of the Transformer encoder to the final output space.
[0124] Collecting multi-source legal data ensures that the information learned by the model is rich and comprehensive, covering various legal scenarios. Secondly, preprocessing the data can improve the data quality and lay a good foundation for model training. The classification task in supervised training enables the model to be familiar with different legal categories and enhances its classification ability. Using the cross-entropy loss function can effectively measure the differences in model predictions, which is beneficial to optimizing performance. The application of the Transformer architecture, through the multi-head self-attention mechanism and the feed-forward neural network, can fully capture the complex semantic relationships in the text, improve the model's understanding and processing ability of legal texts, and finally output more professionalism and accuracy.
[0125] The element analysis module includes: a statistical information unit, a case overview unit, and a judicial advice unit;
[0126] The statistical information unit is used for data analysis, calculating the relationship between case types and geographical distributions, and the changing trend of case processing cycles;
[0127] The case overview unit is used to extract the key elements and core content of cases;
[0128] The judicial advice unit is used to obtain judicial advice similar to the case by using rule matching and similarity algorithms.
[0129] In the statistical information unit, calculating the relationship between case types and geographical distributions and calculating the changing trend of case processing cycles specifically include:
[0130] Among them, the method for calculating the relationship between case types and geographical distributions is:
[0131] Assume there are n case types and m regions. Let A ij represent the number of cases of the i-th case type in the j-th region. By calculation, the proportion of various cases in different regions can be obtained. The formula for the proportion of various cases in different regions is:
[0132]
[0133] Among them, the method for calculating the changing trend of case processing cycles is:
[0134] Obtain the start time and end time of processing for each case, and calculate the processing cycle of each case:
[0135] T k = t end,k - t start,k
[0136] In the formula, T k represents the processing cycle of the case, t start,k represents the start time of processing for each case, and t end,k represents the end time of processing for each case;
[0137] Sort the cases in chronological order. For two adjacent cases k and k + 1, calculate the change rate of the processing cycle:
[0138]
[0139] In the formula, r k represents the change rate of the processing cycle;
[0140] By analyzing the change rate, draw a trend chart of the case processing cycle to visually display the change of the processing cycle and provide data support for subsequent reports.
[0141] The calculation of the relationship between case types and regional distributions by the statistical information unit helps relevant departments master the distribution characteristics of cases in each region, so as to reasonably plan judicial resources and achieve precise management. The analysis of the changing trend of the case processing cycle, presented in a visual chart, enables judicial personnel to quickly understand the dynamics of case handling efficiency, facilitating timely adjustment of strategies and improvement of overall efficiency. The case overview unit efficiently extracts key information, avoiding judicial personnel blindly screening through massive content and significantly improving work efficiency. The judicial advice unit uses algorithms to provide suggestions for similar cases, providing reliable references for judicial personnel and enhancing the quality and accuracy of judicial advice.
[0142] The method for the case overview module to extract key elements and core content from case texts is as follows:
[0143] Use natural language processing technology to perform word segmentation on the case text and divide the text into a sequence of words: W 语 ={w 1 , w 2 , …, w n};
[0144] Through part-of-speech tagging and named entity recognition technology, identify key entity information, where the entity information includes: nouns, verbs, personal names, and place names. Let the identified entity set be E 实 ={e 1 , e 2 , …, e n};
[0145] Use semantic analysis technology to analyze the semantic relationships between words, and determine the subject-predicate-object, attributive-adverbial-complementary relationships between words through dependency syntactic analysis. Let the semantic relationship set be R 语 ={r 1 , r 2 , …, r n};
[0146] Construct the key elements and core content of the case based on entity information and semantic relationships, remove duplicate and ambiguous information, and organize it into a unified data format.
[0147] The word segmentation technology of natural language processing disassembles the long and complex case text for easy sorting, just like straightening out a mess. Part-of-speech tagging and named entity recognition accurately identify key information, leaving no important nouns, verbs, personal names, place names, etc. hidden. Semantic analysis and dependency syntax analysis insight into the internal relationships between words and present the logical structure of the case. Removing duplicate and ambiguous information and unifying the format make the information concise and clear. In this way, judicial personnel can quickly grasp the core of the case, improve work efficiency, and make the judicial process more efficient and smooth.
[0148] The method of obtaining judicial suggestions similar to this case by using rule matching and similarity algorithms is as follows:
[0149] Based on the basic data of the new case and the basic data of historical cases extracted by the information acquisition module, and regarding the basic data as key point vectors;
[0150] By calculating the similarity between the new case and historical cases, find the most similar historical case and obtain its judicial suggestion. The calculation method of the similarity between the new case and historical cases is as follows:
[0151]
[0152] In the formula, cos similarity represents the similarity between the new case and historical cases, V new =(v 1 , v 2 ,…, v n ) represents the key point vector of the new case, and V old =(u 1 , u 2 ,…, u n ) represents the key point vector of the historical case.
[0153] Converting the basic data into key point vectors makes complex information concise and comparable. By calculating the similarity, the most similar case can be accurately located from a large number of historical cases quickly, and then the corresponding judicial suggestion can be obtained. This not only saves the time for judicial personnel to re-conceive suggestions, but also provides them with highly referenceable examples, improving the professionalism and accuracy of judicial suggestions and making the judicial work more efficient and coherent.
[0154] In the template filling module of the report generation stage, the method of selecting a report template is as follows:
[0155] Build a template library, where the templates in the library are classified according to the different requirements of legal fields and case types, and each template is associated with a specific legal field and case type;
[0156] According to the case type and legal field information in the statistical information, select a template from the template library, and the selection function is:
[0157] selectTemplate(T 案 ,F 法 )
[0158] In the formula, T 案 represents the case type, and F 法 represents the legal field. This function returns a template that matches the case type and legal field, and fills the content in the statistical information and case overview into the template according to the format requirements of the template.
[0159] The template library is classified by legal field and case type, and can accurately adapt to different case requirements. With the help of the selection function, according to the case type and legal field information in the statistical information, the matching template can be quickly located, greatly improving the screening efficiency. When filling in the content later, the established format of the template provides clear guidance for information integration, ensuring that the report content is well-organized and formatted, reducing the cumbersome work of manual format adjustment, and helping judicial workers to efficiently produce high-quality reports.
[0160] In the rewriting and generation module in the report generation stage, the method for generating a preliminary case summary is:
[0161] Preprocess each case text to remove irrelevant information, where the irrelevant information includes: stop words, punctuation marks, to obtain the preprocessed text;
[0162] Through the word frequency statistics algorithm, calculate the occurrence frequency of each word in all case texts, and select the words with the highest frequency to form a preliminary case summary. Among them, the expression of the word frequency statistics algorithm is:
[0163] TF-IDF(w) = TF(w) × IDF(w)
[0164] In the formula, TF-IDF(w) represents the term frequency-inverse document frequency, TF(w) represents the term frequency, and IDF(w) represents the inverse document frequency. Among them, the calculation formulas for the term frequency and the inverse document frequency are:
[0165]
[0166] In the formula, count(w, d i ) represents the number of times the word w appears in the text d i , and |d i | represents the text d iThe total number of words, n represents the total number of texts, and |{i:w∈d i , i = 1, 2, …, n}| represents the number of case texts containing the word w.
[0167] Preprocessing to remove stop words and punctuation can effectively purify the text, enabling subsequent analysis to be free from interference by useless information and focusing on the core content. The term frequency statistical algorithm combines term frequency with inverse document frequency, fully considering the occurrence of words in individual texts and their importance in the overall document. The high-frequency words selected in this way are a concentrated reflection of the key information of the case. The preliminary case summary formed based on this can not only accurately extract the key points of the case but also present them in a concise form, saving a large amount of time for judicial personnel to quickly understand the main content of the case and improving the efficiency of judicial work.
[0168] A method for generating a judicial advice report based on a legal large language model, used to implement the system for generating a judicial advice report based on a legal large language model, includes:
[0169] Information acquisition stage: Extract basic data such as case party information, summarize historical cases with the help of a legal large language model, refine judicial experience, and after manual review, transform it into legal advice for a type of case, establish associations of case key points, and obtain judicial advice information.
[0170] Element analysis stage: The statistical information unit analyzes the relationship between case types and geographical distributions; the case overview unit extracts and organizes the key elements and core content of the case; the judicial advice unit uses rule matching and similarity algorithms to find judicial advice similar to this type of case.
[0171] Report generation stage: The template filling module fills a preset template according to the statistical information and case overview; the rewriting and generation module summarizes the case content to generate a case summary and optimizes the expression of judicial advice; the report generation module aggregates and outputs, automatically generates a complete report, adds charts, analysis results, and provides a manual review interface.
[0172] The method for generating a judicial advice report based on a legal large language model can extract data comprehensively and accurately, and can deeply explore the value of historical cases. The element analysis dissects the case from multiple angles and provides a reliable reference. The report generation process is efficient and intelligent, with a standardized format and professional content, and also has a review interface, which not only improves the efficiency of judicial work but also guarantees the quality of the report, providing strong support for judicial decision-making.
[0173] The usage process of the present invention is: extracting data, building a database and refining legal advice; analyzing data, identifying key nodes; filling the template and optimizing to generate a report.
[0174] In summary, the advantages of the present invention are as follows: It effectively solves the problem that the supply of legal resources is difficult to meet the demand. Through the automatic generation system based on the legal large language model, the writing efficiency of judicial recommendation reports is greatly improved. This system can comprehensively and objectively analyze case data, avoiding the subjectivity and inconsistency of manual writing, and ensuring the accuracy and credibility of the reports. At the same time, it can quickly respond to the needs of emergency cases, fully exploit the value of historical data and big data analysis, accurately extract valuable information, and provide strong support for judicial decision-making. In addition, this system also provides a user-friendly interface and a manual review interface, enhancing the usability and flexibility of the reports.
[0175] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only the principles of the present invention. Without departing from the spirit and scope of the present invention, various changes and improvements will occur to the present invention, and all these changes and improvements fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A judicial advice report generation system based on a legal large language model, characterized in that: include: Information acquisition module, element analysis module, template filling module, rewriting generation module and report generation module; The information acquisition module is used to extract basic case data, including: information on the parties involved in the case, the number of cases, the type of case, the time of case filing, the content of the case and the result of case handling, and the establishment of an expert advice library; The element analysis module is electrically connected to the information acquisition module, and the element analysis module is used to analyze and process the acquired information, extract and analyze case information, obtain relevant judicial suggestions, establish a correlation relationship between key points of the case, and provide content support for report generation; The template filling module is electrically connected to the element analysis module, and the template filling module is used to fill the content into the preset report template according to the statistical information and the case overview, and adjust the report format and content; The rewriting generation module is electrically connected to the element analysis module, and the rewriting generation module is used to summarize the case content to generate a case summary, and optimize the content of the judicial suggestion so that it conforms to the case type and the language is standardized and easy to understand; The report generation module is electrically connected to the template filling module and the rewriting generation module. The report generation module is used to summarize the output content of the previous module, automatically generate a complete judicial recommendation report, output a common format report, and provide a user interface and a manual review interface. The common formats include but are not limited to: PDF and Word.
2. According to claim 1, a judicial advice report generation system based on a legal large language model is characterized in that: The method of extracting basic case data and establishing an expert advice database in the information acquisition module is as follows: The method for extracting case basic data is: Based on the judicial system database, the basic data related to the case is obtained through the SQL query statement "SELECT party information, case number, case type, case filing time, case content, processing result FROM case table WHERE related conditions"; The method for establishing the expert advice library is as follows: Based on the judicial system database, a historical case data set is obtained, based on which the historical case data set is D = {d1, d2, ..., d n }; Use the constructed legal language model to analyze each case in the historical case data set and extract the key information set, where the key information set is K i ={k i1 ,k i2 ,…,k im }; Based on the extracted key information set, a preliminary set of judicial suggestions is generated through the legal language model, where the judicial suggestion set is S i ={s i1 ,s i2 ,…,s ip }; The generated preliminary judicial suggestions are manually reviewed to determine their rationality and feasibility, and the suggestions that pass the review are marked using a function, where the review function expression is: f(S i ) If f(S i )=1, it means that the suggestion has passed the review; All suggestions with approval marks are organized into a database to complete the establishment of the expert suggestion database.
3. According to claim 2, a judicial advice report generation system based on a legal large language model is characterized in that: The construction method of the legal big language model is: Collect multi-source legal data, including but not limited to: court judgments, legal regulations, legal academic papers, legal interpretations, legal case analysis, and legal contract texts; Based on the collected multi-source legal data, preprocess the data, including data cleaning, text normalization, and word segmentation, and divide the preprocessed data into a training set and a validation set; The legal language model is constructed through the preprocessed multi-source legal data, and the constructed model is supervised and trained. The supervised training includes: classifying legal texts, classifying judicial documents by case type, and classifying legal provisions by legal department. The model is trained using the training set, and the model performance is evaluated using the validation set. The difference between the model prediction result and the actual result is measured using the cross entropy loss function, where the calculation method of the cross entropy loss function is: In the formula, L represents the loss value, y i represents the true label, p i represents the probability predicted by the model; Among them, the expression of the legal large language model is: Assume that the input sequence is x=(x1,x2,…,x n ), and then the word embedding layer gets the embedding vector E=(e1,e2,…,e n ), through the layer Transformer encoder, the calculation of each layer is as follows: h l =TransformerBlock(h l-1 ) In the formula, h l represents the output of the l-th layer Transformer encoder, h 0 =E means the initial input is the embedding vector; TransformerBlock contains a multi-head self-attention mechanism MultiHeadAttention (h l-1 ) and feedforward neural network FEN(h l-1 ), whose expression is: MultiHeadAttention(h l-1 )=Concat(head1,…,head h )W O FEN(h l-1 )=ReLU(h l-1 W1+b1)W2+b2 in, represents the output of the i-th attention head in the multi-head self-attention mechanism, They represent the weight matrices used to project the input into query, key, and value when the i-th attention head is calculated, respectively. W O represents the weight matrix used to concatenate the outputs of each attention head and perform linear transformation in the multi-head self-attention mechanism, W1 and W2 represent the weight matrices of different layers in the feedforward neural network, and b1 and b2 represent bias terms; The final output is: y =softmax(h L W y +b y ) In the formula, y represents the final output of the model, W y , b y It is used to output h of the last layer Transformer encoder. L The weight matrix and bias terms that map to the final output space.
4. According to claim 1, a judicial advice report generation system based on a legal large language model is characterized in that: The element analysis module includes: statistical information unit, case overview unit and judicial recommendation unit; The statistical information unit is used to perform data analysis, calculate the relationship between case types and geographical distribution, and the changing trend of case processing cycles; The case overview unit is used to extract the key elements and core content of the case; The judicial suggestion unit is used to obtain judicial suggestions similar to the case by using rule matching and similarity algorithms.
5. According to claim 4, a judicial advice report generation system based on a legal large language model is characterized in that: In the statistical information unit, the relationship between case types and regional distribution and the trend of case processing cycle are calculated. include: Among them, the method for calculating the relationship between case type and geographical distribution is: Assume there are n types of cases and m regions. ij It represents the number of cases of the i-th case type in the j-th region. The proportion of each type of case in different regions is calculated. The calculation formula for the proportion of each type of case in different regions is: Among them, the method for calculating the change trend of case processing cycle is: Get the processing start time and end time of each case, and calculate the processing cycle of each case: T k =t end,k -t start,k Where, T k Represents the case processing cycle, t start,k Indicates the processing start time of each case, t end,k Indicates the processing completion time of each case; Sort the cases in chronological order, and for two adjacent cases k and k+1, calculate the rate of change of the processing cycle: In the formula, r k Indicates the rate of change of the processing cycle; By analyzing the change rate, a trend chart of the case processing cycle is drawn to intuitively display the changes in the processing cycle and provide data support for subsequent reports.
6. According to claim 4, a judicial advice report generation system based on a legal large language model is characterized in that: The method by which the case overview module extracts key elements and core content from the case text is as follows: Use natural language processing technology to segment the case text and divide the text into word sequences: W 语 ={w1,w2,…,w n }; Through part-of-speech tagging and named entity recognition technology, key entity information is identified, including nouns, verbs, names, and place names. Let the identified entity set be E 实 ={e1,e2,…,e n }; Using semantic analysis technology, we analyze the semantic relationship between words, and determine the subject-predicate-object, attributive-adverbial-complement relationship between words through dependency syntax analysis. Let the semantic relationship set be R 语 ={r1,r2,…,r n }; Based on entity information and semantic relationships, the key elements and core content of the case are constructed, and duplicate and ambiguous information is removed and organized into a unified data format.
7. According to claim 4, a judicial advice report generation system based on a legal large language model is characterized in that: Using rule matching and similarity algorithms, the method to obtain judicial recommendations similar to this case is: The basic data of new cases and historical cases extracted by the information acquisition module are regarded as key point vectors; By calculating the similarity between new cases and historical cases, the most similar historical case is found and its judicial advice is obtained. The calculation method of the similarity between new cases and historical cases is: In the formula, cos similarity Represents the similarity between new cases and historical cases, V new =(v1,v2,…,v n ) represents the key point vector of the new case, V old =(u1,u2,…,u n ) represents the key point vector of the historical case.
8. According to claim 1, a judicial advice report generation system based on a legal large language model is characterized in that: In the template filling module of the report generation phase, the method for selecting a report template is: Establish a template library, in which the templates are classified according to different needs and case types in the legal field, and each template is associated with a specific legal field and case type; According to the case type and legal field information in the statistical information, select a template from the template library. The selection function is: selectTemplate(T 案 ,F 法 ) Where, T 案 Indicates the case type, F 法 Represents the legal field. This function returns a template that matches the case type and legal field, and fills the content in the statistical information and case overview into the template according to the format requirements of the template.
9. According to claim 1, a judicial advice report generation system based on a legal large language model is characterized in that: In the rewriting generation module of the report generation stage, the method for generating a preliminary case summary is: Preprocess each case text to remove irrelevant information, including stop words and punctuation marks, to obtain a preprocessed text; The frequency of occurrence of each word in all case texts is calculated through the word frequency statistics algorithm, and the words with the highest frequency are selected to form a preliminary case summary. The expression of the word frequency statistics algorithm is: TF-IDF(w)=TF(w)×IDF(w) In the formula, TF-IDF(w) means term frequency-inverse document frequency, TF(w) means term frequency, and IDF(w) means inverse document frequency. The calculation formulas for term frequency and inverse document frequency are: In the formula, count(w,d i ) indicates that word w is in text d i The number of occurrences in |d i | indicates text d i The total number of words in the text, n represents the total number of texts, |{i:w∈d i ,i=1,2,…,n}| represents the number of case texts containing word w.
10. A method for generating a judicial recommendation report based on a legal large language model, characterized in that: A system for generating a judicial advice report based on a legal large language model according to any one of claims 1 to 9, comprising: Information acquisition stage: extract basic case-related data from the system, including case party information, case number, case type, case filing time, case content and case handling results, establish an expert advice database, use the legal big language model to summarize and conclude a large number of historical cases, extract valuable judicial experience, and convert it into legal advice for a type of case after manual review. Establish the relationship between the key points in the case and its correlation, and obtain judicial advice information related to this type of case; Element analysis stage: Data analysis is performed through the statistical information unit to calculate the relationship between case types and geographical distribution. The key elements and core content of the case are extracted through the case overview unit, and information is extracted and sorted from the case text. Through the judicial suggestion unit, rule matching and similarity algorithms are used to identify the relationship between the key nodes in the case and the judicial suggestions, and obtain the judicial suggestions that are most similar to this type of case; Report generation stage: The template filling module fills the content into the preset report template according to the statistical information and case overview. The rewriting generation module summarizes the content of all cases in a certain type of case in a time period, generates a preliminary case summary, and optimizes the language and professional expression of the judicial advice content according to the current case type. The report generation module summarizes the output content of the template filling module and the rewriting generation module, automatically generates a final complete judicial advice report, adds relevant charts and data analysis results, and provides a manual review interface.
Citation Information
Cited By
Judicial decision prediction system based on generative artificial intelligence
CN120746777A