Legal long text abstract generation method based on large language model

By using a legal text summarization method based on a large language model, the professionalism and efficiency issues of legal text summarization in existing technologies are solved, and high-precision, low-cost automatic legal text summarization is achieved, which is applicable to various types of legal texts.

CN121880550APending Publication Date: 2026-04-17NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANTONG UNIV
Filing Date
2025-11-25
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing methods for generating long legal text summaries lack the integration of legal expertise, making them unsuitable for complex and diverse legal texts. This results in low summary production accuracy, weak generalization ability, and high manual processing costs, making it difficult to meet the demand for rapid processing of massive amounts of legal texts.

Method used

A method for generating long legal text summaries based on a large language model is adopted. The text is preprocessed using OCR technology, segmented according to the legal text structure and labeled with core information. Combined with the four elements of criminal law and sentencing template prompts, semantic features are extracted using a pre-trained model, cross-attention feature alignment and Transformer layer processing are performed, the output layer parameters are fine-tuned, and the weighted cross-entropy loss function is used to optimize the generation of the summary.

Benefits of technology

It achieves high-precision legal text summarization, has strong cross-scenario adaptability, significantly reduces manual costs, improves processing efficiency, and the generated summaries meet legal norms and practical needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121880550A_ABST
    Figure CN121880550A_ABST
Patent Text Reader

Abstract

The invention provides a legal long text abstract generation method based on a large language model, and belongs to the technical field of crossing of artificial intelligence and legal text processing. The technical problems that a traditional legal long text abstract generation method is insufficient in semantic understanding depth, weak in field adaptability, poor in cross-scene generalization ability and high in labor cost are solved. The method comprises the following steps: firstly, carrying out cleaning, logic segmentation and crime determination and sentencing core element pre-labeling on a legal text; secondly, constructing a criminal law four-key modular prompt template; then, strengthening the learning of key legal elements; and finally, quantifying the abstract quality through dynamic verification-optimization closed loop, automatically generating a supplementary prompt guide model for iterative correction when the abstract quality exceeds a threshold value, and outputting an abstract conforming to a criminal law theory framework. According to the method, the abstract can be automatically generated, the labor cost is reduced, and efficient support is provided for scenes such as legal document processing and judicial aid decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of artificial intelligence and legal long text processing, specifically involving a method for generating legal long text summaries based on a large language model. Background Technology

[0002] Current methods for generating long legal text summaries mainly rely on two types of approaches: rule matching and template-driven methods, and traditional machine learning models.

[0003] 1. Rule-based matching and template-driven method: This method involves manually defining keyword matching rules, such as identifying the content following "Party A:" or "Party B:" as party information. Alternatively, it involves constructing regular expression templates, such as matching 18-digit ID card numbers, to generate summaries. For example, in contract summary generation, regular expressions can be used to extract information about the court with jurisdiction. However, this method has significant drawbacks: legal texts are flexible in expression; for instance, the same "liability for breach of contract" can be expressed as "compensation for delayed delivery" or "losses incurred for failure to deliver on time," leading to semantic ambiguity. Similarly, "rights stipulated in this clause" may refer to the rights of either the buyer or the seller. Rule-based matching and template-driven methods cannot cover all scenarios and are prone to omissions (e.g., failure to identify the party information corresponding to "supplier" and "purchaser") and errors (e.g., mistakenly classifying "disclaimer clauses" as "rights clauses"), resulting in extremely poor adaptability to complex legal texts (e.g., judicial interpretations containing multiple nested clauses).

[0004] 2. Traditional Machine Learning Models: These models, based on CNNs, LSTMs, etc., are trained using a large amount of labeled legal text data to generate summaries. For example, the model proposed in the paper "Research on Extraction of Controversial Points from Legal Judgments Based on LSTM" (Computer Applications and Software, Vol. 37, 2020) requires training on multiple labeled judgments to achieve the expected F1 score for extracting controversial points. These models lack the integration of legal domain knowledge, cannot understand the specific connotations of professional terms such as "bona fide acquisition" and "apparent agency," and struggle to capture the logical connection between "legal citation and clause validity," resulting in low accuracy when processing highly specialized legal texts. Furthermore, these models are optimized only for a single text type (e.g., only applicable to judgments). When switching to other text types such as contracts or complaints, a large amount of data needs to be re-labeled, leading to weak cross-scenario generalization ability.

[0005] Furthermore, the volume of texts that need to be processed in legal practice is enormous. For example, companies need to review more than 100,000 contracts annually, and courts need to process more than 500,000 judgments annually. Traditional methods rely on manual verification of each sentence to generate summaries, which makes the processing of each complex text time-consuming, labor-intensive, and inefficient, making it difficult to meet the needs of rapid processing of massive amounts of legal texts.

[0006] The core problem with existing technologies is that they do not fully integrate legal expertise and deep semantic understanding capabilities, and cannot adapt to the complexity and diversity of legal texts, resulting in the inability of summary production accuracy, generalization ability and processing efficiency to meet the needs of practical applications. Summary of the Invention

[0007] To address the technical problems in existing methods for generating long legal text summaries, such as the semantic rigidity of rule matching and template-driven methods, the lack of domain knowledge and weak cross-scenario generalization ability of traditional machine learning models, and the high cost and low efficiency of manual processing, this invention provides a method for generating long legal text summaries based on a large language model. This method ensures that the legal text summaries generated by the model not only conform to the professional theoretical framework but also accurately cover key information, thereby meeting the requirements for professionalism and accuracy of summaries in legal practice.

[0008] This invention is achieved through the following measures: a method for generating long legal text summaries based on a large language model, comprising the following steps:

[0009] Step 1. Text Preprocessing Stage: The user inputs a long legal text. For long legal texts of different formats, this invention first extracts the text content using OCR technology or text parsing tools; then, it cleans the text, removing redundant characters and correcting errors in legal terminology; then, it logically segments and standardizes the format according to the legal text structure of "case facts - charges - evidence - grounds for conviction - sentencing factors"; finally, it outputs preprocessed text in a unified format.

[0010] Step 2. Core Information Pre-annotation: The segmented text is annotated according to the "Four Elements of Criminal Law", including: the object of the crime, the objective aspect of the crime, the subject of the crime, the subjective aspect of the crime, and the text fragments related to the "Specific Circumstances of the Special Provisions of Criminal Law", to provide an annotation basis for subsequent knowledge integration.

[0011] Step 3. Four-Element Hint Module: This module employs a three-part structure of "definition + extraction rules + example" to transform the four-element theory into a structured hint that the model can understand, as shown below (using theft as an example):

[0012] Definition of the object of a crime: Social relations protected by criminal law that are infringed upon by a crime, such as property rights and personal rights. Required elements: The specific object of infringement, such as "the victim's cash"; the related crime, such as "theft"; and the relevant provisions of the Criminal Law, such as Article 264. Example: "Theft infringes upon the victim's property rights (Article 264 of the Criminal Law)."

[0013] Definition of the objective aspect of a crime: The objective manifestation of a criminal act (method of the act, harmful result, causal relationship). It needs to extract: type of act, such as "burglary" or "armed assault"; harmful result, such as "stealing XXX yuan" or "causing minor injury"; causal relationship, such as "the harmful act directly caused the result". Example: "Burning down a house and stealing XXX yuan; the act is directly related to the loss of property."

[0014] Definition of a criminal subject: A natural person / entity that commits a crime must meet the requirements of age and criminal responsibility. Required information includes: age (e.g., "25 years old"); criminal responsibility (e.g., "full criminal responsibility"); and special status (e.g., "state employee"). Example: "The defendant is 25 years old and has full criminal responsibility."

[0015] Definition of the subjective aspect of a crime: the mental attitude towards the criminal act (intent / negligence, criminal purpose). Required elements: intent / negligence, such as "direct intent"; criminal purpose, such as "illegal possession". Example: "With the intent to illegally possess, knowingly stealing another person's property."

[0016] Step 4. Specific Provisions Related Circumstances Hint Module: For common crimes, such as theft, intentional injury, and robbery, supplement the specific circumstances related to conviction in the Special Provisions, directly linking them to the criminal law articles.

[0017] Step 5. Sentencing Template Suggestion: Refer to the Criminal Law and sentencing guidelines to construct a standardized template for aggravating / mitigating circumstances:

[0018] Aggravating Circumstances Template: "Circumstance Name (Based on Article X of the Criminal Law): Specific Act (e.g., "Recidivist, committing another crime within 3 years after serving a sentence")".

[0019] Template for mitigating circumstances: "Circumstance name (based on Article X of the Criminal Law): Specific behavior (e.g., "voluntarily surrendering and truthfully confessing")".

[0020] Step 6. Semantic Feature Extraction: The annotated preprocessed text is linearly concatenated with the constructed "four-element prompts + specific circumstances prompts + sentencing template prompts" in sequence to form a model. The input sequence S is expressed by the following formula:

[0021]

[0022] in, The four-element prompt text, For the purpose of dividing the plot, Sentencing template prompts, The text is the preprocessed legal text; the "+" sign indicates a text concatenation operation.

[0023] The input sequence S is then fed into a pre-trained large language model for the legal domain, such as LawGPT. The model's Transformer layer outputs semantic feature vectors. ( =512 or 1024, i.e., 512-1024 dimensional features, where d is set according to the model size (512 dimensions are suitable for lightweight scenarios, and 1024 dimensions are suitable for high-precision requirements). Each dimension in the vector corresponds to a specific dimension of legal semantics. For example, one dimension might correspond to "whether it includes the circumstance of surrender," and another dimension might correspond to "the amount of theft." This semantic feature vector will be used as input for subsequent feature alignment by a cross-attention-based adapter, and will be analyzed and processed by a deep model containing 3-6 Transformer modules for correlation analysis.

[0024] Step 7. Model Parameter Fine-tuning: Freeze 90% of the model's basic parameters, retain the legal knowledge from the pre-trained model, and fine-tune only the output layer parameters. Optimize the output layer using the weighted cross-entropy loss function, as follows:

[0025]

[0026] in, The number of training samples, This refers to the number of legal element categories, such as "completeness of the four elements" and "accuracy of sentencing circumstances." This is the true label of the k-th type of legal element (such as the object of crime, the subject of crime, etc.) in the i-th sample, where 1 indicates that the element exists and 0 indicates that the legal element does not exist; The model predicts the probability of the existence of the k-th type of element; These are the weighting coefficients.

[0027] Step 8. Initial Summary Generation: The adjusted pre-trained model (hereinafter referred to as the "model") is generated based on the semantic feature vectors extracted in Step 3. Generate an initial summary by combining the suggested constraints. Set the generation parameters: temperature coefficient Control output randomness; summary length This is the length of the original text.

[0028] Step 9. Automatic verification and loss calculation:

[0029] Defining the loss of integrity of legal elements Key elements missing from the quantitative summary:

[0030]

[0031] in, The importance coefficient of the element. Labels for real elements. These are the element labels extracted from the abstract.

[0032] This indicates that no element has been omitted.

[0033] Define sentencing template matching loss To assess the difference between sentencing statements and pre-set templates:

[0034]

[0035] in, Sentencing excerpt from the abstract With preset template The number of characters matched. The maximum length of both. This indicates a complete match.

[0036] Step 10. Secondary optimization: If the total loss Generate supplementary prompts, such as "Please provide the age information of the perpetrator," to guide the model to regenerate the summary. Until Output the final summary .

[0037] In practical use, this invention:

[0038] 1) Input any type of criminal law text, and the model will automatically complete the cleaning, segmentation, and preliminary identification of core information.

[0039] 2) Load the preset prompts for the four elements of criminal law, the specific circumstances of the criminal law, and the sentencing template to form the model input sequence.

[0040] 3) The large language model generates a summary within 1-2 minutes, covering the four elements, specific circumstances, and legal basis, and presents the sentencing circumstances in a standardized manner.

[0041] 4) If information is missing, the model will automatically guide the model to perform secondary optimization and finally output a summary that complies with legal regulations.

[0042] 5) Only the prompt content needs to be adjusted to adapt to different types of criminal law texts, without the need for manual intervention in the core process.

[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0044] 1. High accuracy in abstract generation: The accuracy is significantly improved compared to the traditional LSTM model. It can accurately capture the four elements of criminal law, fully cover the key constituent circumstances of crimes in the special provisions, and the expression of sentencing-related elements conforms to the norms of criminal law provisions. It greatly reduces the problem of abstract distortion caused by omission of elements, logical deviation or non-standard expression, and ensures that the generated content is highly consistent with the criminal law theoretical framework and judicial practice requirements.

[0045] 2. Strong cross-scenario adaptability: No need to re-annotate data, only adjust the text type identifier and key element definition in the prompt to adapt to multiple types of legal texts such as contracts, judgments, and complaints, significantly reducing the generalization cost;

[0046] 3. Low labor costs and high processing efficiency: The entire process from text input to summary generation is completed automatically, greatly reducing the time required to process a complex contract and significantly improving efficiency compared to manual processing; automated model extraction reduces a large amount of manual annotation and verification work. Attached Figure Description

[0047] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.

[0048] Figure 1 This is a schematic diagram of the structure of an embodiment of the present invention. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. Of course, the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0050] Example 1

[0051] Sales contract text summary generation: Referring to Figure 1, this invention is a method for generating long legal text summaries based on a large language model. The specific implementation steps in generating sales contract summaries are as follows:

[0052] S1. Data Preprocessing: Input a PDF sales contract and use the Tesseract OCR engine to extract the text, which has a high text extraction accuracy; Text Cleaning: Remove redundant page number markers such as "Page 1 of 5", correct the error of writing "subject matter" as "land object", and delete meaningless spaces; Segmentation and Standardization: Segment the text according to the logic of "Contract Preface - Subject Matter Clause - Price Clause - Performance Clause - Liability for Breach of Contract - Dispute Resolution - Contract End", and convert the text to TXT format with UTF-8 encoding to obtain the preprocessed text.

[0053] S2. Hint Construction: A three-part structure is used to construct the hints: "Domain Knowledge: This text is a sales contract. 'Party Information' must include the names, unified social credit codes, and addresses of both the buyer and seller; 'Subject Matter Clauses' must extract the name, specifications, and quantity of the subject matter; 'Liability for Breach of Contract' must clearly define the circumstances of delayed delivery and substandard quality, as well as the method of calculating compensation; Task Instructions: Extract the elements of party information, subject matter clauses, price clauses, liability for breach of contract, and dispute resolution methods from the text; finally, perform statistical analysis."

[0054] S3. Semantic Feature Extraction: LawGPT-7B was selected as the large language model. During deployment, a high proportion of the model's basic parameters were frozen, and only the output layer was fine-tuned. The concatenated text was input, and 512-dimensional semantic feature representations were extracted. The feature extraction time met the requirements.

[0055] S4. Feature Alignment and Fusion: Extracting structured features from the sales contract: "Contract Header" corresponds to level code "01", "Breach of Contract Clause" corresponds to level code "05", and marking the relationship between Article 595 of the Civil Code and the "Subject Matter Clause"; Cross-Attention Adapter: Using structured features as the query matrix. The semantic features are key matrices Sum matrix The fusion features are calculated using the following formula:

[0056]

[0057] S5. Summary Generation:

[0058] Transformer summarization generation model: It consists of multiple Transformer module layers, each module containing a feedforward neural network with multi-head attention, layer normalization, and GELU activation function;

[0059] The embedding layer maps the fused features to 512-dimensional input features, which are then input into the Transformer module layer. Multi-head attention is used to capture the quantity-amount correspondence between the "subject matter terms" and the "price terms". The feature fusion layer concatenates the output features of the three modules through residual connections.

[0060] S6. Results Output and Model Optimization:

[0061] The projection layer maps the fused features into 5 dimensions (corresponding to 5 categories of elements), outputs a probability distribution through Softmax, takes the type with the highest probability as the element category, and outputs a list of elements: party information, subject matter terms, price terms, liability for breach of contract, and dispute resolution method;

[0062] Model optimization: The model was trained using the cross-entropy loss function, with the AdamW optimizer learning rate set to 3e-5. After multiple iterations, the F1 score was significantly improved on the validation set containing a large number of buy and sell contracts, meeting the usage requirements.

[0063] The structured prompt mechanism for criminal cases breaks down the four elements, specific circumstances, and sentencing factors into structured text prompts, eliminating the need to build a knowledge graph; it transforms abstract criminal law theories into instructions that the model can understand, solving the problem of missing legal logic in the summary; at the same time, it reduces the complexity of technical implementation and facilitates rapid adaptation to different crimes.

[0064] Template-based constraints on sentencing factors: A template for mitigating / aggravating circumstances is constructed based on the Criminal Law and its judicial interpretations, combined with... The loss function ensures that sentencing statements are standardized and linked to legal basis, avoiding ambiguity.

[0065] Lightweight Adaptive Architecture (Layered Transfer Strategy for Legal Pre-trained Models): Freeze 90% of the basic parameters of the pre-trained model (retain legal knowledge), and only fine-tune 10% of the output layer parameters to reduce computational costs. Adapting to different crimes only requires adjusting the prompts.

[0066] Example 2

[0067] This embodiment focuses on the scenario of generating summaries for criminal appeal documents, a special type of criminal law text. It verifies the optimization effect of the method of the present invention compared with traditional methods that incorporate the core logic of general text summarization, as follows:

[0068] This embodiment focuses on 50 criminal appeal documents issued by the People's Procuratorate. These documents cover appeals against theft, intentional injury, fraud, and embezzlement. Each document includes core modules such as the grounds for appeal, flaws in the original judgment, legal basis, and claims, along with complex content such as citations of key passages from the original judgment and explanations of corroborating evidence, meeting the professional and logical requirements of criminal appeal text practice. The experiment is based on a pre-trained large language model in the legal field, and a conventional deep learning experimental environment is constructed.

[0069] The traditional implementation process of integrating general text summarization logic is as follows: It employs general text feature extraction logic to extract key themes, semantic complexity, and paragraph relevance from documents, performing only basic cleaning of redundant symbols and repeated paragraphs without segmentation or pre-annotation specific to the structure and elements of criminal appeal documents; it refers to general dynamic summary generation strategy rules, assigning initial weight coefficients to the basic model based on key themes and semantic complexity, calculating applicability scores and determining input parameters based on real-time model performance indicators, without setting any specific parameter constraints for appeal elements; it introduces a general text quality prediction model, constructing a quality dataset based on historical general legal text summary data, extracting time-series features and establishing a quality correlation matrix through recurrent neural networks, predicting summary quality scores, and adjusting strategy weight parameters; it uses a general iterative optimization algorithm, setting an optimization objective function that includes constraints on summary quality, generation efficiency, and resource consumption, adjusting dynamic weight parameters through a genetic algorithm, and outputting a summary generation strategy that meets convergence conditions, without setting any specific verification dimensions for criminal appeal elements throughout the process.

[0070] The implementation process of the optimized solution of this invention is as follows:

[0071] Based on general text feature extraction, a special processing logic for criminal appeal documents is added, and the text is segmented according to the logic of "appeal subject - appeal object - original trial flaws - legal basis - claims".

[0072] The document pre-marks key elements in the grounds for appeal, such as "errors in fact-finding," "deviations in the application of law," and "unreasonably lenient / severe sentencing," as well as the corresponding legal provisions of the Criminal Procedure Law and the Criminal Law cited. This corrects deviations in the expression of professional legal terminology in the document and enhances the adaptability of the text to criminal appeal practice.

[0073] Using the common weight allocation and applicability scoring mechanism, a new optimized design specifically for criminal appeals has been added. The appeal documents are equipped with modular prompt templates of "core elements of the appeal + legal provisions and demands", and a three-dimensional relational logic of "appeal circumstances - original trial issues - legal basis" is established, so that the model input parameters focus on the extraction needs of core elements of criminal appeals.

[0074] The basic framework of the general quality prediction model is retained, and the dimensions of quality assessment are optimized in a targeted manner. In addition to the general indicators of semantic similarity and grammatical correctness, exclusive indicators such as the completeness of protest elements, the accuracy of the original trial defect location, and the compliance of legal citation are added, and the quality correlation matrix is ​​reconstructed accordingly.

[0075] Based on a general optimization objective function, the weight coefficients were adjusted, increasing the quality weight γ of the core element of criminal appeal to 0.65. At the same time, a lightweight strategy of "freezing most of the underlying parameters and fine-tuning the top-level parameters" was adopted, and a dynamic verification-optimization closed loop was added. Supplementary prompts were automatically generated and iteratively corrected for summaries that did not meet the specific thresholds for criminal appeals. The prompt logic was optimized with a focus on issues such as "vague descriptions of flaws in the original trial" and "mismatch between legal citations and the grounds for appeal".

[0076] The verification results show that although the traditional method, which incorporates general text summarization logic, can generate basic summaries for criminal appeal documents and improves generation efficiency through dynamic strategy adjustments, it still has significant limitations:

[0077] Due to the lack of specific preprocessing and element constraints for criminal appeal documents, the generated summaries frequently suffer from problems such as omission of core appeal elements and deviation in the identification of flaws in the original trial, failing to accurately reflect the core demands of the appeal documents. Furthermore, the quality prediction model is not adapted to the professional characteristics of criminal appeals, lacking assessment of key dimensions such as "compliance of legal citations" and "logic of appeal reasons," resulting in insufficient legal professionalism and practical reference value in the summaries.

[0078] The optimized solution of this invention incorporates the exclusive processing logic of criminal appeal documents into the basic framework of general text summarization. It retains the core advantages of dynamic strategy adjustment, quality prediction and iterative optimization in general methods, while specifically addressing the professional needs of criminal appeal text summarization. The modular prompt template avoids the problems of missing appeal elements and deviation in the location of original trial flaws from the source. The exclusive verification dimension greatly improves the legal accuracy and practical adaptability of the summary, and the lightweight strategy effectively balances the quality of summary generation and the consumption of system resources.

[0079] In summary, the solution of this invention is significantly superior to traditional methods that incorporate general text summarization logic in terms of the completeness of criminal appeal element extraction and the accuracy of locating defects in the original trial. This fully demonstrates the technical superiority of the solution, which is specifically optimized for a particular type of criminal law text, and can better meet the needs of the procuratorate for more precise and professional text summarization in criminal appeal work.

[0080] Example 3

[0081] This embodiment addresses the scenario of generating complex criminal law text summaries involving multiple crimes in judicial practice, verifying the effectiveness of the criminal law text summary generation method based on a large language model as described above:

[0082] Multiple publicly available criminal law practice texts were selected, covering frequently cited crimes such as theft, intentional injury, fraud, and traffic accidents. The text length conforms to the standard range of indictments for multiple charges and consolidated criminal case files, and each text includes various complex sentencing factors to ensure the experimental scenario closely reflects the complexity and diversity of real judicial texts. The experiment is based on a pre-trained large language model in the legal field, and the experimental environment is built using conventional deep learning frameworks and programming languages.

[0083] The experiment was set up with two comparison schemes: the scheme of this invention and the traditional general text summarization scheme.

[0084] The implementation process of the present invention is as follows:

[0085] First, the experimental text was preprocessed specifically for criminal law, segmented according to the judicial logic of "case facts - charges - evidence - reasons for conviction - sentencing factors", and the text fragments related to the four elements of the crime and sentencing factors corresponding to each crime were pre-marked, and deviations in the expression of legal terminology were corrected.

[0086] Subsequently, a modular prompt template for the four elements was constructed, and core element extraction rules were configured specifically for different crimes. The key distinguishing points such as the behavioral characteristics and harmful consequences of each crime were clarified, and a three-dimensional association logic of "crime-elements-corresponding legal provisions" was established.

[0087] In the model optimization phase, a lightweight strategy of "freezing most of the underlying parameters and fine-tuning the parameters of the top output layer and attention layer" is adopted, and the learning weights of crime differentiation elements and sentencing factors are strengthened by using the weighted cross-entropy loss function.

[0088] Finally, a dynamic verification-optimization closed loop is enabled, and quantitative verification thresholds are set. Verification is carried out from three dimensions: completeness of the four elements, distinguishability of crime elements, and compliance of sentencing circumstances. For summaries that do not meet the thresholds, targeted supplementary prompts are automatically generated, and iterative corrections are made until the requirements are met.

[0089] The implementation process of traditional general text summarization schemes is as follows: only generalized cleaning processing is performed on the experimental text to remove redundant symbols and repeated paragraphs, without logical segmentation and element pre-annotation specific to criminal law; general prompt words without specific targets are used to guide the summary generation, without setting crime-specific extraction rules and element constraints; the model adopts a full parameter fine-tuning method, without reinforcement learning mechanisms for core elements of criminal law; a general iterative optimization algorithm is used, and dynamic weight parameters are adjusted through genetic algorithms to output a summary generation strategy that meets the convergence conditions, without setting criminal law-specific verification dimensions throughout the process.

[0090] Through comparative verification of the two sets of schemes, the traditional general scheme has obvious limitations due to the lack of guidance and optimization mechanisms specific to criminal law scenarios: the core elements of different crimes are easily confused, especially the key distinguishing points such as subjective mentality and behavioral characteristics are difficult to define accurately; the description of sentencing factors lacks standardization, cannot accurately associate with the corresponding legal provisions, and is prone to omission of constituent elements; at the same time, the full parameter fine-tuning leads to large resource consumption during model operation and insufficient flexibility in adapting to multi-crime scenarios.

[0091] The present invention achieves targeted sorting and standardization of text elements through criminal law-specific preprocessing; the three-dimensional association logic of modular prompt templates avoids the problem of confusion of crime elements from the source; the lightweight optimization strategy reduces resource consumption while strengthening the learning of key elements; and the dynamic verification-optimization closed loop further ensures the legal professionalism and accuracy of the summary.

[0092] In summary, the method of this invention is significantly superior to traditional general text summarization methods in terms of the accuracy of distinguishing core elements, the standardization of the description of sentencing circumstances, and the efficiency of resource utilization for complex criminal law text summarization scenarios involving multiple crimes. It can better meet the professional and accurate needs of judicial practice for criminal law text summarization.

[0093] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for generating long legal text summaries based on a large language model, characterized in that, Includes the following steps: S1. Perform data preprocessing on long legal texts, extract content using text parsing tools, clean and logically segment the text to obtain preprocessed text in a unified format; S2. Construct modular prompts that integrate legal knowledge, including prompt units based on the theory of the four elements of criminal law, prompts related to specific circumstances, and prompts for sentencing templates; S3. Concatenate the preprocessed text with the modular prompts to form the input sequence S, which is represented as: ; Where P represents various prompt texts, and T represents preprocessed text; S4. Input sequence S into a pre-trained large language model in the legal domain to extract semantic feature vectors, and then align and fuse the semantic features with the structured features of the legal text through an adapter based on cross-attention. S5. Use the Transformer model to perform deep processing on the fused features, capture the correlation between elements, and output the final summary results.

2. The method for generating long legal text summaries based on a large language model according to claim 1, characterized in that, In step S1, the data preprocessing step includes: extracting text content from the original long legal text using OCR technology or text parsing tools, and cleaning, logically segmenting, and standardizing the format of the extracted text content.

3. The method for generating legal long text summaries based on a large language model according to claim 1, characterized in that, In S2, the modular prompts include at least prompt units based on the theory of the four elements of criminal law. Each prompt unit contains the definition of the elements, their legal significance, and the key information to be extracted.

4. The method for generating long legal text summaries based on a large language model according to claim 3, characterized in that, The modular prompts also include specific crime-related prompts, which are directly linked to criminal law provisions and used to supplement specific circumstances related to conviction.

5. The method for generating legal long text summaries based on a large language model according to claim 1, characterized in that, In step S3, the preprocessed text is linearly concatenated with the modular prompt and then input into the large language model.

6. The method for generating legal long text summaries based on a large language model according to claim 1, characterized in that, In S4, the cross-attention-based adapter uses the structured features of the legal text as the query matrix and the semantic features as the key and value matrices, and achieves feature alignment and fusion with the feedforward neural network through 2 to 4 cross-attention layers.

7. The method for generating legal long text summaries based on a large language model according to claim 1, characterized in that, In S5, the Transformer model contains 3-6 Transformer modules to balance model complexity and efficiency, and performs deep processing on the fused features through multi-head attention mechanism, layer normalization and residual connection.

8. The method for generating legal long text summaries based on a large language model according to claim 1, characterized in that, In step S5, the final summary result is output, the probability distribution of the feature type is output through the projection layer, and the maximum probability value is used as the final feature category.

9. The method for generating legal long text summaries based on a large language model according to claim 1, characterized in that, The pre-trained large language model for the legal domain in step S4 is optimized using a weighted cross-entropy loss function, wherein the loss function is: ; Where N is the number of training samples and M is the number of legal element categories. For real labels, To predict probabilities, The weights are used as coefficients; and the model performance is verified by combining multiple types of legal text datasets, and evaluated by accuracy and F1 score to optimize the model's generalization ability and cross-scenario adaptability.

10. The method for generating legal long text summaries based on a large language model according to any one of claims 1-9, characterized in that, The method can adapt to various types of legal texts such as contracts, judgments, and indictments by adjusting the text type identifiers and element definitions in the three major modules (namely, the four-element prompt module, the special provisions and circumstances prompt module, and the sentencing template prompt module).