Automatic electronic file document generation method and system based on prompt learning

Through an automated electronic file generation method based on prompt learning, using large language model and clustering technology, the problem of insufficient efficiency and accuracy of electronic file generation in the existing technology is solved, and document generation with high authenticity and logical consistency is achieved, and labor costs are reduced.

CN119962506APending Publication Date: 2025-05-09SHANGHAI JIAOTONG UNIV

Patent Information

Application Number
CN202510140885.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

It is difficult for the prior art to generate electronic file documents with high authenticity, accuracy and logical consistency, especially in the process of facing different court template styles and high labor cost annotation.

Method used

An automated electronic file generation method based on prompt learning is adopted to realize the automated generation of electronic file documents by collecting judicial data sets, desensitization processing, fine-tuning of large language models, constructing prompt words, clustering processing and backfill templates.

Benefits of technology

It improves the efficiency and accuracy of electronic file documents generation, reduces labor costs, ensures the authenticity and logical consistency of the generated documents, and has broad applicability and versatility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962506A_ABST
    Figure CN119962506A_ABST
Patent Text Reader

Abstract

The invention provides an automatic electronic file generation method and system based on prompt learning, and the method comprises the steps: S1, collecting a judicial data set, and constructing a case element extraction task data set; s2, performing desensitization processing on the case element extraction task data set to obtain a desensitized data set; s3, constructing a fine-tuned large language model by using the desensitization data set; s4, creating an electronic file template used for generating an electronic file program document, and determining target case elements needing to be filled in the electronic file template; s5, constructing cue words, and extracting case element names and entity information of the target case elements; s6, performing clustering processing on the case element names and the entity information to obtain a clustering result; and step S7, backfilling the entity information into the electronic file template to complete automatic generation of the electronic file document. According to the method, the electronic file document with high authenticity, high accuracy and good logic consistency can be quickly generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the large language model technology in the field of artificial intelligence and the field of natural language processing text analysis, and in particular, to an automatic electronic file document generation method and system based on prompt learning. Background Art

[0002] In-depth intelligent analysis of electronic files is an important part and one of the development directions of smart justice. Through in-depth and intelligent analysis of the content of electronic files, it can have a positive impact on the quality and efficiency of current judicial management. However, when using artificial intelligence technology to analyze electronic files, an important prerequisite is to have a large number of high-quality electronic files and accurately annotate them to train artificial intelligence models. But this premise faces two major practical difficulties: first, electronic files contain a large amount of highly sensitive privacy information, and it is usually difficult to obtain complete electronic files when conducting artificial intelligence training; second, the current cognitive tasks of artificial intelligence require a large amount of annotated data, and the annotation of electronic files requires professional judicial knowledge, and the annotation process will bring high labor costs.

[0003] With the rapid development of large language models, using large language model agents to generate and pre-process documents has become an efficient way. However, the illusion problem of large language models and the time cost of their operation have caused obstacles to a certain extent, making it difficult for them to undertake tasks that require high timeliness and accuracy in the intelligent cognition of electronic files.

[0004] In the task of generating electronic files, different courts have different styles of covers, directories, preparation tables, and bottoms, which are mainly reflected in the data requirements, fonts, font sizes, page margins, background colors, and font colors of the templates. Although the traditional implementation method based on the JAVA AWT component library can achieve the generation effect, it requires extensive modification of the source code in the generation process for different templates, which is inconvenient and inflexible.

[0005] In the existing research on generating electronic documents using large language model agents, most of them use a large corpus to pre-train the large language model, and then use a small batch of corpus in the relevant field to fine-tune the large language model agent to generate electronic documents that meet the professional needs of the field. However, there is a hallucination phenomenon in the current large language model, that is, for the specific input of the user, sometimes chaotic and unpredictable output is generated, which makes it difficult to guarantee the accuracy, authenticity and logical consistency of the generated electronic documents.

[0006] When generating electronic case files based on templates, the styles of the cover, directory, preparation table, bottom of the file, etc. filed in different courts are different, mainly reflected in the data requirements, fonts, font sizes, page margins, background colors, and font colors on the templates. Although the traditional implementation method based on the JAVA AWT component library can achieve the generation effect, it requires a large-scale modification of the source code in the face of different templates during the generation process, which is inconvenient and inflexible. In addition, the sorting of the original case information requires a large number of professionals with judicial knowledge, resulting in high labor costs and low generation efficiency. Therefore, it is a very difficult task to generate electronic case files with high authenticity, high accuracy, and good logical consistency on a large scale, and use them as training and test data for the intelligent cognitive network model of electronic cases.

[0007] In the existing electronic case file generation method, the main focus is on the generation of electronic case files during the judicial process and the digitization of paper electronic case files. Through the search of patent documents, it is found that the patent with the authorization announcement number CN110362799B discloses a method, device, computer equipment and storage medium for generating and processing a ruling based on online arbitration. It uses the case identifier of the current case to obtain the corresponding electronic case file, and obtains the corresponding information from the electronic case file and fills it back into the ruling template to generate a ruling document. This method can only generate ruling documents and relies on existing electronic case files, and cannot generate other types of electronic cases.

[0008] In the existing electronic document generation method using a large language model agent, it is mainly aimed at document retrieval tasks and document generation with low accuracy and timeliness requirements. Through the search of patent documents, it was found that the patent with authorization announcement number CN118069815B discloses a large language model feedback information generation method, device, electronic device and medium, which reconstructs the documents in the knowledge base and stores them in the vector database after vectorization, and then uses the large language model to vectorize the user's input and index it in the vector database, thereby achieving efficient indexing. This method is mainly aimed at the indexing task of electronic documents. In addition, the patent with authorization announcement number CN111723564B discloses an event extraction and processing method for electronic files accompanying the case, which obtains the required file data from the electronic file circulation processing platform accompanying the case and stores it in the database, then constructs an event trigger word dictionary, matches the electronic file event description paragraph, and then performs a text preprocessing method, then extracts event attributes, and finally performs event aggregation, aggregates atomic events into main events, and stores them in the event database. This method uses a Transformer-based bidirectional model to extract features, and the extracted content is the event relationship in the electronic file.

[0009] In summary, in view of the above-mentioned problems of the prior art, researching an automated electronic file document generation method and system based on prompt learning has become a key task that needs to be solved urgently. Summary of the invention

[0010] In view of the defects in the prior art, the purpose of the present invention is to provide an automated electronic file document generation method and system based on prompt learning.

[0011] According to the present invention, an automatic electronic file generation method based on prompt learning includes the following steps:

[0012] Step S1, collect judicial data sets, and according to preset standards, filter out a subset of complete original case elements, annotate the subset, and construct a case element extraction task data set;

[0013] Step S2, desensitizing the case element extraction task data set to obtain a desensitized data set;

[0014] Step S3, using the desensitized data set, fine-tuning the large language model that has been pre-trained on a large-scale corpus to obtain a fine-tuned large language model;

[0015] Step S4, creating an electronic case file template for generating electronic case file procedure documents, and specifying the target case elements to be filled in the electronic case file template;

[0016] Step S5, based on the fine-tuned large language model, construct prompt words according to the target case elements and the case element extraction task data set, and then input them into the fine-tuned large language model to extract the case element name and entity information of the target case element;

[0017] Step S6, clustering the case element names and entity information to obtain clustering results;

[0018] Step S7, fill the entity information in the clustering result back into the electronic file template to complete the automatic generation of the electronic file document.

[0019] Preferably, in step S2, the desensitization process is to randomly replace sensitive information, and the sensitive information includes name, address, ID number and telephone number.

[0020] Preferably, in step S4, in the electronic case file template, specific placeholders are used to mark the locations where case elements need to be filled in, and in different electronic case file types, case elements of the same type use the same name and data format to ensure that the generated electronic case file is consistent and standardized.

[0021] Preferably, in step S5, the prompt word is presented in the form of a Prompt sentence, and the case element name and entity information of the target case element are extracted by inputting the Prompt sentence into the fine-tuned large language model.

[0022] Preferably, in step S5, the form of the Prompt statement is: extract the following case elements from the [judicial data set type] content [judicial data set content]: [case element name], replace non-existent elements with "none", and give the case element name and entity information in the form of a JSON string.

[0023] Preferably, step S6 includes the following sub-steps:

[0024] Step S6.1, for the JSON string including the case element name and entity information extracted in step S5, organize the case element name and entity information into a string in the form of "[case element name]: [entity content]", and use string vectorization technology to represent the string as a fixed-length floating-point vector;

[0025] Step S6.2, performing unsupervised clustering on the floating point vector using the K-means clustering algorithm, wherein K centroids are randomly created based on a predefined K value to obtain a clustering result, which includes text and vocabulary vectors.

[0026] Preferably, in step S6.1, the string vectorization technology uses a Transformer-based bidirectional encoding representation, which has been pre-trained on a large-scale Chinese corpus and represents the string as a floating-point vector with a shape of 1×768.

[0027] Preferably, the K-means clustering algorithm of step S6.2 includes the following sub-steps:

[0028] Step S6.2.1, calculate the sum of squares of clustering errors under different K values, and use the elbow rule to determine the optimal K value;

[0029] Step S6.2.2, assign each data point in the floating point vector to the nearest centroid;

[0030] Step S6.2.3, recalculate the location of the centroid by calculating the mean of all data points assigned to the cluster where the centroid is located, thereby reducing the total intra-cluster variance associated with step S6.2.1. The "mean" in K-means refers to finding the arithmetic mean of the data points in the current cluster to find the new centroid location;

[0031] Step S6.2.4, iterate between step S6.2.2 and step S6.2.3 until the cluster assignment of data points no longer changes, and obtain the clustering result.

[0032] Preferably, step S7 includes: for the clustering results obtained in step S6, selecting the "center point vector" with the shortest average distance to the other vectors in the category as the representative of the category in the clustering results of each category; comparing and calculating the distance between the category representative vector and the case element name vector specified in the electronic file template; setting the number of cluster centers to K, then there are K category representative vectors: [RV1, RV2, ..., RV K ], and there are N case element vectors in the electronic file template: [CEV1, CEV2, …, CEV N ]; First calculate the category representative vector RV i The similarity between each case element vector and find the n case element vectors with the highest similarity, recorded as Then for each case element vector CEV j , find the n category representative vectors with the highest similarity, recorded as If for a pair of class representative vectors and case element vectors <RV i ,CEV j >, while meeting RV i ∈R j and CEV j ∈C i , it is considered that this pair of category representative vector and case element vector represents the same case element and corresponding entity; when there are multiple pairs of category representative vectors and case element vectors that meet the requirements, for each case element vector, select the category representative vector that meets the requirements and has the highest similarity; finally, for N case element vectors, generate N <category representative vector, case element vector> pairs; select the case element name with the highest similarity as the correct classification name, and fill in the corresponding entity information into the electronic file template to generate the corresponding electronic file document.

[0033] The present invention also provides an automated electronic dossier generation system based on prompt learning, comprising:

[0034] Module M1 collects judicial data sets, and according to preset standards, selects subsets with complete original case elements, labels the subsets, and constructs case element extraction task data sets;

[0035] Module M2 performs desensitization processing on the case element extraction task data set to obtain a desensitized data set;

[0036] Module M3 uses the desensitized dataset to fine-tune the large language model that has been pre-trained on a large-scale corpus to obtain a fine-tuned large language model;

[0037] Module M4, creating an electronic case file template for generating electronic case file procedural documents, and specifying the target case elements that need to be filled in the electronic case file template;

[0038] Module M5, based on the fine-tuned large language model, constructs prompt words according to the target case elements and judicial data set, and then inputs them into the fine-tuned large language model to extract the case element name and entity information of the target case element;

[0039] Module M6, clustering the case element names and entity information to obtain clustering results;

[0040] Module M7 fills the entity information in the clustering results back into the electronic file template to complete the automatic generation of the electronic file document.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] 1. The present invention predefines the electronic case file Word template and determines the case element name and definition inside, and collects a large number of public judicial data sets, such as judgments, judicial documents, and indictments, as the original real case sources, to provide a rich and real information basis for the generation of electronic case files. Then, through the constructed prompt words, the large language model agent is used to automatically extract the case elements and their corresponding entities in the public judicial data set, which can realize the automated data processing process and improve the efficiency of electronic case file generation.

[0043] 2. In order to avoid the hallucination problem of the large language model agent, the present invention adopts word vectorization to embed the features of the extracted case elements and corresponding entities, and clusters them together with the case element names in the template, ensuring the consistency of the case element names, effectively reducing the uncertainty of the large language model agent extraction, thereby improving the accuracy and reliability of the generated electronic file documents.

[0044] 3. The present invention generates an electronic file document with high authenticity by filling the extracted entity content back into the electronic file document template. This is because the case source is a real case, which makes the case description show considerable authenticity and logical consistency, providing high-quality electronic file documents for judicial practice, which is conducive to the development and advancement of judicial work.

[0045] 4. The present invention uses public judicial data sets as the data source during the generation process, avoiding the risk of privacy leakage, ensuring information security, and meeting the requirements of judicial work for information confidentiality.

[0046] 5. The present invention can automatically mark in the process of backfilling the template without manual marking, which greatly reduces the labor cost, reduces the resource investment caused by the professional judicial knowledge required for manual marking, and improves the overall operational convenience.

[0047] 6. The present invention can generate electronic case files in a variety of different formats and is independent of real electronic case files, so it has wide applicability and versatility. In particular, the present invention focuses on the task of generating highly simulated electronic case files, and can generate documents that are highly similar to actual judicial case files. The present invention uses predefined electronic case file templates, public judicial data sets, and prompt learning projects to extract and fill in case elements, integrating a variety of technologies and data resources, and provides an innovative and efficient method for generating electronic case files. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0049] Figure 1 A flow chart of an automated electronic dossier document generation method based on prompt learning in an embodiment of the present invention;

[0050] Figure 2 It is a Word template for the electronic file document of the court appearance notice in the embodiment of the present invention;

[0051] Figure 3 Schematic diagram of the elbow rule experiment in an embodiment of the present invention;

[0052] Figure 4 This is an example of a court appearance notice generated in an embodiment of the present invention. DETAILED DESCRIPTION

[0053] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0054] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements that are not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "includes a..." do not exclude the existence of other identical elements in the process, method, article or device that includes the elements.

[0055] The present invention proposes an automated electronic case file generation method and system based on prompt learning. The method predefines the Word template of the electronic case file document and the corresponding case elements, and uses a public judicial data set to extract the case elements from the large language model agent through a prompt learning method, and then uses the Bert model to calculate the text vectors of the extracted case element names and entities, and uses the K-means algorithm to cluster these text vectors. By selecting representative elements for each cluster, it is ensured that the extracted case element names and entities are consistent with the case element names pre-set in the Word template of the electronic case file document, and then the case elements are backfilled. The above-mentioned automated electronic case file document generation method and system based on prompt learning can quickly generate electronic case files with high authenticity, high accuracy and good logical consistency.

[0056] Embodiment 1:

[0057] Figure 1 The present invention is a flowchart of an automated electronic dossier document generation method based on prompt learning in an embodiment of the present invention.

[0058] like Figure 1As shown, an automated electronic case file generation method based on prompt learning is used to solve the problem that the existing technology cannot accurately generate highly authentic simulated electronic case file documents, including: first, generating a Word template for the electronic case file document, determining the case element name and data type that needs to be backfilled, and vectorizing the case elements; second, collecting public judicial data sets, and randomly replacing the hidden information therein, such as elements such as ID card number, mobile phone number, and name; then, according to the needs of case elements and the format of the judicial data set, constructing a prompt word project, and sending the constructed prompt words to the large language model agent so that it can extract the required case elements; next, vectorizing the case element names extracted by the large language model agent and their corresponding entities; then, performing unlabeled clustering on the vectorized text, and using the vector with the smallest average distance to the remaining vectors in each class as the representative vector of the class; finally, calculating the distance between the representative vector and the case element vector in the template, selecting the case element vector with the closest distance as the case element classification of all other vectors in the class, and backfilling the corresponding entities into the template to generate an electronic case file document.

[0059] Specifically, the automated electronic dossier generation method based on prompt learning comprises the following steps:

[0060] Step S1, data preparation and annotation: collect judicial data sets, and according to preset standards, filter out a complete subset of original case elements, annotate the subset, and construct a case element extraction task data set.

[0061] Specifically, judicial data sets include judicial documents such as judgments and verdicts. Original case elements refer to various key information related to judicial cases, such as case number, party information, cause of action, verdict, etc. Based on the electronic case file element set standard of the preset cause of action, a complete subset of original case elements is screened out. The integrity standard of case elements is confirmed based on judicial professional knowledge such as element-based trials and relevant legal provisions. For example, for the cause of "private lending disputes", research is conducted based on the General Principles of the Civil Law of the People's Republic of China, the Property Law of the People's Republic of China, the Guarantee Law of the People's Republic of China, the Contract Law of the People's Republic of China, and the Civil Procedure Law of the People's Republic of China, and the subject information is proposed: lender name / name, ID number / unified credit code, contact information, borrower name / name, identity Certificate number / unified credit code, contact information, guarantor information (if any); loan contract elements: loan agreement (loan note, IOU, WeChat record, etc.), loan amount (consistency in uppercase and lowercase), loan purpose (whether it is legal), interest agreement (whether it exceeds 4 times the LPR), repayment period, liquidated damages clause, contract form (oral / written); performance: actual delivery method (cash / transfer / physical), delivery certificate (bank statement, receipt), partial repayment record, interest payment situation, debt transfer / offset situation; default situation: overdue principal, failure to pay agreed interest, unauthorized change of loan purpose (such as agreed Operating actual gambling); Guarantee situation: guarantor's liability (general / joint), mortgage / pledge registration situation, guarantee scope (principal / interest / fees), guarantee period; Defense grounds: debt has been repaid (repayment certificate), expiration of the statute of limitations (last collection time), invalid contract (such as head interest, false litigation), behavioral capacity defects; Evidence materials: original loan notes, transfer vouchers (note loan), collection records (text messages, notarized letters), witness testimony, audio / video evidence; Special circumstances: joint debts of husband and wife (whether used for family life), effectiveness of corporate borrowing (non-financial institution qualifications), debt transfer notice , agreement to settle debts with property; litigation request: calculation of principal and interest (divided into stages: interest within the period, overdue interest), costs of realizing debt rights (attorney fees, preservation fees), assumption of guarantee liability, application for property preservation; procedural matters: jurisdiction of the court (defendant's place of residence / place of contract performance), reasons for interruption of the statute of limitations, conditions for service of announcement, clues to property for execution; risk warning: 11 categories, including risk of evidence of cash delivery, criminal risk of usury (annual interest rate exceeds 36% and the circumstances are serious), legal liability for false litigation, risk of inability to execute, etc., a total of 51 case elements, the subset is labeled under this standard to construct a case element extraction task dataset.

[0062] Step S2, privacy protection processing: desensitizing the case element extraction task data set to obtain a desensitized data set.

[0063] Specifically, desensitization processing is to randomly replace sensitive information to ensure that there is no real privacy information in the data set. Sensitive information includes names, addresses, ID numbers, and phone numbers.

[0064] Step S3, model fine-tuning training: Use the desensitized dataset to fine-tune the large language model that has been pre-trained on a large-scale corpus to obtain a fine-tuned large language model, so as to give the large language model the ability to accurately extract case elements.

[0065] Step S4, template formulation: Create an electronic case template for generating electronic case procedural documents, and specify the target case elements that need to be filled in the electronic case template.

[0066] Specifically, in the electronic case file template, specific placeholders are used to mark the locations where case elements need to be filled in, and in different electronic case file types, the same type of case elements use the same name and data format to ensure that the generated electronic case file documents are consistent and standardized.

[0067] In this embodiment, a Word file template is constructed using Microsoft Word software, and variables are used as placeholders in the locations where data generation and case element backfilling are required. Each case element is marked with the symbol "[]", and the case element name is used as the internal mark. It should be noted that in different types of electronic files, for the same type of entity, the corresponding case element name should be consistent, and the data format should also be consistent.

[0068] Step S5, prompt word formulation and query: Based on the fine-tuned large language model, according to the target case elements and case element extraction task data set, construct prompt words (Prompt), and then input them into the fine-tuned large language model to extract the case element name and entity information of the target case element.

[0069] Specifically, the prompt words are presented in the form of Prompt sentences. By inputting the Prompt sentences into the fine-tuned large language model, the case element names and entity information of the target case elements are extracted.

[0070] For example: "Extract the following case elements from the [judicial data set type] content [judicial data set content]: [case element name], non-existent elements are temporarily replaced and given in JSON format." In the Prompt statement, the content within the symbol "[ ]" is replaceable content. Among them, "judicial data set type" includes common types of public judicial documents, such as judgment documents, verdicts, appeals, etc.; "judicial data set content" is the main text content of the judicial data set collected in step S1; case element name is the case element type contained in the complete subset of the original case elements screened out in step S1. At this time, the fine-tuned large language model can extract the corresponding case elements from the main content of the judgment, and output the extraction results in the form of a JSON string.

[0071] Step S6, clustering the case element names and entity information to obtain clustering results to ensure that they are consistent with the case element names that need to be filled in the electronic file template in step S4, to avoid information confusion or mismatching.

[0072] Specifically, step S6 includes the following sub-steps:

[0073] Step S6.1, for the JSON string including the case element name and entity information extracted in step S5, considering the hallucination problem of the large language model, extraction errors or random fabrications may occur. Therefore, the case element name and entity information are organized into a string in the form of "

Case Element Name

Entity Content

[0074] Specifically, the string vectorization technology uses a bidirectional encoder representation from Transformers (Bert), which has been pre-trained on a large-scale Chinese corpus and represents strings as floating-point vectors with a shape of 1×768.

[0075] Step S6.2, performing unsupervised clustering on the floating point vector using the K-means clustering algorithm, wherein K centroids are randomly created based on a predefined K value to obtain a clustering result, which includes text and vocabulary vectors.

[0076] More specifically, the K-means clustering algorithm in step S6.2 includes the following sub-steps:

[0077] Step S6.2.1, due to the uncertainty and hallucination of the large language model, it is difficult to determine the number of cluster centers. At this time, the elbow method is used for judgment. Specifically, the sum of squared errors (SSE) under different K values ​​is calculated. As the K value increases, the SSE gradually decreases. When K reaches the number of real cluster centers, the downward trend of SSE slows down to form an "elbow" phenomenon. This K value is selected as the real K value.

[0078] Step S6.2.2, assign each data point in the floating point vector to the nearest centroid. Specifically, by minimizing the Euclidean distance between them, that is, if a data point is closer to the centroid of a cluster than any other centroid, then the data point is attributed to that particular cluster.

[0079] Step S6.2.3, recalculate the position of the centroid by calculating the average of all data points assigned to the cluster where the centroid is located, thereby reducing the total intra-cluster variance associated with step S6.2.2. The "mean" in K-means refers to the arithmetic mean of the data points in the current cluster to find the new centroid position.

[0080] Step S6.2.4, iterate between step S6.2.2 and step S6.2.3 until the cluster assignment of data points no longer changes, and obtain the clustering result.

[0081] Step S7, information backfilling and program document generation: fill in the entity information in the clustering results into the electronic file template to complete the automatic generation of the electronic file document.

[0082] Specifically, for the clustering results obtained in step S6, the "center point vector" with the shortest average distance to the other vectors in the category is selected as the representative of the category in the clustering results of each category; the category representative vector and the case element name vector specified in the electronic file template are compared one by one and the distance is calculated; the number of cluster centers is set to K, and there are K category representative vectors: [RV1, RV2, ..., RV K ], and there are N case element vectors in the electronic file template:

[0083] [CEV1,CEV2,…,CEV N ] First, calculate the category representative vector RV i The similarity between each case element vector and find the n case element vectors with the highest similarity, recorded as Then for each case element vector CEV j , find the n category representative vectors with the highest similarity, recorded as If for a pair of class representative vectors and case element vectors <RVi ,CEV j >, while meeting RV i ∈R j and CEV j ∈C i , then it is considered that this pair of category representative vectors and case element vectors represent the same case element and corresponding entity; when there are multiple pairs of category representative vectors and case element vectors that meet the requirements, for each case element vector, select the category representative vector that meets the requirements and has the highest similarity; finally, for N case element vectors, generate N <category representative vector, case element vector> pairs. Select the case element name with the highest similarity as the correct classification name, and fill the corresponding entity information back into the electronic file template to generate the corresponding electronic file document.

[0084] Embodiment 2:

[0085] The present invention also provides an automated electronic file generation system based on prompt learning. The automated electronic file generation system based on prompt learning can be implemented by executing the process steps of an automated electronic file generation method based on prompt learning, that is, those skilled in the art can understand the automated electronic file generation method based on prompt learning as a preferred implementation of the automated electronic file generation system based on prompt learning.

[0086] The automatic electronic dossier generation system based on prompt learning includes:

[0087] Module M1 collects judicial data sets, and according to preset standards, selects subsets with complete original case elements, labels the subsets, and constructs case element extraction task data sets;

[0088] Module M2 performs desensitization processing on the case element extraction task data set to obtain a desensitized data set;

[0089] Module M3 uses the desensitized dataset to fine-tune the large language model that has been pre-trained on a large-scale corpus to obtain a fine-tuned large language model;

[0090] Module M4, creating an electronic case file template for generating electronic case file procedural documents, and specifying the target case elements that need to be filled in the electronic case file template;

[0091] Module M5, based on the fine-tuned large language model, constructs prompt words according to the target case elements and judicial data set, and then inputs them into the fine-tuned large language model to extract the case element name and entity information of the target case element;

[0092] Module M6, clustering the case element names and entity information to obtain clustering results;

[0093] Module M7 fills the entity information in the clustering results back into the electronic file template to complete the automatic generation of the electronic file document.

[0094] Embodiment 3:

[0095] The following is a further description of an embodiment of generating an electronic file document of a civil first-instance appearance notice.

[0096] Step 1: Collect public judicial data sets and select high-quality content for annotation of case element extraction tasks. For example, for civil judgments:

[0097]

[0098]

[0099] The internal case elements are complete, and the case elements are extracted and annotated. The annotation format is:

[0100]

[0101]

[0102] Then, the hidden private information in the judgment text is randomly replaced, and the candidate text content is selected from the random content pool. For example, "plaintiff Zhang" is replaced with "plaintiff Zhang San", and "defendant Li" is replaced with "defendant Li Si". It should be noted that the random replacement here will not affect the logical consistency and authenticity of the content of the judgment.

[0103] Step 2: Use the annotated data set obtained in step 1 to fine-tune the pre-trained large language model, and fine-tune the large language model for the case element extraction task (Case Element-Named Entity Recognition, CE-NER) capability. In the case element extraction task of the large language model used in this patent, in order to facilitate the subsequent use of the case element extraction results, the case element extraction results are represented in JSON format, and the representation format is the same as the electronic file annotation format given in step 1, expressed as Y = [y1, y2, ..., y n ], and after encoding it into a text string, the text encoder represents it as a floating point number in token form. Therefore, when using a large language model for fine-tuning training of case element extraction tasks, it can be divided into the following steps:

[0104] 1) Construct prompt words. For a given text sequence X in the input public data set, construct prompt words to obtain prompt(X). The prompt used in this patent is as follows:

【

[0106] -Role: Legal text element extraction expert

[0107] -Background: Users need to extract key information from legal documents and present them in a specific format. This information is essential for case analysis, legal research or document management.

[0108] -Profile: You are a professional legal text analyst with profound legal knowledge and text parsing capabilities.

[0109] Ability to accurately identify and extract key elements from legal documents.

[0110] -Skills: You have legal expertise, text analysis skills and data formatting capabilities, and are able to extract required information from complex legal texts and organize the output in a specified format.

[0111] -Goals: Accurately and efficiently extract key information from legal documents and present it in a manner that complies with the JSON data format.

[0112] -Constrains: The extracted information must be accurate, the format must conform to the JSON standard, and must not contain any illegal or non-standard content.

[0113] -OutputFormat: Text output in JSON data format, including all specified legal document elements.

[0114] -Workflow:

[0115] 1. Read and understand the content and structure of legal documents.

[0116] 2. Based on the content of the document, identify and extract key information such as case number, case name, court, etc.

[0117] 3. Organize the extracted information in the specified format to ensure compliance with the JSON data standard.

[0118] -Initialization: In the first conversation, please directly output the following: Hello, I am a professional legal text element extraction expert. Please provide the legal document, and I will extract the key information you need and present it in the specified format.

[0119] Please ensure the completeness and accuracy of the document content. 】

[0120] 2) Input the constructed prompt(X) into the pre-trained large language model. Assume that the original parameters of the pre-trained large language model are W, and the parameters of LoRA are A and B, respectively. The dimensions of A and B are much lower than W to ensure that the training is easy to converge. Then for the input sample prompt(X), its output T = (W + AB) · prompt(X) = W · prompt(X) + AB · prompt(X), and obtain the corresponding output text sequence T = [t1, t2, … t n ], where t i Represents the text characters to be output;

[0121] 3) Transform the output text sequence into an entity tag type and calculate the cross entropy loss between each entity tag and the true tag Y as shown in formula (1).

[0122]

[0123] Where L represents the loss function value corresponding to the sequence X, n is the number of entities in the sequence, i represents the order of the entities in the sequence, the maximum is n, k is the number of entity categories, and j represents the jth label of the entity, the maximum is k. represents the true value corresponding to the jth label of the i-th entity in the sequence. If j is the true label, then Otherwise 0. is the confidence corresponding to the jth label of the ith entity, ranging from 0 to 1.

[0124] 4) Use the loss function value calculated by formula (1) to perform reverse gradient iterative update on the LoRA parameters to update the LoRA parameters. After multiple batches of iterative updates, when the loss function tends to be stable and the extraction effect meets the use requirements, stop training and freeze the LoRA parameters.

[0125] Figure 2 This is a Word template for the electronic file document of the appearance notice in the embodiment of the present invention.

[0126] Step 3: Create a Word template for the electronic dossier of the first instance civil appearance notice. The completed electronic dossier document Word template is as follows: Figure 2As shown, it includes 11 case elements: court name, case number, name of litigation agent, name of plaintiff, name of defendant, cause of action, opening date, contractor, contractor phone number, court address, and date of appearance notice. Among them, the format of court name, case number, name of litigation agent, name of plaintiff, name of defendant, cause of action, contractor, contractor phone number, and court address should be in plain text, while the format of opening time and date of appearance notice should be in date format. In the template, each case element is marked with the symbol "[ ]" to facilitate detection and replacement. Then, the bidirectional encoder representation model based on Transformer (Bidirectional Encoder Representations from Transformers, Bert) is used to vectorize these 11 case elements to obtain the case element vector (CEV), which is stored in the vector database for subsequent comparison.

[0127] Step 4: Based on the case elements in the electronic dossier document template of the first-instance civil appearance notice to be generated and the content of the collected public judicial data set, a corresponding prompt statement is constructed to automatically extract case elements from the large language model agent. The constructed prompt is as follows:

[0128] “

[0129] -Role: Legal text element extraction expert

[0130] -Background: Users need to extract key information from legal documents and present them in a specific format. This information is essential for case analysis, legal research or document management.

[0131] -Profile: You are a professional legal text analyst with profound legal knowledge and text parsing capabilities.

[0132] Ability to accurately identify and extract key elements from legal documents.

[0133] -Skills: You have legal expertise, text analysis skills and data formatting capabilities, and are able to extract required information from complex legal texts and organize the output in a specified format.

[0134] -Goals: Accurately and efficiently extract key information from legal documents and present it in a manner that complies with the JSON data format.

[0135] -Constrains: The extracted information must be accurate, the format must conform to the JSON standard, and must not contain any illegal or non-standard content.

[0136] -OutputFormat: Text output in JSON data format, including all specified legal document elements.

[0137] -Workflow:

[0138] 1. Read and understand the content and structure of legal documents.

[0139] 2. Based on the content of the document, identify and extract key information such as case number, case name, court, etc.

[0140] 3. Organize the extracted information in the specified format to ensure compliance with the JSON data standard.

[0141] -Initialization: In the first conversation, please directly output the following: Hello, I am a professional legal text element extraction expert. Please provide the legal document, and I will extract the key information you need and present it in the specified format.

[0142] Please ensure the completeness and accuracy of the document.

[0143] ”

[0144] At this time, the large language model agent can extract the corresponding case elements from the main content of the judgment and give them in the form of a JSON string. The extracted specific content is shown in Table 1, where the four case elements of court name, case number, defendant, and date of filing are correctly extracted, while the court time and venue do not exist in the original judgment, so they are replaced by "None". For the unavailable content, it can be randomly generated through predetermined logical judgments. For example, in the electronic file document of the court notice, according to laws and regulations, the judgment should be made within one month after the prosecution is filed, and the date of the prosecution in the judgment is consistent with the effective date of the judgment. It can be inferred that the court date is June 15, 2023. The venue of the court should have a geographical correlation with the name of the court, which can be supplemented by the preset correspondence between the court and the tribunal.

[0145] Table 1. Example of output results of fine-tuning the large language model

[0146]

[0147] Step 5: For the case elements and corresponding entity relationships generated in step 3, use the Bert model to vectorize the text. Taking the court name as an example, the case elements and entities are combined into a sentence: "[CLS] Court name: Zhangjiagang People's Court, Jiangsu Province [CLS]", where "[CLS]" represents the placeholder for the position of the marked sentence during vectorization. First, use the Chinese Tokenization tool to convert it into floating point form, and then use the Bert model pre-trained with a large-scale Chinese corpus to vectorize the above sentence to obtain a feature vector with a shape of 1×768.

[0148] Step 6: By executing the process of steps 2-4 on the judgment texts in multiple public data sets, multiple vectorized <case elements, entities> pairs can be obtained, and these vectors are clustered unsupervisedly. Here, this patent uses the K-means clustering method, and its specific algorithm is:

[0149] 1) Assign each data point in the dataset to the nearest centroid (minimizing the Euclidean distance between them), which means that a data point is considered to be in a particular cluster if it is closer to the centroid of that cluster than any other centroid.

[0150] 2) K-means then recalculates the centroids by taking the mean of all the data points assigned to that centroid cluster, thereby reducing the total within-cluster variance associated with the previous step. The “mean” in K-means refers to averaging the data and finding new centroids.

[0151] 3) The algorithm iterates between steps 1) and 2) until the data points have no change clusters.

[0152] 4) Due to the uncertainty and hallucination of the large language model, the number of its cluster centers is uncertain. At this time, the elbow method is used to make a judgment and calculate the sum of squared errors (SSE) under different K values. As K increases, SSE will gradually decrease. When K reaches the actual number of cluster centers, the downward trend will suddenly slow down, forming an "elbow" phenomenon. At this time, the corresponding K value is selected as the actual K value.

[0153] Figure 3 Schematic diagram of the elbow rule experiment in an embodiment of the present invention.

[0154] like Figure 3As shown in the figure, the elbow rule is calculated in 590 case information, and it is found that when the K value is 100, there is an obvious downward trend, so the number of cluster centers is determined to be 100. In other words, when the number of cluster centers is 100, the SSE curve shows an obvious downward trend, so the number of cluster centers is selected as 100 for K-means clustering.

[0155] Step 7: For the clustering results of the <case elements, entities> vectors obtained in step 6, find the vector in each cluster that is closest to the other vectors in the class as the representative vector (Representation Vector, RV) of the class, and then compare it with the 8 case element vectors in the civil first-instance appearance notice. The comparison method is as follows: The number of cluster centers is 100, and there are 100 representative vectors in total: [RV1, RV2, ..., RV 100 ], and there are 8 case element vectors in the template: [CEV1, CEV2, …, CEV8]. First, calculate the representative vector RV i The Euclidean distance between each case element vector and the five case element vectors with the smallest Euclidean distance are given, denoted as Then for each case element vector CEV j , find the 5 representative vectors with the smallest Euclidean distance, recorded as If for a pair of representative vectors and case element vectors <RV i ,CEV j >, while meeting RV i ∈R j and CEV j ∈C i , then it is considered that this pair of representative vectors and case element vectors represent the same case element and corresponding entity. When there are multiple pairs of representative vectors and case element vectors that meet the requirements, for each case element vector, the representative vector with the smallest Euclidean distance that meets the conditions is selected. Finally, for the 8 case element vectors, 8 <representative vector, case element vector> pairs are generated.

[0156] Figure 4 This is an example of a court appearance notice generated in an embodiment of the present invention.

[0157] For the <representative vector, case element vector> pairs found in step 7, fill the entities in the <case element, entity> pairs in each cluster represented by the representative vector into the electronic file document Word template to generate an electronic file document. The generated document is as follows: Figure 4 shown.

[0158] The automatic electronic case file generation method and system based on prompt learning proposed in the present invention can simulate and generate electronic case files in the judicial trial process on a large scale based on existing public real data. At the same time, under the premise of protecting privacy, it ensures that the internal information logic of the generated electronic case files is consistent and the elements are complete, providing high-quality data support for the intelligent cognitive model training and testing of electronic case files in the judicial field.

[0159] Those skilled in the art know that, in addition to realizing the system and its various devices, modules, and units provided by the present invention in a purely computer-readable program code, it is entirely possible to realize the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a hardware component, and the devices, modules, and units included therein for realizing various functions can also be regarded as structures within the hardware component; the devices, modules, and units for realizing various functions can also be regarded as both software modules for realizing the method and structures within the hardware component.

[0160] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A method for generating an automated electronic dossier based on prompt learning, characterized in that: The steps include: Step S1, collect judicial data sets, and filter out a subset of complete original case elements according to preset standards, annotate the subset, and construct a case element extraction task data set; Step S2, performing desensitization processing on the case element extraction task data set to obtain a desensitized data set; Step S3, using the desensitized data set, fine-tuning the large language model that has been pre-trained on a large-scale corpus to obtain a fine-tuned large language model; Step S4, creating an electronic case file template for generating an electronic case file procedure document, and specifying the target case elements to be filled in the electronic case file template; Step S5, based on the fine-tuned large language model, construct prompt words according to the target case elements and the case element extraction task data set, and then input them into the fine-tuned large language model to extract the case element name and entity information of the target case element; Step S6, clustering the case element names and the entity information to obtain a clustering result; Step S7, backfilling the entity information in the clustering result into the electronic file template to complete the automatic generation of the electronic file document.

2. The method for generating an automated electronic dossier based on prompt learning according to claim 1, characterized in that: In step S2, the desensitization process is to randomly replace sensitive information, and the sensitive information includes name, address, ID number and telephone number.

3. The method for generating an automated electronic dossier based on prompt learning according to claim 1, characterized in that: In step S4, in the electronic case file template, specific placeholders are used to mark the locations where case elements need to be filled in, and in different electronic case file types, case elements of the same type use the same name and data format to ensure that the generated electronic case file is consistent and standardized.

4. The method for generating an automated electronic dossier based on prompt learning according to claim 1, characterized in that: In the step S5, the prompt word is presented in the form of a Prompt sentence, and the case element name and entity information of the target case element are extracted by inputting the Prompt sentence into the fine-tuned large language model.

5. The method for generating an automated electronic dossier based on prompt learning according to claim 4, characterized in that: In step S5, the form of the Prompt statement is: extract the following case elements from the [judicial data set type] content [judicial data set content]: [case element name], replace non-existent elements with "none", and give the case element name and the entity information in the form of a JSON string.

6. The method for generating an automated electronic dossier based on prompt learning according to claim 5, characterized in that: The step S6 includes the following sub-steps: Step S6.1, for the JSON string including the case element name and entity information extracted in step S5, the case element name and the entity information are sorted into a string in the form of "[case element name]: [entity content]", and the string is represented as a fixed-length floating-point vector using string vectorization technology; Step S6.2, performing unsupervised clustering on the floating point number vector using a K-means clustering algorithm, wherein K centroids are randomly created based on a predefined K value to obtain a clustering result, wherein the clustering result includes text and vocabulary vectors.

7. The method for generating an automated electronic dossier based on prompt learning according to claim 6, characterized in that: In step S6.1, the string vectorization technology adopts a Transformer-based bidirectional encoding representation, which has been pre-trained on a large-scale Chinese corpus and represents the string as a floating-point vector with a shape of 1×768.

8. The method for generating an automated electronic dossier based on prompt learning according to claim 7, characterized in that: The K-means clustering algorithm of step S6.2 includes the following sub-steps: Step S6.2.1, calculate the sum of squares of clustering errors under different K values, and use the elbow rule to determine the optimal K value; Step S6.2.2, assigning each data point in the floating point vector to the nearest centroid; Step S6.2.3, recalculating the position of the centroid by calculating the average of all data points assigned to the cluster where the centroid is located, thereby reducing the total intra-cluster variance associated with step S6.2.1, where the "mean" in K-means refers to finding the arithmetic mean of the data points in the current cluster to find the new centroid position; Step S6.2.4, iterate between step S6.2.2 and step S6.2.3 until the cluster assignment of data points no longer changes, and obtain the clustering result.

9. The method for generating an automated electronic dossier based on prompt learning according to claim 8, characterized in that: The step S7 comprises: for the clustering results obtained in the step S6, selecting the "center point vector" with the shortest average distance to the other vectors in the category in the clustering results of each category as the representative of the category; comparing and calculating the distance between the category representative vector and the case element name vector specified in the electronic file template; setting the number of cluster centers to K, then there are K category representative vectors: [RV1, RV2, ..., RV K ], and there are N case element vectors in the electronic file template: [CEV1, CEV2, …, CEV N ]; First calculate the category representative vector RV i The similarity between each case element vector and find the n case element vectors with the highest similarity, recorded as Then for each case element vector CEV j , find the n category representative vectors with the highest similarity, recorded as If for a pair of class representative vectors and case element vectors <RV i ,CEV j >, while meeting RV i ∈R j and CEV j ∈C i , it is considered that this pair of category representative vector and case element vector represents the same case element and corresponding entity; when there are multiple pairs of category representative vectors and case element vectors that meet the requirements, for each case element vector, select the category representative vector that meets the requirements and has the highest similarity; finally, for N case element vectors, generate N <category representative vector, case element vector> pairs; select the case element name with the highest similarity as the correct classification name, and fill in the corresponding entity information into the electronic file template to generate the corresponding electronic file document.

10. An automated electronic dossier generation system based on prompt learning, characterized in that: include: Module M1 collects judicial data sets, and according to preset standards, selects a complete subset of original case elements, annotates the subset, and constructs a case element extraction task data set; Module M2, desensitizing the case element extraction task data set to obtain a desensitized data set; Module M3, using the desensitized data set, fine-tuning the large language model that has been pre-trained on a large-scale corpus to obtain a fine-tuned large language model; Module M4, creating an electronic case file template for generating electronic case file procedure documents, and specifying the target case elements to be filled in the electronic case file template; Module M5, based on the fine-tuned large language model, constructs prompt words according to the target case elements and the judicial data set, and then inputs them into the fine-tuned large language model to extract the case element name and entity information of the target case element; Module M6, clustering the case element names and the entity information to obtain a clustering result; Module M7, backfills the entity information in the clustering result into the electronic file template to complete the automatic generation of the electronic file document.

Citation Information

Patent Citations

  • Method, device and computer equipment for generating and processing an award based on online arbitration

    CN110362799B

  • A method for event extraction and processing in electronic case files

    CN111723564B

Cited By

  • Method for processing multi-source data based on K-Means aggregation algorithm of large model

    CN120561797A