Case element automatic extraction method and system based on cue word engineering

Through the method based on prompt word engineering, a large language model is used to perform semantic analysis of case text, the problem of insufficient flexibility in case factor extraction in the existing technology is solved, and the factor extraction with high precision and high integrity is achieved, reducing system maintenance costs.

CN120069051APending Publication Date: 2025-05-30BEIJING INST OF COMP TECH & APPL
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510237277.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-02
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing automatic extraction method of case elements is insufficient flexibility in dealing with complex and diverse case texts, making it difficult to accurately capture deep semantics and contextual relationships, especially in complex syntaxes of long texts and legal language, and it is impossible to accurately understand the meaning of the sentence.

Method used

Using a method based on prompt word engineering, a structured prompt word template is designed, and a pre-trained large language model is used to perform semantic analysis of case text, and case elements are automatically identified and extracted. This method includes multiple rounds of prompt word adjustment and extraction to ensure detailed extraction and accuracy of elements.

Benefits of technology

It significantly improves the accuracy and integrity of case factor extraction, reduces dependence on manual labeled data, reduces the cost of system maintenance and expansion, and has good adaptability and automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069051A_ABST
    Figure CN120069051A_ABST
Patent Text Reader

Abstract

The invention relates to a case element automatic extraction method and system based on cue word engineering, and belongs to the crossing field of natural language processing, artificial intelligence and legal information processing. In order to solve the problems of low case element extraction efficiency and insufficient accuracy, the method comprises the following steps: firstly, reading and preprocessing a case file by a system; secondly, designing a structured cue word template according to the case type and elements needing to be extracted; then, the system calls a pre-trained large language model, semantic analysis is carried out through prompt word guidance, case elements are automatically recognized, and an element type list is generated; then, the system dynamically adjusts cue words according to the preliminary recognition result, multiple rounds of information extraction are executed, and the integrity of element extraction is ensured; and finally, integrating the extraction results by the system, generating a structured case description, and outputting the structured case description. The method gives full play to the semantic comprehension ability of the large language model, improves the efficiency and accuracy of case element extraction, and has wide application potential.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - field of natural language processing, artificial intelligence and legal information processing, and particularly relates to a method and system for automatically extracting case elements based on prompt engineering. Background Art

[0002] With the increasing complexity of legal affairs, the large amount of case materials (including transcripts, evidence documents, trial records, etc.) that legal practitioners and institutions need to process has shown an explosive growth. Efficiently and accurately analyzing and processing these texts is an important requirement in current legal information processing. Automatically extracting case elements, that is, identifying key information such as time, place, parties, events, etc. from case texts, has become one of the key means to improve the efficiency and accuracy of legal work.

[0003] Existing methods for extracting case elements mainly include rule - based methods and traditional machine learning methods. Rule - based methods rely on grammar rules and predefined templates written by industry experts. Although such methods have relatively high accuracy in certain specific fields, their limitations are also obvious. First, they lack flexibility. The writing and maintenance of rules highly depend on domain experts and cannot cope with the changes in different case types and fields. Especially when facing diverse and complex case texts, it is difficult to cover the changes in sentence patterns and structures through simple rules. Second, they have insufficient scalability. When new case types or new fields need to be processed, the adaptability of the rules is poor, and the system needs to be readjusted or even rewritten, resulting in high maintenance costs and an inability to respond quickly, thus affecting the timeliness of case processing.

[0004] Compared with rule - based methods, traditional machine learning methods (such as support vector machines, decision trees, etc.) although reduce the dependence on manual rules, these methods strongly rely on a large amount of manually labeled data and still face the following two major problems in practical applications: one is the limitation of feature engineering. Machine learning methods need to rely on experts to manually design features, which makes it difficult to fully capture the deep semantics and context relationships in the text. Especially when dealing with complex legal languages, the extraction accuracy is relatively low. The other is weak generalization ability. Traditional machine learning models often have difficulty in generalizing when dealing with complex and unstructured legal texts, especially when extracting elements across sentences and paragraphs in long texts, showing significant limitations.

[0005] In addition, existing extraction methods based on rules and traditional machine learning not only consume a large amount of computing resources and take a long time to process, but also in actual operation, it is often difficult to ensure the accuracy and recall rate of the extraction results, and problems such as omission of key case elements or excessive redundant information are likely to occur.

[0006] In recent years, prompt word engineering, as an emerging technology, can effectively guide the generation and reasoning process of the model by providing specific prompt words to the large language model. Compared with traditional methods, prompt word engineering combines the powerful language understanding capabilities of large language models (such as GPT, GLM, etc.) and shows significant advantages in dealing with complex semantic understanding and context association problems. After large-scale pre-training, the large language model can understand and generate complex natural language texts, which not only enables it to have a wide range of language generation capabilities, but also can accurately analyze specific case texts through prompt words and extract the required case elements in a targeted manner.

[0007] The combination of prompt word engineering and large language models can significantly improve the accuracy of factor extraction when processing long texts and complex case descriptions, especially in semantic understanding and context processing, overcoming the limitations of existing technologies in flexibility, accuracy and scalability. Compared with traditional methods, this method can not only improve the accuracy and recall of factor extraction, but also greatly reduce the dependence on manually annotated data and reduce the cost of system maintenance and expansion. Therefore, the development of a case factor automatic extraction method and system based on prompt word engineering can better meet the current needs of legal information processing and has important practical significance and application prospects. Summary of the invention

[0008] 1. Technical issues to be resolved

[0009] The technical problem to be solved by the present invention is how to provide a method and system for automatically extracting case elements based on prompt word engineering, so as to solve the problem that the existing automatic case element extraction methods are insufficiently flexible when processing complex and diverse case texts, and have difficulty in accurately capturing the deep semantics and contextual relationships in the text, especially when processing long texts and complex syntax of legal language, and cannot accurately understand the meaning of sentences and cannot meet the needs of actual element extraction.

[0010] (II) Technical solution

[0011] In order to solve the above technical problems, the present invention proposes a method for automatically extracting case elements based on prompt word engineering, which comprises the following steps:

[0012] Step 1: The system receives and reads the case text, including the indictment, transcript, and evidence documents, and performs basic preprocessing to ensure that the input data format is standardized;

[0013] Step 2: Design a structured prompt word template according to the case type and the elements to be extracted, and use the prompt word template to generate prompt words;

[0014] Step 3: The system calls the pre-trained large language model, performs semantic analysis on the case text in combination with the prompt words, and automatically identifies the case element types in the case text;

[0015] Step Four: Under the guidance of the prompt, the model initially identifies target elements from the case text and generates a list of case element types;

[0016] Step Five: Based on the list of case element types generated in Step Four, the system automatically adjusts the prompt to focus on the content extraction of the current case element type;

[0017] Step Six: Based on the case element extraction prompt in Step Five, after multiple extractions, the system completes the detailed extraction for each element type;

[0018] Step Seven: The system integrates the completed element results, retains all the extracted data, and conducts comparison and integration based on the case type and the elements to be extracted;

[0019] Step Eight: The system outputs the processed case elements in a structured format to generate a standardized case report, which includes the case name, case number, list of case elements, and extraction results.

[0020] The present invention also provides a case element automatic extraction system based on prompt engineering. The system includes: a case file data module, a text data processing module, a prompt generation module, an information extraction engine, and an extraction result output module;

[0021] The case file data module is used to obtain case files, and the case files include text data in various formats, including: indictment opinions, requests for approval of arrest, interrogation transcripts, inquiry transcripts, and evidence documents;

[0022] The text data processing module is used to perform text preprocessing. It recognizes the case text through OCR or directly imports it, and conducts normalization operations on it, including: unifying the encoding format, removing special characters, processing redundant spaces, and paragraph separators;

[0023] The prompt generation module is used to design appropriate prompt templates for different types of cases and case elements. The prompt set consists of keywords, phrases, or questions. The prompt template includes: keywords related to case elements and reasonable syntactic structures, and uses the prompt template to generate prompts;

[0024] The information extraction engine uses a large language model combined with prompts to perform in-depth semantic analysis and context correlation analysis on the text; under the guidance of the prompt, the large language model gradually performs semantic analysis on the case text to identify key elements in the text; after initially identifying the elements, it enters the prompt adjustment stage, and dynamically optimizes the prompt according to the identified element list to make it more focused on the element content that has not been fully extracted; through multiple rounds of prompt adjustment, continuously extract key information from the case text and gradually generate more detailed case elements;

[0025] The extraction result output module is used to summarize and integrate all the extracted elements to generate a final element list, which contains detailed information after multiple rounds of extraction and prompt adjustment.

[0026] (III) Beneficial effects

[0027] The present invention provides an automatic case element extraction method and system based on prompt engineering. By guiding the large language model through prompts, the present invention can more accurately capture the deep semantics and context information in the case text, especially when dealing with long texts and complex legal languages, significantly improving the accuracy of element extraction. Compared with traditional machine learning methods, the present invention reduces the dependence on manually labeled data, reduces the manual design requirements for feature engineering, and enhances the automation of the system. Prompt engineering has good adaptability and can be flexibly extended to cases of different types and fields, only by adjusting the prompt template, greatly reducing the system maintenance and update costs. Description of the drawings

[0028] Figure 1 It is a schematic diagram of the overall architecture of the system of the present invention;

[0029] Figure 2 It is a detailed flowchart of the prompt generation module in the present invention;

[0030] Figure 3 It is an example diagram of the classification prompt construction process in the present invention;

[0031] Figure 4 It is an example diagram of the element extraction prompt construction process in the present invention;

[0032] Figure 5 It is a block diagram of the information extraction engine module in the present invention. Detailed implementation manners

[0033] To make the purpose, content and advantages of the present invention clearer, the following further describes the detailed implementation manners of the present invention in conjunction with the drawings and embodiments.

[0034] To solve the above technical problems, the present invention provides an automatic case element extraction method and system based on prompt engineering. The method and system guide the large language model through prompts for directional information extraction, making full use of the semantic understanding ability of the model. The technical solutions specifically include the following steps:

[0035] Step 1: The system receives and reads the case text, including unstructured data such as indictments, transcripts, and evidence documents, and performs basic preprocessing to ensure that the input data format is standardized;

[0036] Step 2: According to the case type and elements to be extracted (such as person elements, location elements, crime elements, etc.), design a structured prompt template, and use the prompt template to generate prompts;

[0037] Step 3: The system calls a pre-trained large language model, combines the prompts, and performs semantic analysis on the case text to automatically identify the types of case elements in the case text;

[0038] Step 4: Under the guidance of the prompts, the model initially identifies the target elements from the case text and generates a list of case element types, including person elements, location elements, crime elements, etc.;

[0039] Step 5: The system automatically adjusts the prompts according to the list of case element types generated in Step 4 to focus on the content extraction of the current case element type;

[0040] Step 6: Based on the case element extraction prompts in Step 5, after multiple extractions, the system completes the detailed extraction for each element type;

[0041] Step 7: The system integrates the completed element results, retains all the extracted data, and performs comparison and integration according to the case type and elements to be extracted;

[0042] Step 8: The system outputs the processed case elements in a structured format to generate a standardized case report, which includes the case name, case number, list of case elements (persons, locations, crime elements, etc.) and extraction results.

[0043] Example 1:

[0044] This specific implementation method takes the realization of automatically extracting key information (persons, locations, crime elements, etc.) from case texts as the core, and fully utilizes the semantic analysis capabilities of prompt engineering and large language models.

[0045] This embodiment describes the overall architecture and working process of the automatic case element extraction system based on prompt engineering. The system includes: a case file data module, a text data processing module, a prompt generation module, an information extraction engine, and an extraction result output module. Its specific architecture is as Figure 1 shown.

[0046] The case file data module is used to obtain case files, which usually contain various formats of text data (such as prosecution opinions, requests for approval of arrests, interrogation transcripts, inquiry transcripts, and evidence documents, etc.). When extracting these case elements, manual analysis is not only time-consuming but also prone to missing key information. Therefore, the system first needs to preprocess these unstructured texts.

[0047] The text data processing module is used to perform text preprocessing. It recognizes text through OCR or directly imports case texts and performs normalization operations on them. This process involves unifying the encoding format, removing special characters, handling extra spaces and paragraph separators to ensure that the text can be smoothly input into the large language model for further analysis. For the noise information in the text (such as headers, footers, page numbers, etc.), the system will automatically detect and delete it. In addition, for long texts, the system will divide them into multiple logical units and preliminarily mark the paragraphs that may contain key information, such as key parts like event descriptions and person information.

[0048] Figure 2 The detailed process of the prompt word generation module is shown, which elaborates on how to generate appropriate prompt words according to the case type and elements to be extracted. The prompt word generation module is the core component of the system, and its main task is to design appropriate prompt word templates for different types of cases and case elements. The specific steps are as follows:

[0049] S21. Define the case type: The user or the system first determines the type of the case (such as civil, criminal, administrative, etc.).

[0050] S22. Identify the target elements: Determine the key elements to be extracted from the case document (such as persons, locations, criminal acts, etc.).

[0051] S23. Load the domain knowledge base: The system loads the legal vocabulary, term library or knowledge graph related to the specified case type, and these resources include common legal terms and entity categories.

[0052] S24. Construct the prompt word template: According to the case type and element requirements, select or construct the corresponding prompt word template. The template may contain combinations of prompt words for different elements, such as "What is the interest rate of this loan?" or "Was the driver drinking at the time of the accident?".

[0053] S25. Generate specific prompt words: The system combines the specific information of the case and uses the template to generate prompt words suitable for the specific situation.

[0054] S26. Optimize the prompt words: According to the context of the case text, the system dynamically adjusts the prompt words to improve the accuracy and efficiency of information extraction.

[0055] S27. Evaluate and fine-tune: Evaluate and adjust the generated prompt words to ensure that they can effectively guide the information extraction process. The evaluation process may include a feedback mechanism for continuously optimizing the quality of the prompt words.

[0056] S28. Output the set of prompt words: Finally, the system outputs the generated set of prompt words to the information extraction engine.

[0057] The set of prompting words consists of keywords, phrases, or questions, aiming to provide precise guidance for the subsequent information extraction process, ensuring that the system can accurately identify and extract the key information in the case.

[0058] Based on the prompting word generation module, the system designs a series of structured prompting word templates for different case types and case elements to be extracted. These templates not only contain keywords related to the case elements but also contain reasonable syntactic structures to ensure that the large language model can accurately perform semantic analysis and element identification.

[0059] The system first conducts a preliminary analysis of the text to identify the case type (such as criminal cases, civil cases, or administrative cases, etc.). The design of the prompting word templates for different case types is different to meet specific requirements.

[0060] In addition, the prompting word design of the present invention is also applicable to the scenario of charge identification. Taking the charge identification of financial crime cases as an example, the prompting word design is as follows:

[0061] "″"

[0062] - Role: Financial crime analyst

[0063] - Background: The user needs to identify specific charges in financial crime cases for better compliance review, risk assessment, or legal litigation.

[0064] - Introduction: You are an experienced financial crime analyst with an in-depth understanding of financial regulations, criminal law, and relevant judicial interpretations, and can accurately identify and explain the charges in various financial crime cases.

[0065] - Skills: You have the capabilities of legal analysis, case research, risk assessment, and compliance review, and can identify potential criminal acts from complex financial transactions.

[0066] - Goal: Provide accurate identification of financial crime charges, help users understand the nature of the case, and provide a basis for further legal actions.

[0067] - Limitation: It must be based on current laws, regulations, and judicial interpretations to ensure the accuracy and legality of charge identification.

[0068] - Output format: Provide a clear list of charges and a brief explanation of each charge, accompanied by relevant legal provisions.

[0069] - Workflow:

[0070] 1. Collect and analyze the financial transaction records and documents related to the case.

[0071] 2. Identify possible criminal acts based on the transaction records and documents.

[0072] 3. Compare with current laws and regulations to determine the corresponding charges.

[0073] 4. Provide a detailed explanation of the charges and relevant legal provisions.

[0074] - Initialization: {Case text content}

[0075] "″"

[0076] The system generates corresponding prompt word templates according to the case type and elements to be extracted (such as unit employment situation, actual business situation, raised funds situation, etc.). For example, for the text classification task of identifying case elements, the system generates the following classification prompt words:

[0077] "″"

[0078] - Role: Text classification expert

[0079] - Background: The user needs to classify the text to determine which predefined element categories it belongs to (such as unit employment situation, actual business situation, raised funds situation, etc.).

[0080] - Introduction: You are a professional text classification expert with rich experience in natural language processing, and can accurately identify and judge the element categories of the text.

[0081] - Skills: You possess the key skills of text analysis, pattern recognition, and machine learning, and can design and apply efficient text classification models.

[0082] - Goal: According to the text content, accurately judge and output the list of element categories to which it belongs.

[0083] - Restriction: The classification result must be accurate and error-free, avoiding misjudgment, and at the same time handling text diversity and ambiguity.

[0084] - Output format: The output should be a list of element categories to which the text belongs.

[0085] - Workflow:

[0086] 1. Preprocess the input text, including word segmentation, stop word removal, etc.

[0087] 2. Analyze the text using the classification model.

[0088] 3. Judge the element category to which the text belongs according to the analysis result.

[0089] 4. Output the list of element categories to which the text belongs.

[0090] - Example:

[0091] - Example 1: For the text "AA serves as the financial director in XX Company", the output category is ["Unit employment situation"].

[0092] - Example 2: The text "The net profit of XX Company increased by 20% last year", and the output category is ["Actual business situation"].

[0093] - Example 3: The text "XX Company plans to raise 100 million US dollars by issuing bonds", and the output category is ["Fund-raising situation"].

[0094] - Example 4: The text "BB serves as the technical director of XX Company, and the sales volume reached 50 million yuan last year", and the output categories are ["Employment situation in the unit", "Actual business situation"].

[0095] - Initialization: {Case text content}

[0096] "″"

[0097] According to the element classification results, the system will dynamically adjust the prompt word template. This adjustment process automatically optimizes the construction of the prompt word based on the complexity of the case text and the structure of the elements. For example, after identifying the element "Fund-raising situation" in the interrogation transcript, the system will automatically update the prompt word to extract more detailed information. The following is an example of the prompt word template for element extraction:

[0098] "″"

[0099] Your task is to carefully read the content of the interrogation transcript, understand its semantics, and extract the corresponding content strictly according to the JSON template format, ensuring that the template structure is not modified.

[0100] Transcript: {Text content}

[0101] JSON template: {Template for the definition of the fund-raising situation element}

[0102] After extraction, check whether the returned result is consistent with the template structure.

[0103] "″"

[0104] Among them, the template for the definition of the fund-raising situation element is as follows:

[0105] {"Fund-raising situation":

[0106] {

[0107] "Total amount of funds raised": "",

[0108] "Number of fund-absorbing objects": "",

[0109] "Fund-raising time": "",

[0110] "Operation mode of raised funds": "",

[0111] "Unit name": ""

[0112] }

[0113] }

[0114] The construction process of the above prompt words is shown in Figure 3 and Figure 4 , which respectively show the construction process of classification prompt words and the construction process of element extraction prompt words. After the prompt words are generated, the system provides them to the information extraction engine, such as Figure 5 shown.

[0115] In the information extraction engine, large language models (such as GPT, GLM, etc.) can understand complex legal language and perform in-depth semantic analysis and context correlation analysis on the text in combination with the prompt words. The model can not only identify the key elements in the case text, but also analyze the logical relationships between the elements, such as the correlation between time and events, and between people and locations.

[0116] Under the guidance of the prompt words, the large language model gradually performs semantic analysis on the case text and identifies the key elements in the text. The system extracts the preliminary case element types from it, such as unit employment situation, actual business situation, raised funds situation, etc., and generates a preliminary element list.

[0117] After the preliminary elements are identified, the system enters the prompt word adjustment stage. According to the identified element list, the system dynamically optimizes the prompt words to make them more focused on the element content that has not been fully extracted. Through multiple rounds of prompt word adjustment, the system continuously extracts key information from the case text and gradually generates more detailed case elements. The element list may include the following key information:

[0118] Unit employment situation: position, scope of duties, employment time, salary, etc.;

[0119] Actual business situation: assets, debts, product names, etc.;

[0120] Raised funds situation: total amount of raised funds, number of people absorbing funds, raising time, etc.

[0121] The system automatically adjusts the prompt word template according to the element recognition results of each round to ensure that all important case elements can be accurately and comprehensively extracted.

[0122] After the element extraction is completed, the system enters the data integration and result verification stage. The extraction result output module summarizes and integrates all the extracted elements to generate a final element list, which contains the detailed information after multiple rounds of extraction and prompt word adjustment.

[0123] To ensure the accuracy and integrity of legal data, the system automatically detects and marks duplicate information, backtracks and marks possible redundant data to ensure data traceability and consistency. Finally, the system outputs the case elements in a structured format (such as JSON, XML, or CSV) for subsequent data processing and case analysis.

[0124] Example 2:

[0125] This example introduces the application of the automatic case element extraction system based on prompt engineering in criminal cases. Taking the interrogation record of a "robbery case" as an example, the automatic processing process of the system is described.

[0126] Step 1, preprocessing of case text. After receiving the indictment, the system first performs preprocessing. The indictment is usually unstructured text with inconsistent formats. The system removes noise information through text cleaning and format standardization, and performs logical chunking according to content features. For example, it divides personal information, event descriptions, evidentiary materials, etc. into different modules for subsequent processing.

[0127] Step 2, generation of prompts. After preprocessing, the system identifies the case type as "criminal case" and generates prompt templates based on the characteristics of the robbery case. These templates target elements such as "criminal act", "crime location", "suspect information", "victim information", "crime time", etc. For example, the system may generate the following prompts:

[0128] "Please extract the criminal act in the case."

[0129] "Please provide the crime location."

[0130] "Please extract the suspect's name and identity information."

[0131] Step 3, semantic analysis and element extraction. Under the guidance of the prompts, the system calls a large language model (such as GPT or GLM) to perform semantic analysis on the text. The model combines the prompts to automatically identify the main elements in the case, such as the description of the "robbery act", "crime location", "suspect name", etc. The element list generated in this stage includes the basic information of the case.

[0132] Step 4, prompt optimization and multi-round extraction. After the initial extraction, the system analyzes the element list to detect possible missing information. If the "specific time when the robbery occurred" is missing, the system will dynamically adjust the prompt template, generate new prompts and perform extraction again. Through multi-round optimization and extraction, the system ensures that all key elements are completely extracted.

[0133] Step 5, data integration and result verification. After the system completes multi-round extraction, it enters the data integration stage, summarizing the elements extracted in each round into a structured list, such as:

[0134] Criminal act: Robbing a bank

[0135] Crime location: A certain street in a certain city

[0136] Suspect information: Suspect A, male, X years old, ID number XXXXXX

[0137] Crime time: X year X month X day

[0138] Step 6, result output. The system outputs the integrated element information in a structured format (such as JSON, XML, CSV), and generates a detailed case element report, including the case name, number, type, and list of main elements.

[0139] The present invention constructs an automatic extraction system for case elements through prompt engineering combined with natural language processing technology. By using a prompt template to guide the large language model to perform layer-by-layer semantic analysis on the case text, the key information in the case is accurately extracted, solving the problems in the existing technical solutions that it is difficult to efficiently process unstructured text and important case elements are easily missed.

[0140] The present invention introduces a prompt dynamic generation technology for long and complex case texts. Combining the case type and element requirements, the prompt template is automatically adjusted to optimize the accuracy of information extraction. In particular, by using multi-round prompt guidance and semantic analysis technology, the integrity and accuracy of case element extraction are significantly improved, and the result is finally output in a structured format, providing efficient and accurate data support for legal practice.

[0141] Example 3:

[0142] An automatic extraction method for case elements based on prompt engineering includes the following steps:

[0143] Step 1: The system receives and reads the case text, including unstructured data such as indictments, transcripts, and evidence files, and performs basic preprocessing to ensure that the input data format is standardized;

[0144] Step 2: Design a structured prompt template according to the case type and elements to be extracted (such as person elements, location elements, crime elements, etc.);

[0145] Step 3: The system calls a pre-trained large language model and performs semantic analysis on the case text in combination with the prompt to automatically identify the case element types in the case text;

[0146] Step 4: Under the guidance of the prompt, the model initially identifies the target elements from the case text and generates a list of case element types, including person elements, location elements, crime elements, etc.;

[0147] Step Five: The system automatically adjusts the prompt words according to the list of case element types generated in Step Four to focus on the content extraction of the current case element type;

[0148] Step Six: Based on the case element extraction prompt words in Step Five, the system completes the detailed extraction for each element type through multiple extractions;

[0149] Step Seven: The system integrates the completed element results, retains all the extracted data, and conducts comparison and integration based on the case type and the elements to be extracted;

[0150] Step Eight: The system outputs the processed case elements in a structured format to generate a standardized case report, which includes the case name, case number, list of case elements (such as people, locations, criminal elements, etc.) and extraction results.

[0151] Among them, the prompt word template is structurally designed based on the case type and the elements to be extracted, and contains specific grammar and keywords to guide the large language model to identify relevant elements in the case text.

[0152] Among them, the large language model is a pre-trained language model, including GPT, GLM, DeepSeek, etc.

[0153] Among them, the automatic adjustment of the prompt words includes adjusting the prompt words based on the preliminary list of case element types generated in Step Four to focus on the content extraction of the current case element type.

[0154] Among them, the multiple extractions are independent extractions for each case element type to ensure that all identified elements are extracted.

[0155] Among them, the integration of all extracted content includes data verification, retaining all extraction results, and conducting comparison and integration based on the case type and elements.

[0156] Among them, the integration includes organizing the integrated case elements into a complete and coherent case description according to the structure of the case file.

[0157] A case element automatic extraction system based on prompt engineering includes:

[0158] A data preprocessing module for receiving and processing case texts and performing basic preprocessing;

[0159] A prompt word generation module for designing a structured prompt word template according to the case type and the elements to be extracted;

[0160] A large language model call module for conducting semantic analysis on the case text in combination with the prompt words to identify case element types;

[0161] A target element recognition module for initially extracting target elements from case texts and generating a list of case element types;

[0162] A prompt adjustment module for automatically adjusting prompts according to the initial extraction results and focusing on the extraction of the current element type;

[0163] A multi-round extraction module for performing multiple extractions based on the adjusted prompts to ensure the detailed extraction of each element type;

[0164] A data integration module for integrating extraction results, removing redundant data, and performing comparison and integration according to case types and elements;

[0165] A result output module for outputting the processed case elements in a structured format, including the case name, case number, case element list, and extraction results.

[0166] Beneficial effects:

[0167] The present invention proposes a method and system for automatically extracting case elements based on prompt engineering. By guiding the large language model through prompts, the present invention can more accurately capture the deep semantics and context information in case texts, especially when dealing with long texts and complex legal languages, significantly improving the accuracy of element extraction. Compared with traditional machine learning methods, the present invention reduces the dependence on manually labeled data, reduces the manual design requirements for feature engineering, and enhances the automation of the system. Prompt engineering has good adaptability and can be flexibly extended to cases of different types and fields, only by adjusting the prompt template, greatly reducing the system maintenance and update costs.

[0168] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principles of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A method for automatically extracting case elements based on prompt word engineering, characterized in that: The method comprises the following steps: Step 1: The system receives and reads the case text, including the indictment, transcript, and evidence documents, and performs basic preprocessing to ensure that the input data format is standardized; Step 2: Design a structured prompt word template according to the case type and the elements to be extracted, and use the prompt word template to generate prompt words; Step 3: The system calls the pre-trained large language model, performs semantic analysis on the case text in combination with the prompt words, and automatically identifies the case element types in the case text; Step 4: Under the guidance of the prompt words, the model preliminarily identifies the target elements from the case text and generates a list of case element types; Step 5: The system automatically adjusts the prompt words based on the case element type list generated in step 4 to focus on the content extraction of the current case element type; Step 6: The system extracts prompt words based on the case elements in step 5. After multiple extractions, it completes detailed extraction for each element type. Step 7: The system integrates the completed factor results, retains all extracted data, and compares and integrates them according to the case type and the factors to be extracted; Step 8: The system outputs the processed case elements in a structured format and generates a standardized case report, which includes the case name, case number, case element list and extraction results.

2. The method for automatically extracting case elements based on prompt word engineering as claimed in claim 1, characterized in that: Elements include: Character elements, place elements, crime elements.

3. The method for automatically extracting case elements based on prompt word engineering as claimed in claim 1, characterized in that: The prompt word template is structured based on the case type and the elements to be extracted, and contains specific grammar and keywords to guide the large language model to identify relevant elements in the case text.

4. The method for automatically extracting case elements based on prompt word engineering as claimed in claim 1, characterized in that: The large language model is a pre-trained language model, including GPT, GLM and DeepSeek.

5. The method for automatically extracting case elements based on prompt word engineering as claimed in claim 1, characterized in that: The automatic adjustment of prompt words includes adjusting the prompt words based on the preliminary case element type list generated in step 4 to focus on content extraction of the current case element type.

6. The method for automatically extracting case elements based on prompt word engineering according to claim 1, characterized in that: The multiple extractions are performed independently for each case element type to ensure that all identified elements are extracted.

7. The method for automatically extracting case elements based on prompt word engineering as claimed in claim 1, characterized in that: The integration of all extracted contents includes data verification, retaining all extraction results, and comparing and integrating them according to case types and elements.

8. The method for automatically extracting case elements based on prompt word engineering as claimed in claim 1, characterized in that: The integration includes organizing the integrated case elements into a complete and coherent case description according to the structure of the case file.

9. A case element automatic extraction system based on prompt word engineering, characterized in that: The system includes: a case file data module, a text data processing module, a prompt word generation module, an information extraction engine and an extraction result output module; The case file data module is used to obtain case files, which include text data in various formats, including: prosecution opinions, arrest requests, interrogation records, inquiry records and evidence documents; The text data processing module is used to perform text preprocessing, recognize the case text through OCR or directly import it, and perform normalization operations on it, including: unifying the encoding format, removing special characters, and processing redundant spaces and paragraph separators; A prompt word generation module is used to design appropriate prompt word templates for different types of cases and case elements. The prompt word set consists of keywords, phrases or questions. The prompt word template includes: keywords related to the case elements and a reasonable syntactic structure. Prompt words are generated using the prompt word template; The information extraction engine uses a large language model combined with prompt words to conduct in-depth semantic analysis and contextual association analysis of the text. Under the guidance of the prompt words, the large language model gradually performs semantic analysis on the case text and identifies the key elements in the text. After the initial identification of the elements, it enters the prompt word adjustment stage and dynamically optimizes the prompt words based on the identified element list to make them more focused on the element content that has not been fully extracted. Through multiple rounds of prompt word adjustment, key information is continuously extracted from the case text, and more detailed case elements are gradually generated. The extraction result output module is used to summarize and integrate all extracted elements to generate the final element list, which contains detailed information after multiple rounds of extraction and prompt word adjustment.

10. The method for automatically extracting case elements based on prompt word engineering according to claim 9, characterized in that: The specific process of the prompt word generation module: S21. Define case type: The user or system first determines the type of case, including: civil, criminal, administrative; S22. Clarify target elements: Identify the key elements that need to be extracted from the case documents; S23, Loading domain knowledge base: The system loads legal vocabulary, term base or knowledge graph related to the specified case type; S24, constructing a prompt word template: selecting or constructing a corresponding prompt word template according to the case type and element requirements; the template includes prompt word combinations for different elements; S25, generating specific prompt words: the system combines the specific information of the case and uses the template to generate prompt words suitable for the specific situation; S26. Optimize prompt words: According to the context of the case text, the system dynamically adjusts the prompt words to improve the accuracy and efficiency of information extraction; S27, Evaluation and fine-tuning: Evaluate and adjust the generated prompt words to ensure that they can effectively guide the information extraction process; the evaluation process includes a feedback mechanism to continuously optimize the quality of the prompt words; S28. Output prompt word set: Finally, the system outputs the generated prompt word set to the information extraction engine.

Citation Information

Cited By

  • Large model-based element type appeal generation method and system

    CN120781849A

  • Knowledge structured extraction method based on multi-modal large model

    CN121031762A

  • Multi-granularity structured representation-based legal case retrieval method and system

    CN122332522A