Legal instrument element extraction and generation method

By constructing a multi-dimensional analytical model and obtaining a general feature library, and combining the case type to determine the specific feature requirements, the problem of insufficient targeting in the extraction of legal document elements in existing technologies has been solved, and accurate, comprehensive, and judicially logical element extraction has been achieved.

CN121301548APending Publication Date: 2026-01-09MUDANJIANG NORMAL UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511605581.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing methods for extracting and generating legal document elements lack a deep integration of adjudication rules and judicial logic, making it difficult to identify implicit related elements and adapt to the personalized needs of different causes of action and fields, resulting in insufficient targeting of extraction results.

Method used

A multi-dimensional analytical model is constructed, which combines legal document judgment rules data and element validity verification data to perform layered analysis, obtain a general element feature library of similar documents in the same field, determine the specific element requirement features based on the cause of action type, and obtain the core element extraction priority through weight adaptation analysis, and finally generate differentiated related element features.

Benefits of technology

It enables precise extraction of elements from legal documents, enhancing the relevance and practicality of the extraction, ensuring that the extraction results conform to legal logic and have solid data support, and adapting to the processing needs of different types of cases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121301548A_ABST
    Figure CN121301548A_ABST
Patent Text Reader

Abstract

The invention discloses an element extraction and generation method for a legal instrument, and relates to the technical field of artificial intelligence, and the key points of the technical scheme comprise the following steps: constructing a multi-dimensional element analysis model based on legal instrument judgment rule data and element validity verification data, performing hierarchical analysis on the fact affirmation content, the legal reasoning content and the text content of the current legal document to be processed through the analysis model to obtain an initial element extraction result; acquiring a general element feature library of similar current to-be-processed legal instruments in the same legal field to obtain a reference element feature set, and obtaining exclusive element demand features in combination with the cause type of the current to-be-processed legal instruments; the method has the effects of focusing on individual requirements of an individual case, not only avoiding extraction deviation caused by separation from industry conventions, but also accurately matching unique element requirements of different causes, and greatly improving the pertinence and practicability of element extraction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, more particularly, it relates to a legal document element extraction generation method. BACKGROUND

[0002] In judicial practice and legal affairs processing, legal documents, as the core carrier of bearing case facts, legal reasoning and adjudication results, the quality of the element extraction directly affects the case trial efficiency, the scientific nature of judicial decision and the depth of legal research; with the continuous advancement of the rule of law, the number of legal documents has increased explosively, covering civil, criminal, administrative and other fields, and the documents issued by different courts in different cases have significant differences in structure, expression and element emphasis, which puts higher requirements on the accuracy, comprehensiveness and efficiency of element extraction. However, the existing legal document element extraction generation method is mostly based on simple keyword matching or single-dimensional rule for element identification, which not only lacks deep integration of legal document adjudication rules and judicial logic, but also can only extract simple elements that are explicit in the text, and it is difficult to identify implicit associated elements in the legal reasoning process, such as the causal relationship between elements, the inclination of the judgment, etc.; At the same time, it does not fully consider the individualized element needs of different cases and different fields of documents, and uses a unified extraction standard, resulting in insufficient pertinence of the extraction results, which cannot adapt to the processing needs of different types of cases such as contract disputes and tort disputes. SUMMARY

[0003] In view of the deficiencies in the prior art, the purpose of the present application is to provide a legal document element extraction generation method.

[0004] To achieve the above-mentioned purpose, the present application provides the following technical solutions: A legal document element extraction generation method, the method comprising the following steps: Based on the legal document adjudication rule data and the element validity verification data, a multi-dimensional element analysis model is constructed, the fact determination content, the legal reasoning content and the text content of the current legal document to be processed are analyzed by the analysis model, and the initial element extraction result is obtained; Obtain a reference element feature set from a general element feature library of the same type of current legal document in the same legal field, and obtain exclusive element demand features based on the case type of the current legal document to be processed; The exclusive element demand features and the reference element feature set are analyzed to obtain the core element extraction priority; Filter the element features matching the case type from the reference element feature set, and take the corresponding inclination of the judgment as the basic element direction; Fuse the basic element direction and the core element extraction priority to obtain differential associated element features; The exclusive element demand characteristics, the analysis model and the differentiated correlation element characteristics are parameter integrated to generate an extraction model, and the current legal document to be processed is extracted and analyzed to obtain optimized element results one and optimized element results two; According to the optimized element results one and the optimized element results two, the initial optimized element results are checked to obtain a final element extraction scheme.

[0005] Preferably, the analysis model is constructed based on the current legal document to be processed and the element validity checking data, specifically including the following steps: Collect the current legal document to be processed and the element validity checking data; The current legal document to be processed of different types, different fields and different levels is collected to establish the element validity checking data; Based on the current legal document to be processed and the element validity checking data, a multi-dimensional element analysis model is constructed, which includes time dimension, case type dimension, regional dimension and legal text dimension. The time dimension is analyzed by analyzing the current legal document data in different time spans to mine the evolution law of legal application and judicial rules. The case type dimension classifies and analyzes the elements of different case types, the regional dimension analyzes the differences in judicial practice in different regions, and the legal text dimension analyzes the application of specific legal texts in documents.

[0006] Preferably, the fact determination content, the legal reasoning content and the text content of the current legal document to be processed are analyzed by the analysis model to obtain the initial element extraction result, specifically including the following steps: The fact determination content of the current legal document to be processed is analyzed by the analysis model from the time, place, person and event dimension to obtain the first analysis data; The legal reasoning content is analyzed by the analysis model from the legal text application, logical reasoning and argument basis dimension to obtain the second analysis data; The multi-level semantics and logical structure of the text content are analyzed by the analysis model to obtain the third analysis data; The first analysis data, the second analysis data and the third analysis data are combined to obtain the initial element extraction result.

[0007] Preferably, the reference element characteristic set is obtained by acquiring the general element characteristic library of the current legal document to be processed in the same legal field, and the exclusive element demand characteristics are obtained by combining the case type of the current legal document to be processed, specifically including the following steps: Collect various legal documents in the same legal field, extract the repetitive and general key elements in the legal documents to form a general element characteristic library; A reference element feature set is obtained by filtering elements related to the current legal document processing needs from the general element feature library; By analyzing the text content through keyword extraction, semantic understanding, and deep learning classification models, the cause of action of the legal document to be processed can be determined. By combining the reference element feature set with the cause of action type of the legal documents to be processed, the unique requirements of each element under different cause of action types are determined to obtain exclusive element requirement features.

[0008] Preferably, a weighted adaptation analysis is performed on the specific element demand characteristics and the reference element feature set to obtain the core element extraction priority, specifically including the following steps: An evaluation matrix of multidimensional features is established for the specific element demand characteristics and the reference element feature set, and the multidimensional features are arranged according to the category dimension and attribute dimension; Among them, the category dimension includes business function category and data attribute category, and the attribute dimension includes the stability and change frequency of the element; Initial weights are assigned to feature elements based on their importance in multidimensional features. The initial weights are dynamically adjusted based on real-time feedback data and business dynamics. The feature elements are then sorted according to the adjusted weights to obtain the core feature extraction priority.

[0009] Preferably, the process involves selecting feature elements from the reference feature set that match the case type, and using their corresponding judicial inclinations as the basic feature directions. This specifically includes the following steps: The reference feature set includes various types of feature elements and their corresponding refereeing tendencies; The first feature is obtained by matching and filtering the cause of action type of the currently pending legal document with the feature features in the reference feature set; Determine the key information involved in the type of cause of action, and compare the key information with the description of legal relationships and common dispute situations associated with each element feature in the reference element feature set to obtain the second feature; When the first feature corresponds to the second feature, it is a feature element; The refereeing tendency corresponding to the feature characteristics is taken as the basic feature direction.

[0010] Preferably, the differentiated related element features are obtained by integrating the direction of basic elements and the extraction priority of core elements, specifically including the following steps: The basic element direction includes extracting the basic data categories of the legal documents to be processed from multiple data sources, performing standardized preprocessing on the basic data categories, eliminating data format differences, filling missing values ​​and constructing a basic data pool; Based on legal knowledge graphs, legal service-related elements such as legal relationships, case types, and legal provisions are extracted to perform semantic analysis on currently pending legal documents and identify core elements in key legal concepts and logical relationships. Based on the needs of the business scenario, weights are assigned to different elements, and core elements are extracted from the basic data pool and the analysis results of the legal documents to be processed in order of weight sorting. By comparing the core elements of different samples, the correlation patterns between the core elements are determined, and the correlation patterns are quantitatively judged to obtain the characteristics of differentiated correlation elements.

[0011] Preferably, the process of extracting and analyzing the currently pending legal documents to obtain optimized element result one and optimized element result two specifically includes the following steps: The first optimized data is obtained by extracting the specific element requirements of the legal documents to be processed based on the extraction model; The second optimized data is obtained by filtering the direct correlation features between text content and differentiated related features based on the analytical model; Optimized element result one is obtained by extracting elements from the factual content, legal reasoning section, and text content of the legal document to be processed using the first and second optimized data. The first matching data is obtained by matching the current legal documents to be processed with the features of differentiated related elements based on the extraction model; The second matching data is obtained by associating the industry and field development trends corresponding to the exclusive element demand characteristics in the current legal documents to be processed with the analytical model; The optimized element result two is obtained by extracting the development trend and potential related elements of the implicit elements in the text content through the first matching data and the second matching data.

[0012] Preferably, the final element extraction scheme is obtained by verifying the initial optimized element results based on optimized element result one and optimized element result two, specifically including the following steps: Obtain the initial feature extraction results and the optimized feature extraction results; The initial and optimized feature extraction results are verified using preset verification rules to obtain the verification results. The initial feature extraction results and the optimized feature extraction results are verified to obtain the verification results. Based on the verification and validation results, the initial feature extraction scheme and the optimized feature extraction scheme are analyzed to obtain the final feature extraction scheme.

[0013] Compared with existing technologies, this invention has the following beneficial effects: By constructing a multi-dimensional analytical model, a foundation for accurate extraction is laid, improving the quality of initial element extraction. Furthermore, by combining legal document judgment rule data with element validity verification data, a multi-dimensional analytical model is built. This model not only follows the inherent rules of judicial judgment but also controls element quality using validity standards. Simultaneously, it performs layered analysis of document fact-finding, legal reasoning, and text content, achieving accurate positioning and preliminary extraction of different types of elements. This avoids element omissions and confusion caused by a lack of systematic model support, ensuring that the initial element extraction results are both consistent with legal logic and have solid data support. By acquiring a general element feature library of similar documents in the same field, a universally applicable reference element feature set is extracted, guaranteeing the industry universality of element extraction. Combined with the current document case type, specific element requirement features are determined, focusing on individual case-specific needs. This avoids extraction deviations caused by deviating from industry norms and accurately matches the unique element requirements of different case types, significantly improving the targeting and practicality of element extraction. Attached Figure Description

[0014] Fig. 1 A schematic diagram illustrating the steps of a method for extracting and generating elements of legal documents proposed in this invention; Fig. 2 This is a schematic diagram illustrating the initial element extraction result obtained in the element extraction and generation method for legal documents proposed in this invention; Fig. 3 This is a schematic diagram illustrating the optimized element result one and optimized element result two obtained in the element extraction and generation method for legal documents proposed in this invention. Detailed Implementation

[0015] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0016] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0017] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places throughout this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment selectively excluded from other embodiments.

[0018] Reference Figs. 1-3 As shown.

[0019] This embodiment further illustrates the method for extracting and generating elements of legal documents proposed in this invention.

[0020] A method for extracting and generating elements of a legal document, comprising the following steps: Based on legal document judgment rules data and element validity verification data, a multi-dimensional element analysis model is constructed. The analysis model is used to perform hierarchical analysis of the factual determination content, legal reasoning content and text content of the legal document to be processed, and the initial element extraction results are obtained. Obtain a reference feature set by acquiring a general feature library of similar legal documents currently pending in the same legal field, and combine it with the cause of action of the legal documents currently pending to obtain specific feature requirements. Weight adaptation analysis is performed on the specific element demand characteristics and the reference element characteristic set to obtain the core element extraction priority; Select the element features that match the case type from the set of reference element features, and use their corresponding judgment tendencies as the basic element directions; By integrating the direction of basic elements and the priority of core element extraction, differentiated related element features are obtained; The exclusive element demand characteristics, analytical model and differentiated related element characteristics are integrated into parameters to generate an extraction model. The current legal documents to be processed are extracted and analyzed to obtain optimized element result one and optimized element result two. The final element extraction scheme is obtained by verifying the initial optimized element results based on optimized element result one and optimized element result two.

[0021] The adjudication rules data covers the regularities of adjudication logic, legal application standards, and key points of fact-finding in judgment documents across different legal fields, providing a rule-based basis for the model to identify key legal elements in the documents. The element validity verification data includes the validity judgment standards for various elements in different scenarios and the characteristics of common erroneous elements, ensuring that the elements extracted by the model possess basic legality and rationality. The analytical model performs layered analysis on the legal documents to be processed: extracting core elements related to the facts of the case from the fact-finding content, such as party information, time and place of the case, course of action, and damage results; analyzing the legal reasoning content, including legal citations, legal relationship analysis, and elements of reasoning for adjudication; and simultaneously analyzing the overall text content of the document to ensure that no potential key elements are missed. Through the layered analysis of the analytical model, initial element extraction results covering multiple types of information are initially obtained. To enhance the relevance and comprehensiveness of element extraction, two types of key element characteristics are acquired simultaneously: First, by retrieving a general element characteristic database of similar legal documents to be processed within the same legal field, a set of universal element characteristics formed during the long-term processing of such documents is selected, i.e., a reference element characteristic set, thereby indicating the common needs for element extraction in similar cases; Second, by combining the cause of action of the currently pending legal documents, such as contract disputes, tort liability disputes, and marriage and family disputes, an in-depth analysis is conducted on the core elements that legal documents under specific cause of action must contain, the elements that are of key concern in judicial practice, and the special elements closely related to the cause of action, thereby determining the exclusive element requirement characteristics, highlighting the individual needs of element extraction for current cases, and ensuring that the extraction direction is highly consistent with the cause of action. Because different elements have varying importance in case handling, a weighted adaptation analysis is performed on the exclusive element requirement characteristics and reference element feature sets obtained in the early stage. Priority weights are assigned based on the closeness of the exclusive element requirement characteristics to the current cause of action and their impact on the case's judgment. Priority weights are determined by combining the frequency of occurrence of each element in similar cases and the degree of judicial attention received. Through comprehensive consideration and adaptation of feature weights, it is clarified which elements are core elements that must be extracted first, forming a priority for core element extraction and providing sequential guidance for subsequent accurate extraction. From the set of reference feature elements, we further screened out feature elements that perfectly matched the cause of action of the legal documents to be processed. These features are not only universal but also highly consistent with the adjudication logic of specific cause of action. Based on the screened feature elements, combined with the past adjudication results, adjudication tendencies, and application practices of relevant legal provisions in similar cases, we extracted the basic direction that should be followed for the extraction of elements under this type of cause of action, namely the basic element direction. This direction defines a compliant and reasonable scope for subsequent element extraction, ensuring that the extracted elements conform to the adjudication orientation in judicial practice and avoiding deviation from the core legal logic. Using the basic elements as the benchmark framework, the high-priority elements in the core element extraction priority are taken as the core nodes within the framework. At the same time, combined with the differentiated information in the reference element feature set and the exclusive element requirement features, distinctive elements that conform to both general rules and meet exclusive needs are added to form differentiated related element features. These features include both common elements of similar cases and individual elements of the current case, while clarifying the importance order and extraction direction of the elements, thus achieving the organic integration of multi-dimensional element information. The specific element requirements are transformed into personalized extraction parameters for the model, and the differentiated related element features are transformed into feature recognition and priority ranking parameters for the model. At the same time, the original hierarchical parsing parameters of the parsing model are retained. Through the collaborative adaptation and logical association between parameters, a specific extraction model for the current legal document to be processed is constructed. The extraction model is then used to conduct a comprehensive extraction analysis of the current legal document to be processed. The first analysis focuses on extracting core elements according to their priority, prioritizing the extraction of high-weight core elements, resulting in optimized element result one. The second analysis focuses on combining the direction of basic elements and the features of differentiated related elements to comprehensively extract all key elements covering both common and individual characteristics, resulting in optimized element result two. These two optimization results, starting from different emphases, respectively ensure the focus and comprehensiveness of element extraction. To ensure the accuracy and completeness of the element extraction results, optimized element result one and optimized element result two were used as the verification basis. The initial element extraction results were comprehensively verified. First, it was checked whether the initial results contained all the core and key elements of optimized result one and optimized result two. If any were missing, they were added in a timely manner. Second, it was verified whether the elements extracted from the initial results were consistent with the elements in the optimized results in terms of content, attributes, and logical relationships. If any contradictions or errors were found, they were corrected. At the same time, the validity verification data of the elements was used to verify whether the supplemented and corrected elements met the validity standards. Through verification, a final element extraction scheme that is comprehensive and accurate, highlights the key points, and conforms to the judicial logic and the needs of the case was obtained.

[0022] The analytical model is constructed based on the current data of legal document judgment rules and the data of element validity verification. The specific steps include: Collect data on the rules governing the adjudication of currently pending legal documents; Collect and establish data for verifying the validity of elements in currently pending legal documents of different types, fields, and levels; Based on the current data on the adjudication rules and the validity verification data of the legal documents to be processed, a multi-dimensional analytical model is constructed. The multi-dimensional elements include time dimension, case type dimension, geographical dimension, and legal provision dimension. The time dimension analyzes the data on the current legal documents to be processed within different time spans to explore the evolution of the application of law and adjudication rules. The case type dimension classifies and analyzes the elements of cases with different causes of action. The geographical dimension analyzes the differences in judicial practice in different regions. The legal provision dimension analyzes the application of specific legal provisions in the documents.

[0023] This process comprehensively collects the adjudication rules contained in the currently pending legal documents, covering various rule elements upon which the judgments were based. For example, in contract dispute documents, it includes the rules followed in determining the validity of contracts, the standards for determining liability for breach of contract, and the adjudication logic for handling contract termination matters; in tort liability dispute documents, it includes the rules for defining the establishment of tortious acts, the basis for allocating the proportion of liability, and the adjudication criteria for determining the scope of compensation. By systematically collecting and organizing the adjudication rules in different documents, direct rule basis is provided for the subsequent model construction. This study extensively collects pending legal documents from three aspects: type, field, and level. It covers civil, criminal, and administrative legal documents, as well as multiple professional fields such as finance, intellectual property, and labor disputes. In terms of level, it includes documents issued by basic-level courts, intermediate courts, higher courts, and even the Supreme Court. During the collection process, various elements in each document are marked for validity. For example, in labor dispute documents, elements regarding the calculation base for overtime pay are marked as valid if they are consistent with local labor laws and judicial practices. In intellectual property documents, elements regarding the determination of originality are marked as invalid if they contain logical contradictions or are inconsistent with legal provisions. Through the collection and validation of a large number of documents of different types, fields, and levels, a comprehensive dataset for element validity verification is gradually constructed, providing a reference standard for subsequent model identification of valid elements. The collected data on adjudication rules and the validity verification data of elements are deeply integrated to build a multi-dimensional element analysis model around four key dimensions: time, case type, region, and legal provisions. In the time dimension, by analyzing document data across different time spans, the evolution of legal application and adjudication rules is explored. For example, comparing documents related to private lending disputes from five and ten years ago reveals the changing patterns of adjudication rules regarding the upper limit of interest protection as relevant judicial interpretations are revised. In the case type dimension, elements of cases with different causes of action are classified and analyzed. For example, divorce disputes and inheritance disputes are analyzed separately. In divorce disputes, the focus is on the division of marital property and child custody, while in inheritance disputes... The disputes focus on elements such as defining the scope of the estate and determining the qualifications of heirs; from a geographical perspective, the differences in judicial practice across different regions are analyzed. For example, in the same traffic accident liability dispute, economically developed and underdeveloped regions show different judicial tendencies regarding the calculation standards for lost wages and the amount of compensation for mental distress; from the perspective of legal provisions, the application of specific legal provisions in documents is analyzed. For example, the relevant provisions on the validity of standard clauses in the Civil Code are analyzed in their application scenarios, conditions, and correlation with judgment results in various contract dispute documents. Through comprehensive analysis and data integration across these four dimensions, a multi-dimensional analytical model capable of accurately analyzing the elements of legal documents is ultimately constructed.

[0024] The analytical model is used to perform hierarchical analysis of the factual findings, legal reasoning, and textual content of the legal document to be processed, resulting in initial element extraction. This process includes the following steps: The first analytical data is obtained by analyzing the factual content of the legal documents to be processed from the dimensions of time, place, people, and event sequence using an analytical model. The second analytical data is obtained by analyzing the legal reasoning content from the dimensions of legal provisions application, logical reasoning, and argumentation basis through an analytical model. The third parsing data is obtained by analyzing the multi-level semantics and logical structure of the text content through a parsing model; The first, second, and third analytical data are combined to form the initial feature extraction result.

[0025] For the factual findings in legal documents, the analytical model focuses on four core dimensions—time, place, people, and event sequence—to conduct in-depth analysis and obtain primary analytical data. Factual findings, as the foundation of legal documents, directly relate to the basis for subsequent legal judgments; therefore, the analytical process must accurately capture key information. For example, in the case of a sales contract dispute, the model extracts key time nodes such as the contract signing time, goods delivery time, payment deadline, and the time the dispute occurred, clarifying the chronological order of events. In the place dimension, it locates the place of contract performance, the place of goods delivery, the locations of both parties, and relevant geographical information, which may affect the determination of the court with jurisdiction and the application of relevant regional regulations. In the people dimension, it accurately identifies the identity information of the plaintiff, defendant, and third parties, including name, gender, occupation, and address, while also distinguishing witnesses, expert witnesses, and other participating entities. In the event sequence dimension, it analyzes the background of the contract signing, the specifications and quantity of the goods, whether there were any defects during delivery, the fulfillment of payment obligations, and the key details of the specific reasons for the dispute between the parties. The information extracted from these four dimensions is integrated to form primary analytical data reflecting the basic facts of the case. Legal reasoning is the core link connecting the facts of a case with the judgment in a legal document. The analysis process needs to penetrate the surface of the text and uncover the inherent legal logic. For example, in a sales contract dispute, in terms of the application of legal provisions, the analysis model identifies the legal provisions related to the sales contract cited in the document, such as clauses concerning the determination of contract validity, liability for breach of contract, and inspection of the subject matter, clarifying the specific application scenario of each legal provision in the case. In terms of logical reasoning, it traces the judge's reasoning path from the facts of the case to the application of law and then to the judgment, analyzing how to derive a conclusion that conforms to the law through fact-finding. For example, how to deduce the conclusion that the seller should bear liability for breach of contract based on the fact that the goods have quality defects, combined with relevant legal provisions. In terms of the basis of argumentation, it extracts the supporting materials such as judicial interpretations, guiding cases, and legal theories cited by the judge in the argumentation process, which are important bases for enhancing the persuasiveness of the judgment. Integrating the information obtained from these three dimensions of analysis forms the second analytical data that reflects the legal logic of the document.

[0026] Because the text not only contains explicit information about fact-finding and legal reasoning, but also hides many potential logical connections and semantic connotations, the analytical model first conducts in-depth semantic mining of the text. In addition to the explicitly stated information, it also captures implicit semantic tendencies, such as the implicit tendency to judge the legality of a party's behavior in the document. Then, from the logical structure level, it analyzes the overall layout of the document and the logical connection between paragraphs, clarifies the internal connection between the fact-finding content, legal reasoning content, and the main text of the judgment, and judges whether the content is consistent and logically self-consistent. Through a comprehensive analysis of the multi-level semantics and logical structure of the text, a third analytical data that can reflect the overall content relevance of the document is formed. The first, second, and third analytical data obtained through hierarchical analysis will be fully integrated to form initial element extraction results that cover multiple aspects of information, including case facts, legal logic, and textual relationships.

[0027] Obtain a reference feature set by acquiring a general feature library of currently pending legal documents within the same legal field, and combine this with the cause of action type of the currently pending legal documents to obtain specific feature requirements. This process includes the following steps: Collect various legal documents within the same legal field, and extract repetitive and common key elements from these documents to form a general element feature library; A reference element feature set is obtained by filtering elements related to the current legal document processing needs from the general element feature library; By analyzing the text content through keyword extraction, semantic understanding, and deep learning classification models, the cause of action of the legal document to be processed can be determined. By combining the reference element feature set with the cause of action type of the legal documents to be processed, the unique requirements of each element under different cause of action types are determined to obtain exclusive element requirement features.

[0028] This involves collecting various legal documents within the same legal field. This field can be categorized into civil, criminal, and administrative categories, or further subdivided into more specific areas, such as contract disputes or tort liability disputes within the civil field. During the collection process, the elements within the documents are systematically reviewed and analyzed, with a focus on identifying key elements that are repetitive and universal. For example, in the field of contract disputes, among numerous sales contracts, lease contracts, and loan contracts, elements such as basic information of the parties, core points of contract signing and performance, information related to the subject matter of the contract, and manifestations of breach of contract frequently appear. These elements are common to documents within this field. By extracting, classifying, and integrating these common elements, and eliminating redundant information, a clear and comprehensive database of common element features is formed, providing fundamental data support for subsequent element selection. Clearly define the core processing objectives and element extraction direction of the documents currently pending. For example, if the document pending is a housing lease contract dispute, the processing needs may focus on the condition of the leased property, rent payment, lease term, and contract termination conditions. Based on this need, select elements closely related to housing lease contract dispute processing from the general element feature library in the field of contract disputes, such as the identity information of the lessor and lessee, the location and condition of the leased property, rent standard and payment method, lease term agreement, the rights and obligations of both parties, breach of contract and liability. The selected elements constitute a reference element feature set for the document pending, which retains the general attributes of the field and initially meets the needs of the current case. By using keyword extraction technology, high-frequency words related to specific causes of action are captured in the documents. For example, in documents concerning housing rental contract disputes, words such as rent, leased property, lease term, and termination of lease appear frequently. By using semantic understanding technology, the core meaning expressed in the document text is analyzed to clarify the focus of the dispute and the nature of the legal relationship. At the same time, a deep learning classification model is combined. Since this model has been trained on a large amount of document data with labeled causes of action, it can automatically identify the type of cause of action based on text features. Therefore, through the synergistic effect of these three technologies, the specific cause of action of the legal document to be processed can be accurately determined, such as clarifying that it is a housing rental contract dispute rather than other types of contract disputes. By combining the selected set of reference element features with the determined specific cause of action, we can conduct an in-depth analysis of the unique requirements for each element under that specific cause of action. Taking housing rental contract disputes as an example, compared with other contract disputes, they have more special and critical requirements regarding the ownership status of the leased property, the division of responsibility for property maintenance, subletting agreements, and the conditions for the return of the rental deposit. By judging the importance and relevance of each element in the set of reference element features under the current cause of action, we can extract the elements that meet the unique requirements of that specific cause of action, and finally form exclusive element requirement features, providing targeted guidance for the subsequent accurate extraction of elements.

[0029] The priority for extracting core elements is determined by weighting and adapting the specific element demand characteristics and the reference element feature set. This process includes the following steps: An evaluation matrix of multidimensional features is established for the specific element demand characteristics and the reference element feature set, and the multidimensional features are arranged according to the category dimension and attribute dimension; Among them, the category dimension includes business function category and data attribute category, and the attribute dimension includes the stability and change frequency of the element; Initial weights are assigned to feature elements based on their importance in multidimensional features. The initial weights are dynamically adjusted based on real-time feedback data and business dynamics. The feature elements are then sorted according to the adjusted weights to obtain the core feature extraction priority.

[0030] By integrating the specific element requirements and reference element feature sets, all element features to be analyzed are identified. A multi-dimensional feature evaluation matrix is ​​constructed using these features as the core. During matrix construction, the multi-dimensional features are systematically arranged according to category and attribute dimensions. The category dimension includes business function categories and data attribute categories. Business function categories are mainly based on the element's role in legal document processing, such as the element's use for basic case information identification, legal relationship analysis, and judgment prediction. Data attribute categories are distinguished according to the element's data type, such as whether the element belongs to text data, numerical data, or date data. The attribute dimension focuses on the element's inherent characteristics, covering stability and frequency of change. Stability reflects the degree to which the element remains fixed in documents across different cases and periods, while frequency of change reflects the frequency with which the element changes due to case type and time factors. Through this categorized arrangement, each element feature finds its corresponding position in the evaluation matrix, clearly presenting its attribute characteristics across different dimensions. Taking the field of contract disputes as an example, among the key characteristics of this field, the element of the parties' names, from a categorical perspective, falls under the business function category of basic case information identification and the data attribute category of textual data. From an attribute perspective, it exhibits high stability, remaining largely unchanged across different stages of the same case with a low frequency of change. In contrast, the element of calculating liquidated damages for breach of contract falls under the business function category of liability determination and the data attribute category of numerical data. It exhibits low stability, changing due to contractual agreements, breach of contract circumstances, and relevant legal revisions, with a higher frequency of change than the element of the parties' names. Through the evaluation matrix, the multidimensional characteristics of these elements can be clearly displayed. For each dimension of the evaluation matrix, corresponding importance assessment criteria are established. In the business function category dimension, importance is determined based on the degree of influence of the element on the core objectives of case handling. For example, elements used for predicting judgment results have a greater impact on case handling decisions and are more important than elements used only for basic information identification. In the data attribute category dimension, importance is determined by combining the difficulty of data processing and its supporting role in subsequent analysis. For example, numerical compensation amount elements are more important than some textual elements because they are directly used for quantitative analysis. In the attribute dimension, elements with high stability and low change frequency are usually more important than elements with poor stability and frequent changes because they can serve as stable references. Based on the comprehensive importance assessment results of each dimension, a corresponding initial weight is assigned to each element feature. The higher the weight value, the higher the basic importance of the element in the extraction process. Real-time feedback data includes the application effects after element extraction and changes in attention to various elements in practice. For example, if the frequency of a certain element in recent cases increases significantly and its impact on the judgment becomes more significant, its weight needs to be adjusted accordingly. Business dynamics include revisions to legal provisions, adjustments to judicial policies, and changes in business processing needs. If revisions to relevant legal provisions lead to a change in the legal meaning of a certain element, or if business needs shift to focus more on a certain type of element, the weight of the corresponding element needs to be adjusted. All feature elements are sorted according to their adjusted weights, with the element with the highest weight being the primary core element to be extracted, and so on, forming the final priority for core element extraction. This provides clear guidance for the subsequent accurate and orderly extraction of legal document elements.

[0031] The process involves selecting feature elements from the reference feature set that match the case type, and using their corresponding judicial tendencies as the basic feature direction. This includes the following steps: The reference feature set includes various types of feature elements and their corresponding refereeing tendencies; The first feature is obtained by matching and filtering the cause of action type of the currently pending legal document with the feature features in the reference feature set; Determine the key information involved in the type of cause of action, and compare the key information with the description of legal relationships and common dispute situations associated with each element feature in the reference element feature set to obtain the second feature; When the first feature corresponds to the second feature, it is a feature element; The refereeing tendency corresponding to the feature characteristics is taken as the basic feature direction.

[0032] The key features are derived from the commonalities of similar documents within the same legal field, covering multiple dimensions such as basic case information, legal relationships, and points of contention. The corresponding judicial inclinations, on the other hand, are based on the judicial practice patterns summarized from the judgments and reasoning of numerous similar documents. For example, when key features appear in a case, what kind of judgment will the court typically make, or the general discretionary logic followed by judicial organs in cases involving those key features. Taking the reference key feature set in the field of loan disputes as an example, it includes the key features of loan principal amount, loan interest agreement, loan delivery voucher, and repayment record. Each key feature corresponds to a specific judicial inclination. For instance, if the loan interest agreement exceeds the statutory limit, the corresponding judicial inclination is usually that the excess portion will not be supported; if the loan delivery voucher is missing, the corresponding judicial inclination is that further evidence of actual loan delivery is required. Based on the cause of action of the legal document currently pending, the system directly matches it with the feature set of reference features to obtain the first feature. Since the cause of action directly determines the nature of the legal relationship and the core scope of the dispute, the matching process will focus on the feature set of features that are highly related to the cause of action. For example, if the cause of action of the document currently pending is a loan dispute, the system will select the feature set of reference features that are directly related to the cause of action, such as the identity information of the lender and borrower, the signing of the loan contract, the interest rate, the flow of loan funds, and the guarantee. The selected feature set of features together constitutes the first feature, which initially defines the range of features that match the current cause of action. Analyzing the key information involved in the current cause of action type, which is the core constituent element of such cases and the focus of judicial practice; taking loan disputes as an example, the key information involved includes whether the loan agreement was established, whether the loan was actually delivered, whether the interest agreement was legal, and whether the loan has been partially or fully repaid. A comprehensive comparison is made between this key information and the descriptions of the legal relationships associated with each element feature in the reference element feature set, as well as common disputed situations. This involves checking whether the descriptions of the legal relationships associated with each element feature are consistent with the legal relationship of the loan, and simultaneously determining whether the common disputed situations corresponding to the element features are frequently occurring points of contention in loan cases, such as whether the interest exceeds the limit or whether the loan was actually delivered. Thus, element features that match the key information of the cause of action are selected to form the second feature. If the feature characteristics obtained from the two screenings are completely overlapping or highly consistent, it means that these feature characteristics are not only directly related to the cause of action type, but also match the legal relationship and dispute situation corresponding to the key information of the cause of action, and have sufficient matching rationality. At this time, these feature characteristics become the final locked feature characteristics. If there are differences between the two, the matching process needs to be checked again to ensure the accuracy of the screening results. Based on the final determined element characteristics, the corresponding judicial tendencies are extracted from the reference element characteristic set. Judgment tendencies are a concentrated reflection of the judicial practice rules of similar cases and can provide clear directional guidance for the extraction of elements from the documents to be processed. For example, the locked element characteristics correspond to the judicial tendencies of not supporting the interest exceeding the limit and needing to supplement evidence for the absence of delivery vouchers. Together, they constitute the basic element direction to ensure that the subsequent element extraction work is highly consistent with the judicial judgment logic.

[0033] By integrating the direction of basic elements and the priority of core element extraction, differentiated related element features are obtained, which specifically includes the following steps: The basic element direction includes extracting the basic data categories of the legal documents to be processed from multiple data sources, performing standardized preprocessing on the basic data categories, eliminating data format differences, filling missing values ​​and constructing a basic data pool; Based on legal knowledge graphs, legal service-related elements such as legal relationships, case types, and legal provisions are extracted to perform semantic analysis on currently pending legal documents and identify core elements in key legal concepts and logical relationships. Based on the needs of the business scenario, weights are assigned to different elements, and core elements are extracted from the basic data pool and the analysis results of the legal documents to be processed in order of weight sorting. By comparing the core elements of different samples, the correlation patterns between the core elements are determined, and the correlation patterns are quantitatively judged to obtain the characteristics of differentiated correlation elements.

[0034] From various data sources related to the legal documents currently pending processing, basic data categories relevant to case handling are selected. These diverse data sources encompass the textual data of the documents themselves, related case files, and judicial statistics. The extracted basic data categories include basic information about the parties involved, case timelines, information on the subject matter of the case, evidence materials, and content closely related to the basic facts of the case. Taking construction contract disputes as an example, the extracted basic data categories might include the identity information of the contracting party and contractor, contract signing and commencement / completion dates, total project cost, amount already paid, and project quality acceptance records. After extraction, the basic data categories undergo standardized preprocessing. To address the issue of inconsistent data formats across different data sources, scattered text, numerical, and date formats are converted into a unified standard format. For example, the formats of project completion dates (e.g., 2023.10.01 and 2023-10-01) in different documents are standardized into a fixed format. For missing data, by reviewing related evidence materials and referring to data completion rules from similar cases, missing contact information for parties and some missing project progress nodes are filled in. Finally, a well-structured and complete basic data pool is formed, providing high-quality basic data for subsequent element extraction. By leveraging legal knowledge graphs, key elements related to legal services are deeply extracted from currently pending legal documents. These graphs integrate a vast amount of legal concepts, relationships, provisions, and case data, clearly presenting the logical connections between various legal elements. Based on this graph, staff extract key elements such as legal relationships, case types, and cited legal provisions from the documents. This allows them to determine the rights and obligations between the parties, clarifying whether the legal relationship is a construction contract or a labor subcontracting relationship; defining the case type, such as a construction contract dispute involving project price settlement; and identifying relevant legal provisions from the Civil Code's Contract Law and the Regulations on the Administration of Construction Project Quality. Simultaneously, semantic analysis technology is used to deeply interpret the document text, identifying key legal concepts such as project changes, final acceptance, and penalties for delayed completion. The logical relationships between these concepts are determined, such as the causal relationship between project changes and project price adjustments, and the liability relationship between unqualified final acceptance and the contractor's maintenance obligations, thus pinpointing elements with core legal significance in the documents. Different business scenarios have different priorities regarding the elements they require. For example, case mediation scenarios focus more on the amount in dispute and the parties' willingness to mediate, while case judgment prediction scenarios place more emphasis on the application of legal provisions and the judgment results of similar cases. Based on the needs of the current specific business scenario, weights are assigned to the basic elements extracted from the basic data pool and the core legal elements obtained from document analysis: higher weights are assigned to elements that are closely related to the business scenario objectives and have a high degree of influence, while lower weights are assigned to elements that have a low degree of relevance and a small impact. Taking the business scenario of construction project contract dispute settlement as an example, the total project cost, the amount already paid, the project change order, and the bill of quantities are directly related to the settlement of the payment, and their weights will be significantly higher than the secondary element of the parties' professional information. After sorting according to the weight, elements are extracted from the basic data pool and the document analysis results in descending order of weight. The elements with the highest weight ranking are the core elements in the current business scenario. Multiple similar case samples were collected, and the core elements of the current pending documents were compared and analyzed with the core elements of other samples. The combination, frequency of occurrence, and mutual influence of the core elements in different samples were observed to determine the correlation patterns between the core elements. For example, in multiple construction contract disputes, it was found that the three core elements of frequent project changes, increased project volume, and project price exceeding the contractual agreement often appeared simultaneously, forming a fixed correlation pattern. There was a clear causal correlation between failure to pay the project progress payment as agreed and the contractor's work stoppage. These correlation patterns were quantitatively judged, and the probability of occurrence and correlation strength indicators of each correlation pattern in the samples were statistically analyzed to distinguish the differences in correlation patterns in different case samples. For example, the correlation strength between project changes and price adjustments was extremely high in some samples, while the correlation strength was weaker in others due to special agreements. Based on the quantitative analysis results, characteristics that both conform to the commonalities of element correlation in similar cases and reflect the individuality of element correlation in the current pending documents were extracted, ultimately forming differentiated correlation element characteristics.

[0035] The analysis of the currently pending legal documents yields two optimized element results: Result 1 and Result 2. The specific steps include: The first optimized data is obtained by extracting the specific element requirements of the legal documents to be processed based on the extraction model; The second optimized data is obtained by filtering the direct correlation features between text content and differentiated related features based on the analytical model; Optimized element result one is obtained by extracting elements from the factual content, legal reasoning section, and text content of the legal document to be processed using the first and second optimized data. The first matching data is obtained by matching the current legal documents to be processed with the features of differentiated related elements based on the extraction model; The second matching data is obtained by associating the industry and field development trends corresponding to the exclusive element demand characteristics in the current legal documents to be processed with the analytical model; The optimized element result two is obtained by extracting the development trend and potential related elements of the implicit elements in the text content through the first matching data and the second matching data.

[0036] By constructing a dedicated extraction model, targeted extraction is carried out on the specific element requirements of the legal documents to be processed, resulting in the first optimized data. The specific element requirements are personalized element requirements determined by the cause of action type. The extraction model has integrated the parameters corresponding to these features, enabling it to accurately locate and extract relevant elements. Taking a property service contract dispute as an example, the specific element requirements include the property fee standard, the agreement on the quality of property services, the duration of the owner's arrears, and the performance of the property company's services. The corresponding information is accurately extracted from the documents, such as the property fee standard of 2 yuan per square meter per month, the security, cleaning and greening services agreed in the contract, the duration of the owner's non-payment of property fees since January 2023, and the property company's failure to maintain public facilities regularly as agreed. The extracted information constitutes the first optimized data, which fits the specific needs of the current case. Differentiated correlation feature characteristics are formed by integrating the direction of basic elements and the priority of core elements. They include the correlation patterns between elements. The analytical model will delve into the document text to uncover elements directly related to the correlation features. Taking property service contract disputes as an example, if there is a pattern in the differentiated correlation feature characteristics that is directly related to the substandard quality of property services and the owner's arrears, the analytical model will filter out the content directly related to this correlation pattern from the document text. Such as the owner's statement that he refuses to pay due to frequent elevator malfunctions, the property company's defense that the owner's arrears are unrelated to service quality, and the contract's stipulations on the owner's rights when services are substandard. The information integration that directly reflects the correlation of elements forms the second optimized data. During the extraction process, the core focus is on the specific elements clearly defined in the first optimized data, supplemented by the related elements reflected in the second optimized data. This covers all key parts of the document. Factual details related to specific elements and related characteristics are extracted from the factual content, such as the specific amount of the owner's arrears and the specific manifestations of the property management company's performance defects. Legal basis citations and liability determination logic related to the elements are extracted from the legal reasoning content, such as citing relevant clauses of the Civil Code regarding property service contracts to analyze the responsibilities of both parties. At the same time, a comprehensive extraction is performed on the entire text content to ensure that no explicit elements related to the first and second optimized data are omitted. Finally, the results are integrated to form a comprehensive and accurate optimized element result one, fully presenting the explicit key elements in the document.

[0037] Differentiated correlation element features include common and unique patterns of element associations in similar cases. The extraction model compares and matches the elements in the current document with these patterns to identify element combinations in the current document that conform to or closely resemble the correlation patterns. Taking property service contract disputes as an example, if the differentiated correlation element features show a pattern that disputes over the quality of property services in older residential communities are mostly related to aging facilities, the extraction model matches the building year of the community, the degree of aging of facilities, and the focus of the property service quality dispute in the current document with this pattern. If the document mentions that the community was built in 2000, that the elevators have been in use for over 20 years and frequently malfunction, and that the main disputes among owners are concentrated on insufficient facility maintenance, the successfully matched element information constitutes the first matching data. The specific elements and their corresponding characteristics correspond to specific industries or fields. The analytical model combines the latest developments in that industry with the judicial practices within that field to uncover elements in the documents that are related to those trends. For property service contract disputes, the property service industry to which it belongs may be trending towards intelligent management, and the judicial practices within that field may place greater emphasis on protecting the owners' right to know. The analytical model extracts relevant information from the documents, such as statements in the documents about the property company's plan to introduce an intelligent access control system, court rulings in similar cases emphasizing that property companies must disclose service costs, and the owners' demands for disclosure of property fee income and expenditure details in the current case. The information related to the industry and field trends forms the second matching data. Based on the element association patterns reflected in the first matching data, and combined with the industry sector trends reflected in the second matching data, information that is not directly stated in the document text but can be deduced is mined. From the development trend of elements, potential trends such as the community's aging facilities and the possible continued increase in disputes over the quality of future property services can be extracted. From potential related elements, implicit elements such as intelligent transformation may become a key factor in resolving property service disputes in the community, and the protection of owners' right to know may affect the outcome of the case can be extracted. The extracted trends and potential elements are integrated to form the second optimized element result, which supplements the deeper information beyond the explicit elements in the document.

[0038] Based on the verification of optimization result one and optimization result two, the final element extraction scheme is obtained by verifying the initial optimized element results, which specifically includes the following steps: Obtain the initial feature extraction results and the optimized feature extraction results; The initial and optimized feature extraction results are verified using preset verification rules to obtain the verification results. The initial feature extraction results and the optimized feature extraction results are verified to obtain the verification results. Based on the verification and validation results, the initial feature extraction scheme and the optimized feature extraction scheme are analyzed to obtain the final feature extraction scheme.

[0039] The synchronous retrieval process yields two core results: 1. Initial element extraction results generated after the initial hierarchical analysis of legal documents by the analytical model, covering basic elements such as factual findings, legal reasoning, and textual content, but potentially lacking elements or having unclear priorities; 2. Optimized element results obtained through two extractions by a dedicated extraction model. Result one focuses on the comprehensive extraction of explicit key elements, while result two focuses on the discovery of implicit trends and potential related elements. Both optimized results are superior to the initial results in terms of element accuracy, completeness, and depth. Taking a loan dispute case as an example, the initial result only extracts basic elements such as the identities of the lender and borrower, the loan amount, and the loan time; optimized result one supplements with dedicated elements such as interest rate agreement standards, loan disbursement vouchers, repayment records, and element-related information; optimized result two uncovers implicit elements such as interest rate agreements approaching the legal limit potentially circumventing legal risks and lenders repeatedly lending money potentially involving professional lending tendencies.

[0040] The pre-defined verification rules are formulated based on industry norms, judicial practice requirements, and data quality standards for extracting elements from legal documents, covering multiple dimensions such as element completeness, accuracy, and compliance. The completeness rule requires that no core elements be omitted. For example, loan cases must include key elements such as the loan agreement, fund delivery, and interest agreement. If the initial result lacks the interest agreement element, it will be judged as incomplete. The accuracy rule requires that the element information be consistent with the original document. If the loan amount of 100,000 yuan is mistakenly written as 10,000 yuan in the initial result, it will be marked as inaccurate. The compliance rule requires that the elements comply with legal provisions. For example, if the interest agreement extracted in the optimized result exceeds the statutory limit, it will be marked as a compliance issue that needs to be checked in detail. By comparing and reviewing the initial result and the optimized result one by one, the missing, incorrect, and compliance defects of elements in various results are identified, forming a detailed verification result. Cross-compare the initial results with the optimized results to verify the consistency and complementarity of the elements. If the identities of the borrowers and lenders and the loan amount are consistent with those in Optimized Result 1, it indicates a high degree of authenticity of the elements. If the interest agreement and delivery voucher elements supplemented by the optimized results are not reflected in the initial results, it is necessary to verify whether the supplemented elements actually exist in the document text. Combining the original legal document and related evidence, the authenticity of disputed or questionable elements is verified. For example, regarding the implicit element in Optimized Result 2 that the lender may be involved in professional lending, it is necessary to check whether there are statements in the document that the lender has lent money to others multiple times, and whether the information of related cases supports this tendency. Through the verification process, the authenticity and reliability of the elements in various results are judged to form the verification results. Based on the verification results, the defects of the initial results and the advantages of the optimized results are identified. For example, if the initial results have missing elements, optimized result one supplements explicit elements, and result two mines implicit elements. Then, combined with the verification results, the true and reliable elements are screened out, and the erroneous information that fails the verification is eliminated. Taking loan disputes as an example, the basic elements of the identities of the lender and borrower that have been verified correctly in the initial results are retained. The explicit elements of the true and valid interest agreement and delivery voucher in optimized result one are integrated, and the implicit trend elements verified in optimized result two are included. At the same time, the erroneous information found in the verification of various results is corrected, and finally a final element extraction scheme with complete elements, accurate information, and both explicit and implicit elements is formed.

[0041] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A method for extracting and generating elements of legal documents, characterized in that, The method includes the following steps: Based on legal document judgment rules data and element validity verification data, a multi-dimensional element analysis model is constructed. The analysis model is used to perform hierarchical analysis of the factual determination content, legal reasoning content and text content of the legal document to be processed, and the initial element extraction results are obtained. Obtain a reference feature set by acquiring a general feature library of similar legal documents currently pending in the same legal field, and combine it with the cause of action of the legal documents currently pending to obtain specific feature requirements. Weight adaptation analysis is performed on the specific element demand characteristics and the reference element characteristic set to obtain the core element extraction priority; Select the element features that match the case type from the set of reference element features, and use their corresponding judgment tendencies as the basic element directions; By integrating the direction of basic elements and the priority of core element extraction, differentiated related element features are obtained; The exclusive element demand characteristics, analytical model and differentiated related element characteristics are integrated into parameters to generate an extraction model. The current legal documents to be processed are extracted and analyzed to obtain optimized element result one and optimized element result two. The final element extraction scheme is obtained by verifying the initial optimized element results based on optimized element result one and optimized element result two.

2. The method for extracting and generating elements of a legal document according to claim 1, characterized in that, The analytical model is constructed based on the current data of legal document judgment rules and the data of element validity verification. The specific steps include: Collect data on the rules governing the adjudication of currently pending legal documents; Collect and establish data for verifying the validity of elements in currently pending legal documents of different types, fields, and levels; Based on the current data on the rules of adjudication and the validity verification data of legal documents to be processed, a multi-dimensional analytical model is constructed. The multi-dimensional elements include time dimension, case type dimension, geographical dimension, and legal provision dimension. The time dimension analyzes the data on the current legal documents to be processed over different time spans to explore the evolution of the application of law and adjudication rules. The case type dimension classifies and analyzes the elements of cases with different causes of action. The geographical dimension analyzes the differences in judicial practice in different regions. The legal provision dimension analyzes the application of specific legal provisions in the documents.

3. The method for extracting and generating elements of legal documents according to claim 2, characterized in that, The analytical model is used to perform hierarchical analysis of the factual findings, legal reasoning, and textual content of the legal document to be processed, resulting in initial element extraction. This process includes the following steps: The first analytical data is obtained by analyzing the factual content of the legal documents to be processed from the dimensions of time, place, people, and event sequence using an analytical model. The second analytical data is obtained by analyzing the legal reasoning content from the dimensions of legal provisions application, logical reasoning, and argumentation basis through an analytical model. The third parsing data is obtained by analyzing the multi-level semantics and logical structure of the text content through a parsing model; The first, second, and third analytical data are combined to form the initial feature extraction result.

4. The method for extracting and generating elements of legal documents according to claim 3, characterized in that, Obtain a reference feature set by acquiring a general feature library of currently pending legal documents within the same legal field, and combine this with the cause of action type of the currently pending legal documents to obtain specific feature requirements. This process includes the following steps: Collect various legal documents within the same legal field, and extract repetitive and common key elements from these documents to form a general element feature library; A reference element feature set is obtained by filtering elements related to the current legal document processing needs from the general element feature library; By analyzing the text content through keyword extraction, semantic understanding, and deep learning classification models, the cause of action of the legal document to be processed can be determined. By combining the reference element feature set with the cause of action type of the legal documents to be processed, the unique requirements of each element under different cause of action types are determined to obtain exclusive element requirement features.

5. The method for extracting and generating elements of a legal document according to claim 4, characterized in that, The priority for extracting core elements is determined by weighting and adapting the specific element demand characteristics and the reference element feature set. This process includes the following steps: An evaluation matrix of multidimensional features is established for the specific element demand characteristics and the reference element feature set, and the multidimensional features are arranged according to the category dimension and attribute dimension; Among them, the category dimension includes business function category and data attribute category, and the attribute dimension includes the stability and change frequency of the element; Initial weights are assigned to feature elements based on their importance in multidimensional features. The initial weights are dynamically adjusted based on real-time feedback data and business dynamics. The feature elements are then sorted according to the adjusted weights to obtain the core feature extraction priority.

6. The method for extracting and generating elements of a legal document according to claim 5, characterized in that, The process involves selecting feature elements from the reference feature set that match the case type, and using their corresponding judicial tendencies as the basic feature direction. This includes the following steps: The reference feature set includes various types of feature elements and their corresponding refereeing tendencies; The first feature is obtained by matching and filtering the cause of action type of the currently pending legal document with the feature features in the reference feature set; Determine the key information involved in the type of cause of action, and compare the key information with the description of legal relationships and common dispute situations associated with each element feature in the reference element feature set to obtain the second feature; When the first feature corresponds to the second feature, it is a feature element; The refereeing tendency corresponding to the feature characteristics is taken as the basic feature direction.

7. The method for extracting and generating elements of a legal document according to claim 6, characterized in that, By integrating the direction of basic elements and the priority of core element extraction, differentiated related element features are obtained, which specifically includes the following steps: The basic element direction includes extracting the basic data categories of the legal documents to be processed from multiple data sources, performing standardized preprocessing on the basic data categories, eliminating data format differences, filling missing values ​​and constructing a basic data pool; Based on legal knowledge graphs, legal service-related elements such as legal relationships, case types, and legal provisions are extracted to perform semantic analysis on currently pending legal documents and identify core elements in key legal concepts and logical relationships. Based on the needs of the business scenario, weights are assigned to different elements, and core elements are extracted from the basic data pool and the analysis results of the legal documents to be processed in order of weight sorting. By comparing the core elements of different samples, the correlation patterns between the core elements are determined, and the correlation patterns are quantitatively judged to obtain the characteristics of differentiated correlation elements.

8. The method for extracting and generating elements of a legal document according to claim 7, characterized in that, The analysis of the currently pending legal documents yields two optimized element results: Result 1 and Result 2. The specific steps include: The first optimized data is obtained by extracting the specific element requirements of the legal documents to be processed based on the extraction model; The second optimized data is obtained by filtering the direct correlation features between text content and differentiated related features based on the analytical model; Optimized element result one is obtained by extracting elements from the factual content, legal reasoning section, and text content of the legal document to be processed using the first and second optimized data. The first matching data is obtained by matching the current legal documents to be processed with the features of differentiated related elements based on the extraction model; The second matching data is obtained by associating the industry and field development trends corresponding to the exclusive element demand characteristics in the current legal documents to be processed with the analytical model; The optimized element result two is obtained by extracting the development trend and potential related elements of the implicit elements in the text content through the first matching data and the second matching data.

9. The method for extracting and generating elements of a legal document according to claim 8, characterized in that, Based on the verification of optimization result one and optimization result two, the final element extraction scheme is obtained by verifying the initial optimized element results, which specifically includes the following steps: Obtain the initial feature extraction results and the optimized feature extraction results; The initial and optimized feature extraction results are verified using preset verification rules to obtain the verification results. The initial feature extraction results and the optimized feature extraction results are verified to obtain the verification results. Based on the verification and validation results, the initial feature extraction scheme and the optimized feature extraction scheme are analyzed to obtain the final feature extraction scheme.

Citation Information

Cited By

  • Intelligent identification system and method applied to legal instruments

    CN121581584A

  • Method and system for automatically extracting and associating legal document key points

    CN121638262A

  • Large model-based legal text generation proofreading method and system

    CN121920329A