A method and system for extracting and enhancing legal text information based on artificial intelligence

By combining semantic segmentation reconstruction, text logic perception and dynamic knowledge fusion methods, the problem of accurate extraction of hierarchical information and logical structures in legal text is solved, and the accurate identification of legal text and logical relationships are realized, which improves the accuracy and completeness of legal information retrieval.

CN120258001BActive Publication Date: 2025-08-22湖南工商大学

Patent Information

Application Number
CN202510745441.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-08-22
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

It is difficult for the existing technology to accurately extract hierarchical information and logical structures in legal texts, especially legal documents with multi-level nested structures, which makes it difficult for information extraction methods to accurately divide clause relationships, affecting the effectiveness of intelligent tasks; traditional methods rely on fixed text formats to perform poorly on non-standardized texts, making it difficult to portray complex legal logical relationships, and static rules are difficult to adapt to the changing role relationships in legal texts.

Method used

A joint framework of semantic segmentation reconstruction, text logic perception and dynamic knowledge fusion is adopted, combined with the conditional random field two-way long and short-term memory model improved by hierarchical label inference and the causal graph convolutional inference neural network, and through multi-channel logical unit enhancement and role-aware triple extraction methods, the precise structure recognition and logical relationship understanding of legal text are achieved.

Benefits of technology

It realizes accurate structural hierarchical identification and logical correlation understanding of legal texts, improves the accuracy and completeness of legal information retrieval, can adapt to non-standardized texts and variable role relationships, and provides reliable logical reasoning support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258001B_ABST
    Figure CN120258001B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for extracting and enhancing legal text information based on artificial intelligence. The method includes data collection and processing, semantic segmentation and reconstruction, text logic perception, dynamic knowledge fusion, and information extraction enhancement. The present invention relates to the field of text information extraction technology, and specifically refers to a method and system for extracting and enhancing legal text information based on artificial intelligence. The method adopts a joint framework of semantic segmentation and reconstruction, text logic perception, and dynamic knowledge fusion to achieve more accurate identification of the structural hierarchy of legal texts, enhance understanding of the logical association of legal clauses, and dynamically integrate knowledge between different legal documents to facilitate cross-reference and intelligent application between different regulations. The method adopts a conditional random field bidirectional long short-term memory model improved by combining hierarchical label reasoning for semantic segmentation and reconstruction. The method adopts a causal graph convolutional inference neural network enhanced by combining dual-channel logic units for text logic perception.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text information extraction, and in particular to an artificial intelligence-based legal text information extraction and enhancement method and system. Background Art

[0002] AI-based enhanced legal text information extraction methods leverage natural language processing (NLP), knowledge graphs, and machine learning technologies to automatically identify key elements in legal documents (such as clauses, parties, and obligations). These methods improve extraction accuracy through semantic understanding, logical reasoning, and dynamic knowledge fusion. These methods aim to efficiently process massive amounts of legal text, addressing the inefficiencies of traditional manual review and reducing the omission or misjudgment of clauses due to subjective oversight. These methods provide structured, analyzable data support for intelligent legal retrieval and contract review, promoting intelligent and standardized legal services.

[0003] However, a core difficulty in the existing legal text information extraction process lies in how to accurately extract the hierarchical information and logical structure in the text. Unlike general text, legal documents typically have a multi-level nested structure, such as laws, regulations, implementation rules, and case analysis. There are cross-references and logical connections between different levels, which makes it difficult for existing information extraction methods to accurately divide the relationship between clauses, affecting the effectiveness of subsequent intelligent tasks. Among existing semantic segmentation and reconstruction methods, the main goal of semantic segmentation is to disassemble complex legal documents according to their inherent logical structure. However, traditional methods usually rely on fixed text formats (such as titles and numbers) when processing, and perform poorly for non-standardized text. Among existing text logic perception methods, traditional text analysis methods mainly rely on keyword matching or static rule-based reasoning models, but these methods have difficulty accurately characterizing complex legal logical relationships. Among existing dynamic knowledge fusion methods, traditional knowledge extraction methods mainly rely on static rules or predefined templates, which makes it difficult to adapt to the changing role relationships in legal texts. Summary of the Invention

[0004] In view of the above situation, in order to overcome the defects of the prior art, the present invention provides an artificial intelligence-based legal text information extraction enhancement method and system. In the existing legal text information extraction process, there is a core difficulty in the processing of legal texts, which is how to accurately extract the hierarchical information and logical structure in the text. Unlike general texts, legal documents usually have a multi-level nested structure, such as laws, regulations, implementation details, case analysis, etc. There are cross-references and logical associations between different levels, which makes it difficult for existing information extraction methods to accurately divide the relationship between clauses, affecting the effectiveness of subsequent intelligent tasks. The technical problem, this solution creatively adopts semantic segmentation reconstruction, text logic perception and The joint frame framework of dynamic knowledge fusion realizes more accurate identification of the structural hierarchy of legal texts, enhances the understanding of the logical relationship between legal clauses, and dynamically integrates the knowledge between different legal documents to build a knowledge network across legal texts, which is convenient for cross-reference and intelligent application between different regulations. In the existing semantic segmentation and reconstruction methods, there is a technical problem that the main goal of semantic segmentation is to disassemble complex legal documents according to their inherent logical structure, but traditional methods usually rely on fixed text formats (such as titles, numbers) when processing, and have poor performance on non-standardized texts. This solution creatively adopts a conditional random field bidirectional long short-term memory model improved by combining hierarchical label reasoning to perform Semantic segmentation and reconstruction, through the combination of conditional random fields and bidirectional long short-term memory networks, improves the recognition accuracy of hierarchical boundaries, avoids misjudgment across chapters and clauses, and realizes intelligent inference of the hierarchical relationship of clauses. Even in the face of legal texts with irregular formats, the hierarchical structure of different clauses can be accurately identified. In view of the technical problem that traditional text analysis methods mainly rely on keyword matching or reasoning models based on static rules in existing text logic perception methods, but these methods are difficult to accurately depict complex legal logical relationships, this solution creatively uses a causal graph convolutional reasoning neural network enhanced by dual-channel logic units for text logic perception, realizing a method based on semantic understanding and The dual-channel modeling of logical reasoning not only considers the literal relationship on the surface of the text, but also combines contextual information to infer the implicit logic between legal clauses, which helps to automatically determine which clauses are interrelated and provide more reliable logical reasoning support for the application of regulations. In view of the technical problem that traditional knowledge extraction methods in existing dynamic knowledge fusion methods mainly rely on static rules or predefined templates and are difficult to adapt to the changing role relationships in legal texts, this solution creatively adopts a role-aware enhanced triple extraction method for dynamic knowledge fusion, enhances the knowledge fusion capability between laws and regulations, automatically integrates relevant content in multiple legal texts, and improves the accuracy and completeness of legal information retrieval.

[0005] The technical solution adopted by the present invention is as follows: The present invention provides an artificial intelligence-based legal text information extraction and enhancement method, which includes the following steps:

[0006] Step S1: data collection and processing;

[0007] Step S2: semantic segmentation reconstruction;

[0008] Step S3: text logic perception;

[0009] Step S4: dynamic knowledge fusion;

[0010] Step S5: Information extraction enhancement.

[0011] Furthermore, in step S1, the data collection and processing is used to construct a multi-source heterogeneous legal text corpus and perform adaptive preprocessing, specifically, obtaining an original data set of legal text information extraction through multimodal data collection, and performing data optimization processing on the original data set of legal text information extraction to obtain an optimized legal text data set;

[0012] The original data set of legal text information extraction includes legislative texts, case texts, contract template texts, legal service material texts, regulatory texts, and annotation material texts;

[0013] The legal text information extracts the data structure of the original data set, including structured text, unstructured text, layout information and scanned image reference data;

[0014] The steps of data optimization processing include:

[0015] Step S11: adversarial data cleaning, specifically introducing a noise detection model based on a generative adversarial network, combining the structured text and the scanned image reference data, identifying text typesetting errors and non-standard typesetting interference, and obtaining cleaned and optimized text data;

[0016] Step S12: enhancing the layout structure, specifically by introducing a structural information analysis method based on layout rule recognition, identifying the title, clause number, paragraph number, and icon information according to the cleansed and optimized text data, and generating layout structure enhanced label annotation data;

[0017] Step S13: enhancing word segmentation, specifically enhancing the label annotation data according to the layout structure, using a domain-optimized word segmentation tool to enhance word segmentation, and obtaining word segmentation reference data;

[0018] The field-optimized word segmentation tool specifically refers to a word segmentation tool that uses a combination of standard SentencePiece extensions and Spacy extensions to optimize custom rules for legal formats;

[0019] Step S14: data optimization processing, specifically, performing data optimization processing through the adversarial data cleaning, the layout structure enhancement, and the enhanced word segmentation to obtain an optimized legal text dataset;

[0020] The optimized legal text dataset includes text information data, structural feature data and logical label data.

[0021] Furthermore, in step S2, the semantic segmentation reconstruction is used for hierarchical legal document content recognition. Specifically, based on the optimized legal text dataset, an improved conditional random field bidirectional long short-term memory model combined with hierarchical label reasoning is used to perform semantic segmentation reconstruction to obtain hierarchical legal text content recognition data, including the following steps:

[0022] Step S21: Joint feature extraction, specifically, performing multi-channel feature extraction through text vectorization and dependency syntax analysis to obtain syntactic, semantic, and structural information of the legal text and obtain sentence-level legal text feature data;

[0023] Step S22: gating fusion feature embedding, specifically constructing a gating mechanism to integrate semantic information, syntactic information, and structural information based on the sentence-level legal text feature data to obtain fused feature data;

[0024] Step S23: Feature segmentation, specifically using a multi-scale window segmentation strategy to extract sentence-level, clause-level, and chapter-level features, and constructing a bidirectional long short-term memory model to predict the legal clause sentence-level segmentation boundaries to obtain initial segmentation data;

[0025] Step S24: hierarchical label inference, specifically, constructing an improved conditional random field model and performing hierarchical label inference based on the initial segmentation data to obtain sentence-level label data;

[0026] The improved conditional random field model is specifically calculated through inference layer conditional probability calculation and hierarchical consistency correction optimization;

[0027] Step S25: generating a structure tree, specifically, constructing a logical hierarchy tree using a recursive nesting method based on the sentence level label data and the initial segmentation data, and correcting the logical hierarchy tree through structural consistency verification to obtain tree-like hierarchical structure data;

[0028] Step S26: semantic segmentation reconstruction, specifically standardizing the format of the initial segmentation data according to the tree-like hierarchical structure data to obtain standard format segmentation data, and automatically extracting the first legal clause of the standard format segmentation data as a hierarchical summary to construct hierarchical legal text content recognition data.

[0029] Furthermore, in step S3, the text logic perception is used to extract the logical relationship of the legal document. Specifically, based on the hierarchical legal text content recognition data, a causal graph convolutional inference neural network enhanced with a dual-channel logic unit is used to perform text logic perception to obtain a legal text logical relationship dataset, including the following steps:

[0030] Step S31: logical unit annotation, specifically, defining the legal logic form and performing logical unit annotation based on the hierarchical legal text content identification data to obtain logically annotated text content data, and constructing a dual-channel pointer network based on the logically annotated text content data to perform logical boundary identification to obtain legal text content logical identification data;

[0031] The dual-channel pointer network includes a text channel and a logic label channel;

[0032] Step S32: Causal graph convolutional reasoning, specifically constructing a multi-source dynamic causal graph through node definition and variable generation; the node definition specifically uses the logical units in the legal text content logic identification data as nodes, constructs multi-source edges, constructs a temporal perception graph attention network, and obtains causal graph convolutional reasoning data through time dependency modeling;

[0033] The multi-source edges include explicit logical edges, implicit logical edges and entity-aligned edges;

[0034] The explicit logical edge refers to the explicit logical relationship where the keywords "shall", "prohibited" and "unless" appear clearly; the implicit logical edge refers to the implicit logical relationship where cross-clause references exist; the entity alignment edge refers to the entity alignment relationship where multiple clauses involve the same legal object;

[0035] Step S33: Logic verification, specifically, generating interference samples based on the causal graph convolutional reasoning data by randomly deleting conditions and randomly replacing subjects, and conducting logic verification of the interference samples by constructing adversarial logic evaluation to obtain logic consistency evaluation data; the logic consistency evaluation data is used to reflect the logic perception performance;

[0036] Step S34: text logic perception, specifically, through the logic unit labeling, the causal graph convolutional reasoning and the logic verification, the text logic perception model is trained and the model is used to perform text logic perception to obtain the legal text logical relationship dataset.

[0037] Furthermore, in step S4, the dynamic knowledge fusion is used to construct the obligation graph. Specifically, based on the legal text logical relationship dataset and the hierarchical legal text content recognition data, a role-aware enhanced triple extraction method is used to perform dynamic knowledge fusion to obtain legal text structured information data, including the following steps:

[0038] Step S41: Constructing a legal knowledge base, specifically, constructing a legal knowledge memory base based on the legal text logical relationship dataset, and embedding legal provisions features through a pre-trained model to obtain an initial knowledge base of legal text information;

[0039] Step S42: differential knowledge update, specifically constructing a differential knowledge update algorithm. When a legal text is updated, only the legal features corresponding to the affected keywords are fine-tuned and trained to optimize the maintainability of the legal knowledge base.

[0040] Step S43: Constructing a triple extraction model for role constraints, specifically by adopting a generative sequence-to-relationship model and introducing a role-aware mechanism to embed legal roles and classify legal obligations, thereby obtaining role logic reference data. Furthermore, based on the role logic reference data, role and legal obligation extraction is performed to obtain role-obligation triple data.

[0041] The role-obligation triplet data includes a subject, a behavior, and an object;

[0042] The subject and object are used to express roles; the behavior is used to express legal obligations;

[0043] Step S44: Adversarial negative sampling enhancement, specifically, based on the role-obligation triple data, generating negative samples to obtain comparison sample data, and introducing adversarial margin distance as a loss function to perform adversarial negative sampling enhancement, optimize the accuracy of the legal obligations in the role-obligation triple data, and obtain obligation-optimized triple data; Step S45: Triple calibration, specifically, based on the obligation-optimized triple data, performing consistency detection through obligation type query analysis, and obtaining role-optimized triple data;

[0044] Step S46: knowledge fusion, specifically, performing knowledge fusion through the construction of the legal knowledge base, the differential knowledge update, the construction of the role-constrained triple extraction model, the adversarial negative sampling enhancement and the triple calibration to obtain legal text structured information data.

[0045] Furthermore, in step S5, the information extraction enhancement is used for comprehensive enhancement of legal text information extraction, specifically combining the hierarchical legal text content identification data, the legal text logical relationship data set and the legal text structured information data to perform information extraction enhancement and obtain legal text comprehensive information extraction reference data.

[0046] The present invention provides an artificial intelligence-based legal text information extraction and enhancement system, which includes an input processing module, an enhanced extraction module, and a reasoning output module;

[0047] The input processing module is used for data collection and processing, obtains an optimized legal text dataset through data collection and processing, and sends the optimized legal text dataset to the enhanced extraction module;

[0048] The enhanced extraction module is used for semantic segmentation and reconstruction, text logic perception and dynamic knowledge fusion, and obtains hierarchical legal text content recognition data, legal text logical relationship data set and legal text structured information data through semantic segmentation and reconstruction, text logic perception and dynamic knowledge fusion, and sends the hierarchical legal text content recognition data, legal text logical relationship data set and legal text structured information data to the reasoning output module;

[0049] The reasoning output module is used for information extraction enhancement, and through information extraction enhancement, comprehensive information extraction reference data of legal texts is obtained.

[0050] The beneficial effects achieved by the present invention using the above scheme are as follows:

[0051] (1) In the existing process of extracting information from legal texts, a core difficulty lies in how to accurately extract the hierarchical information and logical structure in the text. Unlike general texts, legal documents usually have a multi-level nested structure, such as laws, regulations, implementation rules, case analysis, etc. There are cross-references and logical associations between different levels, which makes it difficult for existing information extraction methods to accurately divide the relationship between clauses, affecting the effectiveness of subsequent intelligent tasks. This solution creatively adopts a joint framework of semantic segmentation reconstruction, text logic perception and dynamic knowledge fusion to achieve more accurate identification of the structural hierarchy of legal texts, enhance the understanding of the logical association of legal clauses, and dynamically integrate the knowledge between different legal documents to build a knowledge network across legal texts, which facilitates cross-references and intelligent application between different regulations;

[0052] (2) In the existing semantic segmentation and reconstruction methods, there is a technical problem that the main goal of semantic segmentation is to disassemble complex legal documents according to their internal logical structure, but traditional methods usually rely on fixed text formats (such as titles and numbers) when processing, and have poor performance on non-standardized texts. This solution creatively adopts a conditional random field bidirectional long short-term memory model improved by hierarchical label reasoning to perform semantic segmentation and reconstruction. By combining the conditional random field and the bidirectional long short-term memory network, the recognition accuracy of the hierarchical boundary is improved, and the misjudgment across chapters and clauses is avoided. The hierarchical relationship of the clauses is intelligently inferred. Even when faced with legal texts with irregular formats, the hierarchical structure of different clauses can be accurately identified.

[0053] (3) In view of the technical problem that traditional text analysis methods mainly rely on keyword matching or static rule-based reasoning models in existing text logic perception methods, but these methods are difficult to accurately depict complex legal logical relationships, this solution creatively adopts a causal graph convolutional reasoning neural network enhanced by dual-channel logic units to perform text logic perception, and realizes dual-channel modeling based on semantic understanding and logical reasoning. It not only considers the literal relationship on the surface of the text, but also infers the implicit logic between legal clauses based on contextual information, which helps to automatically determine which clauses are interrelated and provide more reliable logical reasoning support for the application of regulations;

[0054] (4) In view of the technical problem that traditional knowledge extraction methods in existing dynamic knowledge fusion methods mainly rely on static rules or predefined templates and are difficult to adapt to the changing role relationships in legal texts, this solution creatively adopts a triple extraction method enhanced by role perception to perform dynamic knowledge fusion, thereby enhancing the knowledge fusion capability between laws and regulations, automatically integrating relevant content in multiple legal texts, and improving the accuracy and completeness of legal information retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 A flowchart of an artificial intelligence-based legal text information extraction and enhancement method provided by the present invention;

[0056] Figure 2 A schematic diagram of an artificial intelligence-based legal text information extraction and enhancement system provided by the present invention;

[0057] Figure 3 This is a flow chart of data collection and processing in step S1;

[0058] Figure 4 Schematic diagram of the process of semantic segmentation reconstruction in step S2;

[0059] Figure 5 This is a flowchart of text logic perception in step S3;

[0060] Figure 6 Schematic diagram of the process of dynamic knowledge fusion in step S4.

[0061] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention. DETAILED DESCRIPTION

[0062] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0063] Example 1, see Figure 1 The present invention provides an artificial intelligence-based legal text information extraction and enhancement method, which includes the following steps:

[0064] Step S1: data collection and processing;

[0065] Step S2: semantic segmentation reconstruction;

[0066] Step S3: text logic perception;

[0067] Step S4: dynamic knowledge fusion;

[0068] Step S5: Information extraction enhancement.

[0069] By performing the above operations, we can solve the technical problem that in the existing legal text information extraction process, a core difficulty in the processing of legal texts is how to accurately extract the hierarchical information and logical structure in the text. Unlike general texts, legal documents usually have a multi-level nested structure, such as laws, regulations, implementation rules, case analysis, etc. There are cross-references and logical associations between different levels, which makes it difficult for existing information extraction methods to accurately divide the relationship between clauses, affecting the effectiveness of subsequent intelligent tasks. This solution creatively adopts a joint framework of semantic segmentation reconstruction, text logic perception and dynamic knowledge fusion to achieve more accurate identification of the structural hierarchy of legal texts, enhance the understanding of the logical association of legal clauses, and dynamically integrate the knowledge between different legal documents to build a knowledge network across legal texts, which facilitates cross-reference and intelligent application between different regulations.

[0070] Example 2, see Figure 1 、 Figure 2 and Figure 3In step S1, the data collection and processing is used to construct a multi-source heterogeneous legal text corpus and perform adaptive preprocessing, specifically, obtaining a raw data set of legal text information extraction through multimodal data collection, and performing data optimization processing on the raw data set of legal text information extraction to obtain an optimized legal text data set;

[0071] The original data set of legal text information extraction includes legislative texts, case texts, contract template texts, legal service material texts, regulatory texts, and annotation material texts;

[0072] The legal text information extracts the data structure of the original data set, including structured text, unstructured text, layout information and scanned image reference data;

[0073] The steps of data optimization processing include:

[0074] Step S11: adversarial data cleaning, specifically introducing a noise detection model based on a generative adversarial network, combining the structured text and the scanned image reference data, identifying text typesetting errors and non-standard typesetting interference, and obtaining cleaned and optimized text data;

[0075] Step S12: enhancing the layout structure, specifically by introducing a structural information analysis method based on layout rule recognition, identifying the title, clause number, paragraph number, and icon information according to the cleansed and optimized text data, and generating layout structure enhanced label annotation data;

[0076] Step S13: enhancing word segmentation, specifically enhancing the label annotation data according to the layout structure, and using a domain-optimized word segmentation tool to enhance word segmentation to obtain word segmentation reference data;

[0077] The field-optimized word segmentation tool specifically refers to a word segmentation tool that uses a combination of standard SentencePiece extensions and Spacy extensions to optimize legal format custom rules;

[0078] Step S14: data optimization processing, specifically, performing data optimization processing through the adversarial data cleaning, the layout structure enhancement, and the enhanced word segmentation to obtain an optimized legal text dataset;

[0079] The optimized legal text dataset includes text information data, structural feature data and logical label data.

[0080] Example 3, see Figure 1 、 Figure 2 and Figure 4This embodiment is based on the above embodiment. In step S2, the semantic segmentation reconstruction is used for hierarchical legal document content recognition. Specifically, based on the optimized legal text dataset, an improved conditional random field bidirectional long short-term memory model combined with hierarchical label reasoning is used to perform semantic segmentation reconstruction to obtain hierarchical legal text content recognition data, including the following steps:

[0081] Step S21: Joint feature extraction, specifically, performing multi-channel feature extraction through text vectorization and dependency syntax analysis to obtain syntactic, semantic, and structural information of the legal text and obtain sentence-level legal text feature data;

[0082] The calculation formula for the joint feature extraction is:

[0083] ;

[0084] Where, F i is the sentence-level legal text feature data, Concat(·) is the joint feature extraction operation, and W i is the word vector in the sentence, where i is the sentence index and D i is the syntactic node feature of the sentence, S i It is a structural position encoding feature;

[0085] Step S22: Gated fusion feature embedding, specifically, building a gating mechanism, integrating semantic information, syntactic information, and structural information based on the sentence-level legal text feature data, and obtaining fused feature data. The calculation formula is:

[0086] ;

[0087] Where, is the fusion feature data, Is the activation function, specifically the Sigmoid function, W g is the gating weight, F i is the sentence-level legal text feature data, b g is the gate bias term, is the element-wise multiplication representation function;

[0088] Step S23: Feature segmentation, specifically, using a multi-scale window segmentation strategy to extract sentence-level, clause-level, and chapter-level features, and constructing a bidirectional long short-term memory model to predict the legal clause sentence-level segmentation boundaries to obtain initial segmentation data. The calculation formula is:

[0089] ;

[0090] Where H i is the initial segmentation data, BiLSTM(·) is the bidirectional long short-term memory model representation function, Wink (·) is the sliding window operation function, k is the sentence-level window index, is the fused feature data from the ikth sentence-level window to the i+kth sentence-level window;

[0091] Step S24: hierarchical label inference, specifically, constructing an improved conditional random field model and performing hierarchical label inference based on the initial segmentation data to obtain sentence-level label data;

[0092] The improved conditional random field model is specifically calculated through inference layer conditional probability calculation and hierarchical consistency correction optimization, and the calculation formula is:

[0093] ;

[0094] Where P(Y|H) is the output probability of the improved conditional random field model, which is used to express the probability of outputting the sentence-level label Y under the condition of a given feature sequence H. Z(·) is the partition function, which is used as a normalization factor. H is the feature sequence input, which is used to represent the initial segmentation data. exp(·) is the natural base function. The overall score function is a label feature compatibility score function, which is used to represent the current sentence level label y i and the initial segmentation feature H i The degree of compatibility, The overall level perception transfer score function is used to represent the previous sentence level label y i-1 Transfer to the current sentence level label y i The rationality of L is the transfer structure information parameter between hierarchical labels;

[0095] Step S25: Structure tree generation, specifically, by constructing a logical hierarchy tree using a recursive nesting method based on the sentence level label data and the initial segmentation data, and correcting the logical hierarchy tree through structural consistency verification to obtain tree-like hierarchical structure data. The calculation formula is:

[0096] ;

[0097] Where T is the tree-level structure data, RecGen(·) is the recursive tree building function, Y is the sentence-level label data, n is the total number of labels in the sentence-level label data, y1 is the first sentence-level label, y n is the nth sentence level label;

[0098] Step S26: semantic segmentation reconstruction, specifically, standardizing the format of the initial segmented data according to the tree-like hierarchical structure data to obtain standard format segmented data, and automatically extracting the first legal clause of the standard format segmented data as a hierarchical summary to construct hierarchical legal text content recognition data;

[0099] The hierarchical legal text content identification data includes standard format segmentation data and hierarchical summaries, and the calculation formula is:

[0100] ;

[0101] Where, is the standard format segmentation data, Format(·) is the standardization function, T is the tree-like hierarchical structure data, S is the original structure format segmentation data, which is used to represent the structural position encoding features in the sentence-level legal text feature data, Summary is the hierarchical summary, and Head(·) is the summary extraction function.

[0102] By performing the above operations, in order to address the technical problem that in existing semantic segmentation and reconstruction methods, the main goal of semantic segmentation is to disassemble complex legal documents according to their inherent logical structure, but traditional methods usually rely on fixed text formats (such as titles and numbers) when processing, and perform poorly on non-standardized texts, this solution creatively adopts a conditional random field bidirectional long short-term memory model improved by hierarchical label reasoning to perform semantic segmentation and reconstruction. Through the combination of conditional random fields and bidirectional long short-term memory networks, the recognition accuracy of hierarchical boundaries is improved, and misjudgment across chapters and clauses is avoided, and the hierarchical relationship of clauses is intelligently inferred. Even when faced with legal texts with irregular formats, the hierarchical structure of different clauses can be accurately identified.

[0103] Example 4, see Figure 1 、 Figure 2 and Figure 5 This embodiment is based on the above embodiment. In step S3, the text logic perception is used to extract the logical relationship of legal documents. Specifically, based on the hierarchical legal text content recognition data, a causal graph convolutional inference neural network enhanced with dual-channel logic units is used to perform text logic perception to obtain a legal text logical relationship dataset, including the following steps:

[0104] Step S31: logical unit annotation, specifically, defining the legal logic form and performing logical unit annotation based on the hierarchical legal text content identification data to obtain logically annotated text content data, and constructing a dual-channel pointer network based on the logically annotated text content data to perform logical boundary identification to obtain legal text content logical identification data;

[0105] The dual-channel pointer network includes a text channel and a logic label channel;

[0106] The calculation formula for the logic unit annotation is:

[0107] ;

[0108] Where U j is the logical unit embedding vector, j is the logical perception sentence vector index, is a multilayer perceptron function that defines the legal logical form, E j is the semantic vector corresponding to the jth sentence vector in the hierarchical legal text content recognition data, T j is the label type corresponding to the j-th sentence vector in the hierarchical legal text content identification data;

[0109] The calculation formula of the dual-channel pointer network is:

[0110] ;

[0111] Where, is the text channel output, is the logical label channel output, softmax(·) is the classifier function, q is the text pointer query vector, r is the logical pointer query vector, E i is the semantic vector of the i-th sentence, T i is the label type of the i-th sentence, b i is the logical boundary recognition output, argmax(·) is the maximum value function, is the dual-channel weight;

[0112] Preferably, the specific value of the dual-channel weight is 0.6;

[0113] Step S32: Causal graph convolutional reasoning, specifically constructing a multi-source dynamic causal graph through node definition and variable generation; the node definition specifically uses the logical units in the legal text content logic identification data as nodes, constructs multi-source edges, constructs a temporal perception graph attention network, and obtains causal graph convolutional reasoning data through time dependency modeling;

[0114] The multi-source edges include explicit logical edges, implicit logical edges and entity-aligned edges;

[0115] The explicit logical edge refers to the explicit logical relationship where the keywords "shall", "prohibited" and "unless" appear clearly; the implicit logical edge refers to the implicit logical relationship where cross-clause references exist; the entity alignment edge refers to the entity alignment relationship where multiple clauses involve the same legal object;

[0116] The calculation formula for constructing the multi-source dynamic causal graph is:

[0117] ;

[0118] Where G is a multi-source dynamic causal graph, V is a node parameter, E is an edge parameter, and v j is the node index, U j is the logical unit embedding vector, E exp is an explicit logical edge, E imp is an implicit logical edge, E ent is the entity alignment edge;

[0119] The calculation formula for constructing the temporal perception graph attention network is:

[0120] ;

[0121] Where, is the output hidden state of the l+1th layer of the temporal perception graph attention network, which is used to represent the causal graph convolutional reasoning data, ReLU(·) is the nonlinear activation function, t is the multi-source edge type index, is the adjacency matrix of the t-th edge, is the output hidden state of the l-th layer temporal perception graph attention network, is the graph convolution weight of the t-th edge,

[0122] Step S33: Logic verification, specifically, generating interference samples based on the causal graph convolutional reasoning data by randomly deleting conditions and randomly replacing subjects, and performing logic verification on the interference samples by constructing an adversarial logic evaluation loss to obtain logic consistency evaluation data; the logic consistency evaluation data is used to reflect the logic perception performance;

[0123] The calculation formula of the adversarial logic evaluation loss is:

[0124] ;

[0125] Where, is the adversarial logic evaluation loss, It is for the original legal logic pairing The expected calculation function of The whole is the original legal logic pair, The whole is used to express the probability of judging the original legal logical pair as true, where is the logical consistency discriminator identifier, y ab is the result of pairing determination. The overall probability of judging the generated adversarial sample pair as false is expressed as follows: is to generate adversarial sample pairs;

[0126] Step S34: text logic perception, specifically, through the logic unit labeling, the causal graph convolutional reasoning and the logic verification, the text logic perception model is trained and the model is used to perform text logic perception to obtain the legal text logical relationship dataset.

[0127] By performing the above operations, in order to address the technical problem that in existing text logic perception methods, traditional text analysis methods mainly rely on keyword matching or static rule-based reasoning models, but these methods are difficult to accurately portray complex legal logical relationships, this solution creatively adopts a causal graph convolutional reasoning neural network enhanced with dual-channel logic units to perform text logic perception, and realizes dual-channel modeling based on semantic understanding and logical reasoning. It not only considers the literal relationship on the surface of the text, but also infers the implicit logic between legal clauses based on contextual information, which helps to automatically determine which clauses are interrelated and provide more reliable logical reasoning support for the application of regulations.

[0128] Example 5, see Figure 1 、 Figure 2 and Figure 6 In step S4, the dynamic knowledge fusion is used to construct the obligation graph. Specifically, based on the legal text logical relationship dataset and the hierarchical legal text content recognition data, a role-aware enhanced triple extraction method is used to perform dynamic knowledge fusion to obtain legal text structured information data, including the following steps:

[0129] Step S41: Constructing a legal knowledge base, specifically, constructing a legal knowledge memory base based on the legal text logical relationship dataset, and embedding legal provisions features through a pre-trained model to obtain an initial knowledge base of legal text information. The calculation formula is:

[0130] ;

[0131] Where K i is the legal knowledge vector corresponding to the i-th legal clause, which is used to represent the result of legal clause feature embedding and construct the initial knowledge base of legal text information. i is the sentence index, is a pre-trained legal language model function, used to represent the pre-trained model. The pre-trained model specifically uses the LegalBERT model as the base model. i is the original data of the ith syntactic clause, PosEnc(·) position encoding function, Section i is the chapter position parameter;

[0132] Step S42: differential knowledge update, specifically, constructing a differential knowledge update algorithm. When a legal text is updated, only the legal features corresponding to the affected keywords are fine-tuned and trained to optimize the maintainability of the legal knowledge base. The calculation formula is:

[0133] ;

[0134] Where, is the updated legal knowledge vector, Update(·) is the incremental learning parameter, K i is the legal knowledge vector corresponding to the i-th legal clause, is the model parameter of the pre-trained model, Sim(·) is the cosine similarity metric function, L i is the original data of the ith sentence, It is the updated original text of the new law. is the change threshold;

[0135] Step S43: Constructing a triple extraction model for role constraints, specifically by adopting a generative sequence-to-relationship model and introducing a role-aware mechanism to embed legal roles and classify legal obligations, thereby obtaining role logic reference data. Furthermore, based on the role logic reference data, role and legal obligation extraction is performed to obtain role-obligation triple data.

[0136] The calculation formula for the legal role embedding is:

[0137] ;

[0138] Where R k is the legal role embedding data, RoleEnc(·) is the role encoding function, r k is the legal role name, k is the legal role index, Attn(·) is the attention matching function, C k is the contextual content information of the legal role, K is the legal knowledge vector in the legal knowledge base;

[0139] The calculation formula of the role-obligation triple is:

[0140] ;

[0141] Where (s, a, o) is the role-obligation triple, Gen(·) is the Seq2Rel generation model function, and R k is the legal role embedding data, H t is the context feature vector data;

[0142] The role-obligation triplet data includes a subject, a behavior, and an object;

[0143] The subject and object are used to express roles; the behavior is used to express legal obligations;

[0144] Step S44: Adversarial negative sampling enhancement, specifically, generating negative samples based on the role-obligation triple data to obtain comparison sample data, and introducing adversarial margin distance as a loss function to perform adversarial negative sampling enhancement to optimize the accuracy of the legal obligations in the role-obligation triple data to obtain obligation-optimized triple data;

[0145] The calculation formula for the adversarial negative sampling enhancement is:

[0146] ;

[0147] Where, is the adversarial negative sampling loss, is a set of triplet positive samples, max(·) is the maximum value function, It is the margin distance safety boundary parameter, is the semantic distance of the triples, is the semantic distance of the generated negative sample;

[0148] Step S45: Triple calibration, specifically, optimizing triple data based on the obligation, performing consistency check through obligation type query analysis, and obtaining role optimized triple data. The calculation formula is:

[0149] ;

[0150] Where, is the role optimization triple data, Query(a) is the obligation type abstracted from the behavior, and LegalClass(s,o) is the legal behavior set of the behavior role in the legal knowledge base;

[0151] Step S46: knowledge fusion, specifically, performing knowledge fusion through the construction of the legal knowledge base, the differential knowledge update, the construction of the role-constrained triple extraction model, the adversarial negative sampling enhancement and the triple calibration to obtain legal text structured information data.

[0152] By performing the above operations, in order to address the technical problem that traditional knowledge extraction methods in existing dynamic knowledge fusion methods mainly rely on static rules or predefined templates and are difficult to adapt to the changing role relationships in legal texts, this solution creatively adopts a role-aware enhanced triple extraction method to perform dynamic knowledge fusion, thereby enhancing the knowledge fusion capabilities between laws and regulations, automatically integrating relevant content in multiple legal texts, and improving the accuracy and completeness of legal information retrieval.

[0153] Example 6, see Figure 1 and Figure 2 This embodiment is based on the above embodiment. In step S5, the information extraction enhancement is used for comprehensive enhancement of legal text information extraction, specifically combining the hierarchical legal text content identification data, the legal text logical relationship data set and the legal text structured information data to perform information extraction enhancement and obtain legal text comprehensive information extraction reference data.

[0154] Example 7, see Figure 1 and Figure 2 , based on the above embodiment, this embodiment provides an artificial intelligence-based legal text information extraction and enhancement system, including an input processing module, an enhanced extraction module and a reasoning output module;

[0155] The input processing module is used for data collection and processing, obtains an optimized legal text dataset through data collection and processing, and sends the optimized legal text dataset to the enhanced extraction module;

[0156] The enhanced extraction module is used for semantic segmentation and reconstruction, text logic perception and dynamic knowledge fusion, and obtains hierarchical legal text content recognition data, legal text logical relationship data set and legal text structured information data through semantic segmentation and reconstruction, text logic perception and dynamic knowledge fusion, and sends the hierarchical legal text content recognition data, legal text logical relationship data set and legal text structured information data to the reasoning output module;

[0157] The reasoning output module is used for information extraction enhancement, and through information extraction enhancement, comprehensive information extraction reference data of legal texts is obtained.

[0158] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0159] While the embodiments of the present invention have been shown and described, it will be apparent to those skilled in the art that various changes, modifications, substitutions, and alterations can be made to the embodiments without departing from the principles and spirit of the invention.

[0160] The present invention and its embodiments are described above. This description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this and, without departing from the purpose of the present invention, designs structures and embodiments similar to this technical solution without inventiveness, they shall fall within the scope of protection of the present invention.

Claims

1. An artificial intelligence-based legal text information extraction and enhancement method, characterized by: The method comprises the following steps: Step S1: Data collection and processing to obtain an optimized legal text dataset; Step S2: semantic segmentation and reconstruction, using a conditional random field bidirectional long short-term memory model improved by combining hierarchical label reasoning to perform semantic segmentation and reconstruction, and obtain hierarchical legal text content recognition data, specifically including: multi-channel feature extraction to obtain syntactic, semantic and structural information of the legal text; building a gating mechanism to integrate semantic information, syntactic information and structural information; using a multi-scale window segmentation strategy to extract sentence-level, clause-level and chapter-level features, and by building a bidirectional long short-term memory model, predicting the legal clause sentence-level segmentation boundary; building an improved conditional random field model through inference layer conditional probability calculation and hierarchical consistency correction optimization to perform hierarchical label reasoning; using a recursive nesting method to construct a logical hierarchy tree, and correcting the logical hierarchy tree through structural consistency verification; and automatically extracting a hierarchical summary; Step S3: Text logic perception: Using a causal graph convolutional reasoning neural network enhanced with dual-channel logic units, text logic perception is performed to obtain a legal text logical relationship dataset. Specifically, the process includes: labeling logical units and constructing a dual-channel pointer network consisting of a text channel and a logic label channel to identify logical boundaries; constructing multi-source edges including explicit logical edges, implicit logical edges, and entity alignment edges, constructing a temporal perception graph attention network, and performing causal graph convolutional reasoning; verifying logical consistency by generating interference samples; training a text logic perception model and performing text logic perception. Step S4: Dynamic knowledge fusion, using a triple extraction method enhanced by role perception, to perform dynamic knowledge fusion and obtain structured information data of legal texts, including the following steps: constructing a legal knowledge base; constructing a differential knowledge update algorithm to optimize the maintainability of the legal knowledge base; adopting a generative sequence-to-relationship model and introducing a role perception mechanism to embed legal roles and classify legal obligations and extract roles and legal obligations; introducing adversarial margin distance as a loss function to perform adversarial negative sampling enhancement and optimize the accuracy of legal obligations; performing consistency detection through obligation type query analysis; knowledge fusion; Step S5: Enhance information extraction to obtain reference data for comprehensive information extraction of legal texts.

2. The artificial intelligence-based legal text information extraction and enhancement method according to claim 1, characterized in that: In step S1, the data collection and processing is used to construct a multi-source heterogeneous legal text corpus and perform adaptive preprocessing, specifically, obtaining an original dataset of legal text information extraction through multimodal data collection, and performing data optimization processing on the original dataset of legal text information extraction to obtain an optimized legal text dataset; The original data set of legal text information extraction includes legislative texts, case texts, contract template texts, legal service material texts, regulatory texts, and annotation material texts; The steps of data optimization processing include: Step S11: adversarial data cleaning, specifically by introducing a noise detection model based on a generative adversarial network, combining structured text and scanned image reference data, to identify text typographical errors and non-standard typographical interference, and obtain cleaned and optimized text data; Step S12: enhancing the layout structure, specifically by introducing a structural information analysis method based on layout rule recognition, identifying the title, clause number, paragraph number, and icon information according to the cleansed and optimized text data, and generating layout structure enhanced label annotation data; Step S13: enhancing word segmentation, specifically enhancing the label annotation data according to the layout structure, and using a domain-optimized word segmentation tool to enhance word segmentation to obtain word segmentation reference data; The field-optimized word segmentation tool specifically refers to a word segmentation tool that uses a combination of standard SentencePiece extensions and Spacy extensions to optimize legal format custom rules; Step S14: data optimization processing, specifically, performing data optimization processing through the adversarial data cleaning, the layout structure enhancement, and the enhanced word segmentation to obtain an optimized legal text dataset; The optimized legal text dataset includes text information data, structural feature data and logical label data.

3. The artificial intelligence-based legal text information extraction and enhancement method according to claim 2, characterized in that: In step S2, the semantic segmentation reconstruction is used for hierarchical legal document content recognition. Specifically, based on the optimized legal text dataset, an improved conditional random field bidirectional long short-term memory model combined with hierarchical label reasoning is used to perform semantic segmentation reconstruction to obtain hierarchical legal text content recognition data, including the following steps: Step S21: Joint feature extraction, specifically, performing multi-channel feature extraction through text vectorization and dependency syntax analysis to obtain syntactic, semantic, and structural information of the legal text and obtain sentence-level legal text feature data; Step S22: gating fusion feature embedding, specifically constructing a gating mechanism to integrate semantic information, syntactic information, and structural information based on the sentence-level legal text feature data to obtain fused feature data; Step S23: Feature segmentation, specifically using a multi-scale window segmentation strategy to extract sentence-level, clause-level, and chapter-level features, and constructing a bidirectional long short-term memory model to predict the legal clause sentence-level segmentation boundaries to obtain initial segmentation data; Step S24: hierarchical label inference, specifically, constructing an improved conditional random field model and performing hierarchical label inference based on the initial segmentation data to obtain sentence-level label data; The improved conditional random field model is specifically calculated through inference layer conditional probability calculation and hierarchical consistency correction optimization, and the calculation formula is: ; Where P(Y|H) is the output probability of the improved conditional random field model, which is used to express the probability of outputting the sentence-level label Y under the condition of a given feature sequence H. Z(·) is the partition function, which is used as a normalization factor. H is the feature sequence input, which is used to represent the initial segmentation data. exp(·) is the natural base function. The overall score function is a label feature compatibility score function, which is used to represent the current sentence level label y i and the initial segmentation feature H i The degree of compatibility, The overall level perception transfer score function is used to represent the previous sentence level label y i-1 Transfer to the current sentence level label y i The rationality of L is the transfer structure information parameter between hierarchical labels; Step S25: generating a structure tree, specifically, constructing a logical hierarchy tree using a recursive nesting method based on the sentence level label data and the initial segmentation data, and correcting the logical hierarchy tree through structural consistency verification to obtain tree-like hierarchical structure data; Step S26: semantic segmentation reconstruction, specifically standardizing the format of the initial segmentation data according to the tree-like hierarchical structure data to obtain standard format segmentation data, and automatically extracting the first legal clause of the standard format segmentation data as a hierarchical summary to construct hierarchical legal text content recognition data.

4. The artificial intelligence-based legal text information extraction and enhancement method according to claim 3 is characterized by: In step S3, the text logic perception is used to extract the logical relationship of legal documents. Specifically, based on the hierarchical legal text content recognition data, a causal graph convolutional inference neural network enhanced with dual-channel logic units is used to perform text logic perception to obtain a legal text logical relationship dataset, including the following steps: Step S31: logical unit annotation, specifically, defining the legal logic form and performing logical unit annotation based on the hierarchical legal text content identification data to obtain logically annotated text content data, and constructing a dual-channel pointer network based on the logically annotated text content data to perform logical boundary identification to obtain legal text content logical identification data; The dual-channel pointer network includes a text channel and a logical label channel, and the calculation formula is: ; Where, is the text channel output, is the logical label channel output, softmax(·) is the classifier function, q is the text pointer query vector, r is the logical pointer query vector, E i is the semantic vector of the i-th sentence, T i is the label type of the i-th sentence, b i is the logical boundary recognition output, argmax(·) is the maximum value function, is the dual-channel weight; Step S32: Causal graph convolutional reasoning, specifically constructing a multi-source dynamic causal graph through node definition and variable generation; the node definition specifically uses the logical units in the legal text content logic identification data as nodes, constructs multi-source edges, constructs a temporal perception graph attention network, and obtains causal graph convolutional reasoning data through time dependency modeling; The multi-source edges include explicit logical edges, implicit logical edges and entity-aligned edges; The calculation formula for constructing the multi-source dynamic causal graph is: ; Where G is a multi-source dynamic causal graph, V is a node parameter, E is an edge parameter, and v j is the node index, U j is the logical unit embedding vector, E exp is an explicit logical edge, E imp is an implicit logical edge, E ent is the entity alignment edge; Step S33: Logic verification, specifically, generating interference samples based on the causal graph convolutional reasoning data by randomly deleting conditions and randomly replacing subjects, and constructing an adversarial logic evaluator to perform logic verification on the interference samples to obtain logic consistency evaluation data; the logic consistency evaluation data is used to reflect the logic perception performance; Step S34: text logic perception, specifically, through the logic unit labeling, the causal graph convolutional reasoning and the logic verification, the text logic perception model is trained and the model is used to perform text logic perception to obtain the legal text logical relationship dataset.

5. The artificial intelligence-based legal text information extraction and enhancement method according to claim 4 is characterized by: In step S4, the dynamic knowledge fusion is used to construct the obligation graph. Specifically, the dynamic knowledge fusion is performed based on the legal text logical relationship dataset and the hierarchical legal text content recognition data, using the role-aware enhanced triple extraction method to obtain legal text structured information data, including the following steps: Step S41: Constructing a legal knowledge base, specifically, constructing a legal knowledge memory base based on the legal text logical relationship dataset, and embedding legal provisions features through a pre-trained model to obtain an initial knowledge base of legal text information; Step S42: differential knowledge update, specifically constructing a differential knowledge update algorithm. When a legal text is updated, only the legal features corresponding to the affected keywords are fine-tuned and trained to optimize the maintainability of the legal knowledge base. Step S43: Constructing a triple extraction model for role constraints, specifically by adopting a generative sequence-to-relationship model and introducing a role-aware mechanism to embed legal roles and classify legal obligations, thereby obtaining role logic reference data. Furthermore, based on the role logic reference data, role and legal obligation extraction is performed to obtain role-obligation triple data. The role-obligation triplet data includes a subject, a behavior, and an object; The subject and object are used to express roles; the behavior is used to express legal obligations; Step S44: Adversarial negative sampling enhancement, specifically, generating negative samples based on the role-obligation triple data to obtain comparison sample data, and introducing adversarial margin distance as a loss function to perform adversarial negative sampling enhancement to optimize the accuracy of the legal obligations in the role-obligation triple data to obtain obligation-optimized triple data; Step S45: triple calibration, specifically optimizing triple data according to the obligation, performing consistency detection through obligation type query analysis, and obtaining role optimized triple data; Step S46: knowledge fusion, specifically, performing knowledge fusion through the construction of the legal knowledge base, the differential knowledge update, the construction of the role-constrained triple extraction model, the adversarial negative sampling enhancement and the triple calibration to obtain legal text structured information data.

6. The artificial intelligence-based legal text information extraction and enhancement method according to claim 5, characterized in that: In step S5, the information extraction enhancement is used for comprehensive enhancement of legal text information extraction, specifically combining the hierarchical legal text content identification data, the legal text logical relationship data set and the legal text structured information data to perform information extraction enhancement and obtain legal text comprehensive information extraction reference data.

7. An artificial intelligence-based legal text information extraction and enhancement system, for implementing the artificial intelligence-based legal text information extraction and enhancement method according to any one of claims 1 to 6, characterized in that: It includes input processing module, enhanced extraction module and reasoning output module.

8. The artificial intelligence-based legal text information extraction and enhancement system according to claim 7, characterized in that: The input processing module is used for data collection and processing, obtains an optimized legal text dataset through data collection and processing, and sends the optimized legal text dataset to the enhanced extraction module; The enhanced extraction module is used for semantic segmentation and reconstruction, text logic perception and dynamic knowledge fusion, and obtains hierarchical legal text content recognition data, legal text logical relationship data set and legal text structured information data through semantic segmentation and reconstruction, text logic perception and dynamic knowledge fusion, and sends the hierarchical legal text content recognition data, legal text logical relationship data set and legal text structured information data to the reasoning output module; The reasoning output module is used for information extraction enhancement, and through information extraction enhancement, comprehensive information extraction reference data of legal texts is obtained.

Citation Information

Patent Citations

  • Network security information mapping system and method based on hierarchical perception

    CN118300809A

  • Recommendation method and device based on knowledge graph, equipment and storage medium

    CN119149697A

  • Case evidence logic deduction method based on graph neural network

    CN119721117A

Cited By

  • A hierarchical search and graph reasoning-based vertical procedure intelligent analysis method and system

    CN122528897A