A legal document gist automatic extraction and association method and system

By using document fingerprinting algorithms and multi-hop ripple propagation calculations, changes in legal document data are automatically identified and the case knowledge graph is updated. This solves the problem of insufficient identification of dynamic changes in existing systems, improves the real-time performance and accuracy of case management, and reduces the cost of manual backtracking and judicial risks.

CN121638262BActive Publication Date: 2026-04-24JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
Filing Date
2026-02-04
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing legal big data management and case-handling assistance systems cannot automatically identify the impact of changes in underlying evidence on the upper-level legal logic network when faced with dynamic changes in requirements. This results in high costs for manual backtracking and a tendency to produce systemic omissions. Furthermore, they cannot provide timely early warnings and suggestions, thus affecting the real-time performance and accuracy of case management.

Method used

By comparing newly added legal documents with historical documents using a document fingerprint algorithm, the difference in text data fragments is extracted, semantic parsing is performed to generate candidate triple data, which is then mapped to the case knowledge graph. An atomic update operation is performed, and the affected downstream related nodes are identified through multi-hop ripple propagation calculation, outputting the case-related influence domain view data.

Benefits of technology

It enables real-time identification and accurate analysis of changes in legal document data, reduces the cost of manual backtracking, improves the real-time nature and accuracy of case data changes, ensures the dynamic consistency of the case knowledge graph, and avoids judicial errors and resource waste caused by information lag.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121638262B_ABST
    Figure CN121638262B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of document management, and relates to a legal document gist automatic extraction and association method and system. Newly added legal document data of a target case is acquired, preset document fingerprint algorithms are used to automatically compare the newly added legal document data with historical benchmark documents, difference text data segments are extracted, deep semantic analysis is performed on the difference text data segments, candidate triple data containing entities and relations are generated, the candidate triple data is intelligently mapped to a pre-constructed case knowledge graph, graph anchor point nodes with changed data are quickly located, and then, based on the change type of the difference text data segments, atomic update operations are performed on the graph anchor point nodes, including dynamic marking update of node state attributes, and finally, change source node data with the latest state identification is generated. The method effectively solves the problems of dependence on manual work, low efficiency and errors in current legal case data updating.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of document management technology and relates to a method and system for automatically extracting and associating key points of legal documents. Background Technology

[0002] In the field of large-scale commercial legal services, complex cases involving long cycles and multiple volumes have become the norm. The handling of such cases often spans several years, involving not only a massive amount of documentation, but also a constantly evolving and continuously updated set of case documents as the trial progresses, new evidence is added, and legal opinions are revised. Meanwhile, the logical connections between legal facts are the core foundation for maintaining judicial fairness and consistency in judgments. By using knowledge graph technology to structure fragmented document data into a meticulous legal logic network, it provides irreplaceable support for judges and lawyers in determining case facts, assessing risks, and conducting decision analysis.

[0003] It is worth noting that the legal logic network, as a highly rigorous semantic association system, is extremely sensitive to changes in the validity of evidence, fact-finding, and procedural status. In complex cases, any subtle revision of documents (such as minor adjustments to witness testimony, reinterpretations of key contractual clauses, or changes to procedural rulings) often has an impact beyond the text itself, generating a profound "chain reaction" through the legal logic chain. A complex and subtle dynamic balance exists between this change and the consistency of the overall logic of the network: the invalidation or refutation of a core piece of evidence may lead to the collapse of the entire chain of evidence, thereby triggering a reconstruction of higher-level legal inferences and the final judgment.

[0004] However, current legal big data management and case-handling support systems still have significant technical limitations in responding to dynamic changes. Existing technologies lack quantitative assessment and transmission analysis mechanisms for the impact of changes. When the validity of underlying evidence changes, the system cannot automatically identify the downstream related nodes affected (such as affected factual findings or legal conclusions), forcing legal practitioners to expend considerable effort on manual review and rereading. This passive, reactive processing mode not only creates serious cognitive redundancy and a heavy review burden but also prevents the system from providing early warnings when logical conflicts arise, resulting in significant lag and the risk of omissions in the dynamic management of complex cases. Summary of the Invention

[0005] In view of the problems existing in the prior art, the present invention provides a method and system for automatically extracting and associating key points of legal documents to solve the above-mentioned technical problems.

[0006] To achieve the above and other objectives, the technical solution adopted by the present invention is as follows:

[0007] This invention provides a method for automatically extracting and associating key points of legal documents, the method comprising the following steps:

[0008] Acquire new legal document data for the target case, and use a preset document fingerprint algorithm to compare the new legal document data with historical benchmark documents in the case file database to extract the difference text data fragments;

[0009] Semantic parsing is performed on the differential text data fragments to generate candidate triple data containing entities and relations, and the candidate triple data is mapped to a pre-built case knowledge graph to locate the graph anchor node where data changes have occurred.

[0010] Based on the change type of the differential text data fragment, an atomic update operation is performed on the graph anchor node. The atomic update operation includes updating the marker of the node state attribute and generating the change source node data with the latest state identifier.

[0011] Based on the predefined node dependency topology in the case knowledge graph, multi-hop ripple propagation calculation is performed starting from the changed source node data to identify downstream related nodes affected by the changed source node data. Based on the propagation strength parameters stored in the node dependency topology, the predicted confidence values ​​of the downstream related nodes are calculated.

[0012] The output includes case-related impact domain view data containing data on the changed source node, the affected downstream related nodes, and their updated confidence values.

[0013] Another aspect of the present invention provides an automatic extraction and association system for key points of legal documents, the system comprising:

[0014] The document data acquisition and difference comparison module is used to acquire new legal document data for the target case, and uses a preset document fingerprint algorithm to compare the new data with the historical benchmark documents in the case file database to extract the difference text data fragments.

[0015] The semantic parsing and graph anchor point localization module receives differential text data fragments, generates candidate triple data containing entities and relationships through semantic parsing technology, maps the candidate triple data to a pre-built case knowledge graph, and locates the graph anchor point nodes where data changes have occurred.

[0016] The graph node atomic update operation module performs atomic update operations on graph anchor nodes based on the change type of the differential text data fragments, generating change source node data with the latest status identifier;

[0017] The node dependency topology and ripple calculation module, based on the predefined node dependency topology in the case knowledge graph, takes the changed source node data as the starting point, performs multi-hop ripple propagation calculation, identifies the downstream related nodes affected by the changed source node data, and calculates the predicted confidence value of the downstream related nodes according to the transmission strength parameters stored in the node dependency topology.

[0018] The graph evolution domain analysis output module outputs case association impact domain view data, which includes data on the changed source nodes, all affected downstream related nodes, and their updated confidence values.

[0019] As described above, the method and system for automatically extracting and associating key points of legal documents provided by the present invention have at least the following beneficial effects:

[0020] This invention provides a method and system for automatically extracting and associating key points of legal documents. By acquiring new legal document data of a target case, a preset document fingerprint algorithm is used to automatically compare the new legal document data with historical benchmark documents in the case file database, accurately extracting differing text data fragments. Then, deep semantic analysis is performed on the differing text data fragments to generate candidate triple data containing entities and relationships. The candidate triple data is intelligently mapped to a pre-constructed case knowledge graph to quickly locate the graph anchor node where data changes have occurred. Subsequently, based on the change type of the differing text data fragment, atomic update operations are performed on the graph anchor node, including dynamic marking and updating of node status attributes, and finally generating change source node data with the latest status identifier. This effectively solves the problems of reliance on manual labor, low efficiency, and susceptibility to errors in current legal case data updates. On the one hand, from the perspective of automated and intelligent multi-level analysis, it significantly improves the real-time performance and accuracy of case data change identification, greatly reduces the risk of judicial decision-making errors caused by information lag or comparison deviations, further avoids legal application errors and judicial injustices caused by data asynchrony, and can also effectively eliminate the long-term negative impact on case management efficiency and judicial fairness. On the other hand, by ensuring the dynamic and consistent maintenance of the case knowledge graph, it directly avoids analytical distortion and resource waste caused by outdated or delayed data updates, significantly reduces the labor costs and time consumption caused by manual operation, and ensures the accuracy, scientificity, and reliability of case data analysis, thus providing solid technical support for the intelligent transformation of the judicial system.

[0021] This invention is based on a predefined node dependency topology in a case knowledge graph. Starting with the changed source node data, it performs multi-hop ripple propagation calculations, intelligently identifies downstream related nodes affected by the changed source node data, and automatically calculates the predicted confidence values ​​of downstream related nodes based on the transmission strength parameters stored in the node dependency topology. Finally, it outputs a case association influence domain view data containing the changed source node data, the affected downstream related nodes, and their updated confidence values. This further enhances the ability to deeply analyze and predict the impact of case data changes. On the one hand, through multi-hop propagation and systematic reasoning of dependencies, it comprehensively captures the chain reaction and ripple effect of data changes, significantly improving the completeness and reliability of impact assessment. It effectively reduces systemic judicial risks caused by neglecting global correlations due to local updates, further avoiding problems such as unreasonable resource allocation and low case processing efficiency. At the same time, it can also effectively eliminate the negative impact on the accuracy of case correlation analysis and long-term prediction models. On the other hand, it ensures the intelligent evolution and adaptive optimization of the case knowledge graph, directly avoiding insufficient decision support and delayed emergency response caused by static or isolated analysis. It significantly reduces the possibility of case management chaos and judicial errors caused by omission of dependencies or misjudgment of transmission strength, ensuring the forward-looking, systematic, and scientific nature of case impact assessment. Thus, it provides efficient and reliable data-driven support for judicial risk assessment and optimal resource allocation. Attached Figure Description

[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram showing the connections between the steps of the method of the present invention.

[0024] Figure 2 This is a schematic diagram illustrating the logic for determining the change type of differential text data fragments in this invention.

[0025] Figure 3 This is a schematic diagram illustrating the identification operation logic of the present invention for identifying downstream associated nodes affected by changes in source node data.

[0026] Figure 4 This is a schematic diagram showing the connections of the various modules in the system of the present invention. Detailed Implementation

[0027] The following description, in conjunction with the implementation of this invention, is merely an example and illustration of the concept of this invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the inventive concept or exceed the scope defined in these claims, all of which should fall within the protection scope of this invention.

[0028] In existing technologies, document management for complex cases (such as bankruptcy reorganization and intellectual property series cases) often relies on static storage or version overwriting, making it difficult to balance the dynamic evolution of data with logical consistency. Traditional methods, when documents undergo incremental updates, often only identify superficial textual differences, leading to the neglect of the interconnectedness of the legal factual chain, resulting in broken evidence chains or conflicts in factual findings. Existing systems cannot synchronously perceive the erosive effect of underlying document changes on the upper-level legal logic network. Especially during long-term trials when the validity of evidence changes, manual backtracking is extremely costly and prone to systemic omissions, failing to meet the precision requirements of judicial decision-making.

[0029] To address the aforementioned issues, this research reveals a strong coupling between changes in the effectiveness of legal facts and the dependency topology within the graph. An "document change-logic transmission" model is established to automatically calculate the influence domain. The study found that core evidentiary relationships provide strong support for fact-finding, while auxiliary relationships, though less sensitive, have broad coverage. Therefore, a method for dynamically setting transmission strength parameters based on legal logic types is proposed. Further experimental verification introduces atomic update states and multi-hop ripple propagation mechanisms into the dynamic graph maintenance process, forming a closed-loop update system for knowledge consistency.

[0030] Specifically, the detection system first acquires newly added legal document data for the target case, compares it with historical benchmark documents in the case file database using a pre-defined document fingerprint algorithm, and accurately extracts discrepancy text data fragments. Through in-depth semantic analysis of the discrepancy text, entities and relationships are extracted and candidate triples are generated. These are then mapped to a pre-constructed case knowledge graph to accurately locate the graph anchor nodes where data changes have occurred. Based on the change type presented by the discrepancy text (e.g., addition, modification, or negation), the system performs atomic update operations on the anchor nodes, generating change source node data with the latest validity identifier by marking the node's state attributes. Subsequently, with the change source node as the epicenter, the system performs multi-hop ripple propagation calculations based on a predefined node dependency topology, recursively calculating the predicted confidence values ​​of downstream related nodes using the propagation strength parameters stored in the edge attributes. During continuous propagation, the system identifies all affected nodes and outputs a case-related influence domain view data containing the change source, affected nodes, and their dynamic confidence levels. For paths with propagation strength below a threshold, the system automatically terminates the propagation to focus on the core influence area.

[0031] Compared to existing technologies, traditional methods rely on manual comparison and lack a logical transmission mechanism, making them prone to logical errors when documents are frequently updated. This solution innovatively integrates document fingerprinting technology with the ripple propagation mechanism of knowledge graphs, achieving dynamic compensation for the impact of changes in legal facts by establishing a node-dependent topology model. Unlike existing static storage models, this solution can intelligently trigger atomic updates based on document change types and continuously correct the confidence of downstream nodes through a multi-hop propagation algorithm, significantly improving the reliability of logical judgment in complex case-handling environments. Through the above technical solution, this application effectively overcomes the problem of logical disconnect caused by the dynamic evolution of documents, improving the confidence of fact-finding while ensuring real-time data synchronization. The multi-hop ripple propagation mechanism balances the strong transmissibility of core evidence with the weak correlation of auxiliary evidence, and the influence domain view visualization function ensures the integrity of the case's logical chain.

[0032] After introducing the basic concept of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0033] Example 1:

[0034] Please see Figures 1-3 As shown, a method for automatically extracting and associating key points of legal documents includes the following steps:

[0035] The system acquires new legal document data for the target case, and uses a pre-set document fingerprint algorithm to compare the new legal document data with historical benchmark documents in the case file database to extract the difference in text data fragments.

[0036] Preferably, the extracted differential text data fragments include:

[0037] The newly added legal document data is divided into natural paragraphs, the text feature hash value of each natural paragraph is calculated, and the fingerprint sequence of the current document is generated by combining the paragraphs in order.

[0038] Retrieve the pre-stored benchmark fingerprint sequence of historical benchmark documents, compare the current document fingerprint sequence with the benchmark fingerprint sequence item by item, identify the sequence positions where the feature hash values ​​of the two are inconsistent, and generate a change area location index.

[0039] Based on the location index of the changed area, the newly added legal document data is mapped back, the original text content at the corresponding position is extracted, and non-semantic format control characters are removed to obtain the difference text data fragment containing substantive change information.

[0040] In one specific embodiment, the newly added legal document data is divided into natural paragraphs. The system uses regular expressions to identify the combination features of line breaks and indentation characters, and divides the long text into an ordered set of paragraphs. Then, for each natural paragraph... The text feature hash value is calculated by first extracting key feature words from the paragraph using the TF-IDF algorithm. And obtain each feature word Corresponding weights ( This is a dimensionless relative importance coefficient, typically ranging from [0,1], which is then processed by a pre-defined hash function. Each feature word is mapped to an L-bit binary vector (preferably L=64 bits to achieve a balance between computational efficiency and collision resistance), and the "0" bits in the binary vector are converted to "-1", while the "1" bits are kept as "1", thus generating the feature vector.

[0041] Based on this, calculate the paragraph. The final fingerprint The formula is:

[0042] ;

[0043] in, The i-th item in the generated current document fingerprint sequence has the physical meaning of the semantic compression encoding of the paragraph, which is represented as an L-bit binary string; This is a sign function; it takes the value 1 when the accumulated result is greater than 0, and 0 when it is less than or equal to 0.

[0044] Generating the fingerprint sequence of the current document Then, the system retrieves the benchmark fingerprint sequence of historical benchmark documents from the case file database. The system performs a step-by-step numerical comparison. The comparison logic uses the Hamming distance algorithm, calculated as follows:

[0045] ;

[0046] in, This represents the XOR operation, where k represents the k-th bit of the binary string. The system sets a change detection threshold. ,when > At that time, it is determined that a substantial change has occurred in the location. Typically... The default value is 0, which requires that the hash values ​​be completely consistent. This is mainly to take into account the rigor of legal documents. Any change in the negative word "no" or the amount may cause the fingerprint bit to flip, and it must be captured as a difference point.

[0047] Through the above comparison, the system identified all sequence position indices with inconsistent feature hash values. The system then maps this data back to the newly added legal documents, extracting the original text of the corresponding paragraphs. Finally, to ensure the purity of subsequent semantic parsing, the system uses regular expressions to remove non-semantic format control characters from the text, obtaining the final differential text data fragments.

[0048] Semantic parsing is performed on the differential text data fragments to generate candidate triple data containing entities and relations. The candidate triple data is then mapped to a pre-built case knowledge graph to locate the graph anchor node where data changes have occurred.

[0049] Preferably, semantic parsing is performed on the differential text data fragments to generate candidate triplet data containing entities and relations, including:

[0050] The system performs word segmentation and part-of-speech tagging on the differential text data fragments, identifies noun phrases representing parties to the case, evidence names, or legal facts, and generates a set of legal entities.

[0051] Based on the co-occurrence position of each entity in the legal entity set in the differential text data fragment, the dependency syntax structure between entities is analyzed, the verb or preposition structure connecting different entities is extracted, and the associated action description is generated.

[0052] The legal entity set is mapped to subject nodes and object nodes respectively, and the associated action descriptions are mapped to relation edges. The subject nodes, object nodes and relation edges are encapsulated according to the logical structure of subject-predicate-object to generate candidate triplet data.

[0053] In one specific embodiment, the extracted differential text data fragments are input into a pre-trained Word2Vec embedding layer specifically for the legal domain, transforming discrete text words into a dense word vector matrix of dimension d=128. , where n is the sequence length. Subsequently, the word vector matrix is ​​input into a BiLSTM layer for contextual feature encoding. This network contains two LSTM units, one forward and one backward, with 256 hidden layer nodes. The tanh activation function is used to avoid gradient vanishing. The BiLSTM layer outputs the emission probability score matrix P for each word corresponding to various entity labels (such as B-PER party, I-EVI evidence, O non-entity, etc.), where This represents the nonnormalized probability that the i-th word is predicted as the j-th label.

[0054] To resolve dependencies between labels (e.g., "I-PER" must immediately follow "B-PER"), a CRF layer is connected above the BiLSTM layer to construct a scoring formula for the label sequence:

[0055] ;

[0056] Where S(X,y) is the total score of the predicted label sequence y given the input sequence X, which is a dimensionless numerical value; A is the state transition matrix, which is a (k+2)×(k+2) matrix, where k is the number of label categories. Indicates from the label Transferred to The transition probability weights are automatically learned through training. This represents the emission probability output by the BiLSTM layer.

[0057] During the model training phase, the maximum likelihood estimation method is used, and the loss function is defined as the negative log-likelihood function: By minimizing this loss using the Adam optimizer (learning rate set to 0.001), a legal entity recognition model capable of accurately identifying noun phrases such as "plaintiff," "defendant," and "IOU" is trained, generating a set of legal entities. It represents the set of all possible label sequences.

[0058] After acquiring the set of legal entities, the system constructs a dependency parser tree based on the co-occurrence positions of each entity in the differential text data fragments, analyzing the dominance and subordination relationships between entities. To quantify the strength of associations between entities and extract action descriptions, the system calculates a semantic interaction strength index. The calculation formula is:

[0059] ;

[0060] in, They are respectively the subject and the object entities; is the shortest path length between two entity nodes in the dependency syntax tree, which is represented by the number of hops of the edge; α is a preset syntactic decay coefficient, with a suggested value range of [0.5, 0.8] and an optimal value of 0.6. Its physical meaning is to penalize long-distance dependencies, prevent incorrect associations across sentences or long clauses, and ensure that only closely related entities are connected. Let w be the set of words along the shortest path; Weight(w) is the part-of-speech weight of word w along the path, where verbs are set to 1.0, prepositions to 0.8, and others to 0.2. ( When the preset association threshold is 0.4 (default value), the system extracts the core verbs or prepositional phrases (such as "pay" or "signed at") on the path as the description of the associated action.

[0061] Finally, the set of legal entities is mapped to subject nodes S and object nodes O, respectively, and the descriptions of related actions are mapped to relation edges. A triple confidence scoring mechanism is introduced to encapsulate the generated results. The scoring formula is:

[0062] ;

[0063] in, The confidence level of the candidate triplet, with values ​​[0,1]; The normalized probability of entity recognition output by the CRF layer; This is the entity-relationship weight balance coefficient, dimensionless, with a default value of 0.6, indicating that the accuracy of the entity plays a dominant role in the construction of triples.

[0064] It should be added that the construction and training process of the BiLSTM-CRF semantic parsing model used to generate the legal entity set and candidate triplet data in the above embodiments is as follows:

[0065] First, a dedicated legal corpus was constructed. Data was sourced from publicly available civil judgments from the past five years on the China Judgments Online website, with a total of 50,000 high-quality samples extracted. In the preprocessing stage, the corpus was manually annotated using the BIOES annotation system. Defined entity categories included: parties, legal basis, evidence, amount, and time. Compared to traditional BIO annotation, the BIOES system adds "E" and "S" tags, enabling more precise segmentation of long entities.

[0066] The data was split in an 8:1:1 ratio, with 40,000 articles used as the training set, 5,000 as the validation set, and 5,000 as the test set. Before inputting the data into the model, all text was cleaned to remove meaningless special characters, and full-width characters were uniformly converted to half-width characters. Numbers were replaced with special placeholders to reduce vocabulary sparsity.

[0067] The model's specific hierarchical structure is designed as a three-level architecture: "embedding layer - bidirectional coding layer - decoding layer".

[0068] The embedding layer initializes a lookup table of dimension |V|×d, where |V|=20,000 is the vocabulary size and d=128 is the word vector dimension. This layer is not randomly initialized, but rather initialized using weights obtained from pre-training the Word2VecSkip-gram model on the entire unannotated legal corpus, thereby introducing prior semantic knowledge.

[0069] The bidirectional encoding layer consists of two stacked LSTM layers. For the input word vector at time t... Forward LSTM computes hidden states Backward LSTM computes hidden states The final output state is the concatenation of the two. The dimension is 2 × 256 = 512. To prevent overfitting, a Dropout layer is introduced after the LSTM layer output, with a dropout rate set to 0.5.

[0070] The decoding layer, located at the top, is used to learn the constraints between labels. It receives the output score matrix of the BiLSTM and introduces the state transition matrix A, which forces the model to follow hard logical constraints such as "I-PER cannot directly follow O".

[0071] Model training is an iterative process of minimizing the loss function. A negative log-likelihood function is used. For a given training sample... The loss function L is defined as:

[0072] ;

[0073] Where S(X,y) is the total score of the predicted label sequence y given the input sequence X, which is obtained by summing the emission score and the transition score; the second term The logarithm of the normalization factor is efficiently calculated in polynomial time using a forward-backward algorithm, avoiding exponential path traversal.

[0074] The Adam optimizer is chosen because it combines the advantages of Momentum and RMSprop, making it suitable for handling sparse gradients.

[0075] The learning rate is initially set to 0.001. To achieve fine convergence in the later stages of training, a learning rate decay strategy is introduced. If the validation set loss does not decrease after every 5 time points, the learning rate is multiplied by a decay factor of 0.5, with a minimum lower bound of 1e-5.

[0076] In each training iteration, the number of samples input into the neural network for training is set to 64. This value is selected based on a balance test between GPU memory utilization and gradient oscillation amplitude: too small a value will lead to unstable gradient direction, while too large a value will easily get stuck in local optima.

[0077] The gradient clipping threshold is set to 5.0 to prevent the gradient explosion problem commonly encountered in LSTM training. The maximum iteration period is set to 100, and an early stopping mechanism is enabled. If the F1 score on the validation set does not improve for 10 consecutive iterations, training is stopped and the optimal weights are saved.

[0078] To verify the effectiveness of the model trained in this embodiment, a comparative experiment was conducted on the test set. The comparison objects were the traditional CRF model (benchmark A) and the simple BiLSTM model without CRF layers (benchmark B).

[0079] The experimental results are shown in the table below:

[0080] Model Architecture Accuracy Recall rate F1 Harmonic Mean Entity boundary recognition error rate Reference A 82.3% 79.5% 80.9% 12.4% Benchmark B 86.1% 88.4% 87.2% 8.9% This embodiment 92.8% 91.5% 92.1% 3.2%

[0081] Results Analysis: The F1 score of the model in this embodiment reaches 92.1%, which is significantly better than the benchmark model. In particular, the "entity boundary recognition error rate" metric is reduced to 3.2%, which verifies the key role of the CRF layer in handling label dependencies.

[0082] Preferably, the candidate triplet data is mapped to a pre-constructed case knowledge graph, and the graph anchor node where the data change has occurred is located, including the following steps:

[0083] The candidate triplet data is analyzed, and the names of the subject element and the object element are extracted respectively. The node index of the case knowledge graph is then searched to select nodes with matching names as a set of potential matching nodes.

[0084] Obtain the preset attribute features of each node in the potential matching node set, calculate the similarity value between the preset attribute features and the semantic context of the candidate triplet data, and lock the node with the highest similarity value as the unique mapping node.

[0085] Identify the current connection state of the unique mapping node in the case knowledge graph. When there is a difference between the relation data in the candidate triplet data and the current connection state, the unique mapping node is established as the graph anchor node for which a change operation is to be performed.

[0086] In one specific embodiment, the system parses candidate triplet data.<S,R,O> The subject element name S (e.g., "Zhang San") and object element name O (e.g., "loan contract") are extracted separately, and a fast search is performed in the node name database of the case knowledge graph using an inverted index mechanism. To tolerate minor glyph errors that may occur during OCR recognition, the search process does not use absolute exact matching, but rather performs initial screening based on the edit distance ratio. The screening formula is as follows: Where Dist is the Levenstein distance function, which calculates the minimum number of single-character edits required to transform one string into another, and |·| represents the string length. This refers to the name of the subject or object element extracted from the candidate triplet data. This refers to each retrieved node name in the knowledge graph's node name database. The system sets a preset matching threshold λ, recommended to be in the range [0.85, 0.95], preferably 0.90. The logic behind this value setting is: if set too low (e.g., 0.6), it will introduce a large number of irrelevant nodes with the same name, increasing the subsequent computational load; if set too high (e.g., 1.0), it will lack fault tolerance. All All nodes were included in the potential matching node set. .

[0087] Next, for each potential node in the set... The system acquires its preset attribute features (including text descriptions such as ID number mask, historical case details, and associated addresses) and calculates the semantic context similarity between the candidate triplet and the original text paragraph containing it. To quantify this indicator, the system uses a weighted vector space model, with the calculation formula as follows: ;

[0088] in, This is a dimensionless similarity score, with a value range of [0,1]. The context semantic vector of the original paragraph containing the candidate triplet is generated by extracting the average value of the last hidden state from the pre-trained Legal-BERT model, with a dimension of 768. Nodes in the graph The attribute feature aggregation vector is generated by encoding the same BERT model; the first term is the cosine similarity calculation, which represents the degree of semantic matching. (·) is the type constraint indicator function. It takes 1 if the type of the graph node is consistent with the predicted type of the triplet, and 0 otherwise. This is the semantic weight coefficient, preset to 0.7. The value of this parameter is based on the fact that semantic context typically has higher discriminative power than simple type matching, thus assigning it a higher weight. The system locks the node with the highest Sim value in the calculation results that exceeds the confidence lower bound as the unique mapping node.

[0089] Finally, the system queries the current connection status of the unique mapping node in the graph. Specifically, this involves retrieving the set of all outgoing edges E_{out} originating from the unique mapping node. The system compares the relation data R in the candidate triples with the edge labels in the outgoing edge set. If no edge with semantically consistent with R is found in the outgoing edge set (e.g., the original record in the graph was "no relation," but the current triple is "loan"), or if an edge exists but the attribute values ​​are contradictory (e.g., the graph record is "loan amount = 50,000," but the triple is "loan amount = 100,000"), then a substantial data change is determined to have occurred. In this case, the unique mapping node is established as the graph anchor node, and the change type (new edge or attribute update) is marked.

[0090] Based on the change type of the differential text data fragment, an atomic update operation is performed on the graph anchor node. The atomic update operation includes updating the marker of the node state attribute and generating the change source node data with the latest state identifier.

[0091] Preferably, the logic for determining the change type of the differing text data fragments is as follows:

[0092] Keyword extraction is performed on the differential text data fragments. The extracted predicate verbs and adjectives are matched with a pre-set legal negation word library. The frequency and weight of the matched words are calculated to generate a negation semantic score.

[0093] It should be added that generating a negative semantic score includes:

[0094] Part-of-speech tagging is performed on the differential text data fragments. Predicate verbs and adjectives are selected based on the part-of-speech tags, and a candidate keyword set containing the predicate verbs and adjectives is constructed.

[0095] The candidate keyword set is matched item by item with the pre-set legal negation word library to identify the overlapping words. The preset semantic weight values ​​of the words are retrieved from the legal negation word library to generate a weighted matching sequence containing all matching words and their semantic weight values.

[0096] The semantic weight values ​​recorded in the weighted matching sequence are summed to obtain the total weight value, and the total weight value is determined as the negative semantic score.

[0097] Retrieve the historical attribute data of the anchor node in the case knowledge graph, check whether the historical attribute data is empty, and generate a node existence identifier that represents the old and new states of the node;

[0098] The negative semantic score and node existence identifier are input into a preset logical decision tree: if the node existence identifier is empty, it is determined to be a new type; if the node existence identifier is not empty and the negative semantic score is greater than the preset decision threshold, it is determined to be a negative type; otherwise, it is determined to be a correction type, thereby determining the change type of the difference text data fragment.

[0099] Preferably, an atomic update operation is performed on the map anchor node to generate changed source node data, including:

[0100] Based on the change type of the differential text data fragment, the corresponding status change instruction is matched for the map anchor node. If the change type is negative, the invalidation instruction is matched; if the change type is new or correction, the activation instruction is matched.

[0101] The status change instruction is invoked to overwrite the validity attribute field of the graph anchor node, and the current system time is written into the update time field of the graph anchor node, thereby generating the updated node object;

[0102] The updated node object is associated and encapsulated with the difference text data fragment used as the basis for the change, generating change source node data containing the latest validity attributes and time information.

[0103] In one specific embodiment, the system first performs part-of-speech tagging in natural language processing to filter out predicate verbs (such as "revoke" and "terminate") and adjectives (such as "invalid" and "cancelled") with substantial semantic influence, and constructs a candidate keyword set. Subsequently, a pre-set legal negation lexicon is invoked. This lexicon stores common negative terms in the legal field and their corresponding semantic weight values. Its setting logic is based on the legal enforceability of the terms: for example, terms with final negative meaning, such as "void" and "reject," are recommended to have a weight of 0.9-1.0; while terms with temporary or reversible negative meaning, such as "suspend" and "cease," are recommended to have a weight of 0.5-0.7.

[0104] The system combines the candidate keyword set K with the negative keyword library Perform intersection operations and calculate the negative semantic score based on the matching results. The calculation formula is: ;

[0105] in, The final negative semantic score is given in points. (·) is an indicator function, when the keyword The value is 1 if it exists in the dictionary, and 0 otherwise. This is the preset weight for the keyword in the thesaurus.

[0106] In the calculation Subsequently, the system synchronously retrieves the historical attribute data of the map anchor nodes and generates node existence identifiers. (If the historical data is not empty) =1, otherwise 0). Then, and The input is fed into a pre-defined logic decision tree for a three-level branch decision. The decision logic of this decision tree is as follows:

[0107] 1. First-level judgment: Inspection .like =0 indicates that this node does not yet exist in the graph, and the change type is directly determined as "new type".

[0108] 2. Second-level judgment: If =1, then further comparison The preset threshold is recommended to be set to 0.8. This value is set to filter out weak negations or ordinary negative descriptions, ensuring that negation is only triggered when explicit high-weight words such as "withdraw" or "invalid" appear. If a preset threshold is set, the change type will be determined as "negative type".

[0109] 3. Third-level judgment: If =1 and If the value is less than or equal to the preset threshold, the change type is determined to be "correction type".

[0110] Finally, the system performs atomic update operations on the map anchor nodes based on the determined change type. Atomic update refers to an indivisible, minimal write operation to ensure data consistency. The system matches the corresponding state change instruction based on the change type: if it's a "negative type," it matches an invalidation instruction, overwriting the node's "validity attribute field" to False or 0; if it's a "new type" or "correction type," it matches an activation instruction, overwriting the field to True or 1. Simultaneously, the system captures the current system time point and writes it to the node's "update time point field," generating the updated node object. Finally, the updated node object is associated and encapsulated with the differing text data fragments, outputting the change source node data.

[0111] Based on the predefined node dependency topology in the case knowledge graph, multi-hop ripple propagation calculation is performed starting from the changed source node data to identify downstream related nodes affected by the changed source node data. Based on the propagation strength parameters stored in the node dependency topology, the predicted confidence values ​​of the downstream related nodes are calculated.

[0112] Preferably, the predefined logic for the node dependency topology in the case knowledge graph is as follows:

[0113] Extract entity relationship definitions from the case knowledge graph metadata model, and classify the entity relationship definitions into logical support categories and auxiliary association categories according to the causal strength of legal logic, and generate corresponding dependency classification labels;

[0114] The system invokes preset strength mapping rules to match full transmission coefficients for dependency classification identifiers belonging to the logical support category and preset attenuation coefficients for dependency classification identifiers belonging to the auxiliary association category, thereby generating transmission strength parameters that can quantify influence.

[0115] The transmission strength parameter is encapsulated and written into the attribute configuration field of the entity relationship definition, thereby establishing a node dependency topology structure containing explicit data flow rules.

[0116] Preferably, the logic for generating the conduction strength parameter specifically includes:

[0117] The preset strength mapping rule table is searched based on the dependency classification identifier. When the dependency classification identifier indicates a logical support category, the preset full transmission coefficient is locked, and when the dependency classification identifier indicates an auxiliary association category, the preset attenuation coefficient is locked.

[0118] Read the specific numerical content corresponding to the locked full transfer coefficient and attenuation coefficient, and extract the basic weight value from the numerical content.

[0119] The basic weight values ​​are directly assigned and defined as the conduction strength parameter.

[0120] The working process and principle of this application are as follows: The system performs semantic classification of the edge relationships in the graph based on the causal logic depth and probative strength of the legal entities, and retrieves the corresponding transmission coefficients from a pre-set strength mapping rule table. The pre-set logic of this strength mapping rule table is to simulate the principle of "evidence necessity" in legal argumentation: Relationships that constitute essential elements of the facts of the case (such as the support of "contract content" for "liability for breach of contract") are defined as logical support categories and assigned a full transmission coefficient (1.0) to ensure that changes in core facts can be transmitted to related conclusions without loss; while relationships that only play a supplementary role in the background (such as the connection between "plaintiff's place of residence" and "jurisdiction") are defined as auxiliary connection categories and assigned a decay coefficient (such as 0.3~0.6) to simulate the natural decrease in probative strength of related evidence.

[0121] Next, the system locks down specific coefficient values ​​and directly assigns them as transmission strength parameters, thus transforming abstract legal causal logic into a computable weight matrix. The significance of this is that, through this quantitative processing method, it eliminates the subjective ambiguity in judging the relevance of evidence in traditional manual assessments, providing a standardized "energy transfer" basis for subsequent multi-hop ripple propagation algorithms.

[0122] Finally, by using the generated transmission strength parameter as the core operator for propagation attenuation, the system can automatically filter out redundant nodes with low legal relevance in multi-step recursive calculations, preventing irrelevant "information explosion" caused by data changes. This allows the updated influence of the graph to be accurately focused on the fluctuations of the core evidence chain, thereby outputting influence domain view data that truly reflects the changing patterns of case facts.

[0123] In a specific embodiment, the system traverses the entity relationship definitions in the graph metadata model and classifies the relationships into logical support categories and auxiliary association categories based on the strength of causality in legal logic. "Logical support categories" refer to relationships with strong causal linkage (such as "IOU-proof-loan fact" or "judgment-basis-legal provision"), where the downstream relationship inevitably fails if the upstream collapses. "Auxiliary association categories" refer to weak relationships that only provide background information (such as "plaintiff-residence-address" or "party-colleague relationship-witness"), where the relationship only affects credibility rather than existence.

[0124] Based on this, the system generates conduction strength parameters according to preset strength mapping rules. For logical support categories, lock the full transmission coefficient. This implies the lossless transmission of influence; for auxiliary association categories, a preset attenuation coefficient is locked. The recommended value range is [0.3, 0.6], with a preferred value of 0.5, meaning that the influence is halved during transmission. This parameter is directly encapsulated in the edge attributes, forming a weighted directed graph.

[0125] Subsequently, using the changed source node data as the epicenter (hop 0), the system performs multi-hop ripple propagation calculations. For the node in hop k... and its downstream (k+1)th hop's adjacent nodes Calculate the prediction confidence value after downstream nodes are affected. The calculation formula uses a path attenuation and intensity coupling model, and the formula is as follows:

[0126] ;

[0127] in, This is the updated confidence value of the current upstream node. If the node is invalid, it is 0; if it is valid, it is 1 or a specific probability value. The transmission strength parameter stored on the edge from node i to node j is taken from the above full transmission coefficient or the preset attenuation coefficient. The preset distance attenuation factor is dimensionless and has a preset value of 0.9. Its physical meaning is to simulate the energy dissipation during the ripple diffusion process and prevent the influence from infinitely looping and amplifying in the spectrum. k represents the current jump level.

[0128] To prevent computational divergence, the system sets a propagation cutoff threshold. When the calculated change in prediction confidence When the time is right, the propagation of that path is stopped. Through this iterative calculation, the system can accurately identify all downstream related nodes affected by the changed source node data, and write the calculated prediction confidence values ​​into the attributes of these nodes, completing the synchronous correction of the state.

[0129] Preferably, multi-hop ripple propagation calculations are performed starting from the changed source node data to identify downstream related nodes affected by the changed source node data. The specific operation logic includes:

[0130] Construct a propagation queue, write the changed source node data as the first processing item into the propagation queue, and establish a set of downstream related nodes to store the final result;

[0131] Extract the current processing item from the propagation queue, retrieve candidate next-hop nodes that have outgoing edge connections with the current processing item based on the node dependency topology, and retrieve the predefined propagation strength parameters on the connection path;

[0132] The validity of the conduction strength parameter is verified. When the conduction strength parameter meets the preset conduction conditions, the candidate next-hop node is marked as the affected node and written into the downstream associated node set. At the same time, the affected node is added to the end of the propagation queue to trigger the next round of recursive retrieval until the propagation queue is cleared, thereby identifying all the downstream associated nodes.

[0133] Preferably, the logic for constructing the propagation queue includes:

[0134] Allocate a linear storage space with first-in-first-out characteristics in memory, and define the linear storage space as an empty queue container;

[0135] Extract the unique index identifier of the source node data, associate the unique index identifier with a preset saturation confidence value representing the maximum impact weight, and an initial path level value representing the starting propagation distance, and encapsulate the unique index identifier, preset saturation confidence value, and initial path level value into a root propagation instance.

[0136] Write the root propagation instance as the first element to the head of the empty queue container, and confirm the empty queue container containing the root propagation instance as the initialized propagation queue to be processed.

[0137] Preferably, the logic for constructing the numerical validity of the verification conduction strength parameter is as follows:

[0138] The current confidence value is parsed from the current processing item in the propagation queue, and the current confidence value is multiplied with the propagation strength parameter to obtain the predicted confidence value used to characterize the credibility of the downstream node.

[0139] Parse the current path level value contained in the current processing item, perform a single-unit increment count on the current path level value, and generate the next hop path level value representing the node propagation depth.

[0140] The predicted confidence value is compared with the preset minimum attenuation threshold, and the next hop path level value is compared with the preset maximum propagation hop count. When the predicted confidence value is greater than the minimum attenuation threshold and the next hop path level value is less than the maximum propagation hop count, the conduction strength parameter is determined to meet the preset conduction condition.

[0141] Preferably, the logic for calculating the predicted confidence value of downstream associated nodes based on the conduction strength parameters configured in the node-dependent topology includes:

[0142] Read the current confidence value carried by the current processing node, locate the connection path between the current processing node and the downstream associated node, and obtain the conduction strength parameter corresponding to the connection path;

[0143] The current confidence level value is multiplied by the conduction strength parameter to generate a predicted confidence level value;

[0144] The predicted confidence score is written into the attribute field of the downstream associated node to obtain the predicted confidence score of the downstream associated node.

[0145] In one specific embodiment, the process of performing multi-hop ripple propagation calculations starting from the changed source node data and identifying downstream associated nodes affected by the changed source node data is essentially a directed graph energy decay and diffusion process based on a breadth-first search strategy.

[0146] Specifically, to achieve this process, the system first allocates a linear storage space with first-in, first-out (FIFO) characteristics in memory and initializes it as an empty queue container Q. The system extracts the unique index identifier of the changed source node data ( It configures initial propagation state parameters for it, including a preset saturation confidence value representing the weight of maximum influence. It is a dimensionless constant with a fixed value of 1.0, and its physical meaning is that the influence of the source of change is 100% certain; and the initial path hierarchy value represents the initial propagation distance. The value is 0 (unit: hops). The system will This tuple combination is encapsulated as a root propagation instance, and it is pushed as the first element into the head of Q. At this point, the state of Q changes to an initialized propagation queue, and an empty set of downstream associated nodes is created. Used for deduplicating affected objects.

[0147] Next, the system enters a recursive propagation loop: as long as Q is not empty, a currently processed item is popped from the tail of the queue. The system retrieves data based on a pre-built node dependency topology. For all outgoing edges from the starting point, lock the set of candidate next-hop nodes. And retrieve the predefined conduction strength parameters on each connection edge. (This parameter has been assigned a value in the previous steps based on the logical support or auxiliary association category, for example, 1.0 for strong logic and 0.6 for weak association).

[0148] For each candidate next-hop node in the set The system performs rigorous numerical validity checks, which include two dimensions: energy conduction calculation and boundary constraint determination. First, it calculates the prediction confidence score, which characterizes the reliability of downstream nodes. The calculation formula is:

[0149] ;

[0150] in, The confidence level of the current processing node; The conduction strength parameter on the edge; This is the global environmental attenuation factor, a preset dimensionless coefficient. A suggested value range is [0.9, 0.98], with a preferred value of 0.95. (Introduction) The physical significance lies in simulating the natural entropy increase or signal-to-noise ratio decrease during the multi-level transmission of information, preventing the influence from endlessly looping in the closed loop of the graph.

[0151] Simultaneously, the system increments the current path level value to generate the next-hop path level value. .

[0152] The calculation result is then input into a preset logic comparator:

[0153] 1. Intensity threshold determination: [This is a judgment / determination process.] Is it greater than the minimum attenuation threshold? . It is recommended to preset it to 0.1 (dimensionless) to filter out "long-tail noise" whose impact is negligible due to long-link transmission.

[0154] 2. Depth of circuit breaker determination: [This refers to a specific function or feature, which is not directly related to the previous sentence and can be omitted.] Is it less than the maximum propagation hops? . It is recommended to set it to 5, which represents the number of hops. This value is set based on the "six degrees of separation theory" and computational resource constraints to prevent the algorithm from getting stuck in deep traversal traps in ultra-large-scale graphs.

[0155] If and only if > and < If both conditions are met simultaneously, the preset conduction condition is deemed valid. At this point, the system performs a write operation: [The system will then...] Write node In the "Prediction Confidence" attribute field, complete the status update; and update the node. Mark the affected node and store it. At the same time, new propagation instances will be... The node is appended to the end of queue Q to trigger the next round of recursive retrieval. This process continues until Q is empty, thus outputting the complete set of downstream associated nodes.

[0156] The output includes case-related impact domain view data containing data on the changed source node, the affected downstream related nodes, and their updated confidence values.

[0157] Preferably, the output includes case-related impact domain view data containing changed source node data, affected downstream related nodes, and their updated confidence values. The output process is as follows:

[0158] By integrating the change source node data and each downstream related node, extracting the predicted confidence values ​​corresponding to each downstream related node, and mapping the predicted confidence values ​​to the attribute fields of the corresponding nodes, a complete set of influence domain nodes containing complete change status information is constructed.

[0159] Based on the node dependency topology, the directed edge connections between nodes within the entire set of nodes in the influence domain are extracted. The directed edge connections and the entire set of nodes in the influence domain are then combined in a structured manner to generate change influence subgraph data that represents the path and scope of data change diffusion.

[0160] Using preset graphical mapping rules, visual identifier parameters matching the predicted confidence values ​​are assigned to each node in the change impact subgraph data. The visual identifier parameters and the change impact subgraph data are then encapsulated into case-related impact domain view data in a standard description format and output.

[0161] In one specific embodiment, the system first performs data integration and subgraph construction operations. The system will use the changed source node data (denoted as the starting point of propagation) as the propagation origin. ) and the set of downstream associated nodes determined by preceding multi-hop ripple propagation (denoted as ) Perform a union operation to construct a complete set of influence domain nodes containing complete change status information. Based on this, the system traverses the underlying node dependency topology of the case knowledge graph, filtering out nodes whose start and end points both belong to... Generate the influence domain edge set by connecting all directed edges. By and By performing structured composition, the system instantiates an independent change-impact subgraph of data in memory. This subgraph fully preserves the flow path and diffusion range of data changes, and removes irrelevant redundant nodes from the graph.

[0162] Subsequently, in order to visually demonstrate the extent to which the changes disrupted the chain of evidence in the case, the system used preset graphical mapping rules to map each node in the subgraph. The visual label parameters are calculated to match the predicted confidence values. The calculation of these parameters includes two dimensions: warning color level mapping and dynamic scaling of node size.

[0163] First, for the warning color level mapping, the system uses an RGB linear interpolation algorithm based on confidence level, and the calculation formula is as follows:

[0164] ;

[0165] in, Render the calculated node color vector [R,G,B]; For nodes The prediction confidence value, with a range of [0,1]; The preset high-risk warning color vector is recommended to be [255, 50, 50], which represents a state where the confidence level is zero, i.e., completely ineffective. The preset safety state color vector is recommended to be [50, 200, 50], representing a confidence level of 1, i.e., an unaffected state. The physical meaning of this formula is to smoothly transition the node's color between the red and green spectrum based on the degree of confidence deficiency. The lower the confidence level, the more the node's color tends to be warning red.

[0166] Secondly, regarding dynamic scaling of node sizes, in order to highlight critical nodes that are severely affected, the system calculates the rendering radius of the nodes. The calculation formula is: ;

[0167] in, The display radius of the node, in pixels (px). This is the radius of the basic node, preset to 20px, to ensure the basic visibility of the node; This represents the absolute value of the change in node confidence, reflecting the impact of this change on the node. This is the visual magnification factor, dimensionless, with a default recommended value of 1.5. This formula ensures that nodes whose confidence drops sharply (such as key evidence dropping from 1.0 to 0.0) are significantly magnified in the view, while nodes with stable states remain at their basic size.

[0168] Finally, the system will calculate the results. and The changes are written as extended attributes to the corresponding nodes of the change impact subgraph data, and the change impact subgraph data is serialized into case-related impact domain view data in a standard description format. The output operation is then performed for the front-end rendering engine to call.

[0169] Example 2:

[0170] like Figure 4 As shown, an automatic extraction and association system for key points of legal documents is provided. The system includes a document data acquisition and difference comparison module, a semantic parsing and graph anchor point positioning module, a graph node atomic update operation module, a node dependency topology and ripple calculation module, and a graph evolution domain analysis output module.

[0171] The various modules are connected via wired and / or wireless means to enable data transmission between them;

[0172] The document data acquisition and difference comparison module is used to acquire new legal document data for the target case, and uses a preset document fingerprint algorithm to compare the new data with the historical benchmark documents in the case file database to extract the difference text data fragments.

[0173] The semantic parsing and graph anchor point localization module receives differential text data fragments, generates candidate triple data containing entities and relationships through semantic parsing technology, maps the candidate triple data to a pre-built case knowledge graph, and locates the graph anchor point nodes where data changes have occurred.

[0174] The graph node atomic update operation module performs atomic update operations on graph anchor nodes based on the change type of the differential text data fragments, generating change source node data with the latest status identifier;

[0175] The node dependency topology and ripple calculation module, based on the predefined node dependency topology in the case knowledge graph, takes the changed source node data as the starting point, performs multi-hop ripple propagation calculation, identifies the downstream related nodes affected by the changed source node data, and calculates the predicted confidence value of the downstream related nodes according to the transmission strength parameters stored in the node dependency topology.

[0176] The graph evolution domain analysis output module outputs case association impact domain view data, which includes data on the changed source nodes, all affected downstream related nodes, and their updated confidence values.

[0177] It should be noted that the interval and threshold sizes are set for ease of comparison. The size of the threshold depends on the amount of sample data and the base number set by those skilled in the art for each set of sample data, as long as it does not affect the proportional relationship between the parameter and the quantized value. Furthermore, the above formulas are all dimensionless calculations, and the formulas are derived from software simulations using a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0178] It should be understood that, in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0179] It should be understood that determining B based on A does not mean determining B solely based on A; it also means determining B based on A and / or other information.

[0180] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0181] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for automatically extracting and associating key points of legal documents, characterized in that, The method includes the following steps: Acquire new legal document data for the target case, and use a preset document fingerprint algorithm to compare the new legal document data with historical benchmark documents in the case file database to extract the difference text data fragments; Semantic parsing is performed on the differential text data fragments to generate candidate triple data containing entities and relations, and the candidate triple data is mapped to a pre-built case knowledge graph to locate the graph anchor node where data changes have occurred. Based on the change type of the differential text data fragment, perform an atomic update operation on the graph anchor node to generate the change source node data; Based on the predefined node dependency topology in the case knowledge graph, multi-hop ripple propagation calculation is performed starting from the changed source node data to identify downstream related nodes affected by the changed source node data. Based on the propagation strength parameters stored in the node dependency topology, the predicted confidence values ​​of the downstream related nodes are calculated. The output includes case-related impact domain view data containing data on the changed source node, the affected downstream related nodes, and their updated confidence values; The predefined logic of node dependency topology in the case knowledge graph is as follows: Extract entity relationship definitions from the case knowledge graph metadata model, and classify the entity relationship definitions into logical support categories and auxiliary association categories according to the causal strength of legal logic, and generate corresponding dependency classification labels; The preset strength mapping rules are invoked to match the full transmission coefficient for the dependency classification identifier that belongs to the logical support category, and to match the attenuation coefficient with a preset value for the dependency classification identifier that belongs to the auxiliary association category, thereby generating the transmission strength parameter. The transmission strength parameter is encapsulated and written into the attribute configuration field of the entity relationship definition, thereby establishing a node dependency topology. The logic for generating the conduction strength parameter specifically includes: The preset strength mapping rule table is searched based on the dependency classification identifier. When the dependency classification identifier indicates a logical support category, the preset full transmission coefficient is locked, and when the dependency classification identifier indicates an auxiliary association category, the preset attenuation coefficient is locked. Read the specific numerical content corresponding to the locked full transfer coefficient and attenuation coefficient, and extract the basic weight value from the numerical content. The basic weight values ​​are directly assigned and defined as the conduction strength parameters; Starting with the changed source node data, perform multi-hop ripple propagation calculations to identify downstream related nodes affected by the changed source node data. The specific operation logic includes: Construct a propagation queue, write the changed source node data as the first processing item into the propagation queue, and establish a set of downstream related nodes to store the final result; Extract the current processing item from the propagation queue, retrieve candidate next-hop nodes that have outgoing edge connections with the current processing item based on the node dependency topology, and retrieve the predefined propagation strength parameters on the connection path; The validity of the conduction strength parameter is verified. When the conduction strength parameter meets the preset conduction conditions, the candidate next-hop node is marked as the affected node and written into the downstream associated node set. At the same time, the affected node is added to the end of the propagation queue to trigger the next round of recursive retrieval until the propagation queue is cleared, thereby identifying all the downstream associated nodes.

2. The method for automatically extracting and associating key points of legal documents according to claim 1, characterized in that, Mapping candidate triplet data to a pre-built case knowledge graph and locating graph anchor nodes where data changes have occurred includes the following steps: The candidate triplet data is analyzed, and the names of the subject element and the object element are extracted respectively. The node index of the case knowledge graph is then searched to select nodes with matching names as a set of potential matching nodes. Obtain the preset attribute features of each node in the potential matching node set, calculate the similarity value between the preset attribute features and the semantic context of the candidate triplet data, and lock the node with the highest similarity value as the unique mapping node. Identify the current connection state of the unique mapping node in the case knowledge graph. When there is a difference between the relation data in the candidate triplet data and the current connection state, the unique mapping node is established as the graph anchor node for which a change operation is to be performed.

3. The method for automatically extracting and associating key points of legal documents according to claim 1, characterized in that, The logic for determining the change type of differential text data fragments is as follows: Keyword extraction is performed on the differential text data fragments. The extracted predicate verbs and adjectives are matched with a pre-set legal negation word library. The frequency and weight of the matched words are calculated to generate a negation semantic score. Retrieve the historical attribute data of the anchor node in the case knowledge graph, check whether the historical attribute data is empty, and then generate a node existence identifier; The negative semantic score and node existence identifier are input into a preset logical decision tree: if the node existence identifier is empty, it is determined to be a new type; if the node existence identifier is not empty and the negative semantic score is greater than the preset decision threshold, it is determined to be a negative type; otherwise, it is determined to be a correction type, thereby determining the change type of the difference text data fragment.

4. The method for automatically extracting and associating key points of legal documents according to claim 1, characterized in that, The logic for constructing the numerical validity of the verification conduction strength parameter is as follows: The current confidence value is parsed from the current processing item in the propagation queue, and the current confidence value is multiplied with the propagation strength parameter to obtain the predicted confidence value. Parse the current path level value contained in the current processing item, perform a single-unit increment count on the current path level value, and generate the next hop path level value; The predicted confidence value is compared with the preset minimum attenuation threshold, and the next hop path level value is compared with the preset maximum propagation hop count. When the predicted confidence value is greater than the minimum attenuation threshold and the next hop path level value is less than the maximum propagation hop count, the conduction strength parameter is determined to meet the preset conduction condition.

5. The method for automatically extracting and associating key points of legal documents according to claim 4, characterized in that, The logic for calculating the predicted confidence values ​​of downstream associated nodes based on the conduction strength parameters configured in the node-dependent topology includes: Read the current confidence value carried by the current processing node, locate the connection path between the current processing node and the downstream associated node, and obtain the conduction strength parameter corresponding to the connection path; The current confidence level value is multiplied by the conduction strength parameter to generate a predicted confidence level value; The predicted confidence score is written into the attribute field of the downstream associated node to obtain the predicted confidence score of the downstream associated node.

6. The method for automatically extracting and associating key points of legal documents according to claim 1, characterized in that, The output includes case-related impact domain view data containing changed source node data, affected downstream related nodes, and their updated confidence values. The output process is as follows: Integrate the change source node data and each downstream related node, extract the prediction confidence value corresponding to each downstream related node, and map the prediction confidence value to the attribute field of the corresponding node to construct a complete set of nodes in the influence domain; Based on the node dependency topology, the directed edge connections between nodes within the entire set of nodes in the influence domain are extracted. The directed edge connections and the entire set of nodes in the influence domain are then combined in a structured manner to generate change influence subgraph data that represents the path and scope of data change diffusion. Using preset graphical mapping rules, visual identifier parameters matching the predicted confidence values ​​are assigned to each node in the change impact subgraph data. The visual identifier parameters and the change impact subgraph data are then encapsulated into case-related impact domain view data and output.

7. A system for automatically extracting and associating key points of legal documents, characterized in that, It is implemented based on the automatic extraction and association method of key points of legal documents as described in any one of claims 1-6. The system includes: The document data acquisition and difference comparison module is used to acquire new legal document data for the target case, and uses a preset document fingerprint algorithm to compare the new data with the historical benchmark documents in the case file database to extract the difference text data fragments. The semantic parsing and graph anchor point localization module receives differential text data fragments, generates candidate triple data containing entities and relationships through semantic parsing technology, maps the candidate triple data to a pre-built case knowledge graph, and locates the graph anchor point nodes where data changes have occurred. The graph node atomic update operation module performs atomic update operations on graph anchor nodes based on the change type of the differential text data fragments, generating change source node data; The node dependency topology and ripple calculation module, based on the predefined node dependency topology in the case knowledge graph, takes the changed source node data as the starting point, performs multi-hop ripple propagation calculation, identifies the downstream related nodes affected by the changed source node data, and calculates the predicted confidence value of the downstream related nodes according to the transmission strength parameters stored in the node dependency topology. The graph evolution domain analysis output module outputs case association impact domain view data, which includes data on the changed source nodes, all affected downstream related nodes, and their updated confidence values.

Citation Information

Patent Citations

  • Legal document information extraction method, device, computer equipment and storage medium

    CN110516036A

  • Information extraction method based on legal instruments

    CN114036933A