File clause change tracking method and device, equipment and storage medium

By using knowledge graphs for clause-level structured analysis, changes to document clauses can be identified and tracked, solving the problem of low efficiency in manual comparison and achieving efficient and accurate identification and tracking of changed content.

CN122045150APending Publication Date: 2026-05-15INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610199091.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-11
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In existing technologies, when document clauses are changed, manual comparison of old and new versions is required, which is inefficient and prone to errors.

Method used

Based on a pre-built knowledge graph, clause-level structured analysis is performed. By identifying change nodes, their associated nodes, and influencing nodes through the node relationships in the knowledge graph, influence edges are established to achieve clause-level tracking of changed content.

Benefits of technology

This improved the efficiency and accuracy of identifying changes, ensuring the comprehensiveness of the changes and the accuracy of the structured analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122045150A_ABST
    Figure CN122045150A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a file clause change tracking method and device, equipment and a storage medium. The method comprises the steps that clause-level structured analysis is carried out on a first version file and a second version file based on a pre-constructed knowledge graph, change nodes in the knowledge graph are obtained, and the version of the first version file is earlier than that of a second text file; determining an association node candidate set of the change node according to the relationship between the nodes in the knowledge graph every time the change node is determined; determining whether influence nodes of the change nodes exist in the association node candidate set or not according to the similarity of the association nodes and the change nodes; if yes, establishing an influence edge between the change node and the influence node, and determining the influence node as a new change node; and if not, tracking the change content of the second version file along all the influence edges. And the efficiency and comprehensiveness of change content identification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of document analysis technology, and in particular to a method, apparatus, device, and storage medium for tracking document clause changes. Background Technology

[0002] As a carrier of information, documents are widely used in various industries, especially for information with legal characteristics, which is usually recorded, transmitted, and stored in the form of document files, such as contract documents, agreement documents, and rule documents.

[0003] When certain clauses in these documents are changed, it is often necessary to manually compare the old and new versions, manually identify the changes in key clauses and their impact, which is inefficient and prone to errors. Summary of the Invention

[0004] This application provides a method, apparatus, device, and storage medium for tracking changes to document terms, in order to improve the efficiency and accuracy of identifying changed terms.

[0005] In a first aspect, embodiments of this application provide a method for tracking changes to document terms, the method comprising: Based on a pre-built knowledge graph, a clause-level structured analysis was performed on the first and second version files to obtain the change nodes in the knowledge graph. The version of the first file is earlier than the version of the second text file. For each changed node identified, a candidate set of associated nodes is determined based on the relationships between nodes in the knowledge graph. Determine whether there are any nodes affected by the changed nodes in the candidate set of associated nodes based on the similarity between associated nodes and changed nodes; If it exists, establish an influence edge between the changed node and the affected node, and identify the affected node as the new changed node; If not, trace the changes to the second version file along all affected edges.

[0006] Secondly, embodiments of this application provide a document clause change tracking device, the device comprising: The analysis module is used to perform clause-level structured analysis on the first and second version files based on a pre-built knowledge graph to obtain the change nodes in the knowledge graph. The version of the first file is earlier than the version of the second text file. The associated node determination module is used to determine the candidate set of associated nodes for each determined change node based on the relationships between nodes in the knowledge graph. The influence node determination module is used to determine whether there are influence nodes of the changed node in the candidate set of associated nodes based on the similarity between associated nodes and changed nodes. The influence edge establishment module is used to establish influence edges between the changed node and the affected node if they exist, and to identify the affected node as the new changed node. The impact tracking module is used to track changes to the second version file along all impact edges if they do not exist.

[0007] Thirdly, embodiments of this application also provide an electronic device, which includes: One or more processors; Storage device for storing one or more programs. When one or more programs are executed by one or more processors, the one or more processors implement the document terms change tracking method as provided in any embodiment of this application.

[0008] Fourthly, embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the document terms change tracking method as provided in any embodiment of this application.

[0009] Fifthly, embodiments of this application also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the document terms change tracking method as provided in any embodiment of this application.

[0010] The technical solution of this application embodiment performs clause-level structured analysis on a first version file and a second version file based on a pre-constructed knowledge graph to obtain change nodes in the knowledge graph. The version of the first version file is earlier than the version of the second text file. For each change node identified, a candidate set of associated nodes is determined based on the relationships between nodes in the knowledge graph. The similarity between the associated nodes and the change node is used to determine whether there are any influencing nodes in the candidate set of associated nodes. If so, an influence edge is established between the change node and the influencing node, and the influencing node is identified as a new change node. If not, the changes in the second version file are tracked along all influence edges. Based on this, clause-level structured analysis, with clauses as units for subsequent identification and tracking, better conforms to the structure of the document itself, improving the accuracy of identification to a certain extent. Furthermore, by leveraging the node relationship information in the knowledge graph, the tracking of changes is achieved, improving the efficiency and comprehensiveness of changes identification. Attached Figure Description

[0011] Figure 1 A flowchart illustrating the document clause change tracking method provided in Embodiment 1 of this application; Figure 2 This is a schematic diagram of a document clause change tracking device provided in Embodiment 2 of this application; Figure 3This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of this application. Detailed Implementation

[0012] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the application and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present application, not the entire structure.

[0013] Example 1 Figure 1 This is a flowchart illustrating the document clause change tracking method provided in Embodiment 1 of this application, as follows: Figure 1 As shown, the document terms change tracking method provided in this embodiment can be applied to a document terms change tracking platform mounted on devices with data processing capabilities, such as computers. It can be used in conjunction with some application software to achieve a better experience, and can specifically include the following steps: Step 101: Perform clause-level structured analysis on the first and second version files based on the pre-built knowledge graph to obtain the change nodes in the knowledge graph. The version of the first file is earlier than the version of the second text file.

[0014] In this step, the first version file and the second version file are two different versions of the same file. The second version is usually created by modifying some content in the earlier first version file.

[0015] This document can be, but is not limited to, a contract, an agreement, an equipment user manual, a code of conduct, etc. For ease of explanation, this embodiment will use a contract as an example.

[0016] When performing clause-level structured analysis, clause-level structured difference detection can be performed on the first version document and the second version document to obtain the changes in the second version document relative to the first version document; then, the change nodes corresponding to the changes are determined from the pre-built knowledge graph.

[0017] Among them, clause-level structured difference detection can better fit the content structure of the document in this application, making the difference comparison more consistent with the original complete semantics, improving the accuracy of difference detection, and providing a more accurate basis for subsequent tracking.

[0018] It should be noted that the changes can include changes to the terms and conditions and changes to the entities. A terms and conditions structure tree can be constructed for both the first and second version documents. Difference detection can be performed based on the terms and conditions structure trees of the first and second versions to obtain the changed terms and conditions. Legal entity identification can then be performed on the changed terms and conditions to obtain the changed entities.

[0019] For discrepancy detection, a tree-structured long short-term memory network model can be used. A semantic similarity branch for clauses is added to the rows of this network to solve the problem that traditional row-level comparison cannot identify clause reorganizations, thus achieving clause-level discrepancy detection.

[0020] In a specific example, the document can be parsed first, then a contract structure tree can be constructed. The aforementioned added branches can then be used to perform clause-level difference detection, that is, to compare similarity on a clause-by-clause basis. Finally, a structured difference report is generated. This structured difference report contains at least one changed clause, and each changed clause may, but is not limited to, include the change type, location, old text, and new text.

[0021] After obtaining the content of the amended terms, pre-built annotation standards can be used to identify and extract the legal entities involved in each amended term. For example, annotation standards can be defined to include multiple composite entities such as interest rate adjustment clauses and changes in the scope of guarantees. These annotation standards can then be used to identify and extract the entities and their relationships.

[0022] By introducing legal entity identification, the content of the changed terms is linked to the law, providing legal support for tracking the impact of the changes.

[0023] The above process yields the content of the changed terms and the changed entities, and then the corresponding change nodes are determined based on the pre-built knowledge graph.

[0024] It should be noted that the pre-constructed knowledge graph can include nodes and edges. Nodes can include legal entity nodes, judicial interpretation clause nodes, contract participant nodes, and rights and obligations nodes.

[0025] In a specific example, the legal entity node is the contractual terms that have changed, such as the "penalty interest rate clause." The judicial interpretation clause node provides the legal basis, such as a certain article in a certain law. The contractual party node is used to clarify the affected parties, such as customers, banks, guarantors, etc. The rights and obligations node is the direct object that determines the content and direction of the impact, such as the expected repayment obligation and the right to collect interest.

[0026] Edges can be defined with predefined relationships and dynamically created relationships. Predefined relationships constitute the path of influence propagation, such as reference relationships, constraint relationships, causal relationships, dependency relationships, inclusion / composition relationships, conflict / exclusion relationships, substitution relationships, strengthening / weakening relationships, etc.

[0027] In addition, dynamically created relationships refer to the edges created in subsequent processes, which are used to indicate the influence and direction of change, such as causing, aggravating, or mitigating.

[0028] In this step, the relevance between nodes in the knowledge graph and the content of the change clause can be determined one by one. This relevance can be determined based on semantic similarity and keyword matching. For example, weights can be set for semantic similarity and keyword matching respectively, and the sum of the two weighted by these weights can be used as the relevance between a certain node and the content of the change clause.

[0029] In addition, a relevance threshold can be set to identify nodes that exceed the relevance threshold as the change nodes corresponding to the content of the change clause.

[0030] Step 102: For each changed node identified, determine the candidate set of associated nodes of the changed node based on the relationships between nodes in the knowledge graph.

[0031] In the aforementioned process, the change nodes corresponding to the changed content were obtained from the knowledge graph. This step responds to the determination of various change nodes. As long as a change node is determined, this step can be executed.

[0032] Specifically, because knowledge graphs consist of nodes and edges connecting them, the associated nodes of a changed node can be determined by the connections between nodes. A node and another node are connected by an edge, meaning they have a connection.

[0033] In other words, associated nodes are other nodes that are directly or indirectly connected to the changed node through some type of edge. Therefore, this step can first determine the associated nodes that have associated edges with the changed node based on the relationships between nodes in the knowledge graph; the set of all the associated nodes obtained is then determined as the candidate set of associated nodes for the changed node.

[0034] By using a set approach, the changed node can be mapped to all associated nodes, improving the accuracy and comprehensiveness of subsequent processing.

[0035] In the example above, there are 8 predefined relationships, which means there are 8 types of edges. When determining the associated nodes in this step, we can traverse other nodes that have associated edges with the changed node. These associated edges are any of the aforementioned 8 types of edges. Having an associated edge means that the node is directly or indirectly connected to the changed node through the associated edge.

[0036] Other nodes encountered during the traversal are identified as associated nodes. All associated nodes are then included in a pre-defined set, which is the candidate set of associated nodes for the changed node.

[0037] It should be noted that each changed node will obtain a corresponding set of candidate associated nodes, and subsequent processing will be carried out on a unit basis for each changed node. This allows for parallel processing of each changed node to improve processing efficiency.

[0038] Step 103: Determine whether there are any nodes in the candidate set of associated nodes that are affected by the changed nodes based on the similarity between the associated nodes and the changed nodes.

[0039] Since the higher the similarity of nodes, the greater the probability and degree of mutual influence, this step uses similarity as the criterion for determining the influencing nodes in order to conform to this characteristic.

[0040] It should be noted that affected nodes refer to other nodes that will be affected by the changes made by the modified node. To improve the efficiency of determining affected nodes, this step can be performed directly by setting a preset threshold.

[0041] Specifically, we can first determine the similarity between the changed node and each associated node; then determine whether there are associated nodes with a similarity greater than a preset threshold. If so, we determine that the associated nodes with a similarity greater than the preset threshold are the nodes affected by the changed node, and determine that there are nodes affected by the changed node in the candidate set of associated nodes; if not, we determine that there are no nodes affected by the changed node in the candidate set of associated nodes.

[0042] When determining similarity, the aforementioned method of combining semantic similarity and keyword matching can also be used to determine relevance. Specifically, the semantic similarity between the changed node and the related node can be determined first, then the keyword matching degree between the two can be determined, and then a weighted average can be used to obtain the similarity between the related node and the changed node.

[0043] In a specific example, similarity calculation can be performed using the following formula: Sim(a,b)=α*cos(CLS vector)+β*Jaccard(keyword set), where a is a change node, b is an associated node, the CLS vector is the semantic vector of the change node and the associated node, which is a global semantic compression representation of the entire text, and the keyword set is the set of keywords extracted from the change node and the associated node respectively. In addition, α and β are weights. In a specific example, α=0.7, β=0.3, and the preset threshold is set to 0.75.

[0044] Step 104: If it exists, establish an influence edge between the changing node and the influencing node, and determine the influencing node as the new changing node.

[0045] In this step, if there are influencing nodes, in order to further track the possible deeper impacts, the influencing nodes can be identified as new change nodes, and the aforementioned steps 102 and 103 can be continued until there are no more influencing nodes.

[0046] In addition, to characterize the impact method and improve the efficiency of tracking, this step can determine the type of impact of the changed node on the affected node based on the pre-set impact rules; and establish the impact edge between the changed node and the affected node according to the impact type.

[0047] The impact type refers to how the changed node affects the affected node. It can usually be one of three categories: "cause", "aggravate", or "mitigate". Of course, there can be other required types, which can be determined by the developers as needed.

[0048] Step 105: If not, trace the changes to the second version file along all affected edges.

[0049] In this step, since the aforementioned process has already set influence edges for all the changed nodes and their affected nodes, the influence can be traced one by one from the initial changed node through these influence edges.

[0050] During tracking, an impact summary can be generated for each affected edge to reflect the impact of the change.

[0051] Specifically, this can be achieved using constraint-based text generation technology. For example, a summary template can be defined first, namely, "[Clause Name] Change results in [Affected Subject] [Rights and Obligations] [Direction of Change] (Based on [Legal Basis])".

[0052] Among them, [clause name] comes from the specific content of the change clause.

[0053] The [Affected Entities], [Rights and Obligations], and [Direction of Change] are derived from a structured impact chain obtained through knowledge graphs and impact propagation algorithms. This impact chain clearly indicates which entity (such as a customer or guarantor) will be affected by a change in a certain clause, which specific right or obligation (such as the obligation to repay late) will be affected, and whether the impact is positive (mitigation) or negative (aggravation).

[0054] [Legal basis] refers to the relevant legal provisions (such as "Article 392 of the Civil Code") that the system has pre-associated in the knowledge graph, which are automatically extracted when analyzing the impact path.

[0055] In this embodiment, a clause-level structured analysis is performed on the first and second version files based on a pre-constructed knowledge graph to obtain change nodes in the knowledge graph. The version of the first file is earlier than the version of the second file. For each change node identified, a candidate set of associated nodes is determined based on the relationships between nodes in the knowledge graph. The similarity between associated nodes and change nodes is used to determine whether there are any influencing nodes in the candidate set. If so, an influence edge is established between the change node and the influencing node, and the influencing node is identified as a new change node. If not, the changes in the second version file are tracked along all influence edges. Based on this, clause-level structured analysis, with clauses as units for subsequent identification and tracking, better conforms to the structure of the document itself, improving the accuracy of identification to a certain extent. Furthermore, by leveraging the node relationship information in the knowledge graph, the tracking of changes is achieved, improving the efficiency and comprehensiveness of change identification.

[0056] Example 2 Figure 2 This is a schematic diagram of a document clause change tracking device provided in Embodiment 2 of this application. The document clause change tracking device provided in this embodiment can execute the document clause change tracking method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of the method. This device can be implemented in software and / or hardware, such as... Figure 2 As shown, the document clause change tracking device specifically includes: analysis module 201, associated node determination module 202, affected node determination module 203, affected edge establishment module 204, and affected tracking module 205.

[0057] The analysis module is used to perform clause-level structured analysis on the first and second version files based on a pre-built knowledge graph to obtain the change nodes in the knowledge graph. The version of the first version file is earlier than the version of the second text file. The associated node determination module is used to determine the candidate set of associated nodes for each determined change node based on the relationships between nodes in the knowledge graph. The influence node determination module is used to determine whether there are influence nodes of the changed node in the candidate set of associated nodes based on the similarity between associated nodes and changed nodes. The influence edge establishment module is used to establish influence edges between the changed node and the affected node if they exist, and to identify the affected node as the new changed node. The impact tracking module is used to track changes to the second version file along all impact edges if they do not exist.

[0058] Example 3 Figure 3 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of this application, as shown below. Figure 3 As shown, the electronic device includes a processor 310, a memory 320, an input device 330, and an output device 340; the number of processors 310 in the electronic device can be one or more. Figure 3 Taking a processor 310 as an example; the processor 310, memory 320, input device 330, and output device 340 in the electronic device can be connected via a bus or other means. Figure 3 Taking the example of a connection between China and Israel via a bus.

[0059] The memory 320, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the document clause change tracking method in this embodiment of the invention. The processor 310 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 320, thereby implementing the aforementioned document clause change tracking method. Based on a pre-built knowledge graph, a clause-level structured analysis was performed on the first and second version files to obtain the change nodes in the knowledge graph. The version of the first file is earlier than the version of the second text file. For each changed node identified, a candidate set of associated nodes is determined based on the relationships between nodes in the knowledge graph. Determine whether there are any nodes affected by the changed nodes in the candidate set of associated nodes based on the similarity between associated nodes and changed nodes; If it exists, establish an influence edge between the changed node and the affected node, and identify the affected node as the new changed node; If not, trace the changes to the second version file along all affected edges.

[0060] Furthermore, based on the pre-constructed knowledge graph, a clause-level structured analysis is performed on the first and second version files to obtain the change nodes in the knowledge graph, including: Perform clause-level structured difference detection on the first and second version documents to obtain the changes in the second version document relative to the first version document; Identify the change nodes corresponding to the changed content from a pre-built knowledge graph.

[0061] Furthermore, the changes include changes to the terms and conditions and changes to the physical content; Perform clause-level structured difference detection on the first and second version documents to obtain the changes in the second version document relative to the first version document, including: The clause structure trees of the first and second versions of the file are constructed respectively, and the differences between the clause structure trees of the first and second versions are detected to obtain the content of the changed clauses. Legal entity identification is performed on the content of the amended clauses to obtain the content of the amended entity.

[0062] Furthermore, based on the relationships between nodes in the knowledge graph, a candidate set of associated nodes for the changed node is determined, including: Based on the relationships between nodes in the knowledge graph, identify the associated nodes that have associated edges with the changed node; The set of all associated nodes obtained is determined as the candidate set of associated nodes for the changed node.

[0063] Furthermore, based on the similarity between associated nodes and changed nodes, it is determined whether there are any nodes in the candidate set of associated nodes that are affected by the changed nodes, including: Determine the similarity between the changed node and each associated node; Determine whether there are related nodes with a similarity greater than a preset threshold. If so, identify the related nodes with a similarity greater than the preset threshold as the affected nodes of the changed node, and determine whether there are affected nodes of the changed node in the candidate set of related nodes. If not, it is determined that there are no affected nodes in the candidate set of associated nodes.

[0064] Furthermore, establish influence edges between changing nodes and influencing nodes, including: Determine the type of impact of a changed node on other affected nodes based on pre-defined impact rules; Establish influence edges between change nodes and affected nodes based on the influence type.

[0065] The memory 320 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data created based on terminal usage. Furthermore, the memory 320 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 320 may further include memory remotely located relative to the processor 310, which can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0066] Example 4 Embodiment 4 of this application also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform a document clause change tracking method, the method comprising: Based on a pre-built knowledge graph, a clause-level structured analysis was performed on the first and second version files to obtain the change nodes in the knowledge graph. The version of the first file is earlier than the version of the second text file. For each changed node identified, a candidate set of associated nodes is determined based on the relationships between nodes in the knowledge graph. Determine whether there are any nodes affected by the changed nodes in the candidate set of associated nodes based on the similarity between associated nodes and changed nodes; If it exists, establish an influence edge between the changed node and the affected node, and identify the affected node as the new changed node; If not, trace the changes to the second version file along all affected edges.

[0067] Furthermore, based on the pre-constructed knowledge graph, a clause-level structured analysis is performed on the first and second version files to obtain the change nodes in the knowledge graph, including: Perform clause-level structured difference detection on the first and second version documents to obtain the changes in the second version document relative to the first version document; Identify the change nodes corresponding to the changed content from a pre-built knowledge graph.

[0068] Furthermore, the changes include changes to the terms and conditions and changes to the physical content; Perform clause-level structured difference detection on the first and second version documents to obtain the changes in the second version document relative to the first version document, including: The clause structure trees of the first and second versions of the file are constructed respectively, and the differences between the clause structure trees of the first and second versions are detected to obtain the content of the changed clauses. Legal entity identification is performed on the content of the amended clauses to obtain the content of the amended entity.

[0069] Furthermore, based on the relationships between nodes in the knowledge graph, a candidate set of associated nodes for the changed node is determined, including: Based on the relationships between nodes in the knowledge graph, identify the associated nodes that have associated edges with the changed node; The set of all associated nodes obtained is determined as the candidate set of associated nodes for the changed node.

[0070] Furthermore, based on the similarity between associated nodes and changed nodes, it is determined whether there are any nodes in the candidate set of associated nodes that are affected by the changed nodes, including: Determine the similarity between the changed node and each associated node; Determine whether there are related nodes with a similarity greater than a preset threshold. If so, identify the related nodes with a similarity greater than the preset threshold as the affected nodes of the changed node, and determine whether there are affected nodes of the changed node in the candidate set of related nodes. If not, it is determined that there are no affected nodes in the candidate set of associated nodes.

[0071] Furthermore, establish influence edges between changing nodes and influencing nodes, including: Determine the type of impact of a changed node on other affected nodes based on pre-defined impact rules; Establish influence edges between change nodes and affected nodes based on the influence type.

[0072] Of course, the computer-executable instructions provided in the embodiments of this application are not limited to the above-described method operations, but can also perform related operations in the document clause change tracking method provided in any embodiment of this application.

[0073] Based on the above description of the implementation methods, those skilled in the art can clearly understand that this application can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0074] It is worth noting that in the embodiments of the above-mentioned device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of this application.

[0075] Example 5 This embodiment provides a computer program product, including a computer program / instructions, which, when executed by a processor, implements the document clause change tracking method provided in any embodiment of this application. Specifically, the method may include: Based on a pre-built knowledge graph, a clause-level structured analysis was performed on the first and second version files to obtain the change nodes in the knowledge graph. The version of the first file is earlier than the version of the second text file. For each changed node identified, a candidate set of associated nodes is determined based on the relationships between nodes in the knowledge graph. Determine whether there are any nodes affected by the changed nodes in the candidate set of associated nodes based on the similarity between associated nodes and changed nodes; If it exists, establish an influence edge between the changed node and the affected node, and identify the affected node as the new changed node; If not, trace the changes to the second version file along all affected edges.

[0076] Furthermore, based on the pre-constructed knowledge graph, a clause-level structured analysis is performed on the first and second version files to obtain the change nodes in the knowledge graph, including: Perform clause-level structured difference detection on the first and second version documents to obtain the changes in the second version document relative to the first version document; Identify the change nodes corresponding to the changed content from a pre-built knowledge graph.

[0077] Furthermore, the changes include changes to the terms and conditions and changes to the physical content; Perform clause-level structured difference detection on the first and second version documents to obtain the changes in the second version document relative to the first version document, including: The clause structure trees of the first and second versions of the file are constructed respectively, and the differences between the clause structure trees of the first and second versions are detected to obtain the content of the changed clauses. Legal entity identification is performed on the content of the amended clauses to obtain the content of the amended entity.

[0078] Furthermore, based on the relationships between nodes in the knowledge graph, a candidate set of associated nodes for the changed node is determined, including: Based on the relationships between nodes in the knowledge graph, identify the associated nodes that have associated edges with the changed node; The set of all associated nodes obtained is determined as the candidate set of associated nodes for the changed node.

[0079] Furthermore, based on the similarity between associated nodes and changed nodes, it is determined whether there are any nodes in the candidate set of associated nodes that are affected by the changed nodes, including: Determine the similarity between the changed node and each associated node; Determine whether there are related nodes with a similarity greater than a preset threshold. If so, identify the related nodes with a similarity greater than the preset threshold as the affected nodes of the changed node, and determine whether there are affected nodes of the changed node in the candidate set of related nodes. If not, it is determined that there are no affected nodes in the candidate set of associated nodes.

[0080] Furthermore, establish influence edges between changing nodes and influencing nodes, including: Determine the type of impact of a changed node on other affected nodes based on pre-defined impact rules; Establish influence edges between change nodes and affected nodes based on the influence type.

[0081] Note that the above description is merely a preferred embodiment and the technical principles employed in this application. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments. Many other equivalent embodiments may be included without departing from the concept of this application, and the scope of this application is determined by the scope of the appended claims.

Claims

1. A method for tracking changes to document terms, characterized in that, The method includes: Based on a pre-constructed knowledge graph, a clause-level structured analysis is performed on the first and second version files to obtain the change nodes in the knowledge graph. The version of the first version file is earlier than the version of the second text file. For each changed node identified, a candidate set of associated nodes is determined based on the relationships between nodes in the knowledge graph. Based on the similarity between the associated node and the changed node, determine whether there is an influence node of the changed node in the candidate set of associated nodes; If an influence edge exists, establish an influence edge between the changed node and the affected node, and determine the affected node as the new changed node; If not, trace the changes to the second version file along all the aforementioned impact edges.

2. The method according to claim 1, characterized in that, Based on a pre-constructed knowledge graph, a clause-level structured analysis is performed on the first and second version files to obtain the change nodes in the knowledge graph, including: Perform clause-level structured difference detection on the first version file and the second version file to obtain the changes in the second version file relative to the first version file; The change nodes corresponding to the changed content are determined from the pre-constructed knowledge graph.

3. The method according to claim 2, characterized in that, The changes include changes to the terms and conditions and changes to the physical content; The clause-level structured difference detection of the first and second version documents, to obtain the changes in the second version document relative to the first version document, includes: The clause structure trees are constructed for the first version file and the second version file respectively, and the difference detection is performed based on the clause structure trees of the first version file and the second version file to obtain the content of the changed clauses. Legal entity identification is performed on the content of the amended terms to obtain the amended entity content.

4. The method according to claim 1, characterized in that, The step of determining the candidate set of associated nodes of the changed node based on the relationships between nodes in the knowledge graph includes: Based on the relationships between nodes in the knowledge graph, determine the associated nodes that have associated edges with the changed node; The set of all the associated nodes obtained is determined as the candidate set of associated nodes for the changed node.

5. The method according to claim 1, characterized in that, The step of determining whether an influence node of the changed node exists in the candidate set of associated nodes based on the similarity between the associated node and the changed node includes: Determine the similarity between the changed node and each of the associated nodes; Determine whether there are related nodes with a similarity greater than a preset threshold. If so, identify the related nodes with a similarity greater than the preset threshold as the affected nodes of the changed node, and determine that the affected nodes of the changed node exist in the candidate set of related nodes. If not, it is determined that there is no affected node of the changed node in the candidate set of associated nodes.

6. The method according to claim 1, characterized in that, The establishment of the influence edge between the changed node and the affected node includes: The type of impact of the changed node on the affected node is determined based on pre-set impact rules. Establish influence edges between the changed node and the affected node based on the influence type.

7. A document clause change tracking device, characterized in that, The device includes: The analysis module is used to perform clause-level structured analysis on the first version file and the second version file based on a pre-built knowledge graph to obtain the change nodes in the knowledge graph, wherein the version of the first version file is earlier than the version of the second text file. The associated node determination module is used to determine a candidate set of associated nodes for each determined change node based on the relationships between nodes in the knowledge graph. The influence node determination module is used to determine whether there is an influence node of the changed node in the candidate set of associated nodes based on the similarity between the associated node and the changed node; The influence edge establishment module is used to establish an influence edge between the changed node and the influence node if it exists, and to determine the influence node as the new changed node; An impact tracking module is used to track changes to the second version file along all the impact edges if none exist.

8. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the document clause change tracking method as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the document clause change tracking method as described in any one of claims 1-6.

10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the document clause change tracking method as described in any one of claims 1-6.