File difference comparison method and system based on text data structuring

Through the file difference comparison method based on text data structured file data, a global mind map is generated for node matching and difference analysis, which solves the problem that the existing technology cannot clearly display file differences, and realizes multi-level and multi-perspective differences visualization, which is suitable for the difference comparison between complex and regulations documents.

CN120012753APending Publication Date: 2025-05-16HANGZHOU DIANZI UNIVERSITY SHANGYU INSTITUTE OF SCIENCE & ENGINEERING CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411991968.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing file difference comparison methods cannot clearly display the complexity and depth of information, and cannot effectively display the hierarchical and logical relationships between information, especially when comparing differential documents of regulations.

Method used

The file difference comparison method based on text data structure is adopted. By extracting the text content and regulations of the comparison files and benchmark files, a preliminary tree structure is generated, and the global mind map is obtained through the expansion of local mind maps, node matching and difference analysis are performed, and the differences types of different regulations are determined.

Benefits of technology

Through structured and graphical forms, we can deeply understand the differences between different files, provide multi-level and multi-view visualization methods, helping users observe and compare differences between files from macro to micro, and have a wider range of applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012753A_ABST
    Figure CN120012753A_ABST
Patent Text Reader

Abstract

The invention discloses a file difference comparison method and system based on text data structuralization, and the method comprises the steps: respectively extracting the text contents of a comparison file and a reference file and the regulations in the text contents, and generating a preliminary tree structure; expanding basic regulation nodes in the preliminary tree structure to obtain a global mind map; matching nodes in the reference file and the comparison file, comparing text differences of the matched nodes, obtaining a matching value of each regulation, and obtaining a difference degree score according to the matching value; judging the difference type of the nodes of the basic regulations by using the difference score; according to the matching values and the difference types of the different regulations in the comparison file and the reference file, the difference comparison between the comparison file and the reference file is completed, the text difference situation based on the text difference types between the files can be mapped in a structured and graphical form, and a user is helped to deeply understand the multi-aspect difference between the different files.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of computer science and artificial intelligence technology, and specifically relates to a file difference comparison method based on text data structuring. Background Art

[0002] Text structured generation technology converts unstructured or semi-structured text data into structured data, including text preprocessing, entity recognition, relationship extraction and other parts. With the development of machine learning technology, using machine learning models to learn features and patterns from text and extract structured data has become a new mainstream. Traditional machine learning methods are divided into supervised, semi-supervised and unsupervised. Among them, supervised relationship extraction algorithms have high accuracy, but are affected by the quality and quantity of labeled data, and cannot be extended to new relationships; semi-supervised and unsupervised methods have weak dependence on labeled data and are suitable for large-scale open domain relationship extraction, but have low accuracy and limited scalability. As deep learning-based methods begin to dominate the field of relationship extraction, more and more methods begin to use neural networks to automatically extract features from raw text, avoiding tedious manual feature engineering, and showing strong generalization capabilities when processing complex text relationships. According to the order of entity and relationship extraction, relationship extraction methods can be divided into pipeline methods and entity-relationship joint extraction methods, which usually use CNN, RNN and their improved models for relationship extraction.

[0003] Text matching technology is to evaluate and compare the semantic similarity and relevance of two texts through algorithms. Early text matching technology relied heavily on statistical methods. By fine-tuning and adding a simple classification target, text data was converted into numerical representations that are easy for computers to understand, and each word was processed to perform direct text comparisons, which can adapt to different needs and data characteristics. It can be roughly divided into string-based methods, corpus-based methods, sentence semantic interaction matching, etc. Among them, string-based methods usually match by comparing the similarity of characters or strings in the text, and are suitable for simple and direct text comparison tasks; corpus-based methods rely on large-scale corpora to extract similarities and relationships between texts, and achieve matching through statistical and probabilistic models; sentence semantic interaction matching emphasizes the semantic interaction and understanding between sentences, and improves the matching effect by modeling the interaction between sentences.

[0004] Prompt engineering is a key technology to maximize the effectiveness and accuracy of large language models. The output of the model can be significantly affected by modifying the prompts submitted to the large model. Reasonable optimization of the structure and content of the prompt can better unleash the potential of the model and enable the model to generate content that is more consistent with the unique requirements of a given scenario. Thought chain prompts are a method used in prompt engineering to enhance the reasoning ability of the model. In traditional prompt methods, the model may be directly asked to solve a complex problem without providing enough context or reasoning steps, while thought chain prompts guide the model by providing one or more intermediate reasoning steps to help it better understand the problem and construct its answer. It has been proven to significantly improve the complex reasoning and planning capabilities of large language models. However, the existing file difference comparison methods cannot clearly display the complexity and depth of information and the hierarchy and logical relationship between information when performing difference comparison. At the same time, the effect of difference comparison of regulatory documents is poor. Summary of the invention

[0005] The purpose of the present invention is to provide a file difference comparison method and system based on text data structuring.

[0006] In a first aspect, the present invention provides a file difference comparison method based on text data structuring, which comprises the following steps: Step 1: extract the text content of the comparison document and the benchmark document and the regulations in the text content respectively; Step 2: segment the text content according to the regulations, and segment the text content under each regulation to obtain the text segment set under each regulation; take each regulation and its corresponding text segment set as a node, and generate a preliminary tree structure according to the subordinate relationship of the regulations; Step 3: Take the leaf nodes of the preliminary tree structure as basic regulation nodes, generate corresponding local mind maps according to the logical relationship of the text content in the basic regulation nodes, and expand the basic regulation nodes with the local mind maps to obtain the global mind map; Step 4: Match the nodes in the reference file and the comparison file; Step 5: Segment the text content of the two matched groups of basic regulation nodes; identify the text differences between the two matched groups of basic regulation nodes, obtain the matching value of each regulation, and obtain the difference score of each basic regulation node in the global mind map of different files according to the matching value; Step 6: Determine the difference types of different basic rule nodes according to the difference scores; Step 7: Complete the difference comparison between the comparison file and the benchmark file based on the matching values ​​and difference types of different regulations in the comparison file and the benchmark file.

[0007] Preferably, the difference types include deletion, addition and modification; the method for determining the difference type to which the basic rule node belongs is as follows: The difference score of each basic regulation node is normalized. If the difference score of the basic regulation node in the comparison file is greater than or equal to the preset threshold, it is considered that the content of the regulation has been modified; if the difference score of the basic regulation node in the comparison file is less than the preset threshold, it is considered that the comparison file has added the regulation; if there is a basic regulation node in the benchmark file whose difference score is less than the preset threshold, it is considered that the comparison file lacks the regulation.

[0008] Preferably, in step 4, the method for performing node matching is as follows: Step 4-1. Use the semantic model to obtain the semantic vector e of the text content contained in each node in the global mind map of the benchmark file and the comparison file respectively; Step 4-2. Starting from the root node, traverse each node of the global mind map of the benchmark file and the comparison file hierarchically, and obtain the vector similarity S of the nodes in the global mind map of different files according to the semantic vector e of the node; continue to iteratively match the nodes in different files according to the vector similarity S until the node on the upper layer of the basic rule node is matched; Step 4-3. Obtain the node feature h of each basic rule node v in the global mind map graph (v); Step 4-4. Calculate the total similarity of different basic rule nodes in the global mind map G and the global mind map P according to the node characteristics ; Step 4-4. Based on the total similarity Complete the node matching in different files.

[0009] Preferably, the method for obtaining the vector similarity is as follows: in, is the vector similarity of the global mind map of different files at the mth level node; and are the semantic vectors of nodes in the global mind map of the benchmark file and the comparison file respectively; ; M is the number of layers of the global mind map; is the dot product symbol.

[0010] Preferably, for each basic rule node in the comparison file, a basic rule node with the highest vector similarity or total similarity to the basic rule node is selected in the benchmark file as a matching node.

[0011] Preferably, the node feature h graphThe method for obtaining (v) is as follows: in, is the weight of node i; is the feature vector of node i, ; is the semantic vector of node i; is the text embedding information of the neighboring nodes of node i, ; Embed the text information of node j, the neighboring node of node i; is the number of neighboring nodes of node i; The number of nodes in the local mind map corresponding to the basic rule node; ; .

[0012] Preferably, the total similarity The method to obtain is as follows: in, is the node feature of the basic regulation node a in the comparison document; is the node feature of the basic rule node b in the benchmark file; is the semantic vector of the basic rule node a in the comparison document; is the semantic vector of the basic rule node b in the benchmark file; ; ; A and B are the number of basic regulation nodes in the comparison file and the benchmark file respectively.

[0013] Preferably, in the step 2, dependency syntax is used to detect the set of text segments under each regulation respectively. If the subject-predicate structure of the detected text segment is incomplete, the next text segment is merged with the text segment.

[0014] Preferably, in the step 2, regular expressions are used to segment the text content according to the regulations; a sentence segmentation algorithm is used to segment the text content under each regulation, with sentence line breaks as the segmentation standard.

[0015] In the second aspect, the present invention provides a file difference comparison system based on text data structuring, which is used to process the above-mentioned file difference comparison method; the file difference comparison system includes a file management module, a file comparison module, a difference display module and a positioning index module; the file management module is used to manage comparison files and benchmark files; the file comparison module is used to generate a global mind map of benchmark files and comparison files, and compare the benchmark files and comparison files; the difference display module is used to visualize the differences between the comparison files and the benchmark files; the positioning index module is used to locate the original file according to the global mind map.

[0016] The present invention has the following beneficial effects: 1. The present invention performs text matching and difference analysis in the form of mind map node comparison, and maps the text differences between files based on text difference types in a structured and graphical form, which can help users gain an in-depth understanding of the various differences between different files.

[0017] 2. The present invention uses a multi-level, multi-perspective visualization method to assist users in observing and comparing the differences between files at different levels from macro to micro, and has a wider scope of application. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a schematic diagram of the file comparison module and the visualization display module of the file difference comparison system in the present invention.

[0019] Figure 2 It is a flow chart of the file difference comparison method in the present invention.

[0020] Figure 3 It is a schematic diagram of the pre-prompt word in the present invention.

[0021] Figure 4 This is a schematic diagram of the preliminary tree structure constructed in the present invention.

[0022] Figure 5 It is a schematic diagram of the difference between the comparison document and the reference document in the present invention. DETAILED DESCRIPTION

[0023] The present invention will be further described below in conjunction with the accompanying drawings.

[0024] like Figure 1 As shown, a file difference comparison system based on text data structure includes a file management module, a file comparison module, a difference display module and a positioning index module; the file management module is used to manage comparison files and benchmark files; the file comparison module is used to generate a global mind map of benchmark files and comparison files, and compare the benchmark files and comparison files; the difference display module is used to visualize the differences between the comparison files and the benchmark files; the positioning index module is used to locate the original file according to the global mind map.

[0025] like Figure 2 The file difference comparison method used by the file difference comparison system shown includes the following steps: Step 1: Divide the text files that need to be compared into two categories: comparison files and benchmark files; use the file management module to uniformly clean and convert the benchmark files and comparison files, and extract the text content in the benchmark files and comparison files; obtain the regulations in the text content, such as numbers (such as "Article 1", "Article 2", etc.), titles (such as "Chapter 2", "Section 1", etc.) or specific punctuation marks.

[0026] Step 2: Use regular expressions to segment the text content of the regulations, obtain the list of regulations and the text content under each regulation; remove the noise in the text content, such as irrelevant symbols, special characters, etc., and convert the text content into a unified format. Use the sentence segmentation algorithm to process the text content under each regulation, use sentence line breaks as the segmentation standard, and obtain the text segment set under each regulation; considering that in some cases, the end of the sentence may not be very obvious, use dependency syntactic analysis to determine whether the subject-predicate structure of the sentence is complete to assist in judgment.

[0027] Use dependency syntax to detect the text segments in the text segment set. If the subject-predicate structure of the detected text segment is incomplete, the next text segment is combined with the text segment. Take each regulation and its corresponding text segment set as a node, and generate a preliminary tree structure according to the subordinate relationship of the regulations.

[0028] Step 3: Figure 3 and Figure 4 As shown in the figure, the leaf nodes of the preliminary tree structure are used as basic rule nodes; in order to solve the problem that the hierarchy generated by the model is unstable due to the vague definition of the original rule hierarchy in practical applications, the required hierarchy is refined and clear output cases are provided for LLM to obtain more stable and high-quality structured data, and the pre-prompt words of LLM (large language model) are constructed; the pre-prompt words include input examples and JSON output format standards; based on the pre-prompt words, LLM is used to process the basic rule nodes, and the corresponding local mind map is generated according to the logical relationship of the text content in the basic rule nodes, and the basic rule nodes are expanded with the local mind map to obtain the global mind map.

[0029] Step 4: Node matching 4-1. Use the semantic model to encode the text content contained in each node in the global mind map G of the benchmark file and the global mind map P of the comparison file, and generate the corresponding semantic vector e; 4-2. Starting from the root node, traverse each node of the global mind map G and the global mind map P hierarchically, and calculate the vector similarity S of different nodes of the global mind map G and the global mind map P in the mth layer according to the semantic vector e of the node: in, and are the semantic vectors of the global mind map G and the global mind map P at the mth layer node respectively; ; M is the number of layers of the global mind map; is the dot product symbol.

[0030] For each node in the global mind map G, select the node with the highest vector similarity to the node in the global mind map P as the matching node, and continue to iterate the matching until the node on the upper layer of the basic rule node is matched.

[0031] 4-3. Obtain the node feature h of each basic rule node v in the global mind map graph (v), node feature h graph The expression of (v) is: in, is the weight of node i; is the feature vector of node i, ; is the semantic vector of node i; is the text embedding information of the neighboring nodes of node i, ; Embed the text information of node j, the neighboring node of node i; is the number of neighboring nodes of node i; The number of nodes in the local mind map corresponding to the basic rule node; ; .

[0032] 4-4. Calculate the total similarity of different basic rule nodes in the global mind map G and the global mind map P based on node features , whose expression is: in, is the node feature of the basic regulation node a in the comparison document; is the node feature of the basic rule node b in the benchmark file; is the semantic vector of the basic rule node a in the comparison document; is the semantic vector of the basic rule node b in the benchmark file; ; ; A and B are the number of basic regulation nodes in the comparison file and the benchmark file respectively.

[0033] For each basic rule node in the global mind map P, the basic rule node with the highest total similarity to the basic rule node is selected in the global mind map G as a matching node.

[0034] Step 5: Segment the text content of the two matched basic rule nodes. Based on the segmentation results, apply the Patience Diff algorithm to identify the text differences between the two matched basic rule nodes, obtain the matching value of each rule, and obtain the difference score of each basic rule node in the global mind map G and the global mind map P according to the matching value.

[0035] Step 6. Manually summarize the files with manually annotated differences, and divide the annotated difference types into missing (the comparison file does not have the regulation, but the benchmark file does), new (the comparison file has the regulation, but the benchmark file does not), and modified (both documents one and two have the regulation, but the content or structure of the regulation is different between the two files); further combine the text similarity score, divide the difference type into 0-1 intervals for mapping, and use the normalized interval mapping method to associate the difference score with the difference type; if the difference score of the basic regulation node in the comparison file is greater than or equal to the preset threshold, it is considered that the content or structure of the regulation is different between the comparison file and the benchmark file; if the difference score of the basic regulation node in the comparison file is less than the preset threshold, it is considered that the comparison file has added the regulation; if there is a basic regulation node with a difference score less than the preset threshold in the benchmark file, it is considered that the comparison file lacks the regulation.

[0036] Step 7: Mark all regulations in the comparison file and the differences between the contents of the regulations and the benchmark file according to the matching values ​​and difference types of different regulations in the comparison file and the benchmark file, and complete the difference comparison between the comparison file and the benchmark file. The comparison results are as follows: Figure 5 shown.

Claims

1. A file difference comparison method based on text data structure, characterized by: The following steps are involved: Step 1: extract the text content of the comparison document and the benchmark document and the regulations in the text content respectively; Step 2: segment the text content according to the regulations, and segment the text content under each regulation to obtain the text segment set under each regulation; take each regulation and its corresponding text segment set as a node, and generate a preliminary tree structure according to the subordinate relationship of the regulations; Step 3: Take the leaf nodes of the preliminary tree structure as basic regulation nodes, generate corresponding local mind maps according to the logical relationship of the text content in the basic regulation nodes, and expand the basic regulation nodes with the local mind maps to obtain the global mind map; Step 4: Match the nodes in the reference file and the comparison file; Step 5: Segment the text content of the two matched groups of basic regulation nodes; identify the text differences between the two matched groups of basic regulation nodes, obtain the matching value of each regulation, and obtain the difference score of each basic regulation node in the global mind map of different files according to the matching value; Step 6: Determine the difference types of different basic rule nodes according to the difference scores; Step 7: Complete the difference comparison between the comparison file and the benchmark file based on the matching values ​​and difference types of different regulations in the comparison file and the benchmark file.

2. The method for comparing file differences based on text data structuring according to claim 1, characterized in that: The difference types include missing, new and modified. The method for determining the difference type of the basic rule node is as follows: The difference score of each basic regulation node is normalized. If the difference score of the basic regulation node in the comparison file is greater than or equal to the preset threshold, it is considered that the content of the regulation has been modified; if the difference score of the basic regulation node in the comparison file is less than the preset threshold, it is considered that the comparison file has added the regulation; if there is a basic regulation node in the benchmark file whose difference score is less than the preset threshold, it is considered that the comparison file lacks the regulation.

3. The method for comparing file differences based on text data structuring according to claim 1, characterized in that: In the step 4, the method for performing node matching is as follows: Step 4-1. Use the semantic model to obtain the semantic vector e of the text content contained in each node in the global mind map of the benchmark file and the comparison file respectively; Step 4-2. Starting from the root node, traverse each node of the global mind map of the benchmark file and the comparison file hierarchically, and obtain the vector similarity S of the nodes in the global mind map of different files according to the semantic vector e of the node; continue to iteratively match the nodes in different files according to the vector similarity S until the node on the upper layer of the basic rule node is matched; Step 4-3. Obtain the node feature h of each basic rule node v in the global mind map graph (v); Step 4-4. Calculate the total similarity of different basic rule nodes in the global mind map G and the global mind map P according to the node characteristics ; Step 4-4. Based on the total similarity Complete the node matching in different files.

4. The method for comparing file differences based on text data structuring according to claim 3, characterized in that: The method for obtaining the vector similarity is as follows: in, is the vector similarity of the global mind map of different files at the mth level node; and are the semantic vectors of nodes in the global mind map of the benchmark file and the comparison file respectively; ; M is the number of layers of the global mind map; is the dot product symbol.

5. The method for comparing file differences based on text data structuring according to claim 3, characterized in that: For each basic rule node in the comparison file, the basic rule node with the highest vector similarity or total similarity to the basic rule node is selected in the benchmark file as the matching node.

6. The method for comparing file differences based on text data structuring according to claim 3, characterized in that: The node feature h graph The method for obtaining (v) is as follows: in, is the weight of node i; is the feature vector of node i, ; is the semantic vector of node i; is the text embedding information of the neighboring nodes of node i, ; Embed the text information of node j, the neighboring node of node i; is the number of neighboring nodes of node i; The number of nodes in the local mind map corresponding to the basic rule node; ; .

7. The method for comparing file differences based on text data structuring according to claim 3, characterized in that: The total similarity The method to obtain is as follows: in, is the node feature of the basic regulation node a in the comparison document; is the node feature of the basic rule node b in the benchmark file; is the semantic vector of the basic rule node a in the comparison document; is the semantic vector of the basic rule node b in the benchmark file; ; ; A and B are the number of basic regulation nodes in the comparison file and the benchmark file respectively.

8. The method for comparing file differences based on text data structuring according to claim 1, characterized in that: In the step 2, dependency syntax is used to detect the text segment set under each regulation respectively. If the subject-predicate structure of the detected text segment is incomplete, the next text segment is merged with the text segment.

9. The method for comparing file differences based on text data structuring according to claim 1, characterized in that: In the step 2, regular expressions are used to segment the text content according to the regulations; a sentence segmentation algorithm is used to segment the text content under each regulation, with sentence line breaks as the segmentation standard.

10. A file difference comparison system based on text data structure, characterized by: A method for processing a file difference comparison method based on text data structured as described in claim 1; the file difference comparison system comprises a file management module, a file comparison module, a difference display module and a positioning index module; the file management module is used to manage comparison files and reference files; the file comparison module is used to generate a global mind map of reference files and comparison files, and compare the reference files and comparison files; The difference display module is used to visualize the differences between the comparison file and the benchmark file; the positioning index module is used to locate the original file according to the global mind map.