File comparison method, device and equipment and computer readable storage medium
By obtaining the document object model tree of XML files and based on node comparison algorithms, the problem of limitation of accuracy and inefficiency caused by the comparison of file comparison tools in the existing technology is solved, and a more accurate and efficient display of XML files is achieved.
Patent Information
- Application Number
- CN202311466403.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-03
- Publication Date
- 2025-05-06
AI Technical Summary
The existing file comparison tools are compared in behavior units, which have problems such as limited accuracy, structural insensitiveness and inefficiency, especially when processing XML files, it is difficult to accurately display the differences between files.
By obtaining the document object model tree of the XML file to be compared, obtaining the node collection at the same level, and determining the editing plan and the editing value of the scheme based on the node's comparison algorithm, displaying the difference information between the files.
It improves the accuracy and efficiency of file comparison, can better handle the particularity of XML language structure, avoid unnecessary interference, and intuitively display the specific locations of modifications and differences.
Smart Images

Figure CN119940335A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data analysis technology, and in particular to a file comparison method, device, electronic device and computer-readable storage medium. Background Art
[0002] Extensible Markup Language (XML) is a markup language tool for storing and transmitting data. It uses custom tags to identify and organize data, allowing users to customize tags and tag structures to adapt to different application requirements and data formats. XML has the characteristics of scalability, readability and cross-platform. It can be used for a variety of purposes such as data exchange, configuration files and data storage. File comparison is an important research field in the field of computer science, which is used to compare and analyze the differences between two or more files. It plays a role in version control, file synchronization and content analysis. File comparison tools for XML use background technologies such as parsing technology, comparison algorithms and difference extraction.
[0003] The applicant found in the specific implementation process that all the file comparison tools on the market currently compare files in line units, and there is no comparison tool specifically for XML. The traditional method of comparing files in line units has the disadvantages of limited accuracy, insensitivity to structure, and low efficiency. Summary of the invention
[0004] Embodiments of the present application provide a file comparison method, device, electronic device, computer-readable storage medium, and computer program product, which can accurately display the differences between files.
[0005] The technical solution of the embodiment of the present application is implemented as follows:
[0006] The present application provides a file comparison method, the method comprising:
[0007] Acquire a first file and a second file to be compared, parse the first file to obtain a first document object model tree, and parse the second file to obtain a second document object model tree;
[0008] Acquire a first node set from the first document object model tree, and acquire a second node set from the second document object model tree, wherein a first node in the first node set and a second node in the second node set are located at the same level;
[0009] Determine at least one editing scheme between the first node set and the second node set, and determine a scheme editing value corresponding to each editing scheme;
[0010] Determine a minimum scheme editing value from the scheme editing values corresponding to the respective editing schemes;
[0011] When the minimum scheme editing value indicates that there are differences between the first node set and the second node set, difference information between the first file and the second file is displayed based on the target editing scheme corresponding to the minimum scheme editing value.
[0012] The present application provides a file comparison device, the device comprising:
[0013] A file parsing module, used to obtain a first file and a second file to be compared, parse the first file to obtain a first document object model tree, and parse the second file to obtain a second document object model tree;
[0014] an acquisition module, configured to acquire a first node set from the first document object model tree, and acquire a second node set from the second document object model tree, wherein a first node in the first node set and a second node in the second node set are located at the same level;
[0015] A first determining module, configured to determine at least one editing scheme between the first node set and the second node set, and determine a scheme editing value corresponding to each editing scheme;
[0016] A second determining module, configured to determine a minimum scheme editing value from the scheme editing values corresponding to the respective editing schemes;
[0017] The difference display module is used to display the difference information between the first file and the second file based on the target editing scheme corresponding to the minimum scheme editing value when the minimum scheme editing value indicates that there is a difference between the first node set and the second node set.
[0018] An embodiment of the present application provides an electronic device for file comparison, the electronic device comprising:
[0019] A memory for storing computer executable instructions;
[0020] The processor is used to implement the file comparison method provided in the embodiment of the present application when executing the computer executable instructions stored in the memory.
[0021] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the file comparison method provided in the embodiment of the present application when executed by a processor.
[0022] An embodiment of the present application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, the file comparison method provided in the embodiment of the present application is implemented.
[0023] The embodiments of the present application have the following beneficial effects:
[0024] In an embodiment of the present application, when it is necessary to compare the difference between a first file and a second file, the first file and the second file are first parsed respectively to obtain a first document object model tree and a second document object model tree, and then a first node set and a second node set located at the same level are obtained from the first document object model tree and the second document object model tree. Through a node-based comparison algorithm, at least one editing scheme between the first node set and the second node set is determined, and the scheme editing value corresponding to each editing scheme is determined. When it is determined that there is a difference between the first file and the second file based on the minimum scheme editing value, the difference information between the first file and the second file is presented. In an embodiment of the present application, since the comparison is made with node sets located at the same level, that is, the comparison is based on the semantic information of the file, and the differences in formal issues between lines are not considered (extra or less carriage returns and spaces, different order of attributes, etc.), the accuracy of file comparison can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 This is the result of using the Beyond Compare file comparison tool to compare files in rows.
[0026] Figure 2 is a schematic diagram of the architecture of a file comparison system 100 provided in an embodiment of the present application;
[0027] Figure 3 is a schematic diagram of the structure of the terminal 400 provided in an embodiment of the present application;
[0028] Figure 4A This is a schematic diagram of an implementation flow of the file comparison method provided in an embodiment of the present application;
[0029] Figure 4B A schematic diagram of an implementation flow of determining at least one editing scheme provided in an embodiment of the present application;
[0030] Figure 4C A schematic diagram of a process for implementing the determination of the first editor-in-chief value provided in an embodiment of the present application;
[0031] Figure 4D A schematic diagram of a process flow for determining a second editor-in-chief value is provided for an embodiment of the present application;
[0032] Figure 4EA schematic diagram of a process for determining a third editor-in-chief value is provided for an embodiment of the present application;
[0033] Figure 5 A schematic diagram of another implementation flow of the file comparison method provided in an embodiment of the present application;
[0034] Figure 6 A schematic diagram of determining the difference value of two groups of different nodes provided in an embodiment of the present application;
[0035] Figure 7 A schematic diagram of an interface for dragging an XML file to be compared to a comparison tool provided in an embodiment of the present application;
[0036] Figure 8 A schematic diagram of a file comparison result determined by a file comparison method provided by an embodiment of the present application;
[0037] Fig. 9 This is a schematic diagram of file comparison results determined using the Beyond Compare method in the related art. DETAILED DESCRIPTION
[0038] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.
[0039] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0040] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0041] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0042] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0043] 1) Extensible Markup Language (XML): A markup language used to describe and save data. It allows users to customize tags and elements to represent data in specific fields or applications. XML uses a set of tags to define the structure and semantics of data, as well as the relationship between tags, so that data can be exchanged and shared between different systems and platforms.
[0044] 2) Beyond Compare: A professional file and folder comparison tool. It can display the differences between files or folders by comparing file contents, file attributes, and folder structure, and provides advanced merging functions. Beyond Compare can provide flexible configuration options when comparing and merging various file types, and provides a friendly user interface that allows users to intuitively compare and process file differences. See Figure 1 , Figure 1 This is the result of using the Beyond Compare file comparison tool to compare files in rows.
[0045] 3) Document Object Model (DOM) tree: A programming interface for representing and manipulating XML or HTML document structures. It treats each element, attribute, text, and comment in a document as an object and organizes them in a hierarchical manner to form a tree structure. The DOM tree can be accessed and manipulated by programming languages (such as JavaScript), and developers can use the methods and properties provided by DOM to add, delete, modify, and query any part of the document.
[0046] 4) Dynamic Programming Algorithm (DP): A mathematical optimization method for solving multi-stage decision problems. It decomposes the problem into several sub-problems and uses the optimal solutions of the sub-problems to construct a global optimal solution. The dynamic programming algorithm usually improves computational efficiency by creating state transition equations and using memoization technology. In an embodiment of the present application, the dynamic programming algorithm is used to calculate the difference value of the node matching scheme to find the optimal solution that matches as many identical nodes as possible.
[0047] 5) Exhaustive matching: An algorithm or method for searching and matching that determines the best or satisfactory match by trying all possible combinations or choices one by one. In search and matching problems, exhaustive matching traverses the elements, conditions, or rules to be matched and tries all possible combinations or choices to find the best match that meets specific requirements. This matching usually involves comparing, evaluating, or calculating indicators, criteria, or scores for each possible match until the optimal solution or a solution that meets specific conditions is found.
[0048] 6) SIFT4 algorithm (String Similarity-Focused Technique 4, SIFT4): An algorithm for calculating string similarity. It is an improved version of string similarity calculation based on edit distance. It includes steps such as string preprocessing, skip threshold setting, character scanning, character comparison, skip operation and edit distance calculation. The preprocessing operation ensures string consistency and comparability, while the skip threshold setting speeds up the execution. The character scanning compares the characters in the source string and the target string in order and determines the differences based on a series of character comparison rules. When the match fails, the algorithm uses skip operations to try to find a better match. By calculating the number of skip operations, the edit distance calculation obtains the string difference value to measure the string similarity.
[0049] The existing file comparison tools use rows as the basic elements for comparison, which makes it difficult to accurately present the differences between XML structures. For example, the extra spaces between nodes and the order of attribute fields in an XML file are equivalent in XML semantics, but using a row-based comparison method will display these semantically irrelevant modifications. In addition, in actual projects, an XML node may have dozens of attribute fields, and row-based comparison tools find it difficult to display these differences at the same time, so they are inefficient when processing large XML files.
[0050] In order to solve these problems, the present application embodiment proposes a file comparison method. The method uses the DOM tree to represent the structure of the XML document and uses nodes as the basic unit to perform difference calculation, so as to accurately and comprehensively reflect the differences of the XML documents. Compared with the traditional line-based comparison tools, this method can better handle the particularity of the XML language structure, avoid unnecessary interference, and can also more efficiently handle the difference calculation of large XML files.
[0051] The embodiments of the present application provide a file comparison method, device, electronic device, computer-readable storage medium and computer program product, which can provide clearer, more accurate and comprehensive file difference information.
[0052] The following describes an exemplary application of the electronic device provided by the embodiment of the present application. The electronic device provided by the embodiment of the present application can be implemented as various types of user terminals such as a laptop computer, a tablet computer, a desktop computer, a set-top box, a mobile device (for example, a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable gaming device), a smart phone, a smart speaker, a smart watch, a smart TV, and a vehicle-mounted terminal, and can also be implemented as a server. The following describes an exemplary application when the device is implemented as a terminal.
[0053] See also Figure 2 , Figure 2 It is a schematic diagram of the architecture of the file comparison system 100 provided in an embodiment of the present application. To support an exemplary application, the terminal 400 is connected to the server 200 via the network 300. The network 300 may be a wide area network or a local area network, or a combination of the two.
[0054] The terminal 400 is used to display the first file and the second file on the graphical interface 410, and the server 200 is used to store the first file and the second file. The terminal 400 obtains the first file and the second file to be compared from the server 200, and parses the first file and the second file respectively, and obtains the first document object model tree and the second document object model tree accordingly. Then, the first node set is extracted from the first document object model tree, and the second node set is extracted from the second document object model tree. The nodes in the first node set and the second node set are at the same level. Next, at least one editing scheme and the scheme editing value corresponding to each editing scheme are determined. The minimum value is found from the scheme editing values of all editing schemes. When the minimum scheme editing value indicates that there is a difference between the first node set and the second node set, the difference information between the first file and the second file is displayed according to the target editing scheme corresponding to the minimum scheme editing value, so that the user can more accurately understand the modification differences between the files.
[0055] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. The terminal 400 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a car terminal, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.
[0056] See also Figure 3 , Figure 3 is a schematic diagram of the structure of the terminal 400 provided in an embodiment of the present application, Figure 3 The terminal 400 shown includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the terminal 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figure 3 Various buses are labeled as bus system 440 .
[0057] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0058] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0059] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.
[0060] The memory 450 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0061] In some embodiments, memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.
[0062] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0063] A network communication module 452, used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 include: Bluetooth, Wireless Compatibility Certification (WiFi), and Universal Serial Bus (USB), etc.;
[0064] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., display screen, speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripherals and displaying content and information);
[0065] The input processing module 454 is used to detect one or more user inputs or interactions from one of the one or more input devices 432 and translate the detected inputs or interactions.
[0066] In some embodiments, the device provided in the embodiments of the present application can be implemented in software. Figure 3 The file comparison device 455 stored in the memory 450 is shown, which can be software in the form of a program and a plug-in, etc., and includes the following software modules: a file parsing module 4551, an acquisition module 4552, a first determination module 4553, a second determination module 4554 and a difference display module 4555. These modules are logical, and therefore can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.
[0067] In other embodiments, the device provided in the embodiments of the present application can be implemented in hardware. As an example, the device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the file comparison method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application specific integrated circuits (Application Specific Integrated Circuit, ASIC), digital signal processors (Digital Signal Processor, DSP), programmable logic devices (Programmable Logic Device, PLD), complex programmable logic devices (Complex Programmable Logic Device, CPLD), field programmable gate arrays (Field-Programmable Gate Array, FPGA) or other electronic components.
[0068] As mentioned above, the file comparison method provided in the embodiments of the present application can be implemented by various types of electronic devices. Figure 4A , Figure 4A This is a schematic diagram of an implementation flow of the file comparison method provided in the embodiment of the present application, combined with Figure 4A The steps shown are explained, Figure 4A The body of a step is the terminal.
[0069] In step 101, a first file and a second file to be compared are obtained, the first file is parsed to obtain a first document object model tree, and the second file is parsed to obtain a second document object model tree.
[0070] In some embodiments, the first file and the second file can be obtained by the terminal from a local storage space, or from a server or other terminal, or the first file can be obtained from the local storage space of the terminal and the second file can be obtained from a server or other terminal, or the first file can be obtained from a server or other terminal and the second file can be obtained from the terminal locally. The embodiment of the present application does not limit the acquisition source of the first file and the second file. Wherein the first file and the second file are both XML files. After obtaining the first file and the second file to be compared, it is necessary to parse the first file to obtain the first document object model tree, and parse the second file to obtain the second document object model tree. For example, there are two XML files, namely the first file and the second file. The user wants to compare the two files and find the differences in their contents. First, read and load the two files to obtain the file contents. Then, use an XML parser or a related library to parse the first file into a document object model tree, and then parse the second file into another document object model tree.
[0071] Among them, you can, but are not limited to, use built-in libraries such as the javax.xml.parsers package in Java and the xml.dom.minidom module in Python to parse XML files; or use third-party libraries to parse XML files. For example, in Java, commonly used libraries include Jsoup, Xerces, DOM4J, etc., and in Python, commonly used libraries include xml.etree.ElementTree, lxml, etc.; or use an XML parser such as a SAX (Simple API for XML) parser; or use a programming language developed by the developer himself to parse XML files.
[0072] The document object model tree is used to represent the object model of an XML document, and includes different types of nodes such as root nodes, element nodes, text nodes, attribute nodes, and comment nodes. The root node represents the topmost node of the DOM tree and represents the root element of the entire document; the element node represents the tag element in the XML document, and these nodes can have child nodes and attributes; the text node represents the text content in the ML document; the attribute node represents the attribute of the element in the XML document; and the comment node represents the comment content in the XML document. These nodes form a tree representation of a hierarchical structure, and each node represents a specific document element or content. The structures of the first document object model tree and the second document object model tree are basically similar, both of which contain a root node and various types of nodes. The difference between them is that they represent different documents.
[0073] In step 102, a first node set is obtained from a first document object model tree, and a second node set is obtained from a second document object model tree.
[0074] In some embodiments, after parsing the first document object model tree and the second document object model tree, the terminal traverses the first document object model tree and the second document object model tree, taking all nodes at the same level in the first document model tree as a first node set, and taking all nodes at the same level in the second document model tree as a second node set. For example, after reading and loading two XML files and parsing them into corresponding document object model trees, a recursive traversal method is used to traverse the DOM tree corresponding to the first XML file, starting from the root node of the DOM tree corresponding to the first XML file, each node and its child nodes are visited in turn, and the traversed nodes are formed into a first node set. Then, a depth-first traversal method is used to traverse the DOM tree corresponding to the second XML file, starting from the root node of the DOM tree corresponding to the second XML file, each node and its child nodes are visited in turn, and the traversed nodes are formed into a second node set.
[0075] The traversal of the first document object model tree and the second document object model tree may be performed by, but is not limited to, depth-first traversal, breadth-first traversal, pre-order traversal, in-order traversal or post-order traversal.
[0076] In step 103, at least one editing scheme between the first node set and the second node set is determined, and scheme editing values corresponding to each editing scheme are determined.
[0077] In some embodiments, in two given XML files, the first node set and the second node set refer to the set of each node and its child nodes in the DOM trees corresponding to the two files. It is determined that there is at least one editing scheme between the two node sets, and the editing scheme refers to a specific method for modifying, adding or deleting nodes in the first file or the second file. Each editing scheme is assigned a scheme editing value, which is used to calculate the difference value of the node content in the editing scheme.
[0078] In some embodiments, see Figure 4B , Figure 4B A schematic diagram of an implementation flow of determining at least one editing scheme provided in an embodiment of the present application, Figure 4A Step 103 shown may be performed as follows Figure 4B Steps 1031 to 1035 shown in the figure are implemented by combining Figure 4B Specific instructions.
[0079] Step 1031: Based on a dynamic programming algorithm, exhaustively match each first node in the first node set with each second node in the second node set to obtain at least one editing scheme.
[0080] The first node set is edited according to an editing scheme to obtain a second node set; the editing scheme includes at least one of a first node to be deleted, a second node to be added, and modifying a first target node to a second target node.
[0081] In some embodiments, based on a dynamic programming algorithm, each first node in the first node set can be matched with each second node in the second node set by exhaustive matching, thereby obtaining at least one editing scheme. In this process, a dynamic programming algorithm is used to solve the optimal matching problem. Specifically, the first node set is edited according to the editing scheme to obtain the second node set. The editing scheme includes three operations: deleting the first node to be deleted, adding the second node to be added, or modifying the first target node to the second target node.
[0082] Step 1032: When the editing scheme only includes the first node to be deleted, determine a first total editing value for deleting the first node, and determine the first total editing value as the scheme editing value corresponding to the editing scheme.
[0083] In some embodiments, see Figure 4C , Figure 4C A schematic diagram of a process for determining a first editor-in-chief value according to an embodiment of the present application is provided. Figure 4B The step of "determining the first total edit value of deleting the first node" in step 1032 shown in the figure can be performed as follows Figure 4C Steps 321 to 325 shown are implemented as follows:
[0084] Step 321: Determine whether the first node has a first child node.
[0085] In some embodiments, a depth-first search algorithm is used to traverse from the root node to check whether the root node has a left child node. If so, it is determined that the root node has a left child node. Otherwise, the right child node of the root node, the left child node of the right child node, etc. will continue to be recursively traversed until the left child node is found or the traversal ends. In this way, it is possible to accurately determine whether the first node has a first child node, and perform corresponding operations as needed.
[0086] When the first node does not have a first child node, the process proceeds to step 322 ; and when the first node has a first child node, the process proceeds to step 324 .
[0087] Step 322: Obtain the first node edit value for deleting the first node.
[0088] In some embodiments, when obtaining the first node edit value of the first node, in order to avoid repeated calculations, it is first determined whether the first node edit value of deleting the first node is stored in the cache space. When the first node edit value of deleting the first node is not stored in the cache space, it is necessary to determine the first node edit value of deleting the first node and store the first node edit value in the cache space. When the first node edit value of deleting the first node is stored in the cache space, it is not necessary to repeatedly calculate the first node edit value, and the first node edit value can be directly obtained from the cache space. Among them, the first node edit value of deleting the first node needs to use the SIFT4 algorithm to calculate the first node edit value of deleting the first node.
[0089] In some embodiments, when determining to delete the first node edit value of the first node, the first string corresponding to the first node can be obtained based on the first file, and then the edit value of deleting the first string can be determined, and the edit value of deleting the first string can be determined as the first node edit value.
[0090] Step 323: Determine the first node edit value as the first overall edit value.
[0091] In some embodiments, when the first node does not have a first child node, the first total edit value is determined only by the first node edit value, which represents the effect of the first node edit operation of deleting the first node. Specifically, when the first node does not have a first child node, the first total edit value includes the edit operation of deleting the first node. For example, there is an XML file that contains a node tree structure, and the first node is , it has no child nodes. Now we want to determine the first node edit value of the first node that deleted this file. First, we check whether the first node edit value of the first node that deleted the first node is stored in the cache space, and find that the cache space is empty. Then, we get the first string corresponding to the first node from the first file, that is, . Next, the SIFT4 algorithm is used to calculate the edit value of deleting the first string, assuming that the calculated edit value is B. The edit value B is stored in the cache space for future use. Finally, the first node edit value B is determined as the first total edit value. Through the above steps, the first file and the second file to be compared are successfully obtained, and the first node edit value of deleting the first node is determined and used as the first total edit value.
[0092] Step 324: Determine to delete the first child node edit value of the first child node, and determine to delete the first node edit value of the first node.
[0093] In some embodiments, first, obtain the first child node of the first node. In the node tree structure, the first child node is all the child nodes of the first node. When determining to delete the first child node edit value of the first child node, similar to step 322, first determine whether the first child node edit value for deleting the first child node is stored in the cache space. If the first child node edit value for deleting the first child node is stored in the cache space, then directly obtain it from the cache space. If the first child node edit value is not stored in the cache space, then obtain the first string corresponding to the first child node. This string includes but is not limited to the tag name, attributes, and text content information of the node. Then, use the SIFT4 algorithm to calculate the first child node edit value for deleting the first child node.
[0094] When determining to delete the first node edit value of the first node, the implementation process of step 322 may be referred to.
[0095] Step 325: Determine the sum of the first subnode edit value and the first node edit value as the first section total edit value.
[0096] In some embodiments, the first child node edit value is summed with the first node edit value to obtain the first section total edit value, and the first section total edit value is stored for subsequent use. For example, suppose there is an XML file containing a node tree structure, and the first node is <parent>, which has a <child>In order to determine the first child node edit value of the first child node to be deleted, and thereby determine the first node edit value of the first node to be deleted, the following steps may be performed: First, obtain the first node <parent>The first child node of <child>. Next, from the first child node <child>Then, use the SIFT4 algorithm to calculate and delete the first child node. <child>The edit value of the first child node. Assuming that the edit value of the first child node is 1 and the edit value of the first node is 3, the first child node edit value 1 and the first node edit value 3 are calculated to obtain the total edit value of the first section: del( <parent>)+del( <child>)=1+3=4. The total edit value 4 of the first section of this node is stored in the cache space for future use. Through the above steps, it is successfully determined that the first child node is deleted. <child>The first child node edit value, and based on this, it is determined to delete the first node <parent>Then, the sum of the first child node edit value and the first node edit value 4 is determined as the total edit value of the first section.
[0097] Step 1033: When the editing scheme includes only the second node to be added, determine a second total editing value for adding the second node, and determine the second total editing value as the scheme editing value corresponding to the editing scheme.
[0098] In some embodiments, see< / parent> < / child> < / child> < / parent> < / child> < / child> < / child> < / parent> < / child> < / parent> Figure 4D , Figure 4D A schematic diagram of a process for determining a second editor-in-chief value is provided for an embodiment of the present application. Figure 4B The step of "determining to increase the second total edit value of the second node" in step 1033 shown in FIG. 1034 can be performed as follows: Figure 4D Steps 331 to 334 shown are implemented as follows:
[0099] Step 331: Determine whether the second node has a second child node.
[0100] In some embodiments, the method for determining whether the second node has a second child node is similar to the implementation process of step 321. When the second node does not have a second child node, the process proceeds to step 332; when the second node has a second child node, the process proceeds to step 334.
[0101] Step 332: Obtain the second node edit value of the added second node.
[0102] In some embodiments, when obtaining the second node edit value of the second node, in order to avoid repeated calculations, it is first determined whether the cache space stores the second node edit value of the second node. When the cache space does not store the second node edit value of the second node, it is necessary to determine the second node edit value of the second node and store the second node edit value in the cache space. When the cache space stores the second node edit value of the second node, it is not necessary to repeatedly calculate the second node edit value, and the second node edit value can be directly obtained from the cache space. Among them, the second node edit value of the second node needs to be calculated using the SIFT4 algorithm to add the second node edit value of the second node.
[0103] In some embodiments, when determining to increase the second node editing value of the second node, the first string corresponding to the second node can be obtained based on the second file, and then the editing value of the first string can be determined to be increased, and the editing value of the first string can be determined as the second node editing value.
[0104] Step 333: Determine the second node edit value as the second overall edit value.
[0105] In some embodiments, when the second node does not have a second child node, the second total edit value is determined only by the second node edit value. Based on the determination of whether the second node has a second child node, the SIFT4 algorithm is used to calculate and obtain the second node edit value of the added second node or to obtain the edit value from the cache space, and the edit value is determined as the second total edit value.
[0106] Step 334: Determine to increase the second child node edit value of the second child node, and determine to increase the second node edit value of the second node.
[0107] In some embodiments, first, the second child node of the second node is obtained. In the node tree structure, the second child node is all the child nodes under the second node, that is, the second node connected to the second node. When determining the second child node edit value of adding the second child node, similar to step 332, first determine whether the second child node edit value of adding the second child node is stored in the cache space. If the second child node edit value is stored in the cache space, then it can be directly obtained from the cache space. If the second child node edit value is not stored in the cache space, then the first string corresponding to the second child node is obtained. This string includes but is not limited to the tag name, attributes and text content information of the node. Then, the SIFT4 algorithm is used to calculate the second child node edit value of adding the second child node.
[0108] When determining to increase the second node edit value of the second node, the implementation process of step 332 may be referred to.
[0109] Step 335: The sum of the second subnode edit value and the second node edit value is determined as the second section total edit value of the second node.
[0110] In some embodiments, when the edit value of the second subnode and the edit value of the second node are to be added, the edit value of the second subnode and the edit value of the second node are first obtained from step 334, and then the two edit values are summed. The edit values are numbers, and they are directly added. Finally, the sum of the edit value of the second subnode and the edit value of the second node is obtained, and this sum is the second total edit value of the second node.
[0111] Step 1034: When the editing scheme only includes modifying the first target node to the second target node, determine a third total editing value for modifying the first target node to the second target node, and determine the third total editing value as the scheme editing value corresponding to the editing scheme.
[0112] In some embodiments, see Figure 4E , Figure 4E A schematic diagram of a process for determining a third editor-in-chief value is provided for an embodiment of the present application. Figure 4BThe step of "determining to modify the first target node to the third total edit value of the second target node" in step 1034 shown in the figure can be performed as follows: Figure 4E Steps 341 to 3412 shown are implemented as follows:
[0113] Step 341: Determine whether the first target node has a first target child node.
[0114] In some embodiments, when the first target node does not have a first target child node, the process proceeds to step 342 ; and when the first target node has a first target child node, the process proceeds to step 347 .
[0115] Step 342: determine whether the second target node has a second target child node.
[0116] In some embodiments, when the second target node does not have a second target child node, the process proceeds to step 343 ; when the second target node has a second target child node, the process proceeds to step 345 .
[0117] Step 343: Obtain a second character string corresponding to the first target node from the first file, and obtain a third character string corresponding to the second target node from the second file.
[0118] In some embodiments, the first file is opened and the content is read through the file operation function provided by the computer programming language. Then, according to the identifier of the first target node, the second string corresponding to the first target node is obtained in the first file, and the content in the first file is traversed according to the identifier of the first target node to locate the part containing the target node information. After finding the target node, according to the format and structure of the file, the node is parsed using appropriate methods and rules, and the target node is located according to the identifier, and then the second string is extracted and obtained. The specific method can be an XML parser. Then, an operation similar to obtaining the second string is performed to obtain the third string corresponding to the second target node from the second file.
[0119] Step 344: Modify the second character string to the edit value of the third character string, and determine it as the third overall edit value.
[0120] In some embodiments, it is first determined whether an edit value for modifying the second string to the third string is stored in the cache space. If the edit value is stored in the cache space, the edit value is directly obtained from the cache space, and the edit value is determined as the third total edit value. If the edit value is not stored in the cache space, the SIFT4 algorithm can be used to determine the edit value for modifying the second string to the third string, the determined edit value is determined as the third total edit value, and the edit value is stored in the cache space.
[0121] Step 345: Obtain the eleventh character string corresponding to the second target node and the twelfth character string corresponding to the second target sub-node from the second file, and concatenate the eleventh character string and the twelfth character string to obtain a fourth combined character string.
[0122] In some embodiments, the eleventh character string corresponding to the second target node and the twelfth character string corresponding to the second target child node may be obtained from the second file using a process similar to step 343. After obtaining the eleventh character string and the twelfth character string, they may be concatenated using a string concatenation function, that is, the two character strings are concatenated together to generate a fourth combined character string.
[0123] Step 346: Obtain the thirteenth character string corresponding to the first target node from the first file, modify the thirteenth character string to the edit value of the fourth combined character string, and determine it as the third total edit value.
[0124] In some embodiments, a process similar to step 343 can be used to obtain the thirteenth character string corresponding to the first target node from the first file, and then determine to modify the thirteenth character string to an edit value of the fourth combined character string, and determine the edit value as the third total edit value. Among them, the first file can be read, and the structure of the first file can be parsed to traverse the nodes in the file to find a node that matches the first target node. Once the first target node is found, the content of the node will be further parsed to extract the thirteenth character string. The fourth combined character string obtained by concatenating the eleventh character string and the twelfth character string is modified to the edit value of the fourth combined character string, which means that the third total edit value is the edit result of the fourth combined character string, which is used to represent the edit result of the first target node.
[0125] Step 347: Determine whether the second target node has a second target child node.
[0126] In some embodiments, when the second target node does not have a second target child node, the process proceeds to step 348 ; when the second target node has a second target child node, the process proceeds to step 3410 .
[0127] Step 348: Obtain from the first file a fourth character string corresponding to the first target node and a fifth character string corresponding to the first target sub-node, and concatenate the fourth character string and the fifth character string to obtain a first combined character string.
[0128] In some embodiments, when the second target node does not have a second target child node, a fourth string corresponding to the first target node and a fifth string corresponding to the first target child node are obtained from the first file. After obtaining the fourth string and the fifth string, they can be concatenated using a string concatenation function, that is, the two strings are concatenated together to generate a first combined string.
[0129] Step 349: Obtain the sixth character string corresponding to the second target node from the second file, modify the first combined character string to the edit value of the sixth character string, and determine it as the edit value of the third node.
[0130] In some embodiments, it is first determined whether an edit value for modifying the first combination string to the sixth string is stored in the cache space. If the edit value is stored in the cache space, the edit value is directly obtained from the cache space, and the edit value is determined as the third node edit value. If the edit value is not stored in the cache space, the SIFT4 algorithm can be used to determine the edit value for modifying the first combination string to the sixth string, the determined edit value is determined as the third node edit value, and the edit value is stored in the cache space.
[0131] Step 3410: Obtain the seventh character string corresponding to the first target node and the eighth character string corresponding to the first target sub-node from the first file, and concatenate the seventh character string and the eighth character string to obtain a second combined character string.
[0132] In some embodiments, when the first target node has a first target child node, and the second target node has a second target child node, a seventh string corresponding to the first target node and an eighth string corresponding to the first target child node can be obtained from the first file using a process similar to step 349. After obtaining the seventh string and the eighth string, they can be concatenated using a string concatenation function, that is, the two strings are concatenated together to generate a second combined string.
[0133] Step 3411: Obtain the ninth character string corresponding to the second target node and the tenth character string corresponding to the second target sub-node from the second file, and concatenate the ninth character string and the tenth character string to obtain a third combined character string.
[0134] In some embodiments, when the first target node has a first target child node, and the second target node has a second target child node, a ninth character string corresponding to the second target node and a tenth character string corresponding to the second target child node can be obtained from the second file using a process similar to step 349. After obtaining the ninth character string and the tenth character string, they can be concatenated using a string concatenation function, that is, the two character strings are concatenated together to generate a third combined character string.
[0135] Step 3412: Modify the second combined character string to the edit value of the third combined character string, and determine it as the third total edit value.
[0136] In some embodiments, it is first determined whether an edit value for modifying the second combination string into a third combination string is stored in the cache space. If the edit value is stored in the cache space, the edit value is directly obtained from the cache space, and the edit value is determined as the third total edit value. If the edit value is not stored in the cache space, the SIFT4 algorithm can be used to determine the edit value for modifying the second combination string into the third combination string, the determined edit value is determined as the third total edit value, and the edit value is stored in the cache space.
[0137] Step 1035: When the editing scheme includes at least two of a first node to be deleted, a second node to be added, a second node to be added, and modifying the first target node to the second target node, the sum of at least two of the corresponding first total editing value, second total editing value, and third total editing value is determined as the scheme editing value corresponding to the editing scheme.
[0138] In some embodiments, when the editing scheme includes at least two operations of deleting a first node, adding a second node, and modifying a first target node to a second target node, exemplarily, when the editing scheme includes a first node to be deleted and a second node to be added, the sum of the first total editing value and the second total editing value is determined as the scheme editing value corresponding to the editing scheme.
[0139] It should be noted that if there are multiple first nodes to be deleted or multiple second nodes to be added in the editing scheme, then it is necessary to accumulate the first total editing values corresponding to the multiple first nodes to be deleted, or accumulate the second total editing values corresponding to the multiple second nodes to be added. Exemplarily, the editing scheme includes node A to be deleted, node B to be deleted, node C to be added, and node D to be added, then the editing value of the editing scheme is the sum of the first total editing value of deleting node A, the first total editing value of deleting node B, the second total editing value of adding node C, and the second total editing value of adding node D.
[0140] In step 104, a minimum scheme editing value is determined from the scheme editing values corresponding to the various editing schemes.
[0141] In some embodiments, the scheme editing values corresponding to the various editing schemes are compared to determine a minimum scheme editing value, and the minimum scheme editing value may be a real number greater than or equal to 0.
[0142] In step 105, when the minimum edit scheme value indicates that there is a difference between the first node set and the second node set, the difference information between the first file and the second file is displayed based on the target edit scheme corresponding to the minimum edit scheme value.
[0143] In some embodiments, when the minimum scheme edit value is 0, it indicates that there is no difference between the first node set and the second node set. At this time, a new first node set can be obtained from the first document object model tree at the next level of the level where the current first node set is located, and a new second node set can be obtained from the second document object model tree at the next level of the level where the current second node set is located, and then the minimum scheme edit value corresponding to the next level is determined according to the above steps 101 to 104. When the minimum scheme difference value is greater than 0, it indicates that there is a difference between the first node set and the second node set. At this time, based on the target edit scheme corresponding to the minimum scheme edit value, the difference information between the first file and the second file is displayed.
[0144] In some embodiments, when the target editing scheme includes a first node to be deleted, based on the first node identifier of the first node, a character string to be deleted corresponding to the first node identifier is obtained from the first file, and the character string to be deleted is displayed in the first preset display style to display the file content. The first preset display style may include, but is not limited to, the font, font color, font size, background color, etc. of the character string to be deleted. For example, the first preset display style may be a red font color and a yellow background color.
[0145] In some embodiments, when the target editing scheme includes a second node to be added, based on the second node identifier of the second node, the string to be added corresponding to the second node to be added is obtained from the second file, and the string to be added is displayed according to the second preset display style. Exemplarily, the second preset display style may include, but is not limited to, different fonts, font colors, font sizes, background colors, and other settings, and the second preset display style is different from the first preset display style to distinguish between the string to be deleted and the string to be added.
[0146] In some embodiments, when the target editing scheme includes modifying the first target node to the second target node, a first target string corresponding to the first target node is obtained from the first file, and a second target string corresponding to the second target node is obtained from the second file. A first substring is determined from the first target string, and a second substring is determined from the second target string, and the first substring and the second substring are displayed according to a third preset display style, wherein the first substring in the first target string is modified to the second substring to obtain the second target string.
[0147] In some embodiments, when the target editing scheme includes modifying the first target node to the second target node, based on the first target node identifier of the first target node, the first target string corresponding to the first target node is obtained from the first file, and based on the second target node identifier of the second target node, the second target string corresponding to the second target node is obtained from the second file, and then the first substring is determined from the first target string, and the second substring is determined from the second target string, wherein the first substring in the first target string is modified to the second substring, that is, the first substring is the part of the first target string that is different from the second target string. Exemplarily, the first target string is "year = 2019" and the second target string is "year = 2022", then the first substring is "19" and the second substring is "20", and then the first substring and the second substring are displayed according to the third preset display style. Exemplarily, the third preset display style can be to display the first substring and the second substring in red bold form, or to display the background color of the first substring and the second substring as purple.
[0148] By the above-mentioned method, the embodiment of the present application can accurately identify and match each node in the XML file based on the comparison method of the node, so as to accurately determine the difference between the nodes. The embodiment of the present application also adopts a dynamic programming algorithm, can efficiently calculate the difference value between two XML files, and find the matching scheme corresponding to the minimum difference value, strengthen the comparison accuracy, improve the accuracy and speed of the comparison, and make the result more reliable and explainable. At the same time, the embodiment of the present application can intuitively show the specific location of the modification and the difference, so that the user can more easily understand and compare the difference between the files. Whether it is in the case of processing complex XML structures or more attribute fields, the method can provide complete and clear comparison results, improve the user's operating efficiency and experience.
[0149] The following is an explanation of an exemplary application of the embodiments of the present application in a practical application scenario.
[0150] The file comparison method provided in the embodiment of the present application can be applied to any scenario where XML file comparison is required.
[0151] See also Figure 5 , Figure 5 Another implementation flow diagram of the file comparison method provided in the embodiment of the present application is shown below. Figure 5 The steps shown are described in detail.
[0152] Step 301: Obtain and read two XML files, parse the two XML files respectively, and obtain corresponding DOM trees.
[0153] Here, the terminal first obtains two XML files and parses them using the XML parsing library. Through the parsing process, the two XML files are converted into corresponding DOM trees, in which each element is a node and a parent-child relationship is established between nodes. By traversing the DOM tree, the terminal can easily access the attributes and content of each node, thereby obtaining the data in the XML file.
[0154] The XML file itself is a data format with a tree structure, consisting of a series of elements and attributes. These elements are hierarchical in XML, nested and contain sub-elements and text content.
[0155] In order to implement the parsing of XML files, the embodiment of the present application adopts an open source XML parsing library, which can convert XML files into a DOM tree (Document Object Model), in which each element is represented as a node. The nodes are connected through parent-child relationships to form a tree structure. Among them, the top-level node is called the root node, which is the starting point of the entire DOM tree.
[0156] Traversing the DOM tree means visiting each node in the DOM tree in sequence according to certain rules. During the traversal process, the attributes and content of each node can be obtained and manipulated. Common traversal methods are depth-first traversal and breadth-first traversal. In depth-first traversal, the current node is visited first, and its child nodes are visited recursively until there are no child nodes or all child nodes have been visited. Then, it goes back to the parent node and continues to visit other child nodes. In breadth-first traversal, the nodes are visited layer by layer in hierarchical order. In the process of traversing the DOM tree, the type and attributes of the node are used to determine the characteristics of the node, so as to perform corresponding operations. For example, by determining whether the type of the node is an element node, it is determined whether the content of the node needs to be processed. For element nodes, their tag names can be obtained as the names of the nodes, and their attribute values can be obtained.
[0157] In the DOM tree, each node has different types and attributes, as well as child nodes or text content that may be contained. By traversing this DOM tree, you can easily get the data in the XML file, whether it is an element or an attribute. By accessing the attributes and content of each node, you can accurately extract the required data.
[0158] Step 302: Obtain a set of nodes from the corresponding DOM tree, exhaust all matching possibilities at the same XML level, and perform node matching on the two XML files.
[0159] Here, the server obtains a set of nodes from the DOM tree mentioned above, and exhaustively enumerates all possible matching situations in the same XML level to perform node matching operations. When performing a node matching operation, two node sets at the same level are obtained from the DOM trees corresponding to the two XML files, and node matching is performed on the two node sets. Using a node matching algorithm, all possible matching schemes (corresponding to the editing schemes in other embodiments) are exhaustively enumerated, and the difference value (corresponding to the scheme editing value in other embodiments) is calculated to select the scheme with the smallest difference as the final matching result. In this way, it is possible to match as many identical nodes as possible in the XML file while minimizing the degree of difference between the nodes.
[0160] In the node matching algorithm, all matching schemes need to be exhausted first, that is, the nodes at the same level in the two XML files are compared one by one. For each matching scheme, the terminal will calculate the difference value, which can be used to measure the number of matched nodes and the degree of difference in their changes. The smaller the difference value, the more nodes are matched and the smaller the difference between the two XML files.
[0161] In order to calculate the difference value, factors such as node attributes, text content, and child nodes are comprehensively considered. For example, the node tag names, attribute values, text content, and the number and type of child nodes are compared. Based on these comparison results, the terminal assigns a difference score to each pair of matching nodes. The lower the score, the higher the match.
[0162] By traversing all possible matching schemes and calculating the difference value, multiple matching degree scores are obtained, and the scheme with the smallest difference degree is selected as the final node matching result (corresponding to the target editing scheme in other embodiments). In this way, the same nodes can be matched as much as possible, and the difference between two XML files can be minimized. Traditional file comparison algorithms are usually matched in units of behavior, while the node matching algorithm of the embodiment of the application measures the difference value in units of nodes, and the goal is to minimize the sum of the difference values of all nodes, so as to match as many identical nodes as possible.
[0163] like Figure 6 As shown, Figure 6A schematic diagram of determining the difference value of two groups of different nodes provided in an embodiment of the present application. Add represents the difference value generated by adding a node, and del represents the difference value generated by deleting a node. When adding a node, if the node has child nodes, the difference value of adding the node needs to be accumulated when all child nodes of the node are added when calculating the difference value. When deleting a node, if the node has child nodes, the difference value of deleting the node needs to be accumulated when calculating the difference value. If the node no longer contains child nodes, the SIFT4 algorithm will be used to directly calculate the difference value of the string text corresponding to the node.
[0164] For example, Figure 6 As shown, the nodes in the first and third columns are ABADE, the nodes in the second and fourth columns are BADED, Figure 6 Each node in the display has no child nodes.
[0165] The first difference value calculation process corresponds to the following scheme: Node A in the first column and first row matches the nodes in the second column from top to bottom, and successfully matches Node A in the second column and second row. Because the node order has a key influence on the comparison of XML files and cannot be swapped, B can only match from A in the second column and second row. The same is true for other nodes. Finally, Node A in the first column and first row matches Node A in the second column and second row, Node D in the first column and fourth row matches Node D in the second column and third row, and Node A in the first column and fifth row matches Node A in the second column and fifth row. The A node is in the fourth row of the second column. Therefore, according to this matching scheme, if the first column is to be matched to the second column, it is necessary to first add the B node of the second column, and make the node in the first row of the first column the B node of the first column of the second column. Then delete the B node in the second row of the first column and the A node in the third row of the first column, so that there are no other nodes between the A node in the first row of the original first column and the D node in the fourth row. Finally, add the D node in the fifth row of the second column, and modify the node in the first column to become the node in the second column. Therefore, the difference value = add(B)+del(B)+del(A)+add(D).
[0166] The solution corresponding to the calculation process of the second difference value is: the E node in the fifth row of the third column starts to match the nodes in the fourth column from bottom to top, and successfully matches the E node in the fourth row of the fourth column. The same is true for other nodes, and they are matched in sequence. Finally, the B node in the second row of the third column matches the B node in the first row of the fourth column, the A node in the third row of the third column matches the A node in the second row of the fourth column, the D node in the fourth row of the third column matches the D node in the third row of the fourth column, and the E node in the fifth row of the third column matches the E node in the fourth row of the fourth column. Therefore, according to this matching solution, if the third column is to match the fourth column, it is necessary to first delete the A node in the first row of the third column, and then add the D node in the fifth row of the fourth column, that is, add it to the last row of the fourth column. Therefore, the difference value of this solution = del(A)+add(D).
[0167] Step 303: In the node matching algorithm, a dynamic programming algorithm is introduced to calculate the difference value and obtain the minimum difference value.
[0168] Here, a dynamic programming algorithm is introduced into the node matching algorithm to calculate the difference value and obtain the minimum difference value. The goal of the node matching algorithm is to compare the differences between the two groups of XML nodes and determine the best matching scheme. Since trying all matching schemes may result in a huge amount of calculation, in the embodiment of the present application, a dynamic programming algorithm is used to reduce repeated calculations.
[0169] The dynamic programming algorithm is a basic algorithm structure used to avoid repeated calculations in the node matching algorithm. The algorithm has only three lines, corresponding to three operations: add, delete, and modify. These three lines of formulas can be nested and used. They are not global, but are used in specific operations. This formula will be used many times throughout the algorithm. The characteristics of the dynamic programming algorithm allow the intermediate results generated during the calculation process to be cached, so that other matching schemes can directly extract from the cache when calculating the node difference values in the same interval without repeated calculations. This caching technology significantly improves the comparison efficiency.
[0170] Suppose there are two groups of nodes, A1, A2, ..., A m and B1,B2,...,B n , the embodiment of the present application defines a function d(A1, A2, ..., A m , B1, B2, ..., Bn), which represents the difference value between two XML nodes. This function can quantify the degree of difference between nodes. m Indicates that file A has m nodes, B1 to B n It means that file B has n nodes, and function len represents the difference value caused by adding or deleting a single node, as shown in formula (1).
[0171]
[0172] The core idea of the dynamic programming algorithm is to decompose the problem into sub-problems and use the solutions of the sub-problems to construct the overall solution. According to formula (1), d(A1...A m ,B1...B n ) is the minimum of the following three: the difference between deleting node A1 and d(A2...A m ,B1...B n ) and 、 Increase the difference between node B1 and d(A1...A m ,B2...B n ), modify node A1 to the difference between node B1 and d(A2...A m ,B2...B n ). In the next iteration, d(A2...A m ,B1...B n ) can also be understood as determining the minimum value from the following three: the difference between deleting the A2 node and d(A3...A m ,B1...B n ) and add the B1 node and d(A2...A m ,B2...B n ), modify A2 to the difference between B1 and d(A3...A m ,B2...B n ) and…, perform nested iterations in the above manner, and finally determine the minimum difference value. In the dynamic programming algorithm, the calculated intermediate results can be cached to avoid repeated calculations and further improve the matching efficiency.
[0173] Step 304: According to the minimum difference value, a matching solution corresponding to the minimum difference value is obtained.
[0174] Step 305: Output the difference result of the two XML files according to the matching solution corresponding to the minimum difference value.
[0175] Here, after calculating the minimum difference value according to the formula in the dynamic programming algorithm, the embodiment of the present application can find the matching solution corresponding to the minimum difference value by backtracking. Backtracking means to reversely deduce the source of the value (i.e., which operation caused the minimum difference value) and the corresponding matching solution based on the calculated minimum difference value.
[0176] Specifically, in step 303, a dynamic programming algorithm is introduced to calculate the difference value and obtain the minimum difference value. The algorithm systematically explores the difference values of different matching schemes by decomposing the problem into add, delete and modify operations, and utilizing the cache of intermediate results. Finally, the minimum difference value is obtained. According to this minimum difference value, the embodiment of the present application can find the matching scheme corresponding to the minimum difference value by backtracking. The process of backtracking is to start from the last node and reversely deduce the matching scheme of each node according to the calculated minimum difference value. Specifically, find the corresponding operation according to the minimum difference value, such as add, delete or modify, and record the node corresponding to this operation. Then, continue to backtrack to the previous node and repeat this process until backtracking to the starting node. Finally, a sequence is obtained, which represents the matching scheme corresponding to the minimum difference value.
[0177] With this sequence of matching schemes, the difference results of two XML files can be displayed in a more suitable way according to the needs. For example, the operation of each node in the sequence can be executed in sequence to obtain the difference results between the two XML files. For example, if there are "add(A)", "B" and "del(C)" in the sequence, then node A can be added to the original XML file, node B can be retained, and node C can be deleted to display the difference results of the two XML files. Figure 7 A schematic diagram of an interface for dragging an XML file to be compared to a comparison tool provided in an embodiment of the present application, such as Figure 7 As shown, select two XML files with the mouse and drag the files into the comparison tool to see the comparison results. In addition, the embodiment of the present application can also jump to the next difference through shortcut keys, and see the distribution diagram with differences on the far left (the red area on the left indicates the part with differences). Figure 8 A schematic diagram of a file comparison result determined by a file comparison method using an embodiment of the present application is provided, such as Figure 8 As shown, the style of adding, deleting, and modifying nodes in the first file of the embodiment of the present application is highlighted with light gray shadows, and the second file is highlighted with dark gray shadows for the parts where the node contents of the first file and the second file of the embodiment of the present application are different.
[0178] In summary, the matching solution corresponding to the minimum difference value is found through backtracking, and the difference results of the two XML files are output using this matching solution. After calculating the minimum difference value through the formula in the dynamic programming algorithm, the backtracking process helps the application find the operation that leads to the minimum difference value and its corresponding matching solution, and applies this matching solution to the display of the output difference results.
[0179] The node comparison-based file comparison method provided in the embodiment of the present invention uses XML semantic information to perform file comparison and has multiple advantages in result display. First, in the case of condition attributes with many attribute fields, compared with the traditional line comparison-based tools, the tool in the embodiment of the present application can fully display the modified content without the user having to scroll repeatedly to view, thereby improving the user's operation efficiency. However, the modification of the condition attribute cannot be directly seen in Beyond Compare and needs to be dragged to the scroll bar. Fig. 9 As shown, Fig. 9 It is a schematic diagram of the file comparison results determined by the Beyond Compare method in the related art. Secondly, due to the large number of meaningless carriage returns and spaces between lines in the XML format, it is difficult for conventional comparison tools to distinguish which differences are meaningful, and the tool of the embodiment of the present application can perfectly hide such meaningless differences, allowing users to more clearly understand the real differences between files. In addition, the changes in the order of attributes in XML are semantically equivalent, and the tool of the embodiment of the present application can recognize this equivalence relationship, avoiding unnecessary difference prompts caused by changes in the order of attributes. In summary, the file comparison method based on node comparison of the embodiment of the present invention has significant advantages in display effect and user convenience, and can provide users with a more accurate and clear file comparison experience.
[0180] The following is a description of an exemplary structure of the file comparison device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, Figure 3 As shown, the software modules stored in the file comparison device 455 of the memory 450 may include:
[0181] A file parsing module 4551 is used to obtain a first file and a second file to be compared, parse the first file to obtain a first document object model tree, and parse the second file to obtain a second document object model tree;
[0182] An acquisition module 4552 is used to acquire a first node set from the first document object model tree and acquire a second node set from the second document object model tree, wherein a first node in the first node set and a second node in the second node set are located at the same level;
[0183] A first determining module 4553, configured to determine at least one editing scheme between the first node set and the second node set, and determine a scheme editing value corresponding to each editing scheme;
[0184] A second determining module 4554 is used to determine a minimum scheme editing value from the scheme editing values corresponding to the respective editing schemes;
[0185] The difference display module 4555 is used to display the difference information between the first file and the second file based on the target editing scheme corresponding to the minimum scheme editing value when the minimum scheme editing value indicates that there is a difference between the first node set and the second node set.
[0186] In some embodiments, the first determination module 4553 is also used to exhaustively match each first node in the first node set with each second node in the second node set based on a dynamic programming algorithm to obtain at least one editing scheme; wherein the first node set is edited according to the editing scheme to obtain the second node set; the editing scheme includes at least one of a first node to be deleted, a second node to be added, and a first target node to be modified to a second target node.
[0187] In some embodiments, the first determination module 4553 is also used to perform the following operations for each editing scheme: when the editing scheme only includes a first node to be deleted, determine a first total editing value for deleting the first node, and determine the first total editing value as the scheme editing value corresponding to the editing scheme; when the editing scheme only includes a second node to be added, determine a second total editing value for adding the second node, and determine the second total editing value as the scheme editing value corresponding to the editing scheme; when the editing scheme only includes modifying the first target node to the second target node, determine a third total editing value for modifying the first target node to the second target node, and determine the third total editing value as the scheme editing value corresponding to the editing scheme; when the editing scheme includes at least two of a first node to be deleted, a second node to be added, a second node to be added, and modifying the first target node to the second target node, determine the sum of at least two of the corresponding first total editing value, second total editing value, and third total editing value as the scheme editing value corresponding to the editing scheme.
[0188] In some embodiments, the first determination module 4553 is also used to obtain a first node editing value for deleting the first node when the first node does not have a first child node; determine the first node editing value as the first total editing value; when the first node has a first child node, determine the first subnode editing value for deleting the first subnode, and determine the first node editing value for deleting the first node; determine the sum of the first subnode editing value and the first node editing value as the first section total editing value for deleting the first node.
[0189] In some embodiments, the first determination module 4553 is also used to determine the first node edit value for deleting the first node and store the first node edit value in the cache space when the first node edit value for deleting the first node is not stored in the cache space; and to obtain the first node edit value from the cache space when the first node edit value for deleting the first node is stored in the cache space.
[0190] In some embodiments, the first determination module 4553 is further used to obtain, based on the first file, a first string corresponding to the first node to be deleted; and determine the edit value for deleting the first string as the edit value of the first node.
[0191] In some embodiments, the first determination module 4553 is also used to obtain a second node editing value to add the second node when the second node does not have a second child node; determine the second node editing value as the second total editing value; when the second node has a second child node, determine to add the second subnode editing value of the second subnode, and determine to add the second node editing value of the second node; determine the sum of the second subnode editing value and the second node editing value as the second total editing value to add the second node.
[0192] In some embodiments, the first determination module 4553 is also used to determine the second node editing value to be added to the second node and store the second node editing value in the cache space when the second node editing value to be added to the second node is not stored in the cache space; and to obtain the second node editing value from the cache space when the second node editing value to be added to the second node is stored in the cache space.
[0193] In some embodiments, the first determination module 4553 is also used to obtain a second string corresponding to the first target node from the first file and a third string corresponding to the second target node from the second file when neither the first target node nor the second target node has a child node; modify the second string to an edit value of the third string and determine it as the third total edit value.
[0194] In some embodiments, the first determination module 4553 is also used to obtain a fourth string corresponding to the first target node and a fifth string corresponding to the first target subnode from the first file when the first target node has a first target subnode and the second target node does not have a second target subnode; concatenate the fourth string and the fifth string to obtain a first combined string; obtain a sixth string corresponding to the second target node from the second file; modify the first combined string to the edit value of the sixth string, and determine it as the edit value of the third node.
[0195] In some embodiments, the first determination module 4553 is also used to, when the first target node has a first target child node and the second target node has a second target child node, obtain from the first file a seventh character string corresponding to the first target node and an eighth character string corresponding to the first target child node; concatenate the seventh character string and the eighth character string to obtain a second combined character string; obtain from the second file a ninth character string corresponding to the second target node and a tenth character string corresponding to the second target child node; concatenate the ninth character string and the tenth character string to obtain a third combined character string; modify the second combined character string to an edit value of the third combined character string, and determine it as the third total edit value.
[0196] In some embodiments, the difference display module 4555 is also used to, when the target editing scheme includes a first node to be deleted, obtain the string to be deleted corresponding to the first node to be deleted, and display the string to be deleted in a first preset display style; when the target editing scheme includes a second node to be added, obtain the string to be added corresponding to the second node to be added from the second file, and display the string to be added in a second preset display style; when the target editing scheme includes modifying the first target node to the second target node, obtain the first target string corresponding to the first target node from the first file, and obtain the second target string corresponding to the second target node from the second file; determine a first substring from the first target string, and determine a second substring from the second target string, and display the first substring and the second substring in a third preset display style, wherein the first substring in the first target string is modified to the second substring to obtain the second target string.
[0197] An embodiment of the present invention provides an electronic device, including:
[0198] A memory for storing executable instructions;
[0199] The processor is used to implement the file comparison method provided by the embodiment of the present invention when executing the executable instructions stored in the memory.
[0200] The embodiment of the present application provides a computer program product, which includes a computer program or a computer executable instruction, and the computer program or the computer executable instruction is stored in a computer-readable storage medium. The processor of the electronic device reads the computer executable instruction from the computer-readable storage medium, and the processor executes the computer executable instruction, so that the electronic device executes the file comparison method described in the embodiment of the present application.
[0201] The present application embodiment provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will be caused to execute the file comparison method provided by the present application embodiment, for example, FIG. 4A to FIG. 4E The file comparison method shown.
[0202] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0203] In some embodiments, computer executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.
[0204] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).
[0205] As an example, computer executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed at multiple sites and interconnected by a communication network.
[0206] In summary, the node-based file comparison method of the embodiment of the present application adopts node matching and dynamic programming algorithms, which can accurately and efficiently compare XML files, display modified content and hide meaningless differences, while considering semantic information such as attribute order, providing accurate and clear comparison results, and significantly improving user experience and the reliability of comparison effects.
[0207] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent substitutions and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A file comparison method, characterized in that: The method comprises: Acquire a first file and a second file to be compared, parse the first file to obtain a first document object model tree, and parse the second file to obtain a second document object model tree; Acquire a first node set from the first document object model tree, and acquire a second node set from the second document object model tree, wherein a first node in the first node set and a second node in the second node set are located at the same level; Determine at least one editing scheme between the first node set and the second node set, and determine a scheme editing value corresponding to each editing scheme; Determine a minimum scheme editing value from the scheme editing values corresponding to the respective editing schemes; When the minimum scheme editing value indicates that there are differences between the first node set and the second node set, difference information between the first file and the second file is displayed based on the target editing scheme corresponding to the minimum scheme editing value.
2. The method according to claim 1, characterized in that The determining at least one editing scheme between the first node set and the second node set comprises: Based on a dynamic programming algorithm, exhaustively match each first node in the first node set with each second node in the second node set to obtain at least one editing scheme; wherein the first node set is edited according to the editing scheme to obtain a second node set; The editing scheme includes at least one of a first node to be deleted, a second node to be added, and modifying a first target node to a second target node.
3. The method according to claim 2, characterized in that The determining of the scheme editing value corresponding to each editing scheme includes: For each edit scenario, do the following: When the editing scheme includes only a first node to be deleted, determining a first total editing value for deleting the first node, and determining the first total editing value as a scheme editing value corresponding to the editing scheme; When the editing scheme includes only a second node to be added, determining a second total editing value for adding the second node, and determining the second total editing value as a scheme editing value corresponding to the editing scheme; When the editing scheme only includes modifying the first target node to the second target node, determining a third total editing value for modifying the first target node to the second target node, and determining the third total editing value as the scheme editing value corresponding to the editing scheme; When the editing scheme includes at least two of a first node to be deleted, a second node to be added, a second node to be added, and modifying the first target node to the second target node, the sum of at least two of the corresponding first total editing value, second total editing value, and third total editing value is determined as the scheme editing value corresponding to the editing scheme.
4. The method according to claim 3, characterized in that The step of determining to delete the first total edit value of the first node comprises: When the first node does not have a first child node, obtaining a first node edit value for deleting the first node; determining the first node edit value as the first overall edit value; When the first node has a first child node, determining to delete a first child node edit value of the first child node, and determining to delete a first node edit value of the first node; The sum of the first subnode edit value and the first node edit value is determined as the first section total edit value for deleting the first node.
5. The method according to claim 4, characterized in that The obtaining the first node edit value of deleting the first node includes: When the first node edit value for deleting the first node is not stored in the cache space, determining to delete the first node edit value of the first node, and storing the first node edit value in the cache space; When the first node edit value for deleting the first node is stored in the cache space, the first node edit value is obtained from the cache space.
6. The method according to claim 3, characterized in that The determining to increase the second total edit value of the second node comprises: When the second node does not have a second child node, obtaining a second node edit value for adding the second node; determining the second node edit value as the second overall edit value; When the second node has a second child node, determining to increase a second child node editing value of the second child node, and determining to increase a second node editing value of the second node; The sum of the second subnode edit value and the second node edit value is determined to increase the second node's second section total edit value.
7. The method according to claim 6, characterized in that The acquiring and adding a second node edit value of the second node comprises: When the second node edit value for adding the second node is not stored in the cache space, determining to add the second node edit value for the second node, and storing the second node edit value in the cache space; When the cache space stores the second node edit value for adding the second node, the second node edit value is obtained from the cache space.
8. The method according to claim 3, characterized in that The determining of modifying the first target node to a third total edit value of the second target node includes: When neither the first target node nor the second target node has a child node, obtaining a second character string corresponding to the first target node from the first file, and obtaining a third character string corresponding to the second target node from the second file; The second character string is modified to an edit value of the third character string, and is determined as the third total edit value.
9. The method according to claim 8, characterized in that The determining of modifying the first target node to a third total edit value of the second target node includes: When the first target node has a first target child node, and the second target node does not have a second target child node, obtaining a fourth character string corresponding to the first target node and a fifth character string corresponding to the first target child node from the first file; concatenating the fourth character string and the fifth character string to obtain a first combined character string; Acquire a sixth character string corresponding to the second target node from the second file; The first combined character string is modified to the edit value of the sixth character string, and is determined as the edit value of the third node.
10. The method according to claim 8, characterized in that The determining of modifying the first target node to a third total edit value of the second target node includes: When the first target node has a first target child node, and the second target node has a second target child node, obtaining a seventh character string corresponding to the first target node and an eighth character string corresponding to the first target child node from the first file; concatenating the seventh character string and the eighth character string to obtain a second combined character string; Acquire, from the second file, a ninth character string corresponding to the second target node and a tenth character string corresponding to the second target subnode; concatenating the ninth character string and the tenth character string to obtain a third combined character string; The second combined character string is modified to an edit value of the third combined character string, and is determined as the third total edit value.
11. The method according to any one of claims 1 to 6, characterized in that: The displaying the difference information between the first file and the second file based on the target editing scheme corresponding to the minimum scheme editing value includes: When the target editing scheme includes a first node to be deleted, obtaining a character string to be deleted corresponding to the first node to be deleted, and displaying the character string to be deleted in a first preset display style; When the target editing scheme includes a second node to be added, obtaining a character string to be added corresponding to the second node to be added from the second file, and displaying the character string to be added according to a second preset display style; When the target editing scheme includes modifying a first target node into a second target node, obtaining a first target string corresponding to the first target node from the first file, and obtaining a second target string corresponding to the second target node from the second file; A first substring is determined from the first target string, and a second substring is determined from the second target string, and the first substring and the second substring are displayed according to a third preset display style, wherein the first substring in the first target string is modified into a second substring to obtain the second target string.
12. A document comparison device, characterized in that: The device comprises: A file parsing module, used to obtain a first file and a second file to be compared, parse the first file to obtain a first document object model tree, and parse the second file to obtain a second document object model tree; an acquisition module, configured to acquire a first node set from the first document object model tree, and acquire a second node set from the second document object model tree, wherein a first node in the first node set and a second node in the second node set are located at the same level; A first determining module, configured to determine at least one editing scheme between the first node set and the second node set, and determine a scheme editing value corresponding to each editing scheme; A second determining module, configured to determine a minimum scheme editing value from the scheme editing values corresponding to the respective editing schemes; The difference display module is used to display the difference information between the first file and the second file based on the target editing scheme corresponding to the minimum scheme editing value when the minimum scheme editing value indicates that there is a difference between the first node set and the second node set.
13. An electronic device, characterized in that: The electronic device comprises: A memory for storing computer executable instructions; A processor, configured to implement the file comparison method according to any one of claims 1 to 11 when executing the computer executable instructions stored in the memory.
14. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the file comparison method according to any one of claims 1 to 11 is implemented.
15. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the document comparison method according to any one of claims 1 to 11 is implemented.