Method, device and electronic device for tracking difference content between documents

By using a neural network model to identify and fit the elements in the document and establish an element relationship diagram, the problem of low maintenance efficiency when updating documents is solved, efficient tracking of content differences between documents is achieved, and maintenance costs are reduced.

CN114595675BActive Publication Date: 2025-09-26CHINA CONSTRUCTION BANK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210233372.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-10
Publication Date
2025-09-26
Estimated Expiration
2042-03-10

AI Technical Summary

Technical Problem

In the existing technology, when faced with a large number of document updates, the efficiency of maintaining related documents is low, resulting in reduced efficiency and increased costs of system training.

Method used

Through the neural network model, the characters in the document are identified and fitted with features in at least two dimensions, and a feature relationship diagram is established to clearly and intuitively track the differences between documents.

Benefits of technology

Improved document maintenance efficiency and reduced related system maintenance costs, including training costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114595675B_ABST
    Figure CN114595675B_ABST
Patent Text Reader

Abstract

The present application relates to the field of data processing technology, and in particular to a method, device and electronic device for tracking the difference content between documents. The method includes sorting at least two documents according to the time of document release to obtain a document set; wherein the document is a document containing preset content; through a neural network model, the elements in each document in the document set are determined to obtain an element set; wherein the neural network model is used to perform feature recognition and fitting of at least two dimensions of characters in the document, and determine the elements corresponding to the fitting features based on the fitting features obtained by fitting; based on the correspondence between the elements in the element set and the documents, an element relationship graph is obtained; wherein the element relationship graph indicates the existence of any type of elements in different documents. The above method can solve the problem of low efficiency in related document maintenance when facing a large number of document updates in the existing technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method, device, and electronic device for tracking difference content between documents. Background Art

[0002] Every year, from the central government to local governments, and ultimately to the headquarters and branches of enterprises and institutions, a wide range of industry-specific documents are released, such as guidelines, special projects, and business notices. These documents, primarily focused on regulatory content, are constantly updated and often interrelated. For example, the release of a new regulation often triggers changes to other documents within the enterprise or institution, such as guidelines and special projects. These changes require significant human and material resources to maintain (or update) related documents, which in turn reduces the efficiency of regulatory training and increases unnecessary training costs. Therefore, current manual maintenance methods are no longer sufficient to meet relevant business needs. A more efficient maintenance method is needed for these documents. This approach can effectively improve maintenance efficiency and reduce regulatory maintenance costs, including training costs, when faced with large volumes of document updates and discrepancies. Summary of the Invention

[0003] The present application provides a method, device and electronic device for tracking the difference content between documents, which are used to solve the problem of low efficiency in maintaining related documents when a large number of documents are updated in the prior art.

[0004] In a first aspect, the present application provides a method for tracking differences between documents, the method comprising:

[0005] Sorting at least two documents according to their release time to obtain a document set; wherein the documents are documents containing preset content;

[0006] Determining elements in each document in the document set using a neural network model to obtain an element set; wherein the neural network model is used to perform feature recognition and fitting of at least two dimensions on characters in the document, and determining elements corresponding to the fitting features based on fitting features obtained by fitting;

[0007] Based on the correspondence between the elements in the element set and the documents, an element relationship graph is obtained; wherein the element relationship graph indicates the existence of any type of elements in different documents.

[0008] This method uses a neural network model to extract features from chronologically sorted documents and creates a feature relationship graph based on the correspondence between features and documents. This graph clearly and intuitively displays the transformation of any type of feature across related documents, effectively tracking differences between documents. Furthermore, it can quickly and efficiently determine the changes in the business lines associated with these differences, effectively improving the maintenance efficiency and reducing maintenance costs for related documents.

[0009] In one possible implementation, determining the elements in each document in the document set using a neural network model, before obtaining the element set, includes:

[0010] In the neural network model, character features of any dimension in the sample document are identified;

[0011] Based on the identified character features of any dimension, the character features of different dimensions are fitted to obtain fitting features, and the character corresponding to the fitting features and the element corresponding to the character are determined, until the accuracy of the determined element being the preset element corresponding to the character is greater than a set threshold;

[0012] Based on the element categories in the element graph, the elements in the sample document are classified to determine the element category to which the elements belong; until the accuracy of the determined element category being the preset element category to which the elements belong is greater than a classification threshold.

[0013] In a possible implementation, in the neural network model, identifying character features of any dimension in the document includes:

[0014] In the neural network model, a segment of a first length in each sample document in the sample document set is converted into a corresponding document vector; wherein the document vector is extracted based on character features of any dimension of characters in the segment;

[0015] Based on preset rules, the elements in the document vector are identified to determine the character features of any dimension; wherein the preset rules include the meanings represented by different combinations of different elements, and the corresponding relationship between the meanings and the character features.

[0016] In one possible implementation, determining the elements in each document in the document set using a neural network model, before obtaining the element set, includes:

[0017] For each of the segments of the second length in the document, a document identifier is added; wherein the segment consists of at least one character, and the document identifier uniquely identifies the document corresponding to the segment;

[0018] Then, based on the correspondence between the elements in the element set and the documents, an element relationship graph is obtained, including:

[0019] Extracting the document identifier included in each element of each type of element from the element set;

[0020] Based on the correspondence between the document identifier and the document, a correspondence between the elements in the element set and the document is established, thereby obtaining an element relationship graph.

[0021] In a possible implementation manner, obtaining an element relationship graph based on the correspondence between elements in the element set and documents includes:

[0022] Based on the existence situation, determining a difference feature between the at least two documents in the element relationship graph; wherein the difference feature is composed of at least one element;

[0023] The first document of the at least two documents is annotated with the difference feature to obtain key information corresponding to each document; wherein the key information includes whether the element is a new element, the importance of the difference feature, the relevance of the difference feature to the corresponding regulatory content, and the semantic understanding output of the difference feature.

[0024] In a possible implementation manner, after obtaining the document set, the method further includes:

[0025] Determining subsequences in the two adjacent documents in the document set, respectively; wherein the subsequences indicate sentences in the two adjacent documents, in which characters are arranged in a set order;

[0026] Determine the longest subsequence between the two adjacent documents by dynamic programming;

[0027] Comparing any two longest subsequences in the two adjacent documents to determine two target longest subsequences that meet similarity requirements; wherein the longest subsequence indicates the longest subsequence that is consistent with the set subsequence order;

[0028] Determining the different contents between the two target longest subsequences, and marking the different contents as the different contents;

[0029] Based on the difference content, the existence of any type of elements in the element relationship diagram is adjusted.

[0030] In a second aspect, the present application provides a device for tracking differences between documents, the device comprising:

[0031] Sorting unit: used for sorting at least two documents according to the release time of the documents to obtain a document set; wherein the documents are documents containing preset content;

[0032] Model unit: used to determine the elements in each document in the document set through a neural network model to obtain an element set; wherein the neural network model is used to perform feature recognition and fitting of at least two dimensions on the characters in the document, and determine the elements corresponding to the fitting features based on the fitting features obtained by fitting;

[0033] A generating unit is used to obtain an element relationship graph based on the correspondence between the elements in the element set and the documents; wherein the element relationship graph indicates the existence of any type of elements in different documents.

[0034] In a possible implementation manner, the device further includes a training unit, which is specifically used to identify character features of any dimension in the sample document in the neural network model; based on the identified character features of any dimension, fit the character features of different dimensions to obtain fitting features, and determine the characters corresponding to the fitting features, as well as the elements corresponding to the characters, until the accuracy of the elements determined to be the preset elements corresponding to the characters is greater than a set threshold; based on the element categories in the element map, classify the elements in the sample document and determine the element category to which the elements belong; until the accuracy of the element categories determined to be the preset element category to which the elements belong is greater than a classification threshold.

[0035] In one possible implementation, the training unit is specifically used to convert a paragraph of a first length in each sample document in the sample document set into a corresponding document vector in the neural network model; wherein the document vector is extracted based on character features of any dimension of characters in the paragraph; and the elements in the document vector are identified based on preset rules to determine the character features of any dimension; wherein the preset rules include the meanings represented by different combinations of different elements, and the correspondence between the meanings and the character features.

[0036] In one possible implementation, the apparatus further includes an identification unit, specifically configured to add a document identifier to each of the segments of the second length in the document; wherein the segment is composed of at least one character, and the document identifier uniquely identifies the document corresponding to the segment;

[0037] The generation unit is specifically used to extract the document identifier included in each element of each type of element in the element set; based on the correspondence between the document identifier and the document, establish the correspondence between the elements in the element set and the document, and obtain the element relationship graph.

[0038] In a possible implementation manner, the generation unit is further used to determine the difference features between the at least two documents in the element relationship diagram based on the existence situation; wherein the difference features are composed of at least one element; the first document among the at least two documents marks the difference features to obtain key information corresponding to each document; wherein the key information includes whether the element is a new element, the importance of the difference features, the relevance of the difference features to the corresponding regulatory content, and the semantic understanding output of the difference features.

[0039] In a possible implementation, the device further includes an optimization unit, which is specifically used to determine subsequences in the two adjacent documents in the document set respectively; wherein the subsequence indicates a sentence in which characters are arranged in a set order in the two adjacent documents; determine the longest subsequence in the two adjacent documents through dynamic programming; compare any two longest subsequences in the two adjacent documents to determine two target longest subsequences that meet similarity requirements; wherein the longest subsequence indicates the longest subsequence that is consistent with the set subsequence order; determine the different content in the two target longest subsequences, and mark the different content as the different content; based on the different content, adjust the existence of any type of element in the element relationship graph.

[0040] In a third aspect, the present application provides an electronic device, comprising:

[0041] Memory for storing computer programs;

[0042] The processor is configured to execute the computer program stored in the memory to implement the method described in the first aspect and any possible implementation manner.

[0043] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect and any possible implementation manner is implemented.

[0044] In a fifth aspect, the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the method described in the first aspect and any possible implementation manner. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 A flowchart of a method for tracking differences between documents provided in this application;

[0046] Figure 2 A schematic diagram of some element maps provided for this application;

[0047] Figure 3 A schematic diagram of tracking content differences between documents using a neural network model provided by this application;

[0048] Figure 4 A schematic diagram of the structure of a device for tracking differences between documents provided by this application;

[0049] Figure 5 A schematic diagram of the structure of an electronic device for tracking differences between documents provided in this application. DETAILED DESCRIPTION

[0050] To address the low efficiency of document maintenance when a large number of documents are updated in existing technologies, this application proposes a method for tracking differences between documents. This method sorts all documents and feeds them into a neural network model to identify and output document elements. It then establishes a correspondence between elements and documents, forming an element relationship graph. This element relationship graph allows for clear and intuitive identification of document deletions, additions, and modifications, effectively improving document maintenance efficiency and reducing system maintenance costs, including training costs.

[0051] It should be noted that the acquisition, storage, use, and processing of data in the technical solution of this application comply with the relevant provisions of national laws and regulations.

[0052] In order to better understand the above technical solution, the technical solution of the present application is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. In the absence of conflict, the embodiments of the present application and the technical features in the embodiments can be combined with each other.

[0053] Please refer to Figure 1 The present invention provides a method for tracking differences between documents to improve document maintenance efficiency. The method specifically includes the following implementation steps:

[0054] Step 101: Sort at least two documents according to their release time to obtain a document set.

[0055] The document contains preset content.

[0056] Specifically, the impact of the release of a new system is often comprehensive, affecting not only the content of the previous version of the system, but also changes in the content of related business documents. Especially in the financial field, the system adjustment requirements issued by the superior regulatory agency may affect system changes in multiple business lines. In the embodiment of the present application, all documents are sorted according to certain rules to obtain a document set. In this way, the relevant content of the documents in the document set can establish a timeline order. Here, the certain rule is the document release time.

[0057] It should be noted that the solution provided in the embodiment of the present application can be applied to scenarios where system documents need to be maintained. Among them, the system refers to the operating procedures or behavioral guidelines that regulate individual actions with rules (or patterns). System documents refer to documents that have passed the quality system certification and involve enterprise technology, finance, operations, safety, personnel, archives, management and other related content; these documents have file names, file numbers, corporate seals, etc. Generally, there will be overlaps in business, specifications, and related special work content between system documents.

[0058] Step 102: Determine the elements in each document in the document set using a neural network model to obtain an element set.

[0059] The neural network model is used to perform feature recognition and fitting of at least two dimensions on characters in a document, and to determine elements corresponding to the fitting features based on the fitting features obtained from the set.

[0060] Specifically, before using the neural network model, it is necessary to train the neural network model so that the neural network model can output correct content. The following is a detailed description of the training of the neural network model:

[0061] Generally, a neural network model includes an input layer, a hidden layer, and an output layer. The input layer is used to receive samples of two-dimensional visual patterns, specifically, to read characters from sample documents within a sample document set that will be input into the neural network model. After reading, the model enters the hidden layer, which is used to identify character features along any dimension of the characters. When identifying character features, due to the varying lengths of documents, the neural network model first divides each sample document within the sample document set into segments of a unit length, namely, segments of a first length, and converts each segment of the first length into a corresponding document vector. The document vector is extracted based on character features along any dimension of the characters within the segment of the first length. Character features along any dimension can be understood as features extracted from different aspects of the characters. Character features along any dimension can be the structure of each character within the segment of the first length, or the part of speech of each character within the segment of the first length. Then, based on preset rules, the elements in the document vector are identified to determine character features along any dimension. The preset rules may include the meanings represented by different combinations of elements, as well as the corresponding relationships between these meanings and character features.

[0062] Furthermore, after performing multi-dimensional character feature recognition on each first-length paragraph in the document, character features of different dimensions can be fitted to obtain fitting features, and the characters corresponding to the fitting features and the elements corresponding to the characters can be determined, until the elements determined by the neural network model are the preset elements corresponding to the corresponding characters in the sample document, and the accuracy rate is greater than the set threshold, then it is determined that the feature fitting function in the neural network model has been trained and meets the use requirements. Specifically:

[0063] First, the document vectors of character features of different dimensions including the same character can be fitted to determine the feature vector of each document. There can be multiple feature vectors, which are used to respectively indicate the character features of different first-length paragraphs in a sample document. Then, the first element corresponding to the element in each feature vector can be determined based on the mapping relationship between the character features and the elements. It should be noted that the correspondence between elements and elements here is not limited to one-to-one correspondence, and a combination of multiple elements, whether continuous or discontinuous, can correspond to one element. After obtaining the first element, all first elements are compared with the corresponding preset elements to determine whether they are the same, so as to obtain the accuracy of the feature fitting function in the neural network model. Finally, the accuracy is compared with the set threshold. When the accuracy is less than or equal to the set threshold, it is determined that the recognition fitting function does not meet the use requirements, and the corresponding parameters in the hidden layer need to be further debugged; until the accuracy is greater than the set threshold.

[0064] When the accuracy of the feature fitting function of the above-mentioned neural network model is greater than the set threshold, the training of the recognition fitting function is terminated and the training of the classification function is started. The classification function can be specifically performed by a classifier. During execution, the elements in the sample document obtained by the above-mentioned recognition fitting function must be classified based on the element categories in the element map, and the element category to which the element belongs must be determined; until the element category determined is that the accuracy of the preset element category to which the element belongs is greater than the classification threshold. The following describes the element map and the element categories in the element map:

[0065] The element graph described in the embodiment of the present application can be a structured knowledge graph that includes element categories. In this knowledge graph, the concepts of all elements in the system document and their mutual relationships are clearly described. The mutual relationship here can be an inclusion relationship, that is, the relationship between elements and element categories. The element category can be a specific business type, job role, system, etc. For example, if a document includes the elements of British pounds to RMB and US dollars to RMB, then these two elements can be classified into the element category of foreign currency exchange business. Figure 2 is a schematic diagram of some element maps. Figure 2 As shown, the factor map includes relationships between factors and between factor categories. Factor categories can include foreign exchange business, regulations, and work. Factors can include business operations, technical specifications, work specifications, emergency plans, business specifications, and so on.

[0066] After determining the features and the feature categories to which the features belong, the feature categories can be output through the output layer of the neural network model. The output results can be the feature categories of all sample documents in the sample document set, or the feature categories and the features included in the feature categories.

[0067] Furthermore, after the neural network model is trained, it can be used to extract a feature set from the document described in step 101. Specifically, the document set can be input into the neural network model, and the neural network model determines the features included in each document in the document set, as well as the feature categories corresponding to any feature class, and outputs the results to obtain the feature set. The feature set can then include at least one feature category, and each feature category can include at least one feature.

[0068] It should be noted that the neural network model described in the embodiments of the present application mainly refers to a convolutional neural network model. The methods described in the embodiments of the present application are also applicable to other deep learning models and will not be repeated here.

[0069] Step 103: Based on the correspondence between the elements in the element set and the documents, obtain an element relationship graph.

[0070] Among them, the element relationship diagram indicates the existence of any type of elements in different documents.

[0071] Specifically, first, the document identifier included in each element of each type of element in the element set can be extracted. Then, based on the correspondence between the document identifier and the document, the correspondence between the elements in the element set and the document is established to obtain an element relationship graph.

[0072] It's worth noting that the document identifiers are added before the document set is fed into the neural network. Specifically, a document identifier is added for each segment of the second length within each document. Each segment consists of at least one character, and the document identifier uniquely identifies the document corresponding to the segment.

[0073] Furthermore, the determination of the difference between any two documents in the document set can be based on the element relationship diagram in this step 104. Specifically, first, based on the existence of any type of elements in the element relationship diagram, the difference features between at least two documents are determined. The difference feature can be composed of at least one element, and the element can correspond to the same element category or different element categories. Then, in the first document of the at least two documents mentioned above, the difference feature is marked, and the key information corresponding to each document can be obtained. The key information may include whether the element is a new element, the importance of the difference feature, the relevance of the difference feature to the corresponding regulatory content, and the semantic understanding output of the difference feature.

[0074] Furthermore, in order to improve the accuracy of the existence of any type of elements reflected in the element relationship diagram, the existence of elements in the element relationship diagram can also be checked based on the document set determined in step 101. Specifically:

[0075] The difference content can be determined using a longest subsequence-based algorithm or an edit distance-based algorithm;

[0076] The above algorithm based on edit distance may be to determine the difference between two character strings by statistically calculating the number of edits required to convert between the two character strings.

[0077] The algorithm based on the longest subsequence can be: first, determine the subsequences in the two adjacent documents in the document set respectively. The subsequence refers to the sentence in which the characters are arranged in a set order in the above two adjacent documents. Then, the longest subsequence in the above two adjacent documents is determined by dynamic programming. Then, in the above two adjacent documents, any two longest subsequences are compared to determine the two target longest subsequences that meet the similarity requirements. The longest subsequence refers to the longest subsequence that is consistent with the set subsequence order. Finally, the content with differences in the above two target longest subsequences can be determined, and the content with differences can be marked as the difference content. The difference content is then compared with the existence of the elements in the element relationship diagram. That is, based on the difference content, adjustments are made to the existence of any type of elements in the element relationship diagram.

[0078] It should be noted that the longest subsequence-based or edit distance-based algorithms described above for determining differences can also be used independently in conjunction with Natural Language Processing (NLP). NLP technology can be used to intelligently analyze differences, thereby avoiding the costly and inefficient practice of manual proofreading. This allows for efficient determination of differences between at least two documents in a document collection, streamlining the process and improving the efficiency of relevant policy training.

[0079] The following is an example of using the neural network model, please refer to Figure 3 , Figure 3 A schematic diagram of tracking content differences between documents using a neural network model provided in an embodiment of the present application.

[0080] First, all documents are sorted based on the release time to obtain the document set, such as Figure 3 As shown in (a), the order of the documents in the document set is: Document 1_v1, Document 1_v2, Document 2_v1, Document 2_v2, Document 3_v1. Document 1, Document 2, and Document 3 represent different types of institutional documents, each of which is related in content. The "v1" and "v2" after the document name indicate the document version number. For example, Document 1_v2 is an updated version of Document 1_v1.

[0081] Then, after the document set is determined, the document set can be input into the neural network model. In the neural network model, the input layer reads the two-dimensional visual pattern samples, that is, the input layer vectorizes all the characters in the document set according to the model translation rules to obtain a number of document vectors. After that, all document vectors enter the hidden layer, and the hidden layer begins to perform feature recognition on each element in the document vector of the document set. In order to improve the accuracy of the neural network model for character recognition, at least two hidden layers are generally set in the neural network model to extract features of different dimensions (angles) for elements representing the same character. In the embodiment of the present application, the number of hidden layers is set to 4, such as Figure 3 As shown in part (b) of the figure, h1 to h4 represent a hidden layer, which is used to identify character features in different dimensions. Furthermore, after extracting features of different dimensions from the characters in the document set, the character features can be fitted in the output layer to determine the elements corresponding to the fitted features and the element categories to which the elements belong. Figure 3 As shown in part (c) of the figure, each square in part (c) represents a category of features obtained by fitting the character features obtained from h1 to h4. This category of features can include features of the same category but different content. It should be noted that each feature in the feature category obtained in part (c) carries the document identifier, that is, as shown in part (d), each feature category includes an identifier for the document name in part (a).

[0082] Furthermore, after determining the elements in all documents in the document set and the element categories to which the elements belong, the elements and element categories can be output through the neural network model. At the same time, the output content also includes the document identifiers of any type of elements as shown in part (d). In this way, an element relationship graph can be established through the document identifiers of each type of elements. In this element relationship graph, each type of element involved in the content of the document set, the different element contents in each type of element, and the existence of each type of element in each document are included. Figure 3 Part (e) of FIG is a diagram showing a partial element relationship, where 1 indicates that the corresponding document contains an element of the corresponding type, and 0 indicates that the corresponding document does not contain an element of the corresponding type.

[0083] In the element relationship diagram obtained based on the above method, the business line represented by each type of element and the changes in the relevant business line in all (system) documents can be obtained intuitively and efficiently, so as to achieve the purpose of efficiently tracking the differences between documents. In this way, the changes in the business line where the difference content is located can be quickly and effectively determined.

[0084] Based on the same inventive concept, the present invention provides a device for tracking the difference content between documents. Figure 1 The method for tracking the difference content between the documents shown corresponds to the method for tracking the difference content between the documents shown. The specific implementation of the device can be found in the description of the above method embodiment part. The repeated parts will not be repeated here. Figure 4 , the device comprises:

[0085] Sorting unit 401 is used to sort at least two documents according to the release time of the documents to obtain a document set.

[0086] The document is a document containing preset content.

[0087] Model unit 402: used to determine the elements in each document in the document set through a neural network model to obtain an element set; wherein the neural network model is used to perform feature recognition and fitting of at least two dimensions on the characters in the document, and determine the elements corresponding to the fitting features based on the fitting features obtained by fitting.

[0088] The device for tracking the difference content between documents also includes a training unit, which is specifically used to identify character features of any dimension in the sample document in the neural network model; based on the identified character features of any dimension, the character features of different dimensions are fitted to obtain fitting features, and the characters corresponding to the fitting features and the elements corresponding to the characters are determined, until the accuracy of the elements determined to be the preset elements corresponding to the characters is greater than a set threshold; based on the element categories in the element map, the elements in the sample document are classified to determine the element category to which the elements belong; until the accuracy of the element categories determined to be the preset element categories to which the elements belong is greater than a classification threshold.

[0089] The training unit is specifically used to, in the neural network model, convert a paragraph of a first length in each sample document in the sample document set into a corresponding document vector; wherein the document vector is extracted based on the character features of any dimension of the characters in the paragraph; identify the elements in the document vector based on preset rules, and determine the character features of any dimension; wherein the preset rules include the meanings represented by different combinations of different elements, and the correspondence between the meanings and the character features.

[0090] Generating unit 403: used to obtain an element relationship graph based on the correspondence between the elements in the element set and the documents; wherein the element relationship graph indicates the existence of any type of elements in different documents.

[0091] The apparatus for tracking difference content between documents further includes an identification unit, specifically configured to add a document identifier to a segment of the second length in each document; wherein the segment is composed of at least one character, and the document identifier uniquely identifies the document corresponding to the segment;

[0092] The generating unit 403 is specifically used to extract the document identifier included in each element of each type of element in the element set; based on the correspondence between the document identifier and the document, establish the correspondence between the elements in the element set and the document, and obtain the element relationship graph.

[0093] The generation unit 403 is also used to determine the difference features between the at least two documents in the element relationship diagram based on the existence situation; wherein the difference features are composed of at least one element; the first document among the at least two documents marks the difference features to obtain key information corresponding to each document; wherein the key information includes whether the element is a new element, the importance of the difference features, the relevance of the difference features to the corresponding regulatory content, and the semantic understanding output of the difference features.

[0094] The device for tracking the difference content between documents also includes an optimization unit, which is specifically used to determine the subsequences in the two adjacent documents in the document set respectively; wherein the subsequence indicates a sentence in which characters are arranged in a set order in the two adjacent documents; determine the longest subsequence in the two adjacent documents through dynamic programming; compare any two longest subsequences in the two adjacent documents to determine two target longest subsequences that meet the similarity requirements; wherein the longest subsequence indicates the longest subsequence that is consistent with the set subsequence order; determine the different content in the two target longest subsequences, and mark the different content as the different content; based on the different content, adjust the existence of any type of element in the element relationship diagram.

[0095] Based on the same inventive concept as the above-mentioned method for tracking the difference content between documents, an electronic device is also provided in the embodiment of the present application. The electronic device can implement the function of the above-mentioned method for tracking the difference content between documents. Please refer to Figure 5 , the electronic device includes:

[0096] At least one processor 501, and a memory 502 connected to the at least one processor 501. The specific connection medium between the processor 501 and the memory 502 is not limited in the embodiment of the present application. Figure 5 In the example, the processor 501 and the memory 502 are connected via a bus 500. Figure 5 The bus 500 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5The diagram is represented by only one thick line, but this does not mean that there is only one bus or one type of bus. Alternatively, the processor 501 may also be referred to as a controller, without limitation to the name.

[0097] In the embodiment of the present application, the memory 502 stores instructions that can be executed by at least one processor 501. The at least one processor 501 can execute the method for tracking the difference content between documents discussed above by executing the instructions stored in the memory 502. The processor 501 can implement Figure 4 The functions of each module in the device shown.

[0098] Among them, the processor 501 is the control center of the device, which can use various interfaces and lines to connect the various parts of the entire control device, and monitor the device as a whole by running or executing instructions stored in the memory 502 and calling data stored in the memory 502, the various functions of the device and processing data.

[0099] In one possible design, processor 501 may include one or more processing units. Processor 501 may integrate an application processor and a modem processor. The application processor primarily processes the operating system, user interface, and application programs, while the modem processor primarily processes wireless communications. It is understood that the modem processor may not be integrated into processor 501. In some embodiments, processor 501 and memory 502 may be implemented on the same chip. In some embodiments, they may also be implemented on separate chips.

[0100] The processor 501 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method for tracking the difference content between documents disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor.

[0101] The memory 502 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory 502 may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, etc. The memory 502 is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 502 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.

[0102] By programming the processor 501, the code corresponding to the method for tracking the difference content between documents described in the above embodiment can be fixed into the chip, so that the chip can execute the code when running. Figure 1 The steps of the method for tracking the difference content between documents are shown. How to design and program the processor 501 is a technology well known to those skilled in the art and will not be described in detail here.

[0103] Based on the same inventive concept, an embodiment of the present application further provides a storage medium storing computer instructions. When the computer instructions are executed on a computer, the computer executes the method for tracking the difference content between documents discussed above.

[0104] In some possible implementations, various aspects of the method for tracking difference content between documents provided in the present application can also be implemented in the form of a program product, which includes program code. When the program product is run on an apparatus, the program code is used to enable the control device to execute the steps of the method for tracking difference content between documents according to various exemplary implementations of the present application described above in this specification.

[0105] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0106] The program product of the method for tracking differences between documents provided in embodiments of the present invention can be implemented in a portable compact disc read-only memory (CD-ROM) and include program code, and can be run on a computing device. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0107] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0108] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0109] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and the like, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0110] It should be noted that although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to embodiments of the present invention, the features and functions of two or more units described above may be embodied in a single unit. Conversely, the features and functions of a single unit described above may be further divided and embodied by multiple units.

[0111] Furthermore, although the operations of the method of the present invention are described in a particular order in the accompanying drawings, this does not require or imply that these operations must be performed in this particular order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0112] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0113] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0114] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0115] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0116] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A method for tracking differences between documents, characterized in that: The method comprises: Sorting at least two documents according to their release time to obtain a document set; wherein the documents are documents containing preset content; In the neural network model, character features of any dimension in the sample document are identified; Based on the identified character features of any dimension, the character features of different dimensions are fitted to obtain fitting features, and the character corresponding to the fitting features and the element corresponding to the character are determined, until the accuracy of the determined element being the preset element corresponding to the character is greater than a set threshold; Based on the element categories in the element graph, the elements in the sample document are classified to determine the element category to which the elements belong; until the accuracy of the determined element category being the preset element category to which the elements belong is greater than a classification threshold; Determining the elements in each document in the document set using the neural network model to obtain an element set; wherein the neural network model is used to perform feature recognition and fitting on characters in at least two dimensions in the document, and determining the elements corresponding to the fitting features based on the fitting features obtained by fitting; Based on the correspondence between the elements in the element set and the documents, an element relationship graph is obtained; wherein the element relationship graph indicates the existence of any type of elements in different documents.

2. The method according to claim 1, wherein In the neural network model, character features of any dimension in the document are identified, including: In the neural network model, a segment of a first length in each sample document in the sample document set is converted into a corresponding document vector; wherein the document vector is extracted based on character features of any dimension of characters in the segment; Based on preset rules, the elements in the document vector are identified to determine the character features of any dimension; wherein the preset rules include the meanings represented by different combinations of different elements, and the corresponding relationship between the meanings and the character features.

3. The method according to any one of claims 1 to 2, characterized in that Determining the elements in each document in the document set by the neural network model, before obtaining the element set, includes: For each of the segments of the second length in the document, a document identifier is added; wherein the segment consists of at least one character, and the document identifier uniquely identifies the document corresponding to the segment; Then, based on the correspondence between the elements in the element set and the documents, an element relationship graph is obtained, including: Extracting the document identifier included in each element of each type of element from the element set; Based on the correspondence between the document identifier and the document, a correspondence between the elements in the element set and the document is established, thereby obtaining an element relationship graph.

4. The method according to claim 3, wherein The obtaining of an element relationship graph based on the correspondence between elements in the element set and documents includes: Based on the existence situation, determining a difference feature between the at least two documents in the element relationship graph; wherein the difference feature is composed of at least one element; The first document of the at least two documents is annotated with the difference feature to obtain key information corresponding to each document; wherein the key information includes whether the element is a new element, the importance of the difference feature, the relevance of the difference feature to the corresponding regulatory content, and the semantic understanding output of the difference feature.

5. The method according to claim 4, wherein After obtaining the document set, the following steps are also included: Determining subsequences in two adjacent documents in the document set, respectively; wherein the subsequences indicate sentences in the two adjacent documents, in which characters are arranged in a set order; Determine the longest subsequence between the two adjacent documents by dynamic programming; Comparing any two longest subsequences in the two adjacent documents to determine two target longest subsequences that meet similarity requirements; wherein the longest subsequence indicates the longest subsequence that is consistent with the set subsequence order; Determining the different contents between the two target longest subsequences, and marking the different contents as the different contents; Based on the difference content, the existence of any type of elements in the element relationship diagram is adjusted.

6. A device for tracking differences between documents, characterized in that: The device comprises: Sorting unit: used for sorting at least two documents according to the release time of the documents to obtain a document set; wherein the documents are documents containing preset content; A training unit: used to identify character features of any dimension in a sample document in a neural network model; based on the identified character features of any dimension, fit the character features of different dimensions to obtain fitting features, and determine the characters corresponding to the fitting features, as well as the elements corresponding to the characters, until the accuracy of the elements determined to be the preset elements corresponding to the characters is greater than a set threshold; based on the element categories in the element map, classify the elements in the sample document and determine the element categories to which the elements belong; until the accuracy of the element categories determined to be the preset element categories to which the elements belong is greater than a classification threshold; Model unit: used to determine the elements in each document in the document set through the neural network model to obtain an element set; wherein the neural network model is used to perform feature recognition and fitting of at least two dimensions on the characters in the document, and determine the elements corresponding to the fitting features based on the fitting features obtained by fitting; A generating unit is used to obtain an element relationship graph based on the correspondence between the elements in the element set and the documents; wherein the element relationship graph indicates the existence of any type of elements in different documents.

7. The device according to claim 6, characterized in that The training unit is specifically used to convert a paragraph of a first length in each sample document in the sample document set into a corresponding document vector in the neural network model; wherein the document vector is extracted based on the character features of any dimension of the characters in the paragraph; identify the elements in the document vector based on preset rules and determine the character features of any dimension; wherein the preset rules include the meanings represented by different combinations of different elements, and the correspondence between the meanings and the character features.

8. The device according to any one of claims 6 to 7, characterized in that The apparatus further includes an identification unit, specifically configured to add a document identifier to each of the segments of the second length in the document; wherein the segment is composed of at least one character, and the document identifier uniquely identifies the document corresponding to the segment; The generation unit is specifically used to extract the document identifier included in each element of each type of element in the element set; based on the correspondence between the document identifier and the document, establish the correspondence between the elements in the element set and the document, and obtain the element relationship graph.

9. The device according to claim 8, wherein The generation unit is also used to determine the difference features between the at least two documents in the element relationship diagram based on the existence situation; wherein the difference features are composed of at least one element; the first document among the at least two documents marks the difference features to obtain key information corresponding to each document; wherein the key information includes whether the element is a new element, the importance of the difference features, the relevance of the difference features to the corresponding regulatory content, and the semantic understanding output of the difference features.

10. The device according to claim 9, wherein The device also includes an optimization unit, which is specifically used to determine subsequences in two adjacent documents in the document set respectively; wherein the subsequence indicates a sentence in which characters are arranged in a set order in the two adjacent documents; determine the longest subsequence in the two adjacent documents through dynamic programming; compare any two longest subsequences in the two adjacent documents to determine two target longest subsequences that meet similarity requirements; wherein the longest subsequence indicates the longest subsequence that is consistent with the set subsequence order; determine the different content in the two target longest subsequences, and mark the different content as the different content; based on the different content, adjust the existence of any type of element in the element relationship graph.

11. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the method according to any one of claims 1 to 5 when executing the computer program stored in the memory.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

13. A computer program product, characterized in that When the computer program product is run on a computer, the computer is caused to execute the method according to any one of claims 1 to 5.