An electronic official document data security storage method

By constructing directed graphs and clustering analysis based on semantic relationships, the problem of lack of data integrity verification during encryption of electronic document data is solved, and secure storage and tampering detection of electronic document data is realized.

CN118981803BActive Publication Date: 2025-06-10WEI COUNTY ZHICHENG DATA COMMUNICATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411154936.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2025-06-10
Estimated Expiration
2044-08-21

AI Technical Summary

Technical Problem

The prior art lacks a data integrity verification mechanism when encrypting electronic document data, resulting in data being tampered with during transmission or storage, and key leakage or improper management will reduce the reliability of data integrity verification.

Method used

By performing word segmentation on electronic official document data, and constructing a directed graph based on the semantic relationship between word segmentation, semantic parameters and key parameters of the phrases are obtained, and clustered with information vectors, data tampering verification of electronic official document data is achieved.

Benefits of technology

This method can deeply understand the semantic changes in electronic document data, identify key nodes, accurately detect and verify the integrity and consistency of data, and improve the storage security of electronic document data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118981803B_ABST
    Figure CN118981803B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of natural language processing, and particularly relates to a method for secure storage of electronic official document data, including: performing word segmentation processing on the electronic official document data and constructing a directed graph; obtaining the key parameters and key nodes of the nodes by combining the distribution of the word vector similarity relationship and the connection relationship between the nodes; performing graph traversal search on the directed graph by combining the key parameters to obtain the main nodes, and obtaining the information vectors of the main nodes according to the key parameters and word vectors of the main nodes; screening the non-main nodes by using the frequency, key parameters and connection relationship in the directed graph of the non-main nodes to obtain secondary nodes; respectively obtaining the main nodes and secondary nodes of the electronic official document data before and after encryption and performing clustering in combination with the information vectors, so as to perform data tampering verification. The present invention can accurately and reliably verify the integrity of the electronic official document data before and after encryption, and improve the storage security of the electronic official document data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly relates to a method for securely storing electronic official document data. Background Art

[0002] In information-based office work, traditional paper official documents cannot meet the requirements of information-based management and are gradually replaced by electronic official documents. However, this also brings new security challenges, making the secure storage of electronic official documents an important part of information-based secure office work.

[0003] Currently, when encrypting electronic official document data, the encryption algorithms commonly used only provide encryption functions and do not provide data integrity verification. This means that when encrypting data, additional mechanisms may be needed to verify the integrity of the data to prevent the data from being tampered with during transmission or storage.

[0004] However, traditional data verification methods rely on keys to generate and verify authentication codes or signatures. If the key is leaked or mismanaged, the authentication code or signature will become invalid, thereby reducing the reliability of data integrity verification and resulting in insufficient accuracy and reliability of data security verification results. Summary of the Invention

[0005] The present invention provides a method for securely storing electronic official document data to solve the existing problems.

[0006] A method for securely storing electronic official document data according to the present invention adopts the following technical solution:

[0007] An embodiment of the present invention provides a method for securely storing electronic official document data, the method comprising the following steps:

[0008] Obtain electronic official document data, perform word segmentation processing on the electronic official document data, and construct a directed graph according to the semantic relationship between the word segments;

[0009] Obtain the semantic parameters of the word groups in combination with the distribution of the similarity relationship between the corresponding word vectors of the word groups, and obtain the key parameters of the nodes and the key nodes in the directed graph in combination with the semantic parameters of the word groups and the connection relationship between the corresponding nodes of the word groups in the directed graph;

[0010] Perform graph traversal search on the directed graph in combination with the key parameters to obtain a plurality of key node sequences, wherein the key node sequences contain a plurality of main nodes, and obtain the information vectors of the main nodes according to the key parameters and word vectors of the main nodes;

[0011] Screen the non-main nodes by using the frequency, key parameters and connection relationship in the directed graph of the non-main nodes to obtain secondary nodes;

[0012] Obtain the main nodes and secondary nodes of the e - official document data before and after encryption respectively, and perform clustering in combination with the information vector, and verify data tampering of the e - official document data according to the clustering results.

[0013] Further, the method for constructing a directed graph according to the semantic relationship between word segments includes the following specific steps:

[0014] Take any kind of phrase as a node in the directed graph, take the word order relationship of the phrase in the e - official document data as the edge of the phrase in the directed graph, take the word order of any two phrases in the e - official document data as the direction of the corresponding edge of the phrase in the directed graph, and take the frequency of the existence of the same word order as the weight of the corresponding edge. The directed graph contains several nodes, and the corresponding edges between different nodes include out - degree edges and in - degree edges.

[0015] Further, the method for obtaining the semantic parameter of a phrase by combining the distribution of the similarity relationship between the corresponding word vectors of the phrase includes the following specific steps:

[0016] Use the Word2Vec algorithm to convert each phrase in the e - official document data into a word vector;

[0017] For any kind of phrase, obtain the cosine similarity between the corresponding word vectors of any two phrases in the same kind of phrases;

[0018] Obtain the variance of the cosine similarities between the corresponding word vectors of all phrases in any kind of phrases, and record it as the semantic parameter of the kind of phrase.

[0019] Further, the method for obtaining the critical parameter of a node and the critical nodes in the directed graph by combining the semantic parameter of the phrase and the connection relationship between the corresponding nodes of the phrase in the directed graph includes the following specific steps:

[0020] Combine the semantic parameter of the phrase and the connection relationship between the corresponding nodes of the phrase in the directed graph to obtain the critical parameter of the corresponding node of the phrase in the directed graph;

[0021] Record the nodes with critical parameters greater than A as critical nodes, where A is a preset first parameter.

[0022] Further, the method for obtaining several critical node sequences by performing graph traversal search on the directed graph in combination with the critical parameters includes the following specific steps:

[0023] Use the depth - first search algorithm and the breadth - first search algorithm respectively, and perform graph traversal search on the directed graph in combination with the critical parameters of the nodes and the direction of the corresponding edges between the nodes, and obtain the critical node sequences, and the critical node sequences include the depth - critical node sequence and the breadth - critical node sequence.

[0024] Further, obtaining the information vector of the nodes in the key node sequence according to the criticality parameters and word vectors of the nodes in the key node sequence includes the following specific method:

[0025] Denote all the nodes in all key node sequences as main nodes, and obtain the vector formed by the criticality parameters of each main node and the word vectors of the corresponding phrases, which is denoted as the information vector of the main node.

[0026] Further, screening the non-main nodes by using the frequency, criticality parameters of the non-main nodes and their connection relationships in the directed graph to obtain secondary nodes includes the following specific method:

[0027] Obtain the frequency of the corresponding phrases of each non-main node in the electronic official document data, and combine the frequency, criticality parameters of the non-main nodes and their connection relationships in the directed graph to obtain the structure coefficient of the non-main nodes.

[0028] Further, respectively obtaining the main nodes and secondary nodes of the electronic official document data before and after encryption and performing clustering in combination with the information vectors includes the following specific method:

[0029] Obtain the main nodes and secondary nodes in the electronic official document data before and after encryption, and obtain the information vectors of all secondary nodes. The method for obtaining the information vectors of the secondary nodes is the same as that of the main nodes;

[0030] According to the Euclidean distance between the information vectors of the main nodes and secondary nodes and using the DBSCAN clustering algorithm, cluster all the main nodes and secondary nodes of the electronic official document data before and after encryption to obtain several clustering clusters.

[0031] Further, verifying data tampering of the electronic official document data according to the clustering results includes the following specific method:

[0032] Collectively refer to the main nodes and secondary nodes as text structure points, and obtain the non-tampering coefficient according to the number of clustering clusters and the proportion of the number of text structure points in the clustering clusters.

[0033] Further, the specific calculation method of the non-tampering coefficient is:

[0034]

[0035] Among them, represents the non-tampering coefficient of the electronic official document data; represents the number of clustering clusters; represents the total number of all text structure points; represents the number of text structure points in the largest clustering cluster; represents the number of key nodes in the directed graph corresponding to the electronic official document data before encryption; Indicates the number of key nodes in the directed graph corresponding to the encrypted e - official document data; Indicates obtaining the absolute value, Indicates the exponential function with the natural constant as the base.

[0036] The beneficial effects of the technical solution of the present invention are as follows: By constructing a directed graph based on the semantic relationships between word segments, the words and their semantic connections in the text can be structured, representing the semantic relationships and text structures between phrases in the e - official document data. By calculating the semantic parameters of phrases, the semantic characteristics of each phrase are reflected, so as to deeply understand the stability degree of the semantic changes of phrases in the e - official document data. Combining the semantic parameters of phrases and the node connection relationships in the directed graph, the key nodes with the most information and influence in the text can be identified, which helps to determine the main information and structure in the text. By traversing the directed graph and analyzing the key node sequence, the main information flow and important text parts in the text can be found. By analyzing the frequency, key parameters of non - main nodes and their connection relationships in the directed graph, the secondary nodes can be screened out, thus assisting in understanding and supplementing the main information. By obtaining the information vectors of the main and secondary nodes, the description degree of the main information in the e - official document data is improved. By respectively obtaining the main and secondary nodes of the e - official document data before and after encryption and clustering them in combination with the information vectors, the consistency and integrity of the e - official document data before and after encryption can be accurately and reliably detected and verified, improving the storage security of the e - official document data. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following - described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 It is the flowchart of the steps of a method for secure storage of e - official document data according to the present invention;

[0039] Figure 2 It is the schematic diagram of the directed graph structure;

[0040] Figure 3 It is the schematic diagram of the directed graph corresponding to the text data;

[0041] Figure 4 It is the schematic diagram of the directed graph corresponding to the text data after the mutual replacement of "Party A" and "Party B";

[0042] Figure 5 It is the schematic diagram of the directed graph corresponding to the text data after the mutual replacement of "rights" and "obligations". DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, detail the specific implementation manner, structure, features and effects of an electronic official document data security storage method proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.

[0045] The following will specifically describe the specific solution of an electronic official document data security storage method provided by the present invention with reference to the accompanying drawings.

[0046] Please refer to Figure 1 , which shows a step flowchart of an electronic official document data security storage method provided by an embodiment of the present invention. The method includes the following steps:

[0047] Step S001, obtain electronic official document data, perform word segmentation processing on the electronic official document data, and construct a directed graph according to the semantic relationship between the word segments.

[0048] It should be noted that as text data, the electronic official document data contains several types of word groups. The word groups form the text data through the combination of word order relationships. To better represent and quantify the text structure characteristics of the electronic official document data, by converting its data structure, a corresponding graph structure is constructed, which can effectively reflect the text structure of the electronic official document data while reducing the data volume.

[0049] Specifically, to implement an electronic official document data security storage method proposed in this embodiment, it is first necessary to collect electronic official document data, and the specific process is as follows:

[0050] First, obtain the text data in the electronic official document data. When the electronic official document data is in picture format, use OCR technology to extract the text data in the picture.

[0051] Then, use the Jieba word segmentation algorithm to perform word segmentation processing on the electronic official document data. During the word segmentation process, do not consider the text data in the electronic official document data to obtain several word groups, and regard the same word groups as one type of word group to obtain several types of word groups.

[0052] Finally, a directed graph is constructed. Any kind of phrase is used as a node in the directed graph, the word order relationship of the phrases in the electronic official document data is used as the edge in the directed graph, the word order of any two phrases in the electronic official document data is used as the direction of the corresponding edge in the directed graph, and the frequency of the same word order is used as the weight of the corresponding edge.

[0053] It should be noted that the directed graph contains several nodes. The corresponding edges between different nodes include out-degree edges and in-degree edges. For example, Figure 2 As shown in the schematic diagram of the directed graph structure, for node 2, edge 1 and edge 2 are the two out-degree edges of node 2, and edge 3 is the one in-degree edge of node 2. Weights 1, 2, and 3 are the weights of the corresponding edges 1, 2, and 3 respectively. Similarly, it can be known that edge 1 is the in-degree edge of node 1, edge 2 is the in-degree edge of node 3, and edge 3 is the out-degree edge of node 3.

[0054] For example: there is text data: The rights and obligations of Party A: responsible for the personnel management of Party B and responsible for paying Party B's salary; The rights and obligations of Party B: subject to the management of Party A and enjoy the salary treatment.

[0055] After word segmentation processing, the several kinds of word segments included are: Party A, 's, rights, and, obligations, responsible for, Party B, personnel, management, responsible for, paying, salary, accepting, enjoying, treatment. Then the schematic diagram of the corresponding directed graph of this text data is as Figure 3 shown.

[0056] It should be noted that the Jieba word segmentation algorithm is an existing word segmentation method, so it will not be specifically described in this embodiment.

[0057] So far, the electronic official document data and the corresponding directed graph are obtained through the above method.

[0058] Step S002: Combining the distribution of the similarity relationship between the corresponding word vectors of the phrases, obtain the semantic parameters of the phrases. Combining the semantic parameters of the phrases and the connection relationship between the corresponding nodes in the directed graph, obtain the key parameters of the nodes and the key nodes in the directed graph.

[0059] It should be noted that since the edges between different nodes in the directed graph represent the word order relationship of the phrases corresponding to the nodes in the electronic official document data, the network structure formed by the nodes reflects the text structure of the electronic official document data. Therefore, when the integrity is damaged during the encrypted storage process of the electronic official document data, that is, the electronic official document data is tampered with, the text structure of the electronic official document data will change, resulting in a change in the network structure of the directed graph.

[0060] The direction of the edges between nodes in a directed graph reflects the word order of the phrases corresponding to the nodes in the electronic official document data. When the electronic official document data is tampered with, the tampering situations usually include: word order swapping, phrase replacement, insertion, deletion, etc., which will cause changes in the number or word order of the phrases. When the number of phrases changes, the connection relationship between the nodes corresponding to the phrases and other nodes in the directed graph will change. When the word order of the phrases changes, the direction of the edges corresponding to the nodes of the phrases and other nodes will change. Additionally, since nouns, verbs, adjectives, prepositions, and conjunctions in official documents are crucial for understanding and processing the content of official documents, special attention should be paid to the phrases corresponding to the above parts of speech in the directed graph.

[0061] For example: If the text data of the above example is tampered with, such that "Party A" and "Party B" are mutually replaced, that is, it is tampered into: Rights and Obligations of Party B: Responsible for the personnel management of Party A, responsible for paying the salary of Party A; Rights and Obligations of Party A: Subject to the management of Party B, enjoy the salary treatment; as Figure 4 shown in the schematic diagram of the directed graph corresponding to the text data after "Party A" and "Party B" are mutually replaced.

[0062] According to Figure 4 it can be known that the connection relationship between the nodes corresponding to the phrases "Party A" and "Party B" and other nodes has changed significantly compared to before the tampering, and the semantic information expressed by the text data before and after the tampering has also changed greatly.

[0063] However, if the phrases "Rights" and "Obligations" are swapped, the text data becomes: Obligations and Rights of Party A: Responsible for the personnel management of Party B, responsible for paying the salary of Party B; Obligations and Rights of Party B: Subject to the management of Party A, enjoy the salary treatment; Although the connection relationship of the corresponding nodes in the corresponding directed graph will also change (as Figure 5 shown in the schematic diagram of the directed graph corresponding to the text data after "Rights" and "Obligations" are mutually replaced), the semantic information has not changed. Therefore, in the case of the same part of speech, the key degrees of different phrases will also be different.

[0064] It should be noted that in the embodiments of the present invention, considering that due to the different word order relationships or combination relationships between the same type of phrases and other phrases in the electronic official document data, the meanings expressed by the phrases included in the same type of phrases in the electronic official document data may be different. Therefore, by obtaining the similarity relationship between the word vectors corresponding to each phrase in the same type of phrases, the stability degree of the semantics expressed by this type of phrases is further analyzed.

[0065] Step 201, combine the distribution of the similarity relationship between the word vectors corresponding to the phrases to obtain the semantic parameters of the phrases.

[0066] Specifically, as an embodiment, the method for obtaining the semantic parameters of a phrase includes:

[0067] First, use the Word2Vec algorithm to convert each phrase in the electronic official document data into a word vector.

[0068] Then, for any kind of phrase, obtain the cosine similarity between the word vectors corresponding to any two phrases in the phrases of the same kind.

[0069] Finally, obtain the variance of the cosine similarities between the word vectors corresponding to all phrases in any kind of phrase, and denote it as the semantic parameter of the kind of phrase.

[0070] It should be noted that the semantic parameter is used to describe the stability degree of the corresponding semantic information of the corresponding kind of phrase in the electronic official document data. The smaller the semantic parameter, the present invention embodiment takes into account that the combination relationship change of the phrases of the same kind with other phrases in the text data will cause the semantic information to change, and the Word2Vec algorithm can obtain the vector of the phrase by combining the combination or word order relationship of the phrase with other phrases in the text data. Therefore, the present invention embodiment converts the phrase into a word vector to fully quantify and represent the specific meaning of the phrases included in the phrases of the same kind in the electronic official document data.

[0071] It should be noted that the Word2Vec algorithm is an existing word vector conversion algorithm, so it will not be specifically described in this embodiment.

[0072] Step 202: Combine the semantic parameter of the phrase and the connection relationship between the corresponding nodes of the phrase in the directed graph to obtain the key parameter of the corresponding node of the phrase in the directed graph.

[0073] It should be noted that in the electronic official document data, the phrases such as nouns, verbs, and adjectives usually have stable meanings, and the word order relationship or combination relationship with other phrases will not change much, so the corresponding semantic parameters are smaller; in addition, the present invention embodiment takes into account that although the phrases connected before and after some prepositions and conjunctions in the text data will change more, they can accurately represent the association between the front and back phrases, so the role generated in the text structure cannot be ignored. Therefore, the present invention embodiment combines the semantic parameter of the phrase and the connection relationship between the corresponding nodes of the phrase in the directed graph to obtain the key parameter of the corresponding node of the phrase in the directed graph in order to more accurately represent the text structure characteristics of the phrase in the electronic official document data, so as to facilitate subsequent verification of whether the electronic official document data has been tampered with.

[0074] Specifically, as an embodiment, the method for obtaining the key parameter of the corresponding node of the phrase in the directed graph includes:

[0075] Denote the nodes corresponding to any type of phrase in the electronic official document data in the directed graph as target nodes. The specific calculation method for the key parameters of the target nodes is as follows:

[0076]

[0077] Among them, represents the key parameter of the target node; represents the semantic parameter of the phrase corresponding to the target node; represents the semantic parameter of the node corresponding to the th out-degree edge of the target node; represents the th out-degree edge corresponding weight of the target node; represents the number of out-degree edges of the target node; represents the semantic parameter of the node corresponding to the th in-degree edge of the target node; represents the th in-degree edge corresponding weight of the target node; represents the number of in-degree edges of the target node; represents the linear normalization function.

[0078] Step 203: Denote the nodes with key parameters greater than A as key nodes, where A is a preset first parameter.

[0079] It should be noted that, according to experience, the first parameter A is usually preset to 0.7, and it can be adjusted according to the actual situation. This embodiment does not make specific limitations.

[0080] It should be noted that the smaller the semantic parameter of the phrase corresponding to the node, the more stable the semantic information expressed by the phrase corresponding to the node in the electronic official document data. When the weight of the out-degree edge or in-degree edge corresponding to the node is larger, it reflects that the text structure feature of the phrase corresponding to the node in the electronic official document data is more obvious, and the role played by the node in the text structure is greater, that is, the node is more critical. Therefore, the key parameter obtained by combining the semantic parameter of the node with the weights of the out-degree edge and in-degree edge corresponding to the node in this embodiment represents the critical degree of the corresponding node in the directed graph that can be used for subsequent verification of data security. That is, when the phrase corresponding to the node in the electronic official document data is tampered with, the greater the change in the semantic information in the electronic official document data, the higher the accuracy and reliability of the detection result when using the corresponding node for data security verification, and the greater the critical degree of the node corresponding to the phrase, and the greater the key parameter of the node.

[0081] Thus, the key parameters of each node in the directed graph and several key nodes are obtained through the above method.

[0082] Step S003: Perform a graph traversal search on the directed graph in combination with key parameters to obtain several key node sequences. Each key node sequence contains several main nodes. Obtain the information vectors of the main nodes based on the key parameters and word vectors of the main nodes.

[0083] It should be noted that in the embodiments of the present invention, considering that there are semantic or text structure associations between the electronic official document data and the phrases with high key degrees, the association structures formed between these phrases effectively represent the main text structures in the electronic official document data. Therefore, by performing a graph traversal search on the directed graph of the electronic official document data, nodes with obvious text structure features are obtained to facilitate subsequent data tampering verification.

[0084] Step 301: Use the depth-first search algorithm and the breadth-first search algorithm respectively, and perform a graph traversal search on the directed graph in combination with the key parameters of the nodes and the directions of the corresponding edges between the nodes to obtain key node sequences.

[0085] The key node sequences include a depth key node sequence and a breadth key node sequence. The specific acquisition method is as follows:

[0086] As an embodiment, the specific steps for obtaining the depth key node sequence are as follows:

[0087] First, use the key node with the largest key parameter in the directed graph as the starting point, use any key node as the ending point, and perform a traversal search on the directed graph using the depth-first search algorithm.

[0088] Then, during the search process, follow the directions of the corresponding edges between the nodes to perform the search, obtain several search paths from the starting point to any ending point, and obtain the sequence formed by all the nodes on the search path according to the search order, which is recorded as the first key node sequence. Among the key node sequences formed by the starting point and all the ending points respectively, the B first key node sequences with the largest number of key nodes are recorded as the depth key node sequence, where B is a preset second parameter.

[0089] It should be noted that according to experience, the second parameter is usually preset to 20 and can be adjusted according to the actual situation. The embodiments of the present invention do not make specific limitations.

[0090] As an embodiment, the specific steps for obtaining the breadth key node sequence are as follows:

[0091] First, use the breadth-first search algorithm to traverse the nodes in the directed graph to obtain the shortest paths between any two key nodes. During the process of traversing using the breadth-first search algorithm, also follow the directions between the nodes to perform the search.

[0092] Then, obtain the sequence formed by all the nodes on the C shortest paths with the largest number of key nodes passed among the shortest paths between all key nodes, which is denoted as the breadth key node sequence, where C is a preset third parameter.

[0093] It should be noted that according to experience, the third parameter is usually preset to 20, which can be adjusted according to the actual situation, and the embodiments of the present invention do not make specific limitations.

[0094] It should be noted that the embodiments of the present invention consider that since the connection direction between nodes in a directed graph reflects the word order of corresponding phrases in electronic official document data, and the change of word order also represents the situation where the electronic official document data has been tampered with. Therefore, in the process of graph traversal of the directed graph in the embodiments of the present invention, traversal search is performed according to the direction of the corresponding edges between nodes, effectively retaining the word order information between phrases and ensuring the reliability of the verification result when performing data tampering verification later; in addition, the method of following the direction of the corresponding edges between nodes during the search process is specifically, for example: in Figure 3 , it is only possible to traverse and search from "Party A" to "management", and it is not possible to search from "management" to "Party A".

[0095] Step 302, denote all the nodes in all key node sequences as main nodes, and obtain the vector formed by the criticality parameter of each main node and the word vector of the corresponding phrase, which is denoted as the information vector of the main node.

[0096] It should be noted that the method for obtaining the information vector is, for example: the criticality parameter of the main node is 0.8, and the word vectors of the corresponding phrase include: 11, 12, 21, then the information vectors of this main node are respectively: , , .

[0097] Step S004, screen non-main nodes by using the frequency, criticality parameter of non-main nodes and their connection relationships in the directed graph to obtain secondary nodes.

[0098] It should be noted that if only main nodes are used for data tampering verification, false positives and false negatives may occur due to over-reliance on the phrases corresponding to the main nodes. Therefore, the embodiments of the present invention select to screen secondary nodes by analyzing the frequency of non-main nodes in electronic official documents and the graph structure information in the directed graph on the premise of ensuring that the critical text structure can be verified without being tampered with.

[0099] First, obtain the frequency of the phrase corresponding to each non-main node in the electronic official document data, and combine the frequency, criticality parameter of the non-main node and its connection relationship in the directed graph to obtain the structure coefficient of the non-main node.

[0100] Optionally, in one embodiment, the specific calculation method for the structure coefficient of any non-primary node is as follows:

[0101]

[0102] where represents the structure coefficient of the non-primary node; represents the frequency of the phrase corresponding to the non-primary node in the e-official document data; represents the key parameter of the non-primary node; represents the number of out-degree edges of the non-primary node that connect to key nodes; represents the number of in-degree edges of the non-primary node that connect to key nodes; represents the exponential function with the natural constant as the base; represents the linear normalization function.

[0103] It should be noted that since data tamperers may bypass the data tampering verification using the primary node by changing the phrase corresponding to the non-primary node or using synonyms, screening the secondary primary nodes can help capture more subtle tampering traces. These phrases may have important semantic roles in specific contexts; in addition, although there are some nodes in the non-primary nodes whose corresponding phrases may have a low frequency of occurrence in the text, their existence or change may indicate tampering, including these phrases can help reduce false negatives in the verification, that is, detect tampering not detected by the keywords. Therefore, by combining the number of connections of the out-degree edge and the in-degree edge to the primary node, the key parameter of the non-primary node, and the frequency of the phrase corresponding to the non-primary node in the e-official document data, it more accurately reflects the degree of influence of the non-primary node on the overall meaning and text structure of the e-official document data. The larger the structure coefficient of the non-primary node, the greater the degree of influence of the phrase corresponding to the non-primary node in the e-official document data.

[0104] Then, the non-primary nodes with a structure coefficient greater than are used as secondary nodes, where D is a preset fourth parameter.

[0105] Finally, the information vectors of all secondary nodes are obtained, and the method for obtaining the information vectors of the secondary nodes is the same as that of the primary node.

[0106] It should be noted that according to experience, the fourth parameter is usually preset to 0.6, which can be adjusted according to the actual situation, and the embodiments of the present invention do not make specific limitations.

[0107] So far, several secondary nodes and the information vector of each secondary node are obtained through the above method.

[0108] Step S005: Obtain the primary nodes and secondary nodes of the e-official document data before and after encryption respectively, perform clustering in combination with the information vectors, and verify data tampering of the e-official document data according to the clustering results.

[0109] Specifically, in step 501: Obtain the primary nodes and secondary nodes of the e-official document data before and after encryption respectively, and perform clustering.

[0110] First, obtain the primary nodes and secondary nodes in the e-official document data before and after encryption.

[0111] Then, perform clustering on all the primary nodes and secondary nodes of the e-official document data before and after encryption according to the Euclidean distance of the information vectors between the primary nodes and secondary nodes and by using the DBSCAN clustering algorithm to obtain several clustering clusters.

[0112] Step 502: Verify data tampering of the e-official document data according to the clustering results.

[0113] First, collectively refer to the primary nodes and secondary nodes as text structure points, and obtain the non-tampering coefficient according to the number of clustering clusters and the proportion of the number of text structure points in the clustering clusters.

[0114] As an embodiment, the specific calculation method of the so-called non-tampering coefficient is:

[0115]

[0116] Wherein, represents the non-tampering coefficient of the e-official document data; represents the number of clustering clusters; represents the number of all text structure points; represents the number of text structure points in the largest clustering cluster; represents the number of key nodes in the directed graph corresponding to the e-official document data before encryption; represents the number of key nodes in the directed graph corresponding to the e-official document data after encryption; represents taking the absolute value, represents the exponential function with the natural constant as the base.

[0117] It should be noted that the larger the non-tampering coefficient, the greater the probability that the e-official document data has not been tampered with, or in other words, the smaller the degree of tampering of the e-official document data.

[0118] Then, judge whether the e-official document data has been tampered with according to the size of the non-tampering coefficient.

[0119] As an embodiment, the specific judgment method is:

[0120] When the e-official document data has not been tampered with;

[0121] When happens, it prompts that the content of the electronic official document data has changed;

[0122] When happens, a warning is issued to prompt that the electronic official document data has been tampered with, where E is a preset fifth parameter.

[0123] It should be noted that according to experience, E is preset to 0.95 and can be adjusted according to the actual situation. This embodiment does not make specific limitations.

[0124] So far, this embodiment is completed.

[0125] It should be noted that the model used in this embodiment is only used to represent the negative correlation relationship and restrict the result of the model output to be within the interval. In specific implementation, it can be replaced with other models with the same purpose. This embodiment only takes the model as an example for description and does not make specific limitations on it, where refers to the input of the model.

[0126] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for securely storing electronic document data, characterized in that: The method comprises the following steps: Acquire electronic official document data, perform word segmentation on the electronic official document data, and construct a directed graph based on the semantic relationship between the word segments; Combined with the distribution of similarity between word vectors corresponding to the phrase, the semantic parameters of the phrase are obtained. Combined with the semantic parameters of the phrase and the connection relationship between the corresponding nodes of the phrase in the directed graph, the key parameters of the node and the key nodes in the directed graph are obtained. Performing graph traversal search on the directed graph in combination with the key parameters to obtain a number of key node sequences, wherein the key node sequences include a number of main nodes, and obtaining information vectors of the main nodes according to the key parameters and word vectors of the main nodes; The non-primary nodes are screened by using their frequencies, key parameters and connection relationships in the directed graph to obtain secondary nodes. The primary nodes and secondary nodes of the electronic document data before and after encryption are obtained respectively and clustered in combination with the information vector, and data tampering verification is performed on the electronic document data according to the clustering results.

2. According to claim 1, a method for securely storing electronic official document data is characterized in that: The specific method of constructing a directed graph according to the semantic relationship between the segmentations is as follows: Any kind of phrase is taken as a node in a directed graph, the word order relationship of the phrases in the electronic document data is taken as the edge of the phrases in the directed graph, the word order of any two phrases in the electronic document data is taken as the direction of the corresponding edge of the phrases in the directed graph, and the frequency of the same word order is taken as the weight of the corresponding edge. The directed graph contains a number of nodes, and the corresponding edges between different nodes include out-degree edges and in-degree edges.

3. According to claim 1, a method for securely storing electronic official document data is characterized in that: The specific method of obtaining the semantic parameters of the phrase by combining the distribution of similarity relationships between the word vectors corresponding to the phrases is as follows: Use the Word2Vec algorithm to convert each phrase in the electronic document data into a word vector; For any kind of phrase, obtain the cosine similarity between the corresponding word vectors of any two phrases in the same kind of phrase; The variance of the cosine similarity between the word vectors corresponding to all phrases in any type of phrase is obtained, and recorded as the semantic parameter of the type of phrase.

4. According to claim 1, a method for securely storing electronic official document data is characterized in that: The method of combining the semantic parameters of the phrases and the connection relationship between the corresponding nodes of the phrases in the directed graph to obtain the key parameters of the nodes and the key nodes in the directed graph includes the following specific methods: Combining the semantic parameters of the phrase and the connection relationship between the corresponding nodes of the phrase in the directed graph, the key parameters of the corresponding nodes of the phrase in the directed graph are obtained; Nodes with a critical parameter greater than A are recorded as critical nodes, where A is a preset first parameter.

5. According to claim 1, a method for securely storing electronic official document data is characterized in that: The method of performing graph traversal search on the directed graph in combination with the key parameters to obtain a number of key node sequences includes: A depth-first search algorithm and a breadth-first search algorithm are used respectively, and a graph traversal search is performed on a directed graph in combination with key parameters of nodes and directions of corresponding edges between nodes to obtain a key node sequence, wherein the key node sequence includes a depth key node sequence and a breadth key node sequence.

6. According to claim 1, a method for securely storing electronic official document data is characterized in that: The specific method of obtaining the information vector of the node in the key node sequence according to the key parameters and word vectors of the node in the key node sequence includes: All nodes in all key node sequences are recorded as main nodes, and a vector formed by the key parameters of each main node and the word vector of the corresponding phrase is obtained, which is recorded as the information vector of the main node.

7. A method for securely storing electronic official document data according to claim 1, characterized in that: The method of screening the non-primary nodes by using the frequencies, key parameters and connection relationships of the non-primary nodes in the directed graph to obtain the secondary nodes includes: The frequency of each non-main node corresponding to a phrase in the electronic document data is obtained, and the structural coefficient of the non-main node is obtained by combining the frequency, key parameters and connection relationship of the non-main node in the directed graph.

8. A method for securely storing electronic official document data according to claim 6, characterized in that: The specific method of respectively obtaining the primary nodes and secondary nodes of the electronic document data before and after encryption and clustering them in combination with the information vector includes: Obtaining the primary node and the secondary node in the electronic document data before and after encryption, and obtaining the information vectors of all the secondary nodes, wherein the information vectors of the secondary nodes are obtained in the same manner as the information vectors of the primary nodes; According to the Euclidean distance of the information vector between the primary node and the secondary node, the DBSCAN clustering algorithm is used to cluster all the primary nodes and secondary nodes of the electronic document data before and after encryption to obtain several clusters.

9. A method for securely storing electronic official document data according to claim 1, characterized in that: The data tampering verification of the electronic document data according to the clustering result includes the following specific methods: The main nodes and secondary nodes are collectively referred to as text structure points. The non-tampering coefficient is obtained based on the number of clusters and the proportion of text structure points in the clusters.

10. A method for securely storing electronic official document data according to claim 9, characterized in that: The specific calculation method of the untampered coefficient is: in, Indicates the tamper-free coefficient of electronic document data; Indicates the number of clusters; Indicates the number of all text structure points; Indicates the number of text structure points in the largest cluster; Indicates the number of key nodes in the directed graph corresponding to the electronic document data before encryption; Indicates the number of key nodes in the directed graph corresponding to the encrypted electronic document data; Indicates obtaining the absolute value. Represents an exponential function with a natural constant as base.

Citation Information

Patent Citations

  • Chinese short text clustering method

    CN106599029A

  • Classified storage method and system based on financial text data

    CN118227798A