Data flow verification method and device for unstructured data

By obtaining user information and operations, and verifying access rights of unstructured data in combination with user databases, the security issues of unstructured data during transmission and storage of internal and external networks are solved, and flexible control and privacy protection are achieved.

CN120337251AActive Publication Date: 2025-07-18JIANGSU TIANYUAN TENERING CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510400827.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The prior art cannot effectively verify and control the security of unstructured data transmission and storage between internal and external networks, and there is a risk of data leakage and tampering, and user information is easily obtained by hackers and lacks flexible access rights control.

Method used

By obtaining user operations and user information, including user accounts, passwords, hardware tokens and IP addresses, and verifying them in combination with preset user databases, unstructured data is encrypted or decrypted based on the verification results, flexible access control of unstructured data is achieved.

Benefits of technology

It realizes flexible control of unstructured data access rights, reduces potential security risks, protects the privacy and security of unstructured data, and improves data processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337251A_ABST
    Figure CN120337251A_ABST
Patent Text Reader

Abstract

The invention discloses a data flow verification method and device for unstructured data. The method comprises the steps that user operation including writing and / or reading the unstructured data and user information including a user account, a user password, a hardware token and a user IP are obtained; verifying the user account, the user password, the hardware token and the user IP based on a preset user database to obtain a verification result; and based on the verification result, in combination with user operation, performing encryption / decryption operation on the unstructured data so as to complete circulation of the unstructured data. By obtaining the user operation and the user information, verifying the user information in combination with the user database and encrypting / decrypting the unstructured data, circulation of the unstructured data is completed, flexible control over the access permission of the unstructured data is achieved, an authorized user can download and browse the data, potential safety risks are reduced, and the user experience is improved. And the privacy security of the unstructured data is protected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data security, and particularly relates to a method and device for verifying data flow of unstructured data. Background Art

[0002] Unstructured data refers to data with irregular or incomplete data structures, without a predefined data model, and is inconvenient to be represented by a two-dimensional logic table of a database. Power data includes unstructured data, and power data security is of great significance for ensuring the stable operation of the power system. Data leakage and data tampering will greatly affect the normal operation of the power system. When power data is transmitted, the data transmission chain of power data is relatively long and involves the transmission of internal and external network environments.

[0003] Due to the particularity of unstructured data itself, there are relatively large risks of data leakage and data tampering during the transmission and storage of unstructured large files between the enterprise intranet and the enterprise extranet. However, the existing technologies cannot effectively verify whether unstructured data is leaked or tampered with during the transmission and storage between the internal and external networks.

[0004] Patent CN111475845A discloses a system and method for identity authorization access to unstructured data. The system grants access rights to unstructured data in the form of digital identity authentication, hashes the unstructured original data, encrypts the unstructured data with the public key during digital identity registration, and stores the hash value and the public key representing digital identity registration in the blockchain, ensuring the immutability of the hash fingerprint of unstructured data. The key of the unstructured data is encrypted by the off-chain archiving server to ensure the security of unstructured data storage. And the identity access control of unstructured data can be sub-authorized to achieve information sharing under data encryption protection.

[0005] Currently, there is no effective mechanism to ensure immutability and control access rights for the storage and access authorization of unstructured data, and there are various security vulnerabilities. And because the storage of unstructured data is easily obtained by hackers and other personnel through illegal means to obtain complete user information, it poses a great threat to the information security of users.

[0006] Therefore, how to flexibly control the access rights to unstructured data so that only authorized users can download and view the data, in order to reduce potential tampering and security risks and protect the privacy and security of unstructured data, is a problem that needs to be solved currently. Summary of the Invention

[0007] In view of the defects existing in the above-mentioned prior art, the present invention provides a method and device for verifying data flow of unstructured data. The method includes: obtaining a user operation and user information, where the user operation includes writing and / or reading unstructured data, and the user information includes a user account, a user password, a hardware token, and a user IP; verifying the user account, the user password, the hardware token, and the user IP based on a preset user database to obtain a verification result; based on the verification result, combining the user operation, performing an encryption / decryption operation on the unstructured data to complete the data flow of the unstructured data. By obtaining the user operation and user information, verifying the user information in combination with the user database, and performing an encryption and / or decryption operation on the unstructured data based on the obtained verification result, the data flow of the unstructured data is completed, the access rights to the unstructured data are flexibly controlled, enabling only authorized users to download and view the data, reducing potential security risks, and protecting the privacy and security of the unstructured data.

[0008] In a first aspect, the present invention provides a method for verifying data flow of unstructured data, specifically including the following steps:

[0009] Obtain a user operation and user information, where the user operation includes writing and / or reading unstructured data, and the user information includes a user account, a user password, a hardware token, and a user IP;

[0010] Verify the user account, the user password, the hardware token, and the user IP based on a preset user database to obtain a verification result;

[0011] Based on the verification result, combine the user operation, and perform an encryption and / or decryption operation on the unstructured data to complete the data flow of the unstructured data.

[0012] Further, based on the verification result, combine the user operation, and perform an encryption / decryption operation on the unstructured data to complete the data flow of the unstructured data, specifically including:

[0013] If the user operation is to read unstructured data and the verification result is verification passed, perform a decryption operation on the unstructured data to complete the reading of the unstructured data;

[0014] If the user operation is to write unstructured data and the verification result is verification passed, perform an encryption operation on the unstructured data to complete the writing of the unstructured data.

[0015] Further, performing a decryption operation on the unstructured data to complete the reading of the unstructured data specifically includes:

[0016] Based on the user operation, obtain the data ciphertext of the corresponding unstructured data;

[0017] Decrypt the ciphertext of the data in combination with the decryption private key corresponding to the unstructured data to complete the reading of the unstructured data, which is specifically expressed as:

[0018] M1 = (C1 d ) mod n

[0019] where M1 is the unstructured data, C1 is the ciphertext of the data, (n, d) is the decryption private key, and mod is the modulo operation.

[0020] Furthermore, perform a decryption operation on the unstructured data to complete the reading of the unstructured data, which specifically includes:

[0021] Based on the user operation, obtain multiple sub-ciphertexts of the corresponding unstructured data;

[0022] Combine the decryption private key corresponding to the unstructured data to decrypt the multiple sub-ciphertexts of the data to obtain multiple unstructured sub-data;

[0023] Concatenate the multiple unstructured sub-data to complete the reading of the unstructured data.

[0024] Furthermore, perform an encryption operation on the unstructured data to complete the writing of the unstructured data, which specifically includes:

[0025] Judge the volume of the unstructured data to obtain a volume judgment result;

[0026] Based on the volume judgment result, analyze the unstructured data to obtain unstructured intermediate data;

[0027] Obtain the encryption and decryption modulus, and give the Euler's totient function. Combine the random encryption exponent to encrypt the unstructured intermediate data to complete the writing of the unstructured data.

[0028] Furthermore, obtain the encryption and decryption modulus, and give the Euler's totient function. Combine the random encryption exponent to encrypt the unstructured intermediate data to complete the writing of the unstructured data, which specifically includes:

[0029] Randomly obtain two different prime numbers and calculate the product of the two prime numbers to obtain the encryption and decryption modulus;

[0030] Based on the two different prime numbers, combine with the encryption and decryption modulus to determine the Euler's totient function;

[0031] Randomly obtain the encryption exponent, combine with the Euler's totient function, and give the decryption exponent, where 1 < encryption exponent < Euler's totient function, and the encryption exponent and the Euler's totient function are relatively prime;

[0032] Encrypt the unstructured intermediate data based on the encryption and decryption modulus and the encryption exponent, and store the encrypted unstructured intermediate data together with the decryption exponent to complete the writing of the unstructured data.

[0033] Further, based on the volume judgment result, analyze the unstructured data to obtain unstructured intermediate data, specifically including:

[0034] If the volume judgment result is less than the volume threshold, give the hash value of the unstructured data;

[0035] Use the unstructured data and the hash value of the unstructured data as the unstructured intermediate data;

[0036] If the volume judgment result is greater than or equal to the volume threshold, split the unstructured data to obtain multiple unstructured intermediate data.

[0037] Further, split the unstructured data to obtain multiple unstructured intermediate data, specifically including:

[0038] Split the unstructured data according to the punctuation marks in the unstructured data to obtain multiple sub-fragments;

[0039] Based on the order of the sub-fragments, combine multiple consecutive sub-fragments to form candidate sub-texts, where each candidate sub-text includes consecutive m sub-fragments, the text volume of each candidate sub-text is less than the preset sub-text threshold, and the text volume of the candidate sub-text and the (m + 1)-th sub-fragment is greater than the preset sub-text threshold, m ∈ N + ;

[0040] Construct a directed acyclic graph according to the order between the candidate sub-texts, where each node in the directed acyclic graph includes at least one candidate sub-text, and the edge in the directed acyclic graph is the cross-degree of two candidate sub-texts;

[0041] Based on the directed acyclic graph, analyze the forward shortest path from the source point to the sink point and the reverse shortest path from the sink point to the source point in the directed acyclic graph, and determine the forward split point corresponding to the forward shortest path and the reverse split point corresponding to the reverse shortest path;

[0042] Compare the forward shortest path and the reverse shortest path, determine the target split point from the forward split point and the reverse split point, and split the unstructured data to obtain unstructured intermediate data.

[0043] Further, store the encrypted unstructured intermediate data, specifically including: storing the encrypted unstructured intermediate data in a file database, where the file database includes multiple blocks, and each block includes a block header and a block body; the block header is used to store the hash value of the current unstructured data, the hash value of the previous unstructured data, and the timestamp of the write time; the block body is used to store the current unstructured data.

[0044] In a second aspect, the present invention also provides a data transfer and verification device for unstructured data, which adopts the data transfer and verification method for unstructured data as described in any one of the above, including:

[0045] A data acquisition module, configured to acquire user operations and user information, where the user operations include writing and / or reading unstructured data, and the user information includes a user account, a user password, a hardware token, and a user IP;

[0046] A user verification module, configured to verify the user account, the user password, the hardware token, and the user IP based on a preset user database to obtain a verification result;

[0047] A data transfer module, configured to perform encryption and / or decryption operations on the unstructured data based on the verification result in combination with the user operation to complete the transfer of the unstructured data.

[0048] The data transfer and verification method and device for unstructured data provided by the present invention at least include the following beneficial effects:

[0049] (1) By acquiring user operations and user information, verifying the user information in combination with the user database, and performing encryption / decryption operations on the unstructured data based on the obtained verification result to complete the transfer of the unstructured data, flexible control of the access rights to the unstructured data is achieved, enabling only authorized users to download and view the data, reducing potential security risks, and protecting the privacy and security of the unstructured data.

[0050] (2) By splitting the unstructured data with a large volume and encrypting and storing each part of the split unstructured data respectively, while ensuring the security of the unstructured data, the processing efficiency of the unstructured data is also improved. Description of the Drawings

[0051] Figure 1 It is a flowchart of the data transfer and verification method for unstructured data provided by an embodiment of the present invention;

[0052] Figure 2 It is a flowchart of analyzing the verification result and user operation provided by an embodiment of the present invention;

[0053] Figure 3 Flowchart for decrypting unstructured data provided by an embodiment of the present invention;

[0054] Figure 4 Flowchart for encrypting unstructured data provided by an embodiment of the present invention;

[0055] Figure 5 Flowchart for obtaining unstructured intermediate data provided by an embodiment of the present invention;

[0056] Figure 6 Flowchart for splitting unstructured data provided by an embodiment of the present invention;

[0057] Figure 7 Flowchart for encrypting unstructured intermediate data provided by an embodiment of the present invention;

[0058] Figure 8 Schematic diagram of the structure of a file database provided by an embodiment of the present invention;

[0059] Figure 9 Schematic diagram of the relationship between operations and hash values provided by an embodiment of the present invention;

[0060] Figure 10 Block diagram of the structure of a data flow verification device for unstructured data provided by an embodiment of the present invention.

[0061] Among them, 201, data acquisition module; 202, user verification module; 203, data flow module. Detailed implementation manners

[0062] In order to better understand the above technical solutions, the following will describe the above technical solutions in detail in combination with the accompanying drawings of the specification and specific implementation manners. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0063] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. "Plurality" generally includes at least two.

[0064] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the commodity or device comprising said element.

[0065] The present invention provides a method and device for data transfer and verification of unstructured data. The method includes: obtaining user operations and user information, where the user operations include writing to and / or reading unstructured data, and the user information includes a user account, a user password, a hardware token, and a user IP; verifying the user account, the user password, the hardware token, and the user IP based on a preset user database to obtain a verification result; based on the verification result, combining with the user operations, performing encryption and / or decryption operations on the unstructured data to complete the transfer of the unstructured data. By verifying the user information and, after successful verification, encrypting the unstructured data written by the user and decrypting the unstructured data read by the user, intelligent encryption and decryption during the writing and reading of unstructured data are realized, ensuring the security of the use of unstructured data. At the same time, when writing to and reading unstructured data, the volume sizes of the unstructured data corresponding to the read and write operations are recorded, and intelligent operations are performed on the read and write operations of unstructured data of different volumes, accelerating the read and write efficiency of unstructured data. At the same time, the unstructured data is monitored and verified in real time to avoid security problems.

[0066] As Figure 1 shown, an embodiment of the present invention provides a method for data transfer and verification of unstructured data, and the specific steps are as follows:

[0067] S101: Obtain user operations and user information.

[0068] Specifically, the user operations include writing to and / or reading unstructured data, and the user information includes a user account, a user password, a hardware token, and a user IP (internet protocol).

[0069] Among them, a user account is a unique identifier used to identify and manage a user's identity. It is usually composed of a username or user ID and is used to distinguish different users. A user password is a verification credential for the user account and is used to verify the legitimacy of the user's identity. The password is usually a string of characters set by the user himself, including letters, numbers, and special symbols. A hardware token is a security device that enhances authentication. It provides high security by generating one-time passwords and is applicable to scenarios that require high security. The hardware token usually generates a random, one-time password regularly. This password is valid for a short period of time, usually only a few minutes. When the user logs in, he needs to enter this password, and the system will verify whether the password is correct. A user IP is a unique identifier assigned to a user device by the Internet Protocol and is used for network communication and device positioning. It has uniqueness and global universality.

[0070] It can be understood that based on different user requirements, user operations include writing unstructured data and reading unstructured data. Different user operations result in different processing of unstructured data in subsequent processes. Among them, the unstructured data can be a specific power file or a folder composed of multiple data files.

[0071] S102: Based on a preset user database, verify the user account, user password, hardware token, and user IP to obtain a verification result.

[0072] Specifically, when the user operates on unstructured data, obtain the user information and compare the user information with the user information already stored in the preset user database. If any one of the user account, user password, hardware token, and user IP fails to match, the verification result is that the verification fails; if all of the user account, user password, hardware token, and user IP match successfully, the verification result is that the verification passes.

[0073] It can be understood that when none of the four items in the user information match successfully, it means that the user is a stranger who has not read or written unstructured data and there is a certain risk. Therefore, the verification result is that the verification fails. When at least one item in the user information matches successfully, it means that an item of the user information is entered incorrectly and needs to be verified again until all the user information matches successfully and the verification passes.

[0074] By comparing the user information with the information in the user database, the security of the user is improved, and thus the security of the use of unstructured data is ensured.

[0075] S103: Based on the verification result and combined with the user operation, perform encryption and / or decryption operations on the unstructured data to complete the transfer of the unstructured data.

[0076] Further, referring to Figure 2, if the user operation is to read unstructured data and the verification result is verification passed, then perform a decryption operation on the unstructured data to complete the reading of the unstructured data;

[0077] If the user operation is to write unstructured data and the verification result is verification passed, then perform an encryption operation on the unstructured data to complete the writing of the unstructured data.

[0078] Further, performing a decryption operation on the unstructured data specifically includes:

[0079] Based on the user operation, obtain the data ciphertext corresponding to the unstructured data;

[0080] Combine the decryption private key corresponding to the unstructured data to decrypt the data ciphertext to complete the reading of the unstructured data, specifically expressed as:

[0081] M1 = (C1 d ) mod n

[0082] where M1 is the unstructured data, C1 is the data ciphertext, (n, d) is the decryption private key, and mod is the modulo operation.

[0083] In a specific implementation manner, when the user information passes the verification and reads or writes a certain power file, the user needs to input the decryption private key of the power file to be read, and read the data ciphertext C1 corresponding to the power file from the database based on the decryption private key of the power file to be read. Use the decryption private key (n, d) corresponding to the power file to decrypt the data ciphertext C1. Substitute the decryption private key (n, d) and the data ciphertext C1 into the decryption formula to obtain the unstructured data M1 corresponding to the power file, and feedback the decrypted unstructured data M1 to the user to complete the user's reading and viewing of the power file.

[0084] Further, referring to Figure 3 , performing a decryption operation on the unstructured data specifically includes:

[0085] Based on the user operation, obtain multiple sub-data ciphertexts corresponding to the unstructured data;

[0086] Combine the decryption private key corresponding to the unstructured data to decrypt the multiple sub-data ciphertexts to obtain multiple unstructured sub-data;

[0087] Concatenate the multiple unstructured sub-data to complete the reading of the unstructured data.

[0088] In a possible implementation, when the unstructured data is large, during the storage of the unstructured data, the unstructured data will be split, and each of the split parts will be encrypted and stored separately. Therefore, when reading this type of unstructured data, it is necessary to obtain the decryption private keys corresponding to each part, decrypt and splice the sub-data ciphertexts respectively, and finally obtain the complete unstructured data. The decryption process of the sub-data ciphertext is the same as the decryption process of the above data ciphertext, which will not be elaborated here. According to the order of the sub-data ciphertexts, decrypt and splice them in turn. After all the sub-data ciphertexts are decrypted, splice the decrypted sub-data ciphertexts to obtain the unstructured data.

[0089] Judge the integrity of the unstructured data. If the unstructured data is complete, send the complete unstructured data to the user to complete the user's reading of the unstructured data. If the unstructured data is incomplete, add the user IP and hardware token to the blacklist. In a specific example, the integrity detection of the unstructured data can be achieved by comparing the size of the unstructured data. In other examples, the integrity can be verified by calculating the checksum of the unstructured data (such as MD5, SHA-1, SHA-256, etc.). The integrity of the file can also be verified by calculating the hash value of the unstructured data. The hash value is a digital fingerprint with uniqueness and irreversibility.

[0090] Further, referring to Figure 4 , perform an encryption operation on the unstructured data, specifically including:

[0091] Judge the volume of the unstructured data to obtain a volume judgment result;

[0092] Based on the volume judgment result, analyze the unstructured data to obtain unstructured intermediate data;

[0093] Obtain the encryption and decryption modulus, and give the Euler's totient function. Combine with a random encryption exponent to encrypt the unstructured intermediate data to complete the writing of the unstructured data.

[0094] Among them, the Euler's totient function is the number of positive integers less than or equal to n that are relatively prime to n, where n is a positive integer.

[0095] Further, referring to Figure 5 , obtain the unstructured intermediate data, specifically including:

[0096] If the volume judgment result is less than the volume threshold, give the hash value of the unstructured data;

[0097] Take the unstructured data and the hash value of the unstructured data as the unstructured intermediate data;

[0098] If the volume judgment result is greater than or equal to the volume threshold, the unstructured data is split to obtain multiple unstructured intermediate data.

[0099] It can be understood that in order to improve the processing efficiency of unstructured data, and to ensure the security of unstructured data and reduce the risk of data leakage, when writing unstructured data, it is necessary to judge the volume of the unstructured data to obtain a volume judgment result. If the volume of the unstructured data is small, that is, the volume judgment result is less than the volume threshold, the unstructured data can be directly written; if the volume of the unstructured data is large, that is, the volume judgment result is greater than or equal to the volume threshold, the unstructured data needs to be split, and each part of the unstructured data is encrypted and stored respectively, while ensuring the security of the unstructured data, improving the data processing efficiency. The above volume threshold is set according to the actual situation. For example, it is set according to different device computing powers, and no limitation is made on this.

[0100] Furthermore, referring to Figure 6 , the unstructured data is split to obtain multiple unstructured intermediate data, which specifically includes:

[0101] The unstructured data is split according to the punctuation marks in the unstructured data to obtain multiple sub - segments;

[0102] Based on the order of the sub - segments, multiple consecutive sub - segments are combined to form candidate sub - texts. Among them, each candidate sub - text includes m consecutive sub - segments, the text volume of each candidate sub - text is less than the preset sub - text threshold, and the text volume of the candidate sub - text and the (m + 1) - th sub - segment is greater than the preset sub - text threshold, m ∈ N + ;

[0103] According to the order between the candidate sub - texts, a directed acyclic graph is constructed. Among them, each node in the directed acyclic graph includes at least one candidate sub - text, and the edge in the directed acyclic graph is the cross - degree of two candidate sub - texts;

[0104] Based on the directed acyclic graph, analyze the forward shortest path from the source point to the sink point and the reverse shortest path from the sink point to the source point in the directed acyclic graph, and determine the forward split point corresponding to the forward shortest path and the reverse split point corresponding to the reverse shortest path;

[0105] Compare the forward shortest path and the reverse shortest path, determine the target split point from the forward split point and the reverse split point, and split the unstructured data to obtain unstructured intermediate data.

[0106] In a specific implementation manner, the unstructured document is split according to the punctuation marks in the unstructured document to obtain a plurality of sub - segments, and the sub - segments are numbered in sequence, and the numbers of the sub - segments are 1, 2, 3, 4, 5, ……, t. Then, based on a preset sub - text threshold, the sub - segments are recombined to obtain a plurality of candidate sub - texts, and the candidate sub - texts are marked as p1, p2, p3, p4, p5, ……, py in sequence. In a specific example, for the candidate sub - text p1, p1 is composed of sub - segments 1, 2, and 3, and the total text volume of 1, 2, and 3 is less than the preset sub - text threshold, and the total text volume of 1, 2, 3, and 4 is greater than or equal to the preset sub - text threshold, that is, for the candidate sub - text p1, m is 3. There is only one combination method for the candidate sub - text p1. For other candidate sub - texts, there may be multiple combination methods. Taking the candidate sub - text p2 as an example, the candidate sub - text p2 can be composed of sub - segments 2, 3, and 4, or can be composed of sub - segments 3 and 4, or can be composed of sub - segments 4 and 5. It can be understood that in the above example, the combination methods of the candidate sub - text p2 all meet the limitations regarding the sub - text threshold. Through the combination method of the candidate sub - text p2, it can be understood that all other candidate sub - texts may include multiple combination methods. After completing the combination of all candidate sub - texts, a directed acyclic graph is constructed based on all the candidate sub - texts. The source point in the directed acyclic graph is the candidate sub - text p1, and the source point is a node with an in - degree of zero. The nodes connected to the candidate sub - text p1 are different candidate sub - texts p2, and the edge between p1 and p2 is the cross - degree of p1 and p2, that is, the square of the total number of cross - text characters between p1 and p2. For example, if p1 is composed of sub - segments 1, 2, and 3, and p2 is composed of sub - segments 2, 3, and 4, then the cross - text between p1 and p2 is sub - segments 2 and 3, and the square of the total number of characters in sub - segments 2 and 3 is used as the cross - degree between p1 and the corresponding p2. Similarly, after completing the construction of other candidate sub - texts, the construction of the directed acyclic graph is completed. The sink point in the directed acyclic graph is the candidate sub - text py, and the sink point is a node with an out - degree of zero.

[0107] After the construction of the directed acyclic graph is completed, the source point is used as the starting point and the sink point is used as the ending point to find the forward shortest path from the source point to the sink point. Then, the sink point is used as the starting point and the source point is used as the ending point to find the reverse shortest path from the sink point to the source point. Both the forward shortest path and the reverse shortest path are composed of various candidate sub-texts. The end of the cross-text of any two candidate sub-texts in the forward shortest path is used as the forward segmentation point. Similarly, the end of the cross-text of any two candidate sub-texts in the reverse shortest path is used as the reverse segmentation point. Judge the number of forward segmentation points and reverse segmentation points. If the number of forward segmentation points is less than or equal to the number of reverse segmentation points, the unstructured data is segmented according to the forward segmentation points. Conversely, if the number of forward segmentation points is greater than the number of reverse segmentation points, the unstructured data is segmented according to the reverse segmentation points to obtain multiple unstructured intermediate data.

[0108] Further, if the volume judgment result is less than the volume threshold, the hash value of the unstructured data is given;

[0109] The unstructured data and the hash value of the unstructured data are used as unstructured intermediate data.

[0110] In the example provided by the present invention, the hash value of the above unstructured data is calculated by the MD5 hash function. In other embodiments, other hash functions can be used, which are not limited herein.

[0111] Further, with reference to Figure 7 , the unstructured intermediate data is encrypted, specifically including:

[0112] Randomly obtain two different prime numbers and calculate the product of the two prime numbers to obtain the encryption and decryption modulus;

[0113] Based on the two different prime numbers and in combination with the encryption and decryption modulus, determine the Euler's totient function;

[0114] Randomly obtain the encryption exponent, and in combination with the Euler's totient function, give the decryption exponent, where 1 < encryption exponent < Euler's totient function, and the encryption exponent and the Euler's totient function are relatively prime;

[0115] Based on the encryption and decryption modulus and the encryption exponent, encrypt the unstructured intermediate data, and store the encrypted unstructured intermediate data together with the decryption exponent to complete the writing of the unstructured data.

[0116] In a specific embodiment, when the user operation is to write unstructured data, the unstructured data that the user needs to write is obtained. Then, an encryption public key and a decryption private key corresponding to the unstructured data are generated, where the encryption public key includes the encryption and decryption modulus and the encryption exponent, and the decryption private key includes the encryption and decryption modulus and the decryption exponent. The specific steps for determining the encryption public key and the decryption private key are as follows:

[0117] First, randomly select two different prime numbers p and q, and calculate the product of the two prime numbers to obtain the encryption and decryption modulus n for unstructured data. Based on the encryption and decryption modulus n, the Euler's totient function φ(n) corresponding to the unstructured data is given, specifically expressed as: φ(n) = (p - 1) × (q - 1). Randomly select an encryption exponent w, where 1 < w < φ(n) and w is relatively prime to φ(n). Through the Euler's totient function and the encryption exponent, the decryption exponent d is given, specifically expressed as: w × d ≡ 1 (mod φ(n)), where ≡ represents the congruence relation, and mod is the modulo operation, that is, the remainders of w * d and 1 divided by φ(n) are the same. Finally, send the encryption public key (n, w) and the decryption private key (n, d) to the database for storage. Based on the encryption public key, the encryption of the unstructured intermediate data can be completed, and the data ciphertext C1 corresponding to the unstructured intermediate data is obtained, specifically expressed as:

[0118] C1 = (M1 w ) mod n

[0119] where M1 is the unstructured intermediate data, mod is the modulo operation, n is the encryption and decryption modulus, and w is the encryption exponent.

[0120] After completing the encryption of the unstructured intermediate data, store the encrypted unstructured intermediate data in the file database. Among them, referring to Figure 8 , the file database includes multiple blocks, and each block includes a block header and a block body; the block header is used to store the hash value of the current unstructured data, the hash value of the previous unstructured data, and the timestamp of the write time; the block body is used to store the current unstructured data.

[0121] Based on the data transfer and verification method for unstructured data, it also includes the judgment of whether the unstructured data has been tampered with. The specific steps are as follows:

[0122] Obtain the hash value in each block of the file database and record it as the leaf hash value i. i is the number of the hash value in each block, and i is a positive integer;

[0123] Calculate the root hash value of each block. Among them, the root hash value is the sum of the respective leaf hash values in the corresponding block, specifically expressed as: root hash value = leaf hash value 1 + leaf hash value 2 +... + leaf hash value x, where x is the number of leaf hash values in the corresponding block;

[0124] Obtain the historical root hash value and compare the historical hash value with the root hash value;

[0125] If the root hash value is not equal to the historical root hash value, it means that the unstructured file in the file database has been tampered with, and an exception alarm is generated;

[0126] If the root hash value is equal to the historical root hash value, it indicates that the unstructured data in the file database has not been tampered with, and the relevant operations such as monitoring or reading the unstructured data are continued to be completed.

[0127] In other embodiments, if the root hash value is not equal to the historical root hash value, that is, the unstructured file in the file database is tampered with, the backup data is obtained and the unstructured data in the corresponding block is repaired. After the repair is completed, the relevant operations such as monitoring or reading the unstructured data are continued to be completed. At the same time, the user IP corresponding to the user operation and the hardware token are added to the blacklist.

[0128] In a specific example, referring to Figure 9 , Operations 1 to 8 represent writing or modifying unstructured data. If any one of Operations 1 to 8 is successfully executed, the hash value corresponding to the successfully executed operation changes. Taking Operations 1 and 2 as an example, Hash 12 is equal to Hash 1 plus Hash 2. If any one of Hash 1 and Hash 2 changes, then Hash 12 will change; the root hash is equal to Hash 1234 plus Hash 5678. If any one of Operations 1 to 8 is successfully executed, the root hash value will change.

[0129] It should be specifically noted that the historical root hash value is the root hash value in all blocks when no operation is performed on the unstructured data in the file database. The historical root hash value is stored in the file database. If the user normally operates on the unstructured data in the file database, the historical root hash value in the database is updated. When the unstructured data in the file database is tampered with, the corresponding hash value changes, which further causes the blocks to be unable to be linked, thereby further improving the security of the power files.

[0130] Referring to Figure 10 , an embodiment of the present invention provides a data flow verification device for unstructured data, including:

[0131] A data acquisition module 201, configured to acquire user operations and user information, where the user operations include writing to and / or reading unstructured data, and the user information includes a user account, a user password, a hardware token, and a user IP;

[0132] A user verification module 202, configured to verify the user account, the user password, the hardware token, and the user IP based on a preset user database to obtain a verification result;

[0133] A data flow module 203, configured to perform encryption and / or decryption operations on the unstructured data based on the verification result and in combination with the user operation to complete the flow of the unstructured data.

[0134] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the described modules can be referred to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0135] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention. Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and variations.

Claims

1. A data flow verification method for unstructured data, characterized in that, Including: Obtain user operations and user information, where the user operations include writing and / or reading unstructured data, and the user information includes user accounts, user passwords, hardware tokens, and user IPs; Based on a preset user database, verify the user account, user password, hardware token, and user IP to obtain a verification result; Based on the verification result, combined with the user operation, perform encryption and / or decryption operations on the unstructured data to complete the transfer of the unstructured data.

2. The data flow verification method for unstructured data according to claim 1, wherein Based on the verification result, combined with the user operation, perform encryption / decryption operations on the unstructured data to complete the transfer of the unstructured data, specifically including: If the user operation is to read unstructured data and the verification result is verification passed, decrypt the unstructured data to complete the reading of the unstructured data; If the user operation is to write unstructured data and the verification result is verification passed, encrypt the unstructured data to complete the writing of the unstructured data.

3. The data flow verification method for unstructured data according to claim 2, wherein Perform a decryption operation on the unstructured data to complete the reading of the unstructured data, specifically including: Based on the user operation, obtain the data ciphertext corresponding to the unstructured data; Combine the decryption private key corresponding to the unstructured data to decrypt the data ciphertext to complete the reading of the unstructured data, specifically expressed as: M1 = (C1 d ) mod n Where M1 is the unstructured data, C1 is the data ciphertext, (n, d) is the decryption private key, and mod is the modulo operation.

4. The data flow verification method for unstructured data according to claim 2, wherein Perform a decryption operation on the unstructured data to complete the reading of the unstructured data, specifically including: Based on the user operation, obtain multiple sub-data ciphertexts corresponding to the unstructured data; Combine the decryption private key corresponding to the unstructured data to decrypt the multiple sub-data ciphertexts to obtain multiple unstructured sub-data; Concatenate the multiple unstructured sub-data to complete the reading of the unstructured data.

5. The data flow verification method for unstructured data according to claim 2, wherein Perform an encryption operation on the unstructured data to complete the writing of the unstructured data, specifically including: Judge the volume of the unstructured data to obtain a volume judgment result; Based on the volume judgment result, analyze the unstructured data to obtain unstructured intermediate data; Obtain the encryption / decryption modulus, give the Euler's totient function, and combine a random encryption exponent to encrypt the unstructured intermediate data to complete the writing of the unstructured data.

6. The data flow verification method for unstructured data according to claim 5, wherein Obtain the encryption / decryption modulus, give the Euler's totient function, and combine a random encryption exponent to encrypt the unstructured intermediate data to complete the writing of the unstructured data, specifically including: Randomly obtain two different prime numbers and calculate the product of the two prime numbers to obtain the encryption / decryption modulus; Based on the two different prime numbers, combined with the encryption / decryption modulus, determine the Euler's totient function; Randomly obtain an encryption exponent, combine the Euler's totient function, and give a decryption exponent, where 1 < encryption exponent < Euler's totient function, and the encryption exponent and Euler's totient function are relatively prime; Based on the encryption / decryption modulus and the encryption exponent, encrypt the unstructured intermediate data and store the encrypted unstructured intermediate data together with the decryption exponent to complete the writing of the unstructured data.

7. The data flow verification method for unstructured data according to claim 5, wherein Based on the volume judgment result, analyze the unstructured data to obtain unstructured intermediate data, specifically including: If the volume judgment result is less than the volume threshold, the hash value of the unstructured data is given; The unstructured data and the hash value of the unstructured data are used as unstructured intermediate data; If the volume judgment result is greater than or equal to the volume threshold, the unstructured data is split to obtain multiple unstructured intermediate data.

8. The data flow verification method for unstructured data according to claim 7, wherein The unstructured data is split to obtain multiple unstructured intermediate data, which specifically includes: The unstructured data is split according to the punctuation marks in the unstructured data to obtain multiple sub-fragments; Based on the order of sub - segments, a plurality of consecutive sub - segments are combined to form candidate sub - texts, where each candidate sub - text includes m consecutive sub - segments, the text volume of each candidate sub - text is less than a preset sub - text threshold, and the text volume of the candidate sub - text and the (m + 1) - th sub - segment is greater than the preset sub - text threshold, m ∈ N + ; A directed acyclic graph is constructed according to the order between candidate sub-texts. Each node in the directed acyclic graph includes at least one candidate sub-text, and the edge in the directed acyclic graph is the cross-degree of two candidate sub-texts; Based on the directed acyclic graph, analyze the forward shortest path from the source point to the sink point and the reverse shortest path from the sink point to the source point in the directed acyclic graph, and determine the forward split point corresponding to the forward shortest path and the reverse split point corresponding to the reverse shortest path; Compare the forward shortest path and the reverse shortest path, determine the target split point from the forward split point and the reverse split point, and split the unstructured data to obtain unstructured intermediate data.

9. The data flow verification method for unstructured data according to claim 6, wherein The encrypted unstructured intermediate data is stored, which specifically includes: storing the encrypted unstructured intermediate data in a file database. The file database includes multiple blocks, and each block includes a block header and a block body; the block header is used to store the hash value of the current unstructured data, the hash value of the previous unstructured data, and the time stamp of the write time; the block body is used to store the current unstructured data.

10. A data flow verification device for unstructured data, characterized in that, The data flow verification method for unstructured data as described in any one of claims 1-9 is adopted, including: A data acquisition module, which is used to acquire user operations and user information. The user operations include writing and / or reading unstructured data, and the user information includes user accounts, user passwords, hardware tokens, and user IPs; A user verification module, which is used to verify the user account, user password, hardware token, and user IP based on a preset user database to obtain a verification result; A data flow module, which is used to perform encryption and / or decryption operations on the unstructured data based on the verification result and in combination with the user operation to complete the flow of the unstructured data.

Citation Information

Patent Citations

  • A method, system, storage medium, and device for parsing semi-structured data.

    CN114936026A

  • Data ingestion pipeline anomaly detection

    US20210117232A1

  • Relying on discourse analysis to answer complex questions by neural machine reading comprehension

    US20220138432A1

  • Method and system for interpreting inputted information

    US20250045304A1