Safe processing method and system for big data

By constructing a hash tree and hash index mechanism, the problem of difficulty in locating the tampered position in the hash algorithm is solved, rapid tampering detection and positioning is achieved, and the security processing efficiency and anti-tampering capability of text data are improved.

CN120850359APending Publication Date: 2025-10-28CHINA INFORMATION SAFETY RES INST CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511003312.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing hash algorithms are unable to locate tampering locations or restore original content, resulting in low efficiency in data tampering detection.

Method used

By generating identifiers, building hash trees and hash index mechanisms, and using hash functions to perform hierarchical hashing on text data, a unique verifier is generated, and combined with state snapshots and distributed storage, rapid tampering detection and positioning can be achieved.

Benefits of technology

It improves the security processing efficiency of text data, reduces storage burden and computing overhead, can quickly locate tampering locations, shorten processing time, and enhance anti-tampering capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120850359A_ABST
    Figure CN120850359A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of text management, and particularly relates to a security processing method and system for big data, and the method comprises the steps: setting a plurality of common characters, selecting text data needing to be subjected to security processing, traversing the first appearance position of each common character, calculating the character bit order, integrating all the character bit orders, and carrying out security processing on the text data. Generating identifiers, wherein each piece of text data corresponds to one identifier; the method comprises the following steps: creating text data, creating child nodes in one-to-one correspondence with the text data, writing identifiers into the child nodes, selecting a hash function, performing pairwise hash on adjacent child nodes to obtain a plurality of first hash values, creating a plurality of levels of father nodes, writing the first hash values into the father nodes, and mounting the corresponding child nodes into the father nodes. By intercepting the state snapshot, lightweight storage can be carried out on the Hash tree, the storage burden is further reduced, and the safety processing efficiency and tamper-proofing capability of the text data are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of text management technology, and in particular to a secure processing method and system for big data. Background Technology

[0002] Data security processing refers to the use of a series of technical means and management measures to ensure the integrity and security of data during storage and use, and to prevent data from being illegally accessed, leaked, tampered with or lost. Its core includes data encryption, access control, identity authentication and data backup and recovery.

[0003] In existing technologies, tamper-proof verification is generally carried out through hash algorithms, digital signatures, and blockchain. Although hash algorithms can quickly detect whether content has been modified, they cannot locate the location of the tampering or restore the original content.

[0004] Therefore, "how to eliminate the location defect caused by tampering in hash algorithms" is the technical problem that this invention needs to solve. Summary of the Invention

[0005] The purpose of this invention is to provide a secure processing method and system for big data, in order to solve the problem of "how to eliminate the location defect caused by tampering in hash algorithms" mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] A method for securely processing big data, the method comprising:

[0008] Set several commonly used characters, select the text data that needs to be processed for security, traverse the first occurrence position of each commonly used character, calculate the character position, integrate all the character positions, and generate an identifier, where each piece of text data corresponds to an identifier.

[0009] Create child nodes that correspond one-to-one with the text data, write the identifier into the child nodes, select a hash function, perform pairwise hashing on adjacent child nodes to obtain several first hash values, create parent nodes at multiple levels, write the first hash values ​​into the parent nodes, and attach the corresponding child nodes to the parent nodes, continue to perform pairwise hashing on adjacent parent nodes to obtain second hash values, and so on, until a unique hash value is obtained, which is recorded as the verification symbol;

[0010] Create a root node and synchronize the validator to the root node. Attach the parent node to the root node, integrate the child nodes, parent nodes, and root node to generate a hash tree, and extract a state snapshot.

[0011] Identify the storage nodes of the text data, send status snapshots to all storage nodes in parallel, send commonly used characters to preset terminals, and set verification rules;

[0012] It receives user-uploaded call requests and uses a hash index mechanism built based on common characters and status snapshots to determine whether to grant call permissions.

[0013] Furthermore, the step of integrating all character positions to generate the identifier includes:

[0014] Check each identifier sequentially to see if it is unique. If not, insert an additional character into the commonly used characters and update the identifier.

[0015] All identifiers are integrated to generate a check set, and a periodic check mechanism is embedded.

[0016] Furthermore, the step of creating child nodes that correspond one-to-one with the text data and writing the identifier into the child nodes includes:

[0017] Configure a number for each piece of text data, and establish a correspondence between the text data and its child nodes based on the number.

[0018] Calculate the frequency of use for each text data and adjust the order of child nodes in the hash tree according to the frequency of use from highest to lowest.

[0019] Furthermore, the step of integrating child nodes, parent nodes, and root nodes to generate a hash tree includes:

[0020] An authentication policy is integrated into the hash tree, granting permission to modify text data upon successful authentication.

[0021] The modification process is recorded, and the hash tree is adjusted, including growth and breakage.

[0022] Furthermore, the method also includes:

[0023] Through the modification process, incremental data is identified, and a one-to-one correspondence between incremental data and state snapshots is established;

[0024] Based on the aforementioned correspondence, rollback points are created, and a rollback mechanism is generated.

[0025] Furthermore, the step of the storage node for the identified text data sending state snapshots to all storage nodes in parallel includes:

[0026] The storage node is divided into cloud nodes and local nodes. Text data with a usage frequency greater than a threshold is stored in the local node, and the remaining text data is stored in the cloud node.

[0027] Utilize all storage nodes to build a distributed storage architecture.

[0028] Furthermore, the method also includes:

[0029] Collect source data of the call request, wherein the source data includes at least: login IP and user identity, and activate the authentication policy;

[0030] The hash indexing mechanism is used to determine whether the validator has changed. If not, the target data is retrieved and access permissions are granted.

[0031] Furthermore, the system includes:

[0032] The identifier module is used to set several commonly used characters, select the text data that needs to be processed for security, traverse the first occurrence position of each commonly used character, calculate the character position, integrate all the character positions, and generate an identifier, where each piece of text data corresponds to one identifier.

[0033] The verification module is used to create child nodes that correspond one-to-one with the text data, write the identifier into the child node, select a hash function, perform pairwise hashing on adjacent child nodes to obtain several first hash values, create parent nodes at multiple levels, write the first hash values ​​into the parent nodes, and attach the corresponding child nodes to the parent nodes, continue to perform pairwise hashing on adjacent parent nodes to obtain second hash values, and so on, until a unique hash value is obtained, which is recorded as the verification symbol.

[0034] The hash module is used to create the root node, synchronize the validator to the root node, attach the parent node to the root node, integrate the child nodes, parent nodes and the root node to generate a hash tree, extract state snapshots, identify the storage nodes of text data, send the state snapshots to all storage nodes in parallel, send commonly used characters to preset terminals, and set verification rules.

[0035] The judgment module receives the call request uploaded by the user, and uses a hash index mechanism built based on common characters and status snapshots to determine whether to grant call permissions.

[0036] Furthermore, the identification module includes:

[0037] An insertion unit is used to sequentially determine whether each identifier is unique. If not, an additional character is inserted into the commonly used characters to update the identifier.

[0038] The embedding unit is used to integrate all identifiers, generate a check set, and embed a periodic check mechanism.

[0039] Furthermore, the verification module includes:

[0040] The configuration unit is used to configure the number of each text data and establish the correspondence between the text data and the child nodes through the number;

[0041] The calculation unit is used to calculate the frequency of use of each text data and adjust the order of child nodes in the hash tree according to the frequency of use from highest to lowest.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] By defining identifiers, text recognition can be improved while reflecting the overall structural features of the text, facilitating text data retrieval and reducing redundant information storage. By constructing a hash tree, it is possible to quickly determine whether text data has been tampered with without verifying the entire dataset, greatly improving verification efficiency. Furthermore, only the hash tree needs to be stored and maintained, significantly reducing storage burden and computational overhead. When data is tampered with, the location of the tampering can be quickly identified, accurately tracing the source of the problem and shortening the handling time. By capturing state snapshots, the hash tree can be stored in a lightweight manner, further reducing the storage burden and greatly improving the efficiency of secure text data processing and its anti-tampering capabilities. Attached Figure Description

[0044] Figure 1 This is a schematic diagram of a hash tree in the big data security processing method provided in this embodiment of the invention;

[0045] Figure 2 A flowchart illustrating a secure big data processing method provided in an embodiment of the present invention;

[0046] Figure 3 This is a first sub-flowchart of the big data security processing method provided in an embodiment of the present invention;

[0047] Figure 4 This is a second sub-flowchart of the big data security processing method provided in an embodiment of the present invention;

[0048] Figure 5 A third sub-flow diagram of the big data security processing method provided in an embodiment of the present invention;

[0049] Figure 6 A block diagram illustrating the composition of a secure big data processing system provided in an embodiment of the present invention;

[0050] Figure 7 A block diagram illustrating the composition of the identification module in the big data security processing system provided in this embodiment of the invention;

[0051] Figure 8 A block diagram illustrating the composition of the verification module in the big data security processing system provided in this embodiment of the invention;

[0052] Figure 9 It is a block diagram of the composition of the hash module in the big data security processing system provided by the embodiments of the present invention. Specific embodiments

[0053] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0054] In Embodiment 1, Figure 1 and Figure 2 show the implementation process of the big data security processing method provided by the embodiments of the present invention, which will be described in detail below, as follows:

[0055] S100: Set a number of common characters, select the text data that needs to be securely processed, traverse the first occurrence position of each common character, calculate the character position, and integrate all the character positions to generate an identifier, where each text data corresponds to an identifier.

[0056] Set multiple common characters. Common characters are those characters that often appear. For example, "de", "shi", "le", etc.; the text data that needs to be securely processed refers to the text data that needs to be managed to prevent being tampered with; according to the common characters, traverse each text data in turn, find the first occurrence position of each common character in the text data, and record the character position of this position in the text data (that is, the specific sequence number in the entire text data); for example, if the 57th character in the text data is "de", then the character position is 57; integrate all the common characters to generate a set of characteristic position sequences composed of the first occurrence positions of the common characters, that is, the identifier. The identifier can be used for subsequent data verification, anti-tampering identification or traceability processing. In actual use, it may occur that the identifiers of two text data are the same. In this case, the number of common characters needs to be increased to avoid the occurrence of the same identifier.

[0057] On this basis, if the sensitivity of the text data is relatively high, more common characters can be set.

[0058] S200: Create child nodes corresponding one-to-one to the text data, write the identifier into the child nodes, select a hash function, perform pairwise hashing on adjacent child nodes to obtain a number of first hash values, create multiple levels of parent nodes, write the first hash values into the parent nodes, and mount the corresponding child nodes to the parent nodes. Continue to perform pairwise hashing on adjacent parent nodes to obtain a second hash value, and so on until a unique hash value is obtained, denoted as the verification symbol.

[0059] Create child nodes, where each piece of text data corresponds to a child node. Child nodes, parent nodes, and the root node are all logical nodes, not physical processing devices, similar to data packets. Write the identifier of each piece of text data into its corresponding child node. Select a hash function (e.g., SHA-256). Note that this hash function can convert data of any length into a fixed-length hash value, typically 256 bits (32 bytes). Sort all child nodes in descending order of frequency of use, and combine the identifiers of adjacent child nodes. Use the hash function to hash the combined identifier to obtain the first hash value. To reduce computation, only the first 10 bytes of the hash result are used as the first hash value. Create parent nodes, which include multiple levels, such as first-level parent nodes, second-level parent nodes, etc. (as shown in the instruction manual). Figure 1 As shown, the first hash value is written to the first-level parent node, and the two child nodes that participated in generating the first hash value are attached to this parent node, establishing a parent-child relationship. The parent-child relationship is mainly used to show the connection between the hash value and the text data. Further, in the same way, the two adjacent sets of parent nodes are hashed again to obtain the second hash value. The second hash value is written to the second-level parent node, and the two parent nodes that participated in generating the second hash value are attached to the second-level parent node. The above hashing and attachment process is repeated, building upwards layer by layer until a unique hash value is obtained. This hash value is defined as the verification character.

[0060] By using the hierarchical hashing method described above, the verification symbol is generated recursively. When a certain text data changes, the corresponding character position will also change, which in turn triggers changes in the first hash value, the second hash value, and the verification symbol. Therefore, it is only necessary to record and maintain the verification symbol to verify whether the text data has been tampered with.

[0061] S300: Creates a root node and synchronizes the validator to the root node, mounts the parent node to the root node, integrates the child nodes, parent nodes and root node, generates a hash tree, extracts a state snapshot, identifies the storage node of the text data, distributes the state snapshot to all storage nodes in parallel, sends commonly used characters to the preset terminal, and sets the verification rules.

[0062] Create a root node, which is the parent node's parent node. Write the verification symbol into the root node. In practical applications, the root node not only stores the verification symbol but also manages all child and parent nodes to prevent structural disorder. Integrate all child, parent, and root nodes to generate a hash tree. A hash tree is a tree-like data structure that can be simply viewed as a binary tree. The child and parent nodes in a hash tree are the same as the child nodes in a binary tree, and the root node is the same as the parent node in a binary tree. However, unlike a binary tree, a hash tree changes according to changes in the text data. These changes include the breaking and growth of parent-child relationships.

[0063] Extract a state snapshot from the hash tree, containing the root node's hash value, parent-child relationships, and their corresponding structural information. Based on the actual storage situation of the text data, determine the storage nodes for the text data; each storage node is the unit or device actually responsible for data storage, such as a storage server or hard disk drive. Send the state snapshot to all storage nodes and send commonly used characters to a preset terminal, which is the device terminal of the storage node or text data administrator. In real-world scenarios, the maintenance and management of text data requires authorization. Therefore, when the storage system composed of storage nodes receives a user's maintenance or management request, it verifies the user's identity using authentication rules: the user must correctly input commonly used characters.

[0064] S400: Receives user-uploaded call requests, constructs a hash index mechanism based on common characters and status snapshots, and determines whether to grant call permissions.

[0065] Upon receiving a user-uploaded request, the system authenticates the user, determines the text data to be accessed, and activates the hash indexing mechanism. This mechanism compares the validator in the state snapshot with the real-time updated validator in the hash tree. If they match, the text data has not been tampered with; otherwise, the text data has been altered. Based on the hash value changes in the hash tree, the system traces the parent-child relationships to quickly locate the child node where the text data has changed. This child node is defined as an abnormal node, and administrators perform data recovery, tracking, and reporting operations.

[0066] In Example 2, Figure 3 The implementation flow of the big data security processing method provided by the embodiment of the present invention is shown below. The steps of integrating all character positions and generating identifiers are described in detail below:

[0067] S101: Sequentially determine whether each identifier is unique. If not, insert an additional character into the commonly used characters and update the identifier.

[0068] Determine if two identical identifiers exist. If they do, insert additional characters into the common characters to increase the number of common characters, and update the identifier.

[0069] S102: Integrate all identifiers, generate a check set, and embed a periodic check mechanism.

[0070] Establish a correspondence between text data and identifiers, integrate all identifiers to obtain a check set, set an update cycle, and periodically update the check set according to a periodic check mechanism; for example, the periodic check mechanism is: after each user call is completed, determine whether the identifier has changed, and if so, update the identifier and check set of the text data.

[0071] In Example 3, Figure 4 The implementation flow of the secure big data processing method provided by the embodiment of the present invention is shown below. The steps of creating child nodes that correspond one-to-one with the text data and writing the identifier into the child nodes are described in detail below:

[0072] S201: Configure a number for each text data, and establish a correspondence between the text data and its child nodes based on the number.

[0073] Each piece of text data is assigned a number based on its type and size. Alternatively, text data can be numbered sequentially. The specific numbering rules are determined by the text data administrators, establishing a connection between the numbers, text data, and child nodes.

[0074] S202: Calculate the usage frequency of each text data and adjust the order of child nodes in the hash tree according to the usage frequency from highest to lowest.

[0075] Calculate the frequency of use for each text data item, and adjust the order of child nodes according to their frequency of use from highest to lowest (see instruction manual). Figure 1 In the middle, from left to right, the frequency of use of child nodes increases sequentially; the advantage of this approach is that it can further improve the query efficiency of tampered text data, and prioritize the comparison of the left path.

[0076] In Example 4, Figure 5 The implementation flow of the secure big data processing method provided by an embodiment of the present invention is illustrated below. The steps of integrating child nodes, parent nodes, and root nodes to generate a hash tree are described in detail below:

[0077] S301: Integrate the authentication policy into the hash tree, and grant permission to modify text data after successful authentication.

[0078] Authentication policies can include password verification or identity verification. If the text data remains unchanged and common characters are not modified, the verification token will also remain unchanged. Therefore, by regularly comparing the verification token, changes to the verification token can be detected in a timely manner, further improving the security of the text data. After a user passes the authentication policy, the content and common characters in the text data can be modified.

[0079] S302: Record the modification process and adjust the hash tree, wherein the adjustment includes: growth and breakage.

[0080] When a user modifies the content of text data, the specific location of the modification, the content before and after the modification, and the timestamp of the modification are recorded. In actual operation, in addition to modification, text data may also be added or deleted. "Growth" refers to the process of adding new child nodes to the existing hash tree when new text data is added, and the hash value change is propagated upwards from the new child nodes, gradually updating the corresponding parent nodes until the root node. "Break" occurs when the text data is completely deleted, the original hash relationship between the nodes is broken, the corresponding child nodes are removed, and the hash values ​​of their corresponding parent nodes and higher levels are recalculated.

[0081] In Example 5, unlike Example 1, the method further includes:

[0082] Through the modification process, incremental data is identified, and a one-to-one correspondence between incremental data and state snapshots is established;

[0083] Based on the aforementioned correspondence, rollback points are created, and a rollback mechanism is generated.

[0084] The system identifies the changed data in each modification process and generates incremental data. Incremental data can be simply understood as revision data reflecting the evolution of text data from one state to another. Based on this, the state snapshot is updated after each user call. A corresponding rollback point is created for each state snapshot, recording the increments in each modification process. A rollback mechanism is generated, allowing incremental data to be applied backwards based on the differences between the current state and historical snapshots when needed, enabling gradual rollback of the text data. By setting incremental data, only the original text data needs to be stored, eliminating the need to fully store the modified text data after each modification, significantly reducing data storage requirements.

[0085] In Example 6, Figure 5 The implementation flow of the secure big data processing method provided by an embodiment of the present invention is illustrated below. The steps of the storage node for identifying text data and sending the status snapshot to all storage nodes in parallel are described in detail below:

[0086] S303: Divide the storage node into cloud nodes and local nodes, store text data with a usage frequency greater than a threshold in the local nodes, and store the remaining text data in the cloud nodes.

[0087] All storage nodes are divided into remote nodes and local nodes. The cloud nodes mainly undertake the complex text data storage and management functions, while the local nodes are deployed close to the user or at the edge, with fast response capabilities. When the usage frequency of a certain text data exceeds a preset threshold, it is identified as high-frequency data and is stored in the local node first to reduce access latency and improve processing efficiency. The remaining text data is stored in the cloud nodes.

[0088] S304: Utilize all storage nodes to build a distributed storage architecture.

[0089] A distributed storage architecture is built using local nodes and cloud nodes to store text data in a distributed manner. Simply put, a distributed storage architecture is an architecture that distributes text data across multiple servers or devices (i.e., storage nodes) for separate storage.

[0090] In Example 7, unlike Example 1, the method further includes:

[0091] Collect source data of the call request, wherein the source data includes at least: login IP and user identity, and activate the authentication policy;

[0092] The hash indexing mechanism is used to determine whether the validator has changed. If not, the target data is retrieved and access permissions are granted.

[0093] To further optimize the authentication strategy, multiple difficulty levels can be set. For example, the difficulty levels can be divided into easy, medium, and hard. The source data of the call request is collected. If the user's login IP is a frequently used IP, the authentication is successful, and the easy difficulty level authentication strategy is activated. After receiving the user's call request, the verification token is recalculated to determine whether the verification token has changed. If it has not changed, it means that the text data has not been tampered with during storage. The text data that the user needs to call is determined, defined as the target data, and the user is granted access permission to call the target data.

[0094] Figure 6 This diagram illustrates the structural block diagram of a big data security processing system provided in an embodiment of the present invention. The big data security processing system 1 includes:

[0095] The identifier module 11 is used to set several commonly used characters, select the text data that needs to be processed for security, traverse the first occurrence position of each commonly used character, calculate the character position, integrate all the character positions, and generate an identifier, wherein each piece of text data corresponds to an identifier.

[0096] The verification module 12 is used to create child nodes that correspond one-to-one with the text data, write the identifier into the child node, select a hash function, perform pairwise hashing on adjacent child nodes to obtain several first hash values, create parent nodes at multiple levels, write the first hash values ​​into the parent nodes, and attach the corresponding child nodes to the parent nodes, continue to perform pairwise hashing on adjacent parent nodes to obtain second hash values, and so on, until a unique hash value is obtained, which is recorded as the verification symbol.

[0097] Hash module 13 is used to create a root node, synchronize the validator to the root node, mount the parent node to the root node, integrate the child nodes, parent nodes and root node to generate a hash tree, extract a state snapshot, identify the storage node of the text data, send the state snapshot to all storage nodes in parallel, send common characters to the preset terminal, and set the verification rules.

[0098] The judgment module 14 is used to receive the call request uploaded by the user, and construct a hash index mechanism based on common characters and status snapshots to determine whether to grant call permission.

[0099] Figure 7 This diagram illustrates the structural composition of a secure big data processing system provided in an embodiment of the present invention. The identification module 11 includes:

[0100] The insertion unit 111 is used to sequentially determine whether each identifier is unique. If not, it inserts an additional character into the commonly used characters and updates the identifier.

[0101] Embedding unit 112 is used to integrate all identifiers, generate a check set, and embed a periodic check mechanism.

[0102] Figure 8 This diagram illustrates the structural block diagram of a secure big data processing system provided in an embodiment of the present invention. The verification module 12 includes:

[0103] Configuration unit 121 is used to configure the number of each text data, and establish the correspondence between text data and child nodes through the number;

[0104] The calculation unit 122 is used to calculate the frequency of use of each text data and adjust the order of child nodes in the hash tree according to the frequency of use from largest to smallest.

[0105] Figure 9This diagram illustrates the structural composition of a secure big data processing system provided in an embodiment of the present invention. The hash module 13 includes:

[0106] Open unit 131 is used to integrate an authentication policy into the hash tree, and grant permission to modify text data after successful authentication.

[0107] Recording unit 132 is used to record the modification process and adjust the hash tree, wherein the adjustment includes: growth and breakage;

[0108] Storage unit 133 is used to divide the storage node into cloud node and local node, store text data with a usage frequency greater than a threshold in the local node, and store the remaining text data in the cloud node.

[0109] Building unit 134 is used to build a distributed storage architecture using all storage nodes.

[0110] The identification module 11 is mainly used to complete step S100, the verification module 12 is mainly used to complete step S200, the hash module 13 is mainly used to complete step S300, and the judgment module 14 is mainly used to complete step S400.

[0111] The insertion unit 111 is mainly used to complete step S101, and the embedding unit 112 is mainly used to complete step S102.

[0112] Configuration unit 121 is mainly used to complete step S201, and calculation unit 122 is mainly used to complete step S202;

[0113] The open unit 131 is mainly used to complete step S301, the record unit 132 is mainly used to complete step S302, the storage unit 133 is mainly used to complete step S303, and the construction unit 134 is mainly used to complete step S304.

[0114] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0115] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

[0116] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for securely processing big data, characterized in that, The method includes: Set several commonly used characters, select the text data that needs to be processed for security, traverse the first occurrence position of each commonly used character, calculate the character position, integrate all the character positions, and generate an identifier, where each piece of text data corresponds to an identifier. Create child nodes that correspond one-to-one with the text data, write the identifier into the child nodes, select a hash function, perform pairwise hashing on adjacent child nodes to obtain several first hash values, create parent nodes at multiple levels, write the first hash values ​​into the parent nodes, and attach the corresponding child nodes to the parent nodes, continue to perform pairwise hashing on adjacent parent nodes to obtain second hash values, and so on, until a unique hash value is obtained, which is recorded as the verification symbol; Create a root node and synchronize the validator to the root node. Attach the parent node to the root node, integrate the child nodes, parent nodes, and root node to generate a hash tree, and extract a state snapshot. Identify the storage nodes of the text data, send status snapshots to all storage nodes in parallel, send commonly used characters to preset terminals, and set verification rules; It receives user-uploaded call requests and uses a hash index mechanism built based on common characters and status snapshots to determine whether to grant call permissions.

2. The method for securely processing big data according to claim 1, characterized in that, The step of integrating all character positions to generate the identifier includes: Check each identifier sequentially to see if it is unique. If not, insert an additional character into the commonly used characters and update the identifier. All identifiers are integrated to generate a check set, and a periodic check mechanism is embedded.

3. The method for securely processing big data according to claim 1, characterized in that, The steps of creating child nodes that correspond one-to-one with the text data and writing identifiers into the child nodes include: Configure a number for each piece of text data, and establish a correspondence between the text data and its child nodes based on the number. Calculate the frequency of use for each text data and adjust the order of child nodes in the hash tree according to the frequency of use from highest to lowest.

4. The method for securely processing big data according to claim 3, characterized in that, The steps of integrating child nodes, parent nodes, and root nodes to generate a hash tree include: An authentication policy is integrated into the hash tree, granting permission to modify text data upon successful authentication. The modification process is recorded, and the hash tree is adjusted, including growth and breakage.

5. The method for securely processing big data according to claim 4, characterized in that, The method further includes: Through the modification process, incremental data is identified, and a one-to-one correspondence between incremental data and state snapshots is established; Based on the aforementioned correspondence, rollback points are created, and a rollback mechanism is generated.

6. The method for securely processing big data according to claim 3, characterized in that, The step of the storage node for the identified text data sending state snapshots to all storage nodes in parallel includes: The storage node is divided into cloud nodes and local nodes. Text data with a usage frequency greater than a threshold is stored in the local node, and the remaining text data is stored in the cloud node. Utilize all storage nodes to build a distributed storage architecture.

7. The method for securely processing big data according to claim 4, characterized in that, The method further includes: Collect source data of the call request, wherein the source data includes at least: login IP and user identity, and activate the authentication policy; The hash indexing mechanism is used to determine whether the validator has changed. If not, the target data is retrieved and access permissions are granted.

8. A secure big data processing system, characterized in that, The system includes: The identifier module is used to set several commonly used characters, select the text data that needs to be processed for security, traverse the first occurrence position of each commonly used character, calculate the character position, integrate all the character positions, and generate an identifier, where each piece of text data corresponds to one identifier. The verification module is used to create child nodes that correspond one-to-one with the text data, write the identifier into the child node, select a hash function, perform pairwise hashing on adjacent child nodes to obtain several first hash values, create parent nodes at multiple levels, write the first hash values ​​into the parent nodes, and attach the corresponding child nodes to the parent nodes, continue to perform pairwise hashing on adjacent parent nodes to obtain second hash values, and so on, until a unique hash value is obtained, which is recorded as the verification symbol. The hash module is used to create the root node, synchronize the validator to the root node, attach the parent node to the root node, integrate the child nodes, parent nodes and the root node to generate a hash tree, extract state snapshots, identify the storage nodes of text data, send the state snapshots to all storage nodes in parallel, send commonly used characters to preset terminals, and set verification rules. The judgment module receives the call request uploaded by the user, and uses a hash index mechanism built based on common characters and status snapshots to determine whether to grant call permissions.

9. The big data secure processing system according to claim 8, characterized in that, The identification module includes: An insertion unit is used to sequentially determine whether each identifier is unique. If not, an additional character is inserted into the commonly used characters to update the identifier. The embedding unit is used to integrate all identifiers, generate a check set, and embed a periodic check mechanism.

10. The big data secure processing system according to claim 8, characterized in that, The verification module includes: The configuration unit is used to configure the number of each text data and establish the correspondence between the text data and the child nodes through the number; The calculation unit is used to calculate the frequency of use of each text data and adjust the order of child nodes in the hash tree according to the frequency of use from highest to lowest.

Citation Information

Cited By

  • Full-process management method and system for investment project

    CN121258122A