Deep learning enhanced mass data security noise reduction method
By performing binary tree rotation and reorganization on the target data, combined with a unique identifier method, the problem of leakage risk still exists after the noise-reduced data is stolen is solved, thus achieving data security and ease of use.
Patent Information
- Application Number
- CN202511115830.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-21
AI Technical Summary
In existing technologies, even after noise reduction, data still faces the risk of leakage if account permissions are stolen, and data security cannot be effectively guaranteed.
By periodically acquiring the target data and its subsequent data, splitting them into odd and even arrays and mapping them into a binary tree, performing rotation operations and recombining them, and combining them with a unique identifier to form mixed data, and then performing reverse restoration in a specified peripheral device, the data is ensured to be obfuscated and unauthorized access is prevented.
Even if an account is stolen, the data cannot be read correctly, thus ensuring data security and ease of use and preventing data leakage.
Smart Images

Figure CN120995498A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of data security noise reduction, and specifically relates to a deep learning enhanced massive data security noise reduction method. BACKGROUND
[0002] Big data has become the core driving force of modern social and industrial development. Massive data contains great value, and through data analysis and mining (such as machine learning and statistical analysis), decisions can be optimized, efficiency can be improved, new knowledge can be discovered, and user experience can be improved. The massive data collected in the real world usually contains noise (measurement error, transmission interference, outliers, missing values, etc.). Noise can seriously interfere with subsequent data analysis and model training effects, leading to incorrect conclusions or inefficient models. Therefore, data noise reduction is a key step in data preprocessing and a prerequisite for releasing data value.
[0003] The patent with publication number CN116955934A discloses a network transmission data noise reduction method, device, computing equipment and storage medium. The method comprises: acquiring data to be transmitted, performing data analysis and processing on the data to be transmitted to divide data attributes; for picture data attributes, inputting picture data into a deep learning model for noise recognition, and outputting first noise reduction data after data noise reduction according to the recognition result; for voice data attributes, dividing noise signals according to clustering operation results, and outputting second noise reduction data after removing noise according to the division results; for text data attributes, outputting third noise reduction data after data noise reduction processing of the text data; and performing data network transmission after integrating and outputting the first noise reduction data, the second noise reduction data and the third noise reduction data. Through data preprocessing, type division and corresponding noise reduction methods for different types of applications, data noise reduction for network transmission is more accurately realized, and data transmission accuracy is improved.
[0004] However, data often contains a large amount of sensitive information, and data leakage incidents occur frequently, not only causing huge economic losses, but also seriously infringing on personal privacy, damaging corporate reputation, and even threatening national security. After noise reduction processing, although part of the noise is removed, the core structure and pattern are often highly correlated with the original sensitive data. In the case of stolen account permissions, after the attacker obtains the noise reduction data, the attacker may use data reconstruction attacks, correlation analysis, model inversion and other technologies to partially or completely infer the original sensitive information.
[0005] Therefore, it is a difficult problem to ensure the security of the noise reduction data and the original data. After the account information with reading permission is stolen, complete data leakage may occur. In order to ensure the security of the noise reduction data, a solution needs to be provided. SUMMARY
[0006] The present application aims to at least solve one of the technical problems existing in the prior art;
[0007] To this end, the present application proposes a deep learning enhanced massive data security denoising method, comprising:
[0008] Periodically obtain the backward data after the denoising model trained on the target data with noise is denoised;
[0009] The target data and its corresponding backward data for one period are split into two equal halves, the data with odd sequence number after splitting is combined with the backward data with odd sequence number to form a data group Cj, and the remaining data in the target data and the remaining data in the forward and backward data form a data group Uj, j = 1,..., n;
[0010] Map Cj and Uj to binary tree one and binary tree two according to the number of j, then perform a rotation operation on the corresponding binary tree one or binary tree two of Cj or Uj, which will disrupt the position of each node in the rotated data group and ensure the shortest rotation path, and then obtain the disrupted Cj or Uj, read the row and column numbers to obtain C1j or U1j, and place them in the order of row before column after, and then combine them with the undisturbed Uj or Cj to obtain mixed data, keeping the undisturbed Cj in front, and there is a unique identifier in the combined data.
[0011] Further, the target data and backward data for one period are marked as Bi and Hi, i = 1,..., m;
[0012] For each Bi, split Bi into two equal parts to obtain two segments, and label all Bi segments in order as Cj, and similarly obtain Uj corresponding to Hi, j = 1,..., n, n = 2m, i.e. Bm is split into C2m-1 and C2m;
[0013] Update Cj and Uj: combine Cj and Uj in the order of odd j in Cj first and odd j in Uj last, then update Cj, and combine Cj and Uj in the order of even j in Cj first and even j in Uj last, then update Uj, the first half of the updated Cj and the first half of the Uj can form Bi, and the remaining data can form Hi.
[0014] Further, the binary tree one construction method for Cj and Uj is:
[0015] Map Cj and Uj to binary tree one and binary tree two in the order of row number first and column number second from small to large.
[0016] Further, the binary tree one or the binary tree two is rotated, and the specific mode is as follows:
[0017] The binary tree one or the binary tree two is rotated, so that all nodes are not in the original node position, and the rotation path is the shortest.
[0018] Further, the rotation path includes a plurality of rotation operations in sequence, the rotation operation includes a rotation node and a rotation direction, the rotation node refers to which node is rotated around, and the rotation direction includes left rotation and right rotation; the shortest rotation path refers to the minimum number of rotations.
[0019] Further, when the disturbed Cj or Uj is combined with the undisturbed Uj or Cj, there is a unique identifier in the middle.
[0020] Further, the specific mode of rotating the binary tree one or the binary tree two is as follows:
[0021] If the number of target data before splitting is odd, the binary tree one is rotated, otherwise the binary tree two is rotated.
[0022] Further, when the binary tree one or the binary tree two is rotated, the binary tree one corresponding to Cj is rotated.
[0023] Further, when the binary tree one or the binary tree two is rotated, the binary tree two corresponding to Uj is rotated.
[0024] Further, when the user needs to restore the data, the updated binary tree one and the binary tree two are obtained, and then the restoration algorithm stored in the specified peripheral device is used to restore the data, and the restoration algorithm is as follows:
[0025] The updated binary tree one and the rotation operation set are obtained;
[0026] The mixed data is split by using the unique identifier, and is mapped to the nodes of the updated binary tree one and the binary tree two, the updated binary tree one is inversely restored by using the rotation operation set, the mixed data corresponding to the split data is mapped to the binary tree one according to the order of the odd bits, and the even bits are mapped to the binary tree two, so that the original target data and the backward data are obtained, the target data is the first half, and the backward data is the second half.
[0027] Compared with the prior art, the beneficial effects of the present application are as follows:
[0028] The application periodically acquires backward data after the target data with noise is processed by the trained denoising model for denoising.
[0029] Cj, Uj are mapped into binary tree one and binary tree two according to the bit number of j, then rotation operation is performed on the corresponding binary tree one or binary tree two of Cj or Uj, the data is rearranged after being disturbed, and binary tree one is inversely restored in the specified peripheral device, so as to realize data confusion, even if the account is stolen, the correct permission corresponding data cannot be read, so that the application is simple and effective, and easy to use. BRIEF DESCRIPTION OF DRAWINGS
[0030] Fig. 1 The flowchart of the application;
[0031] Fig. 2 The binary tree one of the application is shown. DETAILED DESCRIPTION
[0032] The technical solutions of the application will be described in detail below with reference to the embodiments. Obviously, the described embodiments are only part of the embodiments of the application, not all. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.
[0033] Embodiment one:
[0034] Please refer to Figs. 1-2 The application provides a deep learning enhanced massive data security denoising method;
[0035] As an embodiment of the application, it specifically includes the following steps:
[0036] Periodically acquire backward data after the target data with noise is processed by the trained denoising model for denoising.
[0037] The target data and the corresponding backward data of a period are split into two equal halves, the target data after splitting is combined with the backward data with odd sequence number, forming data group Cj, and the data in the remaining target data is combined with the remaining data in the backward data, forming data group Uj, j=1,...,n.
[0038] Cj, Uj are mapped to binary tree one and binary tree two according to the number of bits of j, then a rotation operation is performed on the binary tree one or binary tree two corresponding to Cj or Uj, the position of each node in the rotated data group is disturbed and the shortest rotation path is ensured, and after the disturbance, Cj or Uj is obtained, and after reading the row and column numbers, C1j or U1j is obtained, which is read in the order of row first and column last, and is combined with Uj or Cj which is not disturbed to obtain mixed data, and Cj which is disturbed or not disturbed is kept in front, and there is a unique identifier in the combined data.
[0039] Embodiment two:
[0040] As embodiment two of the present application, it specifically includes:
[0041] Step one: first, a batch of target data to be processed is obtained, and the target data is data that needs to be denoised;
[0042] Step two: then the trained denoising model is used to process the target data, and the denoised data is obtained after processing, which is marked as backward data;
[0043] Step three: periodically obtain the target data and the backward data, and mark the target data and the corresponding backward data as a group of goal data groups, and obtain a plurality of groups of goal data groups obtained in each period;
[0044] Step four: then select a plurality of goal data groups in a period;
[0045] Step five: process the plurality of goal data groups in the period, and the processing method is:
[0046] First, assign all target data to a record identifier, and mark it as Bi, i=1,...,m, indicating that there are m target data, and list it as a target data group;
[0047] Mark the backward data corresponding to the target data Bi as Hi, i=1,...,m, and Hi and Bi are in one-to-one correspondence;
[0048] Then, Bi and Hi are divided into two equal parts, which means that each data in Bi and Hi is split into two half data, one before and one after;
[0049] Then, Bi and Hi after splitting are sequentially marked as Cj and Uj, j=1,...,n, where n=2m; during this process, C1 and C2 can be combined to form B1, C3 and C4 can be combined to form B2, and so on; U1 and U2 can be combined to form H1, U3 and U4 can be combined to form H2, and so on;
[0050] Then reconstruct Cj and Uj, keep the data corresponding to the odd number value of j in Cj, and extract the rest of the data, then extract the data corresponding to the odd number value of j in Uj, keep the even data, and place the data corresponding to the odd number value of j in Uj in front of the data corresponding to the odd number value of j in Cj to form a new Cj;
[0051] Keep the data corresponding to the even number value of j in Cj in front of the data corresponding to the even number value of j in Uj to form a new Uj;
[0052] Update Cj and Uj to obtain updated Cj and Uj, j = 1,..., n, at this time Cj and Uj can obtain Bi and Hi, the first half of Cj and Uj can obtain Bi, and the last half can obtain Hi;
[0053] Step six: Construct a binary tree for Cj and Uj, the specific construction method is as follows:
[0054] Map j in Cj to the nodes of the binary tree, and map it to the binary tree in the order of 1, 2, 3,..., n, the mapping order is first according to the row number from small to large, and the specific arrangement of each row is according to the value from small to large, which is mapped to each column in the order of small to large;
[0055] As shown in Fig. 2 , when there are 14 data, the binary tree is constructed as shown in the figure, and the value of n can only be an even number, so the specific example of 14 is shown in Fig. 2 ;
[0056] Similarly, the same processing is performed on Uj to obtain a binary tree two;
[0057] Step seven: Then perform a rotation operation on the binary tree one, and the rotation operation requirements are as follows:
[0058] First, rotate the binary tree one to ensure that all nodes are not in the original node position, and the rotation path is the shortest, the rotation path includes a number of sequential rotation operations; the rotation operation includes rotating the node and rotating the direction, rotating the node means rotating around which node, rotating the direction includes left rotation and right rotation; the shortest rotation path means the least number of rotations;
[0059] All nodes not in the original node position means that the row and column numbers of the node in the binary tree after rotation are different from the row and column numbers of the node in the binary tree before rotation, at least one of the row or column is different;
[0060] After obtaining the rotated binary tree 1, it is restored to a new C1j, j=1, ..., n, according to the priority of row number first and column number last, in ascending order. At this time, C1j is different from each of Cj. Then, the binary tree 1 is re-identified according to the new C1j. Here, the rotation nodes in each rotation operation involved in the rotation path will be updated according to the new C1j.
[0061] Then C1j and Uj are merged pairwise. During the merging process, there is a readable unique identifier between C1j and Uj, resulting in a mixed data of the target data and the encrypted data after further processing.
[0062] Preserve the mixed data, the set of rotation operations, and the updated binary tree one and binary tree two;
[0063] Step 8: When a user needs to restore data, the updated binary tree one and binary tree two will be obtained. Then, the data will be restored using a restoration algorithm stored in a specified peripheral device. The restoration method is as follows:
[0064] Obtain the updated binary tree and the set of rotation operations;
[0065] Simultaneously, the mixed data will be split using a unique identifier and mapped to the nodes of the updated binary tree 1 and binary tree 2. The updated binary tree 1 will be reversed using a set of rotation operations. The corresponding split mixed data will be mapped to the restored binary tree 1 according to the order of the odd-numbered positions, and the even-numbered positions will be mapped to binary tree 2. The data will be recombined to obtain the original target data and the backward data. The target data is the first half, and the backward data is the second half.
[0066] Specifically, this method involves splitting the mixed data to obtain an odd-numbered ordinal number. After restoring the binary tree, the ordinal numbers are shuffled and restored to their original positions. However, the ordinal numbers in the original positions still correspond to the positions of the odd-numbered split mixed data in the original target data. Therefore, it is possible to recombine and restore the data.
[0067] Example 3:
[0068] As a third embodiment of this application, this embodiment is based on embodiment two, except that:
[0069] Then, when performing a rotation operation on binary tree one, this embodiment does not perform a rotation operation on binary tree one, but instead performs a rotation operation on binary tree two. The requirements for the rotation operation are as follows:
[0070] First, the binary tree two is rotated to ensure that all nodes are not in the original node position, and the rotation path is the shortest, the rotation path includes a number of sequential rotation operations; the rotation operation includes the rotation node and the rotation direction, the rotation node refers to which node is rotated around, and the rotation direction includes left rotation and right rotation;
[0071] All nodes are not in the original node position refers to the number of rows and columns of the rotated node in the binary tree two is different from the number of rows and columns of the node in the binary tree two before rotation, at least one of the rows or columns is different;
[0072] The rotated binary tree two is obtained, which is restored to a new U1j, j=1,...,n in the order of row number priority and column number priority from small to large, at this time U1j and Uj are different; and the binary tree two is re-identified according to the new U1j, at this time the rotation node in each rotation operation involved in the rotation path is updated according to the new U1j;
[0073] Then, the Cj and U1j are merged two by two, and there is a readable unique identifier between Cj and U1j during the merging process, to obtain the mixed data after the target data and the backward processing encryption;
[0074] The mixed data, the rotation operation set, the binary tree one and the updated binary tree two are reserved.
[0075] The restoration principle in step eight is for the binary tree two.
[0076] Embodiment four:
[0077] As embodiment four of the present application, the binary tree one or the binary tree two is rotated, depending on the number of target data before splitting, when it is odd, the binary tree one is rotated, otherwise the binary tree two is rotated.
[0078] Some data in the above formula are calculated by removing the dimension and taking the numerical value, the formula is obtained by software simulation of a large amount of collected data to obtain a formula closest to the real situation; the preset parameters and the preset threshold in the formula are set by the person skilled in the art according to the actual situation or obtained by a large amount of data simulation.
[0079] The above embodiments are only used to illustrate the technical method of the present application and are not limited, although the present application is described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical method of the present application.
Claims
1. A secure noise reduction method for massive datasets enhanced by deep learning, characterized in that, include: The trained denoising model performs denoising on noisy target data and periodically acquires the denoised backward data. For a given period, the target data and its corresponding backward data are split into two equal halves. The target data with an odd sequence number after splitting is combined with the backward data to form a data group Cj. The remaining data are then grouped into data groups Uj with the target data first and the backward data last, where j=1, ..., n. Map Cj and Uj to binary tree one and binary tree two respectively according to the number of bits of j. Then, perform a rotation operation on the binary tree one or binary tree two corresponding to Cj or Uj, shuffling the position of each node in the rotated data group while ensuring the shortest rotation path. After shuffling, obtain the shuffled data group. Read its row and column numbers to obtain C1j or U1j. Read and place them in the order of row first and column last, and combine them with the unshuffled Uj or Cj to obtain mixed data. Keep the shuffled or unshuffled Cj first. The combined data contains a unique identifier in the middle.
2. The method for secure noise reduction of massive data using deep learning enhancement as described in claim 1, characterized in that, The target data and backward data for a given period are labeled Bi, Hi, i=1, ..., m, respectively. For each Bi, Bi is divided into two equal parts to obtain two segments. All segments of Bi are labeled as Cj in order. Similarly, Uj corresponding to Hi is obtained, j=1, ..., n, n=2m. That is, Bm is divided into C2m-1 and C2m. Update Cj and Uj: Combine Cj with data where j is odd first and Uj with data where j is odd last, and update Cj with data where j is even first and Uj with data where j is even last, and update Uj with Uj. The first half of the updated Cj and the first half of the updated Uj can be combined to form Bi, and the remaining data can be combined to form Hi.
3. The method for secure noise reduction of massive data using deep learning enhancement as described in claim 1, characterized in that, The method for constructing a binary tree for Cj and Uj is as follows: Map Cj and Uj to binary trees in ascending order of row number and then column number, respectively, to obtain binary tree one and binary tree two.
4. The method for secure noise reduction of massive data using deep learning enhancement as described in claim 3, characterized in that, Perform a rotation operation on binary tree one or binary tree two, specifically as follows: Rotate either binary tree 1 or binary tree 2 to ensure that all nodes are not in their original positions and that the rotation path is the shortest.
5. A method for secure noise reduction of massive data using deep learning enhancement as described in claim 4, characterized in that, The rotation path includes several sequential rotation operations. Each rotation operation includes a rotation node and a rotation direction. The rotation node refers to the node around which the rotation is performed, and the rotation direction includes left rotation and right rotation. The shortest rotation path means the fewest number of rotations.
6. A method for secure noise reduction of massive data using deep learning enhancement as described in claim 1, characterized in that, When a scrambled Cj or Uj is combined with an unscrambled Uj or Cj, a unique identifier exists in the middle.
7. A method for secure noise reduction of massive data using deep learning enhancement as described in claim 1 or 4, characterized in that, The specific method for performing a rotation operation on binary tree one or binary tree two is as follows: If the number of target data before splitting is odd, rotate binary tree one; otherwise, rotate binary tree two.
8. A method for secure noise reduction of massive data using deep learning enhancement as described in claim 1 or 4, characterized in that, When performing a rotation operation on binary tree one or binary tree two, the rotation operation is performed on binary tree one corresponding to Cj.
9. A method for secure noise reduction of massive data using deep learning enhancement as described in claim 1 or 4, characterized in that, When performing a rotation operation on binary tree 1 or binary tree 2, the rotation operation is performed on binary tree 2 corresponding to Uj.
10. A method for secure noise reduction of massive data using deep learning enhancement as described in claim 1, characterized in that, When a user needs to restore data, the updated binary tree one and binary tree two are obtained, and then the data is restored using a restoration algorithm stored in a specified peripheral device. The restoration algorithm is as follows: Obtain the updated binary tree and the set of rotation operations; Simultaneously, the mixed data will be split using a unique identifier and mapped to nodes in the updated binary tree 1 and binary tree 2. The updated binary tree 1 will be reversed using a set of rotation operations. The split mixed data will be mapped to the restored binary tree 1 according to the order of odd-numbered positions, and even-numbered positions will be mapped to binary tree 2. The data will then be recombined to obtain the original target data and the backward data, with the target data being the first half and the backward data being the second half.
Citation Information
Patent Citations
Network transmission data noise reduction method and device, computing equipment and storage medium
CN116955934A