Dynamic Data Masking Method, Device and Storage Medium

By dividing the numerical domain into data segments and using monotonic incremental functions and pseudo-random number generation rules, the problem of missing sequence relationships in numerical data desensitization is solved, and dynamic order-keeping desensitization is achieved to ensure data security and calculation efficiency.

CN112257111BActive Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011268765.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-13
Publication Date
2025-07-11
Estimated Expiration
2040-11-13

AI Technical Summary

Technical Problem

Existing data desensitization technology cannot effectively realize sequence-keeping desensitization of numerical data, resulting in the data losing its original sequence relationship after desensitization, affecting the fairness and accuracy of the data.

Method used

The dynamic numerical desensitization method is adopted to divide the positive real number domain into data segments and generate desensitization rules using monotonic incremental function and pseudo-random number generator to ensure that the desensitization value maintains the sequential relationship with the source data, and at the same time, the randomness and calcification technology are used to enhance the randomness and calcification.

Benefits of technology

Dynamic order-keeping desensitization of numerical data is realized, ensuring the consistency and rationality of the desensitization results, enhancing the security and reliability of the data, preventing data from being predicted, and with high computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112257111B_ABST
    Figure CN112257111B_ABST
Patent Text Reader

Abstract

The present application relates to a dynamic numerical desensitization method, device and storage medium. The method includes: receiving a numerical desensitization request, which includes a source to-be-desensitized numerical value and a service identifier; obtaining a desensitization rule corresponding to the service identifier; representing the positive real number domain with M data segments, and dividing each data segment into a second preset number of data slices connected end to end; determining the position loc of the source to-be-desensitized numerical value ij , where the position loc ij represents that the source to-be-desensitized numerical value is located in the j-th data slice of the i-th data segment; determining the maximum desensitized value corresponding to the data slice before the j-th data slice in the i-th data segment as the first desensitized numerical value; determining the second desensitized numerical value based on the monotonically increasing function corresponding to the j-th data slice in the desensitization rule; and determining the sum of the first desensitized numerical value and the second desensitized numerical value as the desensitized numerical value. The present application can achieve order-preserving desensitization for any numerical value in the real number domain on the basis of dynamic numerical desensitization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data desensitization, and particularly to a dynamic numerical desensitization method, device and storage medium. Background Art

[0002] With the advent of the big data era, the huge commercial value hidden in big data has been gradually explored. However, it also brings great challenges to the protection of sensitive information. In order to prevent the leakage of sensitive information, during the data sharing process, data desensitization technology is usually used to desensitize the source data to achieve data disguise. The existing data desensitization technology mainly focuses on string desensitization, which realizes data disguise by directly masking or removing some content of the source data. In this desensitization method, the content of the source data has been changed after desensitization. For example, a mobile phone number becomes "134***216" after desensitization. The same applies to names / ID numbers / bank card numbers, etc.

[0003] However, for numerical data, it is often desired that, without revealing the true value, the desensitized data can show the general trend and basic situation of the data, that is, it is required that the desensitized data is not only of numerical type, but also the desensitized data maintains the same order as the source data to be desensitized. For example, if the number of fans of Anchor A is more than that of Anchor B, if the desensitized data is not a specific value, it will not be able to reflect the relationship between the number of fans of Anchor A and Anchor B; or if the desensitized data indicates that the number of fans of Anchor B exceeds that of Anchor A, it will not only be misleading but also unfair. Obviously, the string desensitization method will no longer be applicable, and the order-preserving desensitization of numerical data has become an urgent problem to be solved. Summary of the Invention

[0004] The present application provides a dynamic numerical desensitization method, device and storage medium, which can, when desensitizing the source numerical data to be desensitized, make the desensitized numerical value maintain the same order as the source numerical data to be desensitized, so as to realize the order-preserving desensitization of numerical data.

[0005] On the one hand, the present application provides a dynamic numerical desensitization method, and the method includes:

[0006] Receiving a numerical desensitization request, where the numerical desensitization request includes a source numerical data to be desensitized and a service identifier, and the service identifier represents a unique identifier of the service to which the source numerical data to be desensitized belongs;

[0007] Obtaining a desensitization rule corresponding to the service identifier, where the desensitization rule at least includes second metadata, and the second metadata includes a first preset number of monotonically increasing functions, and the value range of each monotonically increasing function is between 0 and 1;

[0008] The positive real number domain is represented by M data segments, and each of the data segments is divided into a second preset number of data slices connected end to end, where each of the data segments is the first data slice of the next data segment, and M is a positive integer that is infinite;

[0009] Determine the position loc of the source value to be desensitized ij , the position loc ij represents that the source value to be desensitized is located in the j-th data slice of the i-th data segment;

[0010] Determine the first desensitized value corresponding to the source value to be desensitized as the maximum desensitized value corresponding to the data slice preceding the j-th data slice in the i-th data segment;

[0011] Based on the monotonically increasing function corresponding to the j-th data slice in the second metadata, determine the second desensitized value corresponding to the source value to be desensitized;

[0012] Determine the sum of the first desensitized value and the second desensitized value as the desensitized value corresponding to the source value to be desensitized.

[0013] On the other hand, a dynamic value desensitization device is provided, and the device includes:

[0014] A request receiving module, configured to receive a value desensitization request, where the value desensitization request includes a source value to be desensitized and a service identifier, and the service identifier represents a unique identifier of the service to which the source value to be desensitized belongs;

[0015] A rule obtaining module, configured to obtain a desensitization rule corresponding to the service identifier, where the desensitization rule at least includes second metadata, the second metadata includes a first preset number of monotonically increasing functions, and the value range of each of the monotonically increasing functions is between 0 and 1;

[0016] A data segment division module, configured to represent the positive real number domain by M data segments, and divide each of the data segments into a second preset number of data slices connected end to end, where each of the data segments is the first data slice of the next data segment, and M is a positive integer that is infinite;

[0017] A position determination module, configured to determine the position loc of the source value to be desensitized ij , the position loc ij represents that the source value to be desensitized is located in the j-th data slice of the i-th data segment;

[0018] A first desensitized value determination module, configured to determine the first desensitized value corresponding to the source value to be desensitized as the maximum desensitized value corresponding to the data slice preceding the j-th data slice in the i-th data segment;

[0019] The second desensitization value determination module is configured to determine a second desensitization value corresponding to the source value to be desensitized based on the monotonically increasing function corresponding to the j-th data slice in the second metadata;

[0020] The desensitization value generation module is configured to determine the sum of the first desensitization value and the second desensitization value as the desensitized value corresponding to the source value to be desensitized.

[0021] On the other hand, a computer storage medium is provided, in which at least one instruction or at least one program segment is stored, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the dynamic numerical desensitization method as described above.

[0022] This application uses the same desensitization rule for the same service to ensure the consistency of desensitization results, and realizes the dynamic desensitization processing of numerical data; uses the maximum desensitized value corresponding to the data slice before the source value to be desensitized as the base value of the desensitized value, and realizes the order-preserving desensitization of the source values to be desensitized in different data slices; through a monotonically increasing function with a value range between 0 and 1, realizes the order-preserving desensitization of the source values to be desensitized in the same data slice, thereby realizing the order-preserving desensitization of any value in the real number domain. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 is a schematic diagram of an application scenario of a dynamic numerical desensitization method provided by an embodiment of the present application.

[0025] Figure 2 is a flowchart of a dynamic numerical desensitization method provided by an embodiment of the present application.

[0026] Figure 3 is an example diagram of data segment division provided by an embodiment of the present application.

[0027] Figure 4 is another example diagram of data segment division provided by an embodiment of the present application.

[0028] Figure 5 is a flowchart of another dynamic numerical desensitization method provided by an embodiment of the present application.

[0029] Figure 6It is a schematic flowchart of determining the maximum desensitization value corresponding to a data slice provided by an embodiment of the present application.

[0030] Figure 7 It is a schematic flowchart of determining the share value of a data slice provided by an embodiment of the present application.

[0031] Figure 8 It is an example diagram of the maximum desensitization value corresponding to a data slice provided by an embodiment of the present application.

[0032] Figure 9 It is an example diagram of a scatter plot of a desensitization rule provided by an embodiment of the present application.

[0033] Figure 10 It is a schematic flowchart of determining the second desensitization value provided by an embodiment of the present application.

[0034] Figure 11 It is an example diagram of a scatter plot of another desensitization rule provided by an embodiment of the present application.

[0035] Figure 12 It is an example diagram of a scatter plot of another desensitization rule provided by an embodiment of the present application.

[0036] Figure 13 It is an example diagram of performing order-preserving desensitization provided by an embodiment of the present application.

[0037] Figure 14 It is a schematic block diagram of a structure of a dynamic numerical desensitization device provided by an embodiment of the present application.

[0038] Figure 15 It is a schematic block diagram of a structure of another dynamic numerical desensitization device provided by an embodiment of the present application.

[0039] Figure 16 It is a schematic block diagram of a structure of a maximum desensitization value determination module provided by an embodiment of the present application.

[0040] Figure 17 It is a schematic diagram of the hardware structure of a device for implementing the method provided by an embodiment of the present application. Detailed implementation manners

[0041] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data calculation, storage, processing, and sharing.

[0042] Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important support. The back-end services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the back-end system for logical processing. Data at different levels will be processed separately, and various industry data requires the support of a powerful system. This can only be achieved through cloud computing.

[0043] The solution of the embodiment of the present application relates to the technical field of big data in cloud technology. Big data refers to a data set that cannot be captured, managed, and processed by conventional software tools within a certain time range. It is a massive, high-growth-rate, and diverse information asset that requires a new processing mode to have stronger decision-making power, insight discovery ability, and process optimization ability. With the advent of the cloud era, big data has attracted more and more attention. Big data requires special technologies to effectively process a large amount of data tolerated over time. Technologies applicable to big data include massively parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the Internet, and scalable storage systems.

[0044] With the arrival of the big data era, the huge commercial value hidden in big data has been gradually explored. However, at the same time, it has also brought huge challenges to the protection of sensitive information, such as ID numbers, transaction amounts, and incomes, etc. To prevent the leakage of sensitive information, during the data sharing process, data desensitization technology is usually used to desensitize the source data to achieve data camouflage.

[0045] The embodiment of the present application aims at the problem in the prior art that the order-preserving desensitization of numerical data cannot be achieved, and proposes a dynamic numerical desensitization method. To make the purpose, technical solution, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0046] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0047] First, the relevant terms involved in the embodiments of this specification are explained as follows:

[0048] Data desensitization: Data desensitization refers to the deformation of certain sensitive information through desensitization rules to achieve reliable protection of sensitive privacy data.

[0049] Numerical desensitization: Desensitization is performed on numerical data. The difference between numerical desensitization and traditional desensitization is that traditional desensitization directly masks or removes part of the original data, while numerical desensitization results in a complete value. For example, if the monthly income is 1,000 yuan, it may become 1,334 yuan after desensitization, so it is also called numerical desensitization and disguise.

[0050] Sequence preservation: This means that while desensitizing the data, the desensitization result is also guaranteed to be consistent with the source data order. For example, for 4 source data a, b, c and d, if a <b<c<d的关系,对应脱敏结果分别为A,B,C和D,同样符合A<B<C<D的关系。

[0051] Anti-predictability: It means that even if the data user has the ability to construct massive amounts of data, it is difficult to predict the source data corresponding to this batch of data. On the contrary, predictability means that the data user may infer the data relationship by constructing data->observing results and then deduce the desensitization rules, thereby obtaining the source data corresponding to the desensitized data, which makes desensitization completely meaningless. Of course, due to the order-preserving property, it is very easy to predict the value between two values ​​that are not far apart through the idea of ​​squeezing. This is an inherent situation of order-preserving desensitization and does not affect the establishment of the anti-predictability of this application.

[0052] Boundedness: The masked result will not differ too much from the source value, and the result will be guaranteed to be within a certain reasonable range. The purpose of numerical masking is to allow data users to see the general trend and basic situation of the data, but the real value cannot be revealed. If the masked result is too outrageous and completely unusable, it deviates from the original intention of masking.

[0053] Computable: It can be computed for any source data over the entire real number domain, and for any source data, the computation can be completed within a reasonable time, such as within 1 ms.

[0054] Static desensitization: It means that the data set is determined, and then a one-time desensitization process is performed on this data set. In this case, usually only one delivery is required, and subsequent situations do not need to be considered.

[0055] Dynamic desensitization: It means that the data set may be continuously changing, such as new additions, changes, or deletions. Dynamic desensitization cannot rely on the relevance of the data set. For multiple deliveries, the stability of the desensitization results must also be ensured.

[0056] Pseudo-random number generator: Generates a series of "seemingly random" numbers that conform to the statistical characteristics of random numbers but are not random. For a given "random seed", the generated number sequence must be the same.

[0057] 0-1 function: It refers to a function whose domain is [0, 1], range is [0, 1], and is monotonically increasing.

[0058] Exponential periodic function: It refers to a function with a setting similar to a periodic function, such as sin(x), but the period occurs according to "exponentials". For example, in the decimal system, the function graphs of 0 - 10, 0 - 10 2 and 0 - 10 3 as well as 0 - 10 4 are very similar, and this similarity exists in an exponential form, similar to a kind of "fractal".

[0059] Please refer to Figure 1 , which shows a schematic diagram of the application scenario of a dynamic numerical desensitization method provided by an embodiment of the present application. As Figure 1 shown, the implementation environment may include at least one data provider 01, a big data center sharing and exchange platform 02, and at least one data user 03.

[0060] Specifically, the big data center sharing and exchange platform 02 is a general data intermediate layer platform, which provides a place for data sharing between the data provider 01 and the data user 03, and when the data provider 01 has a need to perform order-preserving desensitization on a certain numerical data, it performs numerical desensitization processing on this numerical data and provides the data obtained from the desensitization processing to the data user 03 for use.

[0061] The following introduces the dynamic numerical desensitization method of the present application with the big data center sharing and exchange platform (abbreviated as the sharing and exchange platform) as the execution subject. Figure 2It is a schematic flowchart of a dynamic data masking method provided by an embodiment of the present application. This specification provides the method operation steps as described in the embodiment or flowchart, but based on routine or non-creative labor, more or fewer operation steps may be included. The step sequence listed in the embodiment is only one way among the execution sequences of numerous steps and does not represent the only execution sequence. When the actual system or server product executes, it can be executed in the order of the method shown in the embodiment or the accompanying drawings, or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing). Specifically, as Figure 2 shown, the method may include:

[0062] S201, receiving a data masking request, where the data masking request includes a source data value to be masked and a service identifier, and the service identifier represents the unique identifier of the service to which the source data value to be masked belongs.

[0063] When the shared exchange platform performs data masking on the source data value to be masked, it identifies the source data value to be masked with the service identifier to distinguish which service the source data value to be masked belongs to. Different services have different data masking rules, and the same data masking rules are used under the same service. Among them, the service refers to the real content represented by the source data value to be masked. For example, whether the source data value to be masked represents income, age, exam score, or number of fans, etc.

[0064] S202, obtaining a data masking rule corresponding to the service identifier, where the data masking rule at least includes second metadata, and the second metadata includes a first preset number of monotonically increasing functions, and the value range of each monotonically increasing function is between 0 and 1.

[0065] In the embodiment of the present application, the data masking rule is pre-generated by the shared exchange platform based on the service identifier, and one service identifier corresponds to one data masking rule. Specifically, the shared exchange platform converts the service identifier into a magic number and uses this magic number as a random seed; based on the random seed, a data masking rule corresponding to the service identifier is obtained by using a preset pseudo-random number generator. In specific implementation, if biz_id is used to represent the service identifier, the random seed can be obtained by using the magic() function, that is, the random seed seed can be expressed as seed = magic(biz_id).

[0066] The preset pseudo-random number generator generates a data masking rule g(x) based on the idea of a random function and a fractal structure to construct a masking function. The generated data masking rule g(x) has a monotonically increasing property. Using g(x) to perform data masking on the source data value to be masked x is actually mapping the source data value to be masked x into a target value y so that the target value y maintains the same order as x. The mapping relationship can be expressed by the following formula:

[0067] x → f(x) → y

[0068] In practical applications, to achieve order-preserving desensitization, the desensitization rule f(x) can be any monotonically increasing function, that is, directly using a monotonically increasing function to desensitize the source value x to be desensitized. However, the anti-predictability of this method is weak. Of course, it is also possible to perform iterative calculations one by one based on generating random numbers, but the computability of this method is poor. Especially when dealing with the decimal region and when the source data x is very large, it may lead to the inability to complete the calculation within a reasonable time.

[0069] In the embodiments of the present application, the desensitization rule uses the idea of accumulation and monotonically increasing functions to achieve order-preserving desensitization. The main idea is to divide the positive real number domain into several data intervals. After determining the interval where the source value to be desensitized is located, use the maximum desensitized value that can be represented by the intervals before the position of the source value to be desensitized as the base value, and then use a monotonically increasing function to calculate the calculation problem in the same interval.

[0070] When dealing with the calculation problem in the same interval, a monotonically increasing function is used to calculate the proportion value of the source value to be desensitized in this interval, so as to calculate a desensitized value in this interval. Since the proportion value is a number between 0 and 1, the preset pseudo-random number generator normalizes these first preset monotonically increasing functions, limits the value range of each monotonically increasing function to be between 0 and 1, and stores the results in the second metadata for subsequent desensitization processing. Wherein, the first preset value is a positive integer greater than or equal to 1.

[0071] In some embodiments, the preset pseudo-random number generator can also directly use 0-1 functions as the monotonically increasing functions stored in the second metadata, and these 0-1 functions can be the same or different. In specific implementation, functions such as the deformed logistic function, the real exponential power function, or the composite function of logistic and tan can be selected as the 0-1 functions.

[0072] Among them, the deformed logistic function can be expressed by the following formula:

[0073]

[0074] The real exponential power function can be expressed by the following formula:

[0075] g(x) = x t , t ∈ R +

[0076] The composite function of logistic and tan can be expressed by the following formula:

[0077]

[0078] S203. Represent the positive real number domain with M data segments, and divide each of the data segments into a second preset number of data pieces connected end to end, where each of the data segments is the first data piece of the next data segment, and M is a positive integer that is infinitely large.

[0079] The embodiments of this application perform order-preserving desensitization based on the idea of accumulation. Therefore, it is first necessary to divide the positive real number domain into data segments (sections). Each section corresponds to a data interval, and the previous section is used as the first data piece (part) of the next section. Then, the second preset value is a positive integer greater than or equal to 2, that is, each section is divided into at least 2 parts.

[0080] In practical applications, the termination value of each section can be any specified value or a randomly generated value, as long as it satisfies that each section is within the numerical interval of the next section. However, if it is any specified value or a randomly generated value, it may take a long time during the calculation process. Especially for the processing of random number sequences, the calculation time-consuming will increase with the increase of the input, and this method does not have computability. And if, for the sake of computability, a periodic function is used for processing, it will be very easy to predict. Data users generate sequences of different scales and can always find this periodicity, and then deduce the function formula.

[0081] The embodiments of this application use an exponential periodic function combined with a fixed point to solve the above computability problem. The exponential periodic function is implemented using the first metadata randomly generated in the desensitization rule. The first metadata is a positive integer greater than or equal to 2. The termination value of each section is a value with the first metadata as the base and the number of segments of the section as the exponent, that is, the termination value of the kth (1≤k≤M) section is a value with the first metadata as the base and k as the exponent.

[0082] As Figure 3 and Figure 4 shown, they respectively show the data segment division cases when the first metadata is 10 and 97. In Figure 3 each section starts with a value of 0, the previous section is used as the first part of the next section, each section contains 10 data pieces connected end to end, and each part has an equal numerical interval; for example, the 4th section is the first part of the 5th section, and the numerical intervals of the first part and the second part of the 5th section are both 10 4 In Figure 4Among them, the termination value of each section is a value with base 97. Each section contains 97 parts. The second section is the first part of the third section. The numerical ranges of the second part of the third section and the first part are both 97. 2 。

[0083] It should be noted that Figure 3 and Figure 4 each part in each section in is equally divided. However, in actual implementation, it is not limited to the equally divided situation. Moreover, the preset pseudo-random number generator can randomly select any positive integer greater than or equal to 2 as the first metadata. However, the larger the first metadata, the larger the numerical range that each data segment needs to represent, and too large data will also affect the processing speed. If the first metadata is regarded as a number system, number systems such as binary, octal, and decimal are also relatively common in applications and are easy to predict. Therefore, the preset pseudo-random number generator can randomly select a number within the set number system range as the first metadata. For example, the set number system range can be in the range of 13 to 97.

[0084] S204, determine the position loc of the source value to be desensitized ij , the position loc ij represents that the source value to be desensitized is located in the j-th data slice of the i-th data segment.

[0085] In some application scenarios, the source value to be desensitized can not only be a positive number but also a negative number. For example, if a customer's expenditure within a certain period is 1000 yuan, it is usually represented by -1000. Therefore, the sharing and exchange platform determines which part of which section the source value to be desensitized is located according to the absolute value of the source value to be desensitized. Since each section is the first part of the next section, when determining the position of the source value to be desensitized, it will inevitably fall into the first part of a certain section. At this time, the smaller section will be preferentially selected as the final position.

[0086] For example, if the source value to be desensitized is 9800, according to Figure 3 the division of data segments in, it is located in the 10th part of the 4th section and also in the 1st part of the 5th section. Then it is finally determined that 9800 is located in the 10th part of the 4th section.

[0087] S205, determine the maximum desensitized value corresponding to the previous data slice of the j-th data slice in the i-th data segment as the first desensitized value corresponding to the source value to be desensitized.

[0088] In the embodiment of the present application, the maximum desensitization value corresponding to each data slice represents the maximum desensitized value that can be obtained after desensitizing the source value to be desensitized located in the data slice. Then, on the number line, the maximum desensitization value corresponding to the previous data slice of the j-th data slice can be regarded as a fixed point, and the desensitization values corresponding to the points after the fixed point are all calculated by adding a certain offset based on the fixed point. As Figure 5 shown, before step S205, the method further includes:

[0089] S501. Determine the maximum desensitization value corresponding to each data slice in the i-th data segment.

[0090] The shared exchange platform will pre-determine the maximum desensitization value corresponding to each data slice in the i-th data segment, and use the maximum desensitization value corresponding to the previous data slice of the j-th data slice as the base value of the desensitized value corresponding to the source value to be desensitized, so that the desensitized value corresponding to the source value to be desensitized in the subsequent data segment is always greater than the desensitized value corresponding to the source value to be desensitized in the previous data segment. For example, for the source value to be desensitized S1 and the source value to be desensitized S2, where S1 < S2, and S1 is located in the 3rd data slice of the 4th data segment, and S2 is located in the 6th data slice of the 4th data segment, then the desensitized value D1 corresponding to the source value to be desensitized S1 must also be less than the desensitized value D2 corresponding to the source value to be desensitized S2.

[0091] Referring to Figure 6 shown in, step S501 may specifically include:

[0092] S5011. For each of the M data segments, based on the termination value of the data segment, determine the fixed value of the data segment.

[0093] In the embodiment of the present application, the fixed value of each data segment is obtained by mapping the termination value of the data segment. The fixed value can be regarded as the maximum desensitization value corresponding to the data segment. Of course, the termination value can also be directly determined as the fixed value, but this method is easily predictable. To enhance randomness, in the embodiment of the present application, through the method of partition scaling, the third metadata in the desensitization rule is used to scale the termination value of each data segment to obtain the fixed value of the data segment. Among them, the third metadata represents a scaling sequence with an average value of 1, and the scaling sequence includes at least one scaling value.

[0094] Correspondingly, step S5011 may specifically include: for each of the M data segments, taking the remainder obtained by dividing the number of segments corresponding to the data segment by the length of the third metadata as the first index value; using the first index value as the index in the third metadata to obtain the target scaling value; and determining the product of the target scaling value and the termination value of the data segment as the fixed value of the data segment.

[0095] Taking the third metadata (denoted as zoom) as {0.8, 0.6, 1.6} as an example, since the length of the third metadata is 3, if the first metadata is 10, for the first data segment, its termination value is 10 1 , the number of segments is 1, and the remainder of 1 % 3 is 3, then the target scaling value is zoom[3] = 1.6, and the fixed value of the first data segment is 1.6 * 10 1 . For the second data segment, its termination value is 10 2 , the number of segments is 2, and the remainder of 2 % 3 is 2, then the target scaling value is zoom[2] = 0.6, and the fixed value of the second data segment is 0.6 * 10 2 = 60. Similarly, the fixed value of each data segment can be obtained.

[0096] According to the above description of calculating the fixed value, if n represents the nth data segment and system represents the first metadata, then the fixed value of the nth data segment can be expressed by the following formula:

[0097] zoom[n % len(zoom)] * system n

[0098] where len(zoom) represents the length of zoom.

[0099] Then, according to the above calculation method, when there is only one scaling value 1 in zoom, the termination value of each data segment is the fixed value of that data segment.

[0100] S5012. Obtain the share number of each data slice in the ith data segment.

[0101] Under the determined desensitization rule, the number of data slices included in each data segment is also determined, and the share number of each data slice can be set to be the same, but this conventional method is easily predicted.

[0102] In order to reduce the predicted probability, in the embodiments of the present application, the weights, i.e., the share numbers, occupied by each data slice during the data desensitization process are different. The shared exchange platform can read the fourth metadata in the desensitization rule to obtain the share number of each data slice in the i-th data segment. Wherein, the fourth metadata is a permutation randomly selected from a target permutation, and the target permutation is obtained by permuting all integers from 1 to the second preset value according to a preset permutation rule.

[0103] Specifically, the obtaining of the share number of each data slice in the i-th data segment may include: determining the value corresponding to the q-th bit in the fourth metadata as the share number of the q-th data slice in the i-th data segment.

[0104] For example, if the second preset value is 3 and the preset permutation rule is a full permutation, then all integers from 1 to 3 (1, 2, and 3) are fully permuted, and there are 6 obtained target permutations, namely {1, 2, 3}, {1, 3, 2}, {2, 1, 3}, {2, 3, 1}, {3, 1, 2}, and {3, 2, 1}. Then, a permutation {2, 1, 3} is randomly selected from these 6 permutations as the fourth metadata. Then, the value 2 corresponding to the 1st bit in {2, 1, 3} is the share number of the 1st data slice, the value 1 corresponding to the 2nd bit is the share number of the 2nd data slice, and the value 3 corresponding to the 3rd bit is the share number of the 3rd data slice.

[0105] In some embodiments, the fourth metadata may also be a randomly generated data sequence that contains the second preset value of numerical values, and each numerical value corresponds to the share number of each data slice.

[0106] S5013. Determine the share value of each data slice in the i-th data segment based on the fixed value of the i-th data segment and the share number of each data slice in the i-th data segment.

[0107] According to the definition of the data segment, each data segment is the first data slice of the next data segment, and the termination value of each data segment is a fixed value. Therefore, the total share value of each data segment can be determined based on the fixed value of each data segment, and then the share value of each data slice can be determined based on the share number of each data slice in this data segment. However, for the first data segment, its first data slice is not any data segment, so special processing is required when determining the total share value.

[0108] Specifically, as shown in Figure 7 Step S5013 may include:

[0109] S50131. Determine whether the i-th data segment is the first data segment.

[0110] If the i-th data segment is the first data segment, then perform the following steps:

[0111] S50132. Determine the fixed value of the i-th data segment as the first total share value.

[0112] For the first data segment, since the first data slice in this data segment is not a certain data segment and the starting value of this data segment is zero, then all data slices in this data segment share a numerical range as large as the fixed value of this data segment. Therefore, the fixed value of this data segment is determined as the first total share value.

[0113] S50133. Determine the sum of the share numbers of each data slice in the i-th data segment as the first total share number.

[0114] S50134. Divide the first total share value by the first total share number to obtain the first share value.

[0115] S50135. For each data slice in the i-th data segment, determine the product of the share number of the data slice and the first share value as the share value of the data slice.

[0116] If the i-th data segment is not the first data segment, then perform the following steps:

[0117] S50136. Determine the difference between the fixed value of the i-th data segment and the fixed value of the (i - 1)-th data segment as the second total share value.

[0118] For non-first data segments, since the first data slice in this data segment is the previous data segment, then in this data segment, all data slices except the first data slice share a numerical range between the fixed values of the two data segments. Therefore, the difference between the fixed value of this data segment and the fixed value of the previous data segment is determined as the second total share value.

[0119] For example, if fixed[i] represents the fixed value of the i-th data segment and fixed[i - 1] represents the fixed value of the (i - 1)-th data segment, then the second total share value is fixed[i] - fixed[i - 1].

[0120] S50137. Determine the sum of the share numbers of all data slices except the first data slice in the i-th data segment as the second total share number.

[0121] S50138. Divide the second total share value by the second total share number to obtain the second share value.

[0122] S50139. For the first data slice of the \(i\)-th data segment, determine the fixed value of the \((i - 1)\)-th data segment as the share value of the first data slice.

[0123] S501310. For the \(p\)-th (\(p\neq1\)) data slice of the \(i\)-th data segment, determine the product of the share number of the \(p\)-th data slice and the second share value as the share value of the \(p\)-th data slice.

[0124] For example, assume that the total share value of the \(i\)-th data segment is 1500, and the fourth metadata is \(\{3, 5, 4, 1, 2\}\). If the \(i\)-th data segment is the first data segment, then the 5 data slices in the first data segment share the total share value of 1500, and the total share number is the sum of the share numbers of each data slice, that is, \(3 + 5 + 4 + 1 + 2 = 15\). The first share value can be obtained as \(1500 / 15 = 100\). Further, the share values of each data slice can be determined as \(3\times100 = 300\), \(5\times100 = 500\), \(4\times100 = 400\), \(1\times100 = 100\), \(2\times100 = 200\) in sequence.

[0125] If the \(i\)-th data segment is not the first data segment, then the 4 data slices except the first data slice in the \(i\)-th data segment share the total share value of 1500, and the total share number is the sum of the share numbers of the 4 data slices except the first data slice, that is, \(5 + 4 + 1 + 2 = 12\). The second share value can be obtained as \(1500 / 12 = 125\). Further, the share value of the first data slice can be determined as the fixed value of the \((i - 1)\)-th data segment, and the share values of the other data slices are \(5\times120 = 600\), \(4\times120 = 480\), \(1\times120 = 120\), \(2\times120 = 240\) in sequence.

[0126] S5014. For the first data slice of the \(i\)-th data segment, determine the share value of the first data slice as the maximum desensitization value corresponding to the first data slice.

[0127] S5015. For the \(m\)-th (\(m\neq1\)) data slice of the \(i\)-th data segment, determine the sum of the maximum desensitization value corresponding to the \((m - 1)\)-th data slice and the share value of the \(m\)-th data slice as the maximum desensitization value corresponding to the \(m\)-th data slice.

[0128] For example, taking the case where the \(i\)-th data segment is not the first data segment as an example, as Figure 8As shown, the share values of other data parts in the above example are 5 * 120 = 600, 4 * 120 = 480, 1 * 120 = 120, 2 * 120 = 240 in sequence. Assuming the fixed value of the (i - 1)-th data segment is 500, that is, the share value of the first part is 500, then the maximum desensitization value corresponding to the first part is 500, the maximum desensitization value corresponding to the second part is 500 + 600 = 1100, the maximum desensitization value corresponding to the third part is 1100 + 480 = 1580, the maximum desensitization value corresponding to the fourth part is 1580 + 120 = 1700, and the maximum desensitization value corresponding to the fifth part is 1700 + 240 = 1940.

[0129] As Figure 9 shown therein, it is an example scatter plot of the i-th section (where x ranges from 0 to 1e 10 ) in the desensitization rule f(x). The horizontal axis represents the source value to be desensitized, and the vertical axis represents the desensitized value. In Figure 9 it, the (i - 1)-th section (shown as section[i - 1] in the figure) is the first part of the i-th section (shown as part[1] in the figure). The fixed value of the i-th section (shown as fixed[i] in the figure) is the maximum desensitization value corresponding to the i-th section, and the fixed value of the (i - 1)-th section (shown as fixed[i - 1] in the figure) is the maximum desensitization value corresponding to the (i - 1)-th section, that is, all parts in the i-th section except the first part share the share of the difference between fixed[i] and fixed[i - 1]. The points corresponding to the two values of fixed[i] and fixed[i - 1] in the coordinate axis can both be regarded as fixed points. The desensitized value corresponding to the source value to be desensitized in the i-th section can be regarded as the target point. As the source value to be desensitized is different, the corresponding target point moves between the two fixed points.

[0130] S206. Based on the monotonically increasing function corresponding to the j-th data slice in the second metadata, determine the second desensitized value corresponding to the source value to be desensitized.

[0131] The first desensitized value ensures the order-preserving property of the source values to be desensitized in different parts. For the source values to be desensitized in the same part, in order to achieve the order-preserving property, the embodiments of the present application use the first preset number of monotonically increasing functions in the second metadata to achieve it.

[0132] Referring to Figure 10 shown therein, step S206 may include:

[0133] S2061. Determine the position label of the source value to be desensitized in the j-th data slice.

[0134] For two values in the same part, their positions in this part are different. In this application, position labels are used to mark the position of the source value to be desensitized in the j-th data slice.

[0135] Specifically, take the difference between the source value to be desensitized and the starting value of the j-th data slice as the first value; take the difference between the ending value and the starting value of the j-th data slice as the second value; and determine the ratio of the first value to the second value as the position label of the source value to be desensitized in the j-th data slice.

[0136] For example, if the source value to be desensitized is 15800, then according to the Figure 3 data segment division in, 15800 is located in the 2nd part of the 5th section. The ending value of the 2nd part is 20000 and the starting value is 10000. Then the position label of 15800 in the 2nd part can be expressed as:

[0137]

[0138] S2062. Use the position label as the input of the monotonically increasing function corresponding to the j-th data slice, and obtain the proportion value of the source value to be desensitized in the j-th data slice.

[0139] In the embodiments of this application, each section contains a second preset number of parts, and the second metadata includes a first preset number of monotonically increasing functions. The first preset value and the second preset value can be the same or different. When they are different, the sharing and exchange platform can randomly select a monotonically increasing function from the second metadata, or sequentially select a monotonically increasing function in order as the monotonically increasing function corresponding to the j-th data slice; when they are the same, in addition to the above two methods, the sharing and exchange platform can also correspond each monotonically increasing function to each part one by one. For example, the first monotonically increasing function in the second metadata is used as the monotonically increasing function corresponding to the 1st part, the second monotonically increasing function in the second metadata is used as the monotonically increasing function corresponding to the 2nd part, and so on. Of course, there can be other selection methods, which are not specifically limited in the embodiments of this application.

[0140] Suppose the monotonically increasing function corresponding to the j-th data slice is a real exponential power function, and t is 1, that is, the monotonically increasing function corresponding to the j-th data slice is g(x) = x. Then the proportion value of the source value to be desensitized in the j-th data slice is 0.58.

[0141] S2063, determine the product of the occupancy ratio and the share value of the j-th data slice as the second desensitized value corresponding to the source value to be desensitized.

[0142] For example, if the share value of the j-th data slice is 240, then the second desensitized value is 0.58 * 240. It can be understood that even if two source values to be desensitized are in the same data slice, if these two source values to be desensitized are the same, then the finally obtained second desensitized values will necessarily be equal; if these two source values to be desensitized are different, due to different position tags in this data slice, combined with the monotonically increasing characteristic of the 0-1 function, the second desensitized values corresponding to these two source values to be desensitized will also maintain the same order.

[0143] S207, determine the sum of the first desensitized value and the second desensitized value as the desensitized value corresponding to the source value to be desensitized.

[0144] By taking the sum of the first desensitized value and the second desensitized value as the desensitized value corresponding to the source value to be desensitized, not only does the desensitized value have the same order as the source value to be desensitized, but also the result is ensured to be within a certain reasonable range. That is to say, the value desensitized by the desensitization rule f(x) has certain characteristics such as chaos or self-similarity.

[0145] Such as Figure 11 and Figure 12 shown, which respectively show the scatter plot examples of f(x) when the x value ranges are 0 to 1e 6 and 0 to 1e 7 for different services. It can be seen from the figure that for the same two values to be desensitized, due to different services, the corresponding desensitized values may be different. And in the same service, the desensitized values always have order preservation, and the desensitized values differ from the source values within a certain range.

[0146] Next, take a certain live streaming platform service as an example, where it is necessary to open data such as the host's gift income, the number of fans, and the popularity number for the studio or the trade union to conduct data analysis and business intelligence, and elaborate on the dynamic data desensitization method of the present application. Since in the initial cooperation stage, this live streaming platform is not willing to provide real data, but cannot be too outrageous in disguise, the sharing and exchange platform needs to perform order-preserving desensitization on these data to ensure fairness. For example, if the number of fans of host A is more than that of host B, if the disguised data shows that the number of fans of host B exceeds that of host A, it will affect the legitimate interests and evaluation of host A, which is obviously unfair.

[0147] Suppose that in the desensitization rule corresponding to the number of fans obtained by the shared exchange platform, the first metadata (denoted by system) is 10; the second metadata (denoted by z1funcs) includes 10 0-1 functions; the third metadata (denoted by zoom) has only 1 scaling value, that is, the termination value of each section is the fixed value of that section; the fourth metadata (denoted by perm) is a permutation randomly selected from the target permutation obtained by permuting all integers from 1 to 10, and perm is {10, 3, 4, 6, 2, 1, 8, 5, 7, 9}.

[0148] The shared exchange platform represents the positive real number domain with M sections, as Figure 3 shown. The initial value of each section is zero, and the termination value is the value with base 10 and exponent equal to the section number where the section is located. Each section contains 10 parts, and the difference between the termination value and the start value of each part in each section is equal, that is, each section is divided into 10 equal parts. The shared exchange platform then uses the desensitization rule corresponding to the number of fans to desensitize the number of fans of two anchors respectively.

[0149] In an example, suppose the number of fans extracted by the shared exchange platform are x1 = 25768 and x2 = 14726 respectively.

[0150] For the number of fans x1 = 25768 of the first anchor, first determine that x1 is located in the 3rd part of the 5th section. Since the fixed value fixed[5] of the 5th section is 10^ 5 (100000), and the fixed value fixed[4] of the 4th section is 10^ 4 (10000), then the total share value of the 5th section is fixed[5] - fixed[4] = 90000. Then the total number of shares in the 5th section, that is, the second total number of shares, is:[[]]

[0151] sum(perm[2:system]) = 3 + 4 + 6 + 2 + 1 + 8 + 5 + 7 + 9 = 45

[0152] Next, the share value represented by each share, that is, the second share value, is: 90000 / 45 = 2000. The share values share of each part calculated therefrom can be respectively represented as:[[]]

[0153] share(part[1]) = fixed(4) = 10000

[0154] share(part[2]) = perm[2] * 2000 = 3 * 2000 = 6000

[0155] share(part[3]) = perm[3] * 2000 = 4 * 2000 = 8000

[0156] share(part[4]) = perm[4] * 2000 = 6 * 2000 = 12000

[0157] share(part[5]) = perm[5] * 2000 = 2 * 2000 = 4000

[0158] share(part[6]) = perm[6] * 2000 = 1 * 2000 = 2000

[0159] share(part[7]) = perm[7] * 2000 = 8 * 2000 = 16000

[0160] share(part[8]) = perm[8] * 2000 = 5 * 2000 = 10000

[0161] share(part[9]) = perm[9] * 2000 = 7 * 2000 = 14000

[0162] share(part

[10] ) = perm

[10] * 2000 = 9 * 2000 = 18000

[0163] Since each part is connected end to end, that is, the termination value of the previous part is the starting value of the next part, the termination value of part[2] is:

[0164] 10000 + share(part[2]) = 10000 + 3 * 2000 = 16000

[0165] And so on, as Figure 13 shown, the numerical range of each part can be calculated. The maximum value (termination value) of the numerical range of each part is the maximum desensitization value corresponding to that part. Then the first desensitization value corresponding to x1 can be obtained as: fixed[4] + share(part[2:2]) = 16000.

[0166] And x1 = 25768 is located in the 3rd part. The original range of the 3rd part is [20000, 30000]. Then the position label of x1 in the 3rd part is:

[0167]

[0168] Assume that the monotonically increasing function g(x) = x corresponding to the 3rd part. Then, according to the position label 0.5768, the obtained occupancy ratio is also 0.5768. Correspondingly, the second desensitization value corresponding to x1 can be obtained as 0.5768 * share(part[3]) = 4614.4. In summary, the desensitized value corresponding to x1 can be obtained as 16000 + 4614.4 = 20614.4.

[0169] For the number of fans x2 = 14726 of the second anchor, it is determined that x2 is located in the 2nd part of the 5th section. Since the same desensitization rule as x1 is adopted, that is, the first metadata, the second metadata, the third metadata, and the fourth metadata are all the same, the maximum desensitization value corresponding to each part is the same as Figure 13 also the same. Then, the first desensitization value corresponding to x2 can be obtained as: fixed[4] = 10000.

[0170] And x2 = 14726 is located in the 2nd part, and the original interval of the 2nd part is [10000, 20000]. Then the position label of x2 in the 2nd part is:

[0171]

[0172] Assume that the monotonically increasing function corresponding to the 2nd part is g(x) = x. Then, according to the position label 0.4762, the obtained occupancy ratio is also 0.4762. Correspondingly, the second desensitization value corresponding to x2 can be obtained as 0.4762 * share(part[2]) = 2857.2.

[0173] In summary, the desensitized value corresponding to x2 = 14726 can be obtained as 10000 + 2857.2 = 12857.2. It can be found that the desensitized value corresponding to x1 is also greater than the desensitized value corresponding to x2.

[0174] In another example, assume that the number of fans extracted by the shared exchange platform is x3 = 29863 and x4 = 23654 respectively. Both x3 and x4 are located in the 3rd part of the 5th section. The maximum desensitization value corresponding to each part is also the same as Figure 13 the same. The first desensitization value corresponding to x3 and the first desensitization value corresponding to x4 are both the maximum desensitization value corresponding to the 2nd part, that is, 16000.

[0175] But the position label of x3 in the 3rd part is:

[0176]

[0177] The position label of x4 in the 3rd part is:

[0178]

[0179] Correspondingly, the second desensitized value corresponding to x3 is: 0.9863 * share(part[3]) = 7890.4, and the second desensitized value corresponding to x4 is: 0.3564 * share(part[3]) = 2851.2. Obviously, the desensitized value corresponding to x3 is also greater than the desensitized value corresponding to x4.

[0180] It is not difficult to see from the above example that after the sharing and exchange platform performs desensitization processing on the number of fans of the live streamers on the live streaming platform, the number of fans of each live streamer can still maintain the same order as before desensitization, that is, the order-preserving desensitization of the number of fans in the live streaming scenario is achieved; moreover, the desensitized value will not deviate too much from the value before desensitization, and the result will be ensured to be within a certain reasonable range, and the anti-predictability makes it difficult for data users to deduce the real data.

[0181] The embodiment of the present application also provides a dynamic numerical desensitization device. Refer to Figure 14 shown in

[0182] A request receiving module 1410, configured to receive a numerical desensitization request, where the numerical desensitization request includes a source value to be desensitized and a service identifier, and the service identifier represents a unique identifier of the service to which the source value to be desensitized belongs;

[0183] A rule obtaining module 1420, configured to obtain a desensitization rule corresponding to the service identifier, where the desensitization rule at least includes second metadata, and the second metadata includes a first preset number of monotonically increasing functions, and the value range of each monotonically increasing function is between 0 and 1;

[0184] A data segment division module 1430, configured to represent the positive real number domain with M data segments, and divide each data segment into a second preset number of data slices connected end to end, where each data segment is the first data slice of the next data segment, and M is a positive integer of infinity;

[0185] A position determination module 1440, configured to determine the position loc ij of the source value to be desensitized, and the position loc ij represents that the source value to be desensitized is located in the jth data slice of the ith data segment;

[0186] A first desensitized value determination module 1450, configured to determine the maximum desensitized value corresponding to the previous data slice of the jth data slice in the ith data segment as the first desensitized value corresponding to the source value to be desensitized;

[0187] A second desensitization value determination module 1460, configured to determine a second desensitization value corresponding to the source value to be desensitized based on the monotonically increasing function corresponding to the j-th data slice in the second metadata;

[0188] A desensitization value generation module 1470, configured to determine the sum of the first desensitization value and the second desensitization value as the desensitized value corresponding to the source value to be desensitized.

[0189] In some embodiments, as Figure 15 shown, the apparatus may further include:

[0190] A maximum desensitization value determination module 1480, configured to determine a maximum desensitization value corresponding to each data slice in the i-th data segment.

[0191] Specifically, as Figure 16 shown, the maximum desensitization value determination module 1480 may include:

[0192] A fixed value determination unit 1481, configured to, for each of the M data segments, determine a fixed value of the data segment based on the termination value of the data segment;

[0193] A share number acquisition unit 1482, configured to acquire the share number of each data slice in the i-th data segment;

[0194] A share value determination unit 1483, configured to determine a share value of each data slice in the i-th data segment based on the fixed value of the i-th data segment and the share number of each data slice in the i-th data segment;

[0195] A first determination unit 1484, configured to, for the first data slice of the i-th data segment, determine the share value of the first data slice as the maximum desensitization value corresponding to the first data slice;

[0196] A second determination unit 1485, configured to, for the m-th (m≠1) data slice of the i-th data segment, determine the sum of the maximum desensitization value corresponding to the (m - 1)-th data slice and the share value of the m-th data slice as the maximum desensitization value corresponding to the m-th data slice.

[0197] In some embodiments, the desensitization rule further includes third metadata, where the third metadata represents a scaling sequence with an average value of 1, and the scaling sequence includes at least one scaling value.

[0198] Correspondingly, the fixed value determination unit 1481 may include:

[0199] An index value calculation unit, for each of the M data segments, taking the remainder obtained by dividing the number of segments corresponding to the data segment by the length of the third metadata as the first index value;

[0200] A scaling value calculation unit, for using the first index value as an index in the third metadata to obtain a target scaling value;

[0201] A fixed value calculation unit, for determining the product of the target scaling value and the termination value of the data segment as the fixed value of the data segment.

[0202] In some embodiments, the desensitization rule further includes fourth metadata, which is a permutation randomly selected from a target permutation, and the target permutation is obtained by permuting all integers from 1 to the second preset value according to a preset permutation rule.

[0203] Correspondingly, the share number obtaining unit 1482 may include:

[0204] A share number calculation unit, for determining the value corresponding to the q-th bit in the fourth metadata as the share number of the q-th data slice of the i-th data segment.

[0205] In some embodiments, the share value determining unit 1483 may include:

[0206] A data segment judgment unit, for judging whether the i-th data segment is the first data segment;

[0207] A first total share value determining unit, for determining the fixed value of the i-th data segment as the first total share value;

[0208] A first total share number determining unit, for determining the sum of the share numbers of each data slice in the i-th data segment as the first total share number;

[0209] A first share value determining unit, for dividing the first total share value by the first total share number to obtain a first share value;

[0210] A first result determining unit, for each data slice in the i-th data segment, determining the product of the share number of the data slice and the first share value as the share value of the data slice.

[0211] In some embodiments, the share value determining unit 1483 may further include:

[0212] A second total share value determining unit, for determining the difference between the fixed value of the i-th data segment and the fixed value of the (i - 1)-th data segment as the second total share value;

[0213] The second total share number determining unit is configured to determine the sum of the share numbers of all data slices except the first data slice in the \(i\)-th data segment as the second total share number;

[0214] The second share value determining unit is configured to divide the second total share value by the second total share number to obtain a second share value;

[0215] The second result determining unit is configured to, for the first data slice of the \(i\)-th data segment, determine the fixed value of the \((i - 1)\)-th data segment as the share value of the first data slice;

[0216] The third result determining unit is configured to, for the \(p\)-th (\(p\neq1\)) data slice of the \(i\)-th data segment, determine the product of the share number of the \(p\)-th data slice and the second share value as the share value of the \(p\)-th data slice.

[0217] In some embodiments, the second desensitized value determining module 1460 may include:

[0218] The position label determining unit is configured to determine the position label of the source value to be desensitized in the \(j\)-th data slice;

[0219] The ratio occupation determining unit is configured to use the position label as the input of a monotonically increasing function corresponding to the \(j\)-th data slice to obtain the ratio occupation of the source value to be desensitized in the \(j\)-th data slice;

[0220] The fourth result determining unit is configured to determine the product of the ratio occupation and the share value of the \(j\)-th data slice as the second desensitized value corresponding to the source value to be desensitized.

[0221] It should be noted that, when the device provided in the above embodiments implements its functions, only the above division of each functional module is used for illustration. In actual applications, the above functions may be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the method embodiments belong to the same concept, and the specific implementation process thereof can be found in the method embodiments and will not be elaborated here.

[0222] An embodiment of the present application further provides a dynamic value desensitization device, where the device includes a processor and a memory, and at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or at least one program segment is loaded and executed by the processor to perform the dynamic value desensitization method provided in the above method embodiments.

[0223] Furthermore, Figure 17FIG. shows a schematic diagram of the hardware structure of a device for implementing the method provided by the embodiments of the present application. The device may participate in forming or include the device or system provided by the embodiments of the present application. As Figure 17 shown, the device 17 may include one or more (shown as 1702a, 1702b, ……, 1702n in the figure) processors 1702 (the processor 1702 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 1704 for storing data, and a transmission device 1706 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 17 the structure shown is only illustrative and does not limit the structure of the above electronic device. For example, the device 17 may further include more or fewer components than Figure 17 shown in Figure 17 or have a different configuration from

[0224] It should be noted that the above one or more processors 1702 and / or other data processing circuits are generally referred to as "data processing circuits" in this document. The data processing circuit may be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the device 17 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is used for a processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0225] The memory 1704 may be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method described in the embodiments of the present application. The processor 1702 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned dynamic numerical desensitization method. The memory 1704 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, a flash memory, or other non-volatile solid-state memories. In some instances, the memory 1704 may further include a memory remotely disposed relative to the processor 1702, and these remote memories may be connected to the device 17 through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0226] The transmission device 1706 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the communication provider of device 17. In one example, the transmission device 1706 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station to communicate with the Internet. In one example, the transmission device 1706 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0227] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of device 17 (or the mobile device).

[0228] An embodiment of the present application also provides a computer storage medium, in which at least one instruction or at least one segment of program is stored. The at least one instruction or at least one segment of program is loaded and executed by a processor to implement the dynamic data masking method provided by the above method embodiment.

[0229] Optionally, in this embodiment, the above computer storage medium can be located in at least one of multiple network servers in a computer network. Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media that can store program codes such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.

[0230] An embodiment of the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer storage medium. The processor of the dynamic data masking device reads the computer instructions from the computer storage medium, and the processor executes the computer instructions, so that the dynamic data masking device executes the dynamic data masking method provided by the above method embodiment.

[0231] As can be seen from the embodiments of the dynamic data masking method, device, equipment, and storage medium provided by the present application above, the present application uses the same data masking rule for the same service to ensure the consistency of the data masking results and realizes the dynamic data masking of numerical data; the maximum masked value corresponding to the data slice before the source data to be masked is used as the base value of the masked value, realizing the order-preserving data masking of the source data to be masked in different data slices; through a monotonically increasing function with a value range between 0 and 1, the order-preserving data masking of the source data to be masked in the same data slice is realized; the method of using an exponential periodic function and a fixed value is used to solve the problem that it can be calculated within the entire real number range; the methods of using random number systems, full permutations, and partition scaling are used to strengthen randomness and realize anti-predictability, making it difficult for data users to deduce the source real data from the masked value, providing security for data providers; the masked value will not deviate too much from the source value and will ensure that the masked value is within a certain reasonable range, realizing boundedness.

[0232] In summary, the dynamic data masking method of the present application can not only realize order-preserving data masking for any value in the real number domain on the basis of dynamic data masking, but also realize characteristics such as boundedness, anti-predictability, and computability.

[0233] It should be noted that: the above sequence of the embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above specific embodiments of this specification have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be executed in a different order from that in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0234] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device and the electronic equipment, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0235] The above description has fully disclosed the specific implementation manners of the present application. It should be pointed out that any modification made by those skilled in the art to the specific implementation manners of the present application does not depart from the scope of the claims of the present application. Correspondingly, the scope of the claims of the present application is not limited to the foregoing specific implementation manners.

Claims

1. A dynamic data masking method, characterized in that, Including: Receiving a numerical desensitization request, which includes a source numerical value to be desensitized and a service identifier, where the service identifier represents a unique identifier of the service to which the source numerical value to be desensitized belongs; Obtaining a desensitization rule corresponding to the service identifier, where the desensitization rule at least includes second metadata, and the second metadata includes a first preset number of monotonically increasing functions, and the value range of each monotonically increasing function is between 0 and 1; Representing the positive real number domain with M data segments, and dividing each data segment into a second preset number of data slices connected end to end, where each data segment is the first data slice of the next data segment, and M is a positive integer of infinity; Determine the position loc of the source value to be desensitized ij , where the position loc ij represents that the source value to be desensitized is located in the j-th data slice of the i-th data segment; Determining the maximum desensitization value corresponding to the previous data slice of the j-th data slice in the i-th data segment as the first desensitization value corresponding to the source numerical value to be desensitized; if j = 2, the maximum desensitization value corresponding to the previous data slice of the j-th data slice is the share value of the first data slice in the i-th data segment, if j ≥ 2, the maximum desensitization value corresponding to the previous data slice of the j-th data slice is the sum of the maximum desensitization value corresponding to the (j - 2)-th data slice and the share value of the (j - 1)-th data slice, the share value is determined based on the fixed value of the i-th data segment and the share number of each data slice in the i-th data segment, the fixed value is obtained by mapping the termination value of the i-th data segment, and the share number is the weight occupied by the data slice in the data desensitization process; determining the second desensitization value corresponding to the source numerical value to be desensitized based on the monotonically increasing function corresponding to the j-th data slice in the second metadata; the second desensitization value is the product of the occupancy ratio and the share value of the j-th data slice, and the occupancy ratio is obtained by using the position tag as the input of the monotonically increasing function corresponding to the j-th data slice, and the position tag is used to mark the position of the source numerical value to be desensitized in the j-th data slice; Determining the sum of the first desensitization value and the second desensitization value as the desensitized value corresponding to the source numerical value to be desensitized.

2. The method according to claim 1, wherein Before determining the maximum desensitization value corresponding to the previous data slice of the j-th data slice in the i-th data segment as the first desensitization value corresponding to the source numerical value to be desensitized, it further includes the step of determining the maximum desensitization value corresponding to each data slice in the i-th data segment; The determining the maximum desensitization value corresponding to each data slice in the i-th data segment includes: For each of the M data segments, determining the fixed value of the data segment based on the termination value of the data segment; Obtaining the share number of each data slice in the i-th data segment; Determining the share value of each data slice in the i-th data segment based on the fixed value of the i-th data segment and the share number of each data slice in the i-th data segment; For the first data slice of the i-th data segment, determining the share value of the first data slice as the maximum desensitization value corresponding to the first data slice; For the m-th data slice of the i-th data segment, the sum of the maximum desensitization value corresponding to the (m - 1)-th data slice and the share value of the m-th data slice is determined as the maximum desensitization value corresponding to the m-th data slice, where m ≠ 1.

3. The method according to claim 2, wherein The desensitization rule further includes third metadata, and the third metadata represents a scaling sequence with an average value of 1, and the scaling sequence includes at least one scaling value; Correspondingly, for each of the M data segments, based on the termination value of the data segment, determining the fixed value of the data segment includes: For each of the M data segments, taking the remainder of dividing the number of segments corresponding to the data segment by the length of the third metadata as the first index value; Using the first index value as the index in the third metadata to obtain the target scaling value; Determining the product of the target scaling value and the termination value of the data segment as the fixed value of the data segment.

4. The method according to claim 2, wherein The desensitization rule further includes fourth metadata, and the fourth metadata is a permutation randomly selected from a target permutation, and the target permutation is obtained by permuting all integers from 1 to the second preset value according to a preset permutation rule; The obtaining the share number of each data slice in the i-th data segment includes: Determining the value corresponding to the q-th bit in the fourth metadata as the share number of the q-th data slice in the i-th data segment.

5. The method according to claim 2, wherein The determining the share value of each data slice in the i-th data segment based on the fixed value of the i-th data segment and the share number of each data slice in the i-th data segment includes: Determining whether the i-th data segment is the first data segment; If the i-th data segment is the first data segment, then determining the fixed value of the i-th data segment as the first total share value; Determining the sum of the share numbers of each data slice in the i-th data segment as the first total share number; Dividing the first total share value by the first total share number to obtain the first share value; For each data slice in the i-th data segment, determining the product of the share number of the data slice and the first share value as the share value of the data slice.

6. The method according to claim 5, wherein The method further includes: If the i-th data segment is not the first data segment, then determining the difference between the fixed value of the i-th data segment and the fixed value of the (i - 1)-th data segment as the second total share value; Determining the sum of the share numbers of all data slices except the first data slice in the i-th data segment as the second total share number; Dividing the second total share value by the second total share number to obtain the second share value; For the first data slice of the i-th data segment, determining the fixed value of the (i - 1)-th data segment as the share value of the first data slice; For the p-th data slice of the i-th data segment, determining the product of the share number of the p-th data slice and the second share value as the share value of the p-th data slice, where p ≠ 1.

7. The method according to claim 1, characterized in that The determining the second desensitized value corresponding to the source value to be desensitized based on the monotonically increasing function corresponding to the j-th data slice in the second metadata includes: Determine the position label of the source value to be desensitized in the j-th data slice; Use the position label as the input of the monotonically increasing function corresponding to the j-th data slice to obtain the proportion value of the source value to be desensitized in the j-th data slice; Determine the product of the proportion value and the share value of the j-th data slice as the second desensitized value corresponding to the source value to be desensitized.

8. The method according to claim 1, wherein The desensitization rule further includes first metadata, and the first metadata is a positive integer greater than or equal to 2; Correspondingly, the termination value of the k-th data segment is a value with the first metadata as the base and k as the exponent, where 1 ≤ k ≤ M.

9. A dynamic numerical desensitization device, characterized in that, The device includes: A request receiving module, configured to receive a numerical desensitization request, where the numerical desensitization request includes a source value to be desensitized and a service identifier, and the service identifier represents a unique identifier of the service to which the source value to be desensitized belongs; A rule acquisition module, configured to acquire a desensitization rule corresponding to the service identifier, where the desensitization rule at least includes second metadata, the second metadata includes a first preset number of monotonically increasing functions, and the value range of each monotonically increasing function is between 0 and 1; A data segment division module, configured to represent the positive real number domain with M data segments, and divide each data segment into a second preset number of data slices connected end to end, where each data segment is the first data slice of the next data segment, and M is a positive integer of infinity; A location determination module for determining the location loc of the source value to be desensitized ij , where the location loc ij represents that the source value to be desensitized is located in the j-th data slice of the i-th data segment; A first desensitized value determination module, configured to determine the maximum desensitized value corresponding to the previous data slice of the j-th data slice in the i-th data segment as the first desensitized value corresponding to the source value to be desensitized; if j = 2, the maximum desensitized value corresponding to the previous data slice of the j-th data slice is the share value of the first data slice in the i-th data segment, if j ≥ 2, the maximum desensitized value corresponding to the previous data slice of the j-th data slice is the sum of the maximum desensitized value corresponding to the j-2-th data slice and the share value of the j-1-th data slice, the share value is determined based on the fixed value of the i-th data segment and the share number of each data slice in the i-th data segment, the fixed value is obtained by mapping through the termination value of the i-th data segment, and the share number is the weight occupied by the data slice during the data desensitization process; A second desensitized value determination module, configured to determine the second desensitized value corresponding to the source value to be desensitized based on the monotonically increasing function corresponding to the j-th data slice in the second metadata; the second desensitized value is the product of the proportion value and the share value of the j-th data slice, and the proportion value is obtained by using the position label as the input of the monotonically increasing function corresponding to the j-th data slice, and the position label is used to mark the position of the source value to be desensitized in the j-th data slice; A desensitized value generation module, configured to determine the sum of the first desensitized value and the second desensitized value as the desensitized value corresponding to the source value to be desensitized.

10. The device according to claim 9, characterized in that, The device further includes a maximum desensitized value determination module, configured to determine the maximum desensitized value corresponding to each data slice in the i-th data segment; The maximum desensitized value determination module includes: A fixed value determination unit, which is configured to determine a fixed value of each of the M data segments based on the termination value of the data segment; A share number acquisition unit, which is configured to acquire the share number of each data slice in the i-th data segment; A share value determination unit, which is configured to determine the share value of each data slice in the i-th data segment based on the fixed value of the i-th data segment and the share number of each data slice in the i-th data segment; A first determination unit, which is configured to determine the share value of the first data slice of the i-th data segment as the maximum desensitization value corresponding to the first data slice; A second determination unit, which is configured to determine the sum of the maximum desensitization value corresponding to the (m - 1)-th data slice and the share value of the m-th data slice as the maximum desensitization value corresponding to the m-th data slice for the m-th data slice of the i-th data segment, where m ≠ 1.

11. The device according to claim 10, characterized in that, The desensitization rule further includes third metadata, and the third metadata represents a scaling sequence with an average value of 1, and the scaling sequence includes at least one scaling value; Correspondingly, the fixed value determination unit includes: An index value calculation unit, which is configured to, for each of the M data segments, use the remainder obtained by dividing the number of segments corresponding to the data segment by the length of the third metadata as the first index value; A scaling value calculation unit, which is configured to use the first index value as an index in the third metadata to obtain a target scaling value; A fixed value calculation unit, which is configured to determine the product of the target scaling value and the termination value of the data segment as the fixed value of the data segment.

12. The device according to claim 10, wherein The desensitization rule further includes fourth metadata, and the fourth metadata is a permutation randomly selected from a target permutation, and the target permutation is obtained by permuting all integers from 1 to the second preset value according to a preset permutation rule; The share number acquisition unit includes: A share number calculation unit, which is configured to determine the value corresponding to the q-th bit in the fourth metadata as the share number of the q-th data slice of the i-th data segment.

13. The device according to claim 10, characterized in that, The share value determination unit includes: A data segment judgment unit, which is configured to judge whether the i-th data segment is the first data segment; A first total share value determination unit, which is configured to, if the i-th data segment is the first data segment, determine the fixed value of the i-th data segment as the first total share value; A first total share number determination unit, which is configured to determine the sum of the share numbers of each data slice in the i-th data segment as the first total share number; A first share value determination unit, which is configured to divide the first total share value by the first total share number to obtain a first share value; A first result determination unit, which is configured to, for each data slice in the i-th data segment, determine the product of the share number of the data slice and the first share value as the share value of the data slice.

14. The device according to claim 13, characterized in that, The share value determination unit further includes: A second total share value determination unit, which is configured to, if the i-th data segment is not the first data segment, determine the difference between the fixed value of the i-th data segment and the fixed value of the (i - 1)-th data segment as the second total share value; A second total share number determining unit, configured to determine, as the second total share number, the sum of the share numbers of all data slices except the first data slice in the i-th data segment; A second share value determining unit, configured to divide the second total share value by the second total share number to obtain a second share value; A second result determining unit, configured to, for the first data slice of the i-th data segment, determine the fixed value of the (i - 1)-th data segment as the share value of the first data slice; A third result determining unit, configured to, for the p-th data slice of the i-th data segment, determine the product of the share number of the p-th data slice and the second share value as the share value of the p-th data slice, where p ≠ 1.

15. The device according to claim 9, characterized in that, The second desensitized value determining module includes: A position label determining unit, configured to determine the position label of the source value to be desensitized in the j-th data slice; An occupancy ratio determining unit, configured to use the position label as the input of a monotonically increasing function corresponding to the j-th data slice to obtain the occupancy ratio of the source value to be desensitized in the j-th data slice; A fourth result determining unit, configured to determine the product of the occupancy ratio and the share value of the j-th data slice as the second desensitized value corresponding to the source value to be desensitized.

16. The device according to claim 9, characterized in that, The desensitization rule further includes first metadata, and the first metadata is a positive integer greater than or equal to 2; Correspondingly, the termination value of the k-th data segment is a value with the first metadata as the base and the k as the exponent, where 1 ≤ k ≤ M.

17. A computer storage medium, characterized in that, At least one instruction or at least one program segment is stored in the computer storage medium, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the dynamic numerical desensitization method according to any one of claims 1 - 8.

18. A dynamic numerical desensitization device, characterized in that, The device includes a processor and a memory, at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the dynamic numerical desensitization method according to any one of claims 1 - 8.

19. A computer program product, characterized in that, The computer program product includes computer instructions, the computer instructions are stored in a computer storage medium, a processor of the dynamic numerical desensitization device reads the computer instructions from the computer storage medium, and the processor executes the computer instructions, so that the dynamic numerical desensitization device executes the dynamic numerical desensitization method according to any one of claims 1 - 8.

Citation Information

Patent Citations

  • System and method for order-preserving encryption for numeric data

    US20050147240A1

  • Sensitive information processing method, device, server and security determination system

    WO2016034068A1