Data desensitization method, device, equipment, medium and product
By using generative adversarial networks to perform feature processing and model iteration on financial data, desensitized data that conforms to the distribution of real financial data is generated, solving the privacy protection problem of tabular financial data and realizing the security of data in sharing and analysis.
Patent Information
- Application Number
- CN202510980768.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-10-28
AI Technical Summary
Existing data anonymization technologies struggle to generate anonymized data that accurately reflects the distribution of real financial data when processing tabular and highly correlated financial data, leading to a higher risk of leakage of personal privacy and trade secrets.
Generative adversarial networks are used to normalize the features of financial data and iterate the model. The generator is trained by using the cross-entropy loss function and noise parameter update strategy to generate desensitized data that conforms to the distribution of real financial data.
The generated anonymized data can effectively protect personal privacy and business secrets when shared, tested, or analyzed, ensuring that legally protected information is not leaked.
Smart Images

Figure CN120850320A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data processing, and in particular to a data desensitization method, apparatus, equipment, medium, and product. Background Technology
[0002] With the rapid development of artificial intelligence technology, its applications have spread across various industries. Particularly in the financial industry, massive amounts of transaction data, market data, and customer data are generated daily. This data exhibits diverse characteristics and demands extremely high levels of privacy and security. Data anonymization is a technique that uses technical means to transform, replace, or delete sensitive data to protect privacy and confidential information. However, traditional data anonymization techniques are less effective at processing structured financial data that is highly correlated and often consists of tabular data.
[0003] Therefore, for tabular and highly correlated financial data, how to conduct comprehensive feature processing and model iteration to generate anonymized data that conforms to the distribution of real financial data, and ensure that personal privacy, trade secrets or legally protected information are not leaked when the data is shared, tested or analyzed, is an urgent problem to be solved. Summary of the Invention
[0004] This invention provides a data anonymization method, apparatus, device, medium, and product for comprehensive feature processing and model iteration of tabular and highly correlated financial data, thereby generating anonymized data that conforms to the distribution of real financial data, ensuring that personal privacy, trade secrets, or legally protected information are not leaked when the data is shared, tested, or analyzed.
[0005] According to one aspect of the present invention, a data anonymization method is provided, comprising:
[0006] In response to the request for data anonymization of financial data, a pre-stored financial data table is determined, and the financial data table is subjected to normalized feature processing to obtain the normalized feature vector corresponding to the financial data table.
[0007] Determine the values of the generated discrete variables corresponding to each target sample in the financial data table, and based on the values of the generated discrete variables, determine the loss function of the preset generative adversarial network based on the cross-entropy loss function;
[0008] Based on the preset objective function, loss function, preset noise parameter update strategy, and model parameter update strategy of the generative adversarial network, the preset generative adversarial network is iteratively trained, and the trained generative adversarial network is used to generate de-identified data corresponding to financial data in order to respond to data de-identification requests.
[0009] According to another aspect of the present invention, a data desensitization apparatus is provided, comprising:
[0010] The vector determination module is used to respond to the data anonymization request for financial data, determine the pre-stored financial data table, and perform normalization feature processing on the financial data table to obtain the normalized feature vector corresponding to the financial data table.
[0011] The function determination module is used to determine the values of the generated discrete variables corresponding to each target sample in the financial data table, and based on the values of the generated discrete variables, determine the loss function of the preset generative adversarial network based on the cross-entropy loss function.
[0012] The data anonymization module is used to iteratively train a preset generative adversarial network (GAN) based on its objective function, loss function, preset noise parameter update strategy, and model parameter update strategy. It then uses the trained GAN to generate anonymized data corresponding to the financial data in response to data anonymization requests.
[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data desensitization method according to any embodiment of the present invention.
[0014] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the data desensitization method according to any embodiment of the present invention.
[0015] According to another aspect of the present invention, a computer program product is also provided, the computer program product including a computer program that, when executed by a processor, implements the data desensitization method of any embodiment of the present invention.
[0016] The technical solution of this invention, in response to a data anonymization request for financial data, firstly determines a pre-stored financial data table and performs normalized feature processing on the financial data table to obtain a normalized feature vector corresponding to the financial data table; secondly, it determines the values of the generated discrete variables corresponding to each target sample in the financial data table, and based on the values of the generated discrete variables, determines the loss function of a preset generative adversarial network (GAN) based on the cross-entropy loss function; thirdly, iteratively trains the preset GAN based on its objective function, loss function, preset noise parameter update strategy, and model parameter update strategy, and uses the trained GAN to generate anonymized data corresponding to the financial data, thus responding to the data anonymization request. By performing comprehensive feature processing and model iteration on tabular and highly correlated financial data, anonymized data that conforms to the distribution of real financial data can be generated, ensuring that personal privacy, trade secrets, or legally protected information are not leaked when the data is shared, tested, or analyzed.
[0017] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart of a data desensitization method provided in Embodiment 1 of the present invention;
[0020] Figure 2 This is a flowchart of a data desensitization method provided in Embodiment 2 of the present invention;
[0021] Figure 3 This is a structural block diagram of a data desensitization device provided in Embodiment 3 of the present invention;
[0022] Figure 4 This is a schematic diagram of the structure of the electronic device provided in Embodiment 4 of the present invention. Detailed Implementation
[0023] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0024] It should be noted that the terms "first," "second," "target," "candidate," and "alternative," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices. The acquisition, storage, use, and processing of data in the technical solutions of this application comply with relevant laws and regulations.
[0025] Example 1
[0026] Figure 1 This is a flowchart of a data anonymization method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where comprehensive feature processing and model iteration of financial data are performed based on generative adversarial networks to generate anonymized data that conforms to the distribution of real financial data. This method can be executed by a data anonymization device, which can be implemented in hardware and / or software. The data anonymization device can be configured in electronic devices, such as... Figure 1 As shown, the data anonymization method includes:
[0027] S101. In response to the request for data anonymization of financial data, determine the pre-stored financial data table and perform normalization feature processing on the financial data table to obtain the normalized feature vector corresponding to the financial data table.
[0028] Financial data refers to data that has undergone anonymization to achieve protection. A financial data table is a pre-stored data table associated with financial data, which can store historical financial data similar to the financial data. A financial data table can contain at least one piece of financial data, i.e., one target sample. Each target sample can correspond to at least one target feature. Specifically, each row of the financial data table corresponds to one target sample, and each column corresponds to one target feature of each target sample. A normalized feature vector is a vector containing the encoded features and modal normalized features corresponding to the target samples in the financial data table.
[0029] Optionally, in response to a request to anonymize financial data, the data type of the financial data can be determined, and further matched in a pre-defined database to identify the corresponding financial data table.
[0030] Optionally, in response to a data anonymization request, the operations of S101-S103 of this embodiment can be executed to perform iterative training of a preset generative adversarial network and generate anonymized data. Alternatively, iterative training can be performed periodically to obtain a trained generative adversarial network. When a data anonymization request for financial data is detected, the trained generative adversarial network can be directly used to generate anonymized data corresponding to the financial data. This invention does not impose any limitations on this.
[0031] Optionally, the financial data table is subjected to normalized feature processing to obtain the normalized feature vector corresponding to the financial data table. This includes: determining at least two modalities associated with the financial data table, and for each target sample in the financial data table, assigning the target sample to the modality with the highest probability to obtain the modality normalized feature corresponding to the target sample; determining at least two target features of the target sample recorded in the financial data table, and based on a preset coding strategy, determining the coding features corresponding to the discrete features and continuous features in the target features respectively; and determining the normalized feature vector corresponding to the financial data table based on the coding features and modality normalized features corresponding to each target sample in the financial data table.
[0032] Here, "modality" refers to Gaussian mode. Target features can be continuous or discrete. Modality-normalized features refer to the normalization of continuous target features in the target sample to obtain the desired features; the preset encoding strategy can be one-hot encoding.
[0033] Optionally, a specific mode normalization technique can be used to transform the multimodal distribution into k Gaussian models in a way that maximizes the expectation based on the variational Gaussian mixture model. The posterior probability of the evaluation sample of the continuous feature in these k Gaussian models is calculated, and finally the sample point is assigned to the Gaussian model with the highest probability. Specifically, each Gaussian model corresponds to one mode.
[0034] For example, the maximum probability mode t corresponding to the target sample can be determined based on the following formula:
[0035]
[0036] Where N represents the Gaussian distribution function, μ k Let σ represent the mean of the k-th mode. k π represents the standard deviation of the k-th mode. k Let represent the weight of the k-th mode, and satisfy the formula . K represents the number of modalities, k' refers to the modalities associated with the financial data table, πk' represents the average weight of the K modalities, and x i,j Denotes the target sample, μ k′ σk′ represents the average of the means of the K modes, and σk′ represents the average of the standard deviations of the K modes.
[0037] Optionally, after selecting the most suitable t-th mode for the target sample, the target sample can be normalized within this mode based on the following formula to obtain the corresponding mode-normalized features:
[0038]
[0039] Where, x i,j Let μ represent the target sample, where i represents the i-th feature of the target sample, and j represents the j-th data entry in the financial data table, i.e., the row number of the target sample in the financial data table. t Let σ represent the mean value corresponding to the t-th mode. t This represents the standard deviation corresponding to the t-th mode.
[0040] For continuous features, one-hot encoding is first used to characterize which Gaussian mode represents the target sample, thus obtaining the encoded feature corresponding to the continuous feature. Specifically, if the Gaussian model with the highest probability assigned to the sample point is the third Gaussian model, then the corresponding encoded feature can be determined as [0,0,1], that is, the value of the item selected from the distribution is 1.
[0041] For example, based on the coding features and modal normalization features corresponding to each target sample in the financial data table, the normalized feature vector corresponding to each target sample in the financial data table can be determined using the following formula:
[0042]
[0043] Where, x i ′,j refers to the normalized feature vector corresponding to the target sample, d i,j Let N be the one-hot encoding of the i-th discrete feature of the j-th target sample. cN represents the number of continuous features. d Let αi,j represent the number of discrete features, αi,j represent the modal normalized feature corresponding to the i-th continuous feature of the j-th target sample, and βi represent the modal normalized feature. ,j This represents the encoded feature corresponding to the i-th continuous feature of the j-th target sample. This indicates the addition operation of corresponding items in a matrix.
[0044] S102. Determine the values of the generated discrete variables corresponding to each target sample in the financial data table, and based on the values of the generated discrete variables, determine the loss function of the preset generative adversarial network based on the cross-entropy loss function.
[0045] Here, the generated discrete variables are discrete variables selected for each target sample in the financial data table. The pre-defined loss function of the Generative Adversarial Network (GAN) includes the loss function of the discriminator and the loss function of the generator in the GAN.
[0046] Optionally, determining the values of the generated discrete variables corresponding to each target sample in the financial data table includes: selecting the corresponding generated discrete variables with equal probability for each target sample, and performing probability modeling on all values of the generated discrete variables to obtain the probability of each value of the generated discrete variables; and determining the values of the generated discrete variables corresponding to the target samples based on the probability of each value of the generated discrete variables.
[0047] For example, the discrete variable m can be generated using the following formula. i :
[0048]
[0049] Where i takes values in the range 1,...,N d , where k represents the range of values for the selected discrete variable. N d D represents the number of discrete features. d Let |D| represent the d-th discrete feature. d | represents the number of discrete variables generated.
[0050] Optionally, based on the values of discrete variables, a loss function for the pre-defined generative adversarial network is determined using the cross-entropy loss function. This includes: determining the final condition variable based on the values of the generated discrete variables corresponding to each target sample to construct the loss function of the discriminator in the generative adversarial network; constructing the cross-entropy loss between the features of the real discrete variables and the features of the generated discrete variables based on the values of the generated discrete variables, the values of the real discrete variables, and the pre-defined weight coefficients; and constructing the loss function of the generator in the generative adversarial network based on the cross-entropy loss, the final condition variable, and the ground distance.
[0051] For example, the values of the discrete variables corresponding to each target sample can be summed up according to the following formula to obtain the final condition variable:
[0052]
[0053] Among them, the selected discrete variable mi′ takes the value of 1, while the values of the other variables are all 0. d This represents the number of discrete features. Specifically, the final condition variable refers to the set of lengths of all discrete variables. If there are two discrete variables m1 and m2, m1 has 3 possible values and m2 has 2 possible values, then during the calculation of the mask variable, if m2 is selected as the second value, the final condition variable is m1 + m2, which is [0,0,0] + [0,1]. Since m1 is not selected, all values of m1 are 0.
[0054] For example, the loss function of the discriminator in a generative adversarial network can be as follows:
[0055]
[0056] Among them, E z~p(z) [D(G(z,cond))] represents the Wasserstein distance between the final condition variable and the discriminator output value. This represents the ground motion distance corresponding to the generated discrete variable. G represents the generator, D represents the discriminator, and the data in parentheses are the input data for the generator and discriminator. z is random noise, p(z) represents the distribution of random noise z, which is usually a normal distribution, x is the target sample, and Pr is the actual data distribution of the target sample.
[0057] For example, the loss function for the generator in a generative adversarial network can be as follows:
[0058]
[0059] Among them, E z~p(z) [D(G(z,cond))] represents the geodynamic distance between the final condition variable and the discriminator output value. The generator minimizes the Wasserstein distance Ε. z~p(z) [D(G(z,cond))] is used to narrow the gap with the true distribution, where λ is the weighting coefficient. For the characteristics of the true discrete variable d i With generating discrete feature variables m i Cross-entropy loss between them.
[0060] It should be noted that, in order for the generator to generate discrete variable distribution samples that are as similar as possible to the training data, in addition to random noise, a final condition variable is also needed to guide the training direction of the generator.
[0061] S103. Based on the preset objective function, loss function, preset noise parameter update strategy, and model parameter update strategy of the pre-set generative adversarial network, iteratively train the preset generative adversarial network, and use the trained generative adversarial network to generate de-identified data corresponding to financial data in order to respond to data de-identification requests.
[0062] The pre-defined generative adversarial network (GAN) includes a generator and a discriminator. The pre-defined noise parameter update strategy refers to an adaptive gradient noise mechanism introduced to better protect privacy data and accelerate network model training. Specifically, noise parameters are updated in each round of the model's iterative training process.
[0063] Optionally, during the iterative training of the preset generative adversarial network, the target noise parameters corresponding to the current training round are determined based on the preset noise parameter update strategy, and the updated model parameters corresponding to the current training round are determined based on the model parameter update strategy; the objective function of the preset generative adversarial network is constructed, and the preset generative adversarial network is iteratively trained based on the objective function, loss function, target noise parameters and updated model parameters of the preset generative adversarial network.
[0064] Optionally, during the iterative training of the preset generative adversarial network, the target noise parameter corresponding to the current training round can be determined based on the initial noise parameters, generator loss value, discriminator loss value, preset loss difference weight, preset adaptive decay coefficient, current training round, and total training rounds. That is, the target noise parameter corresponding to the current training round is determined based on the preset noise parameter update strategy.
[0065] The generator loss value refers to the loss value obtained by substituting the output of the previous training round into the generator loss function, and the discriminator loss value refers to the loss value obtained by substituting the output of the previous training round into the discriminator loss function.
[0066] For example, the target noise parameters corresponding to the current training round can be determined based on the following formula:
[0067] σ t =σ init ·β t / T +α·|Loss G -Loss D |
[0068] Where, σ t This refers to the target noise parameter, Loss GThis refers to the generator loss value. D This refers to the discriminator loss value. σ init Let be the initial noise parameter, β be the adaptive decay coefficient, t and T represent the current training round and the total number of training rounds, respectively, and α be the weight of the loss difference. Using the above formula, as the number of training rounds increases, the initial noise can be gradually reduced, improving the training stability of the model. Simultaneously, when the difference between the generator and discriminator losses is too large, noise can be increased to balance adversarial training.
[0069] Optionally, during the iteration of the generative adversarial network (GAN) model, the target pruning gradient can be determined based on the model gradient of the GAN in the current training round and the preset gradient threshold; the updated model parameters of the GAN model can be determined based on the target pruning gradient, learning rate, identity matrix, preset gradient threshold and target noise parameters, that is, the updated model parameters corresponding to the current training round are determined based on the model parameter update strategy.
[0070] The updated model parameters θ of the generative adversarial network (GAN) model include the model parameters of the generator and the discriminator. The generator's model parameters define how to transform low-dimensional random noise z into high-dimensional data samples, with the goal of making the generated anonymized data resemble real financial data as closely as possible. The discriminator's model parameters define how to calculate the authenticity score of the input samples. The goal is to assign high scores to real samples and low scores to generated fake samples.
[0071] For example, assuming the model gradient corresponding to the current training epoch is g, to avoid the gradient being too large and severely affecting model training, gradient clipping is required. Specifically, the target clipping gradient can be determined based on the following formula:
[0072]
[0073] Among them, g clip The gradient is clipped to the target, C is the gradient threshold, and ||g||2 is the L2 norm of the model gradient g of the generative adversarial network in the current training epoch.
[0074] For example, based on the target pruning gradient, learning rate, identity matrix, preset gradient threshold, and target noise parameters, the model parameters of the generative adversarial network model can be updated using the following formula to determine the updated model parameters:
[0075]
[0076] Where θ1 represents the updated model parameters corresponding to the current training epoch t, θ0 represents the model parameters corresponding to the previous training epoch, and g clip The gradient is clipped for the target, C is the gradient threshold, η is the learning rate, I is the identity matrix, and σ is the gradient threshold. tThis refers to the target noise parameter corresponding to the current training round t.
[0077] Optionally, a pre-defined objective function for the generative adversarial network is constructed, including: constructing the hidden layer structure of the generator based on the final condition variable, the pre-defined first activation function, the pre-defined second activation function, the normalization layer, and the fully connected layer; adding noise to the generator's output based on the reparameterization function; and constructing the generator's objective function based on the hidden layer structure. The discriminator's objective function is constructed based on the final condition variable, the pre-defined neuron inactivation strategy, the pre-defined third activation function, and the fully connected layer.
[0078] Optionally, the final condition variable can be determined based on the values of discrete variables generated for each target sample.
[0079] For example, the objective function of the generator can be represented by the following formula:
[0080]
[0081] Among them, h i The following characters represent different hidden layers: `cond` is the input conditional parameter; `BN` and `FC` are normalized and fully connected layers, respectively; `ReLU` and `tanh` are common activation functions; and `Gumbel` is a reparameterization function that adds Gumbel noise to the output to solve the gradient problem in discrete sampling, allowing the gradient to propagate back and update the network parameters. `z` represents random noise, and `Di` represents the i-th generated discrete feature. FC a->b This indicates that the feature length of the fully connected layer changes from the input a to b. The meanings of different subscripts of FC in the above formula are similar and will not be elaborated here.
[0082] For example, the objective function of the discriminator can be represented by the following formula:
[0083]
[0084] Among them, h i Different hidden layers are represented by FC (Fully Connected). To address the mode collapse problem commonly encountered in generative adversarial networks (GANs), the discriminator performs multi-sample processing to improve the diversity of samples generated by the generator, preventing the generator from repeatedly generating the same samples. n represents the number of concatenated samples. `drop` is a commonly used neuron inactivation strategy during model training, `leaky` is a common activation function, and `D(·)` is the output of the discriminator network, distinguishing between real and generated data. `r1` refers to the target sample; `r1` is target sample 1, `cond1` is the final condition variable 1, and `n` is typically 10. FC a->b This indicates that the feature length of the fully connected layer changes from the input a to b. The meanings of different subscripts of FC in the above formula are similar and will not be elaborated here.
[0085] The technical solution of this invention, in response to a data anonymization request for financial data, firstly determines a pre-stored financial data table and performs normalized feature processing on the financial data table to obtain a normalized feature vector corresponding to the financial data table; secondly, it determines the values of the generated discrete variables corresponding to each target sample in the financial data table, and based on the values of the generated discrete variables, determines the loss function of a preset generative adversarial network (GAN) based on the cross-entropy loss function; thirdly, iteratively trains the preset GAN based on its objective function, loss function, preset noise parameter update strategy, and model parameter update strategy, and uses the trained GAN to generate anonymized data corresponding to the financial data, thus responding to the data anonymization request. By performing comprehensive feature processing and model iteration on tabular and highly correlated financial data, anonymized data that conforms to the distribution of real financial data can be generated, ensuring that personal privacy, trade secrets, or legally protected information are not leaked when the data is shared, tested, or analyzed.
[0086] Example 2
[0087] Figure 2 This is a flowchart of a data de-identification method provided in Embodiment 2 of the present invention; based on the above embodiments, this embodiment provides a preferred example of generating de-identified data using a generative adversarial network, specifically, as follows: Figure 2 As shown, the method includes the following steps:
[0088] S201. Determine at least two modalities associated with the pre-stored financial data table, and for each target sample in the financial data table, assign the target sample to the modality with the highest probability to obtain the modality normalization feature corresponding to the target sample.
[0089] S202. Determine at least two target features of the target sample recorded in the financial data table, and determine the coding features corresponding to the discrete features and continuous features in the target features based on the preset coding strategy.
[0090] S203. Based on the coding features and modal normalization features corresponding to each target sample in the financial data table, determine the normalized feature vector corresponding to the financial data table.
[0091] S204. For each target sample, select the corresponding generated discrete variable with equal probability, and perform probability modeling on all values of the generated discrete variable to obtain the probability of each value of the generated discrete variable.
[0092] S205. Based on the probability of each value of the generated discrete variable, determine the value of the generated discrete variable corresponding to the target sample.
[0093] S206. Based on the values of the discrete variables corresponding to each target sample, determine the final condition variables to construct the loss function of the discriminator in the generative adversarial network.
[0094] S207. Based on the values of the generated discrete variables, the values of the real discrete variables, and the preset weight coefficients, construct the cross-entropy loss between the features of the real discrete variables and the features of the generated discrete variables, and combine the final condition variable and the ground motion distance to construct the loss function of the generator in the generative adversarial network.
[0095] S208. During the iterative training of the preset generative adversarial network, the target noise parameters corresponding to the current training round are determined based on the preset noise parameter update strategy, and the updated model parameters corresponding to the current training round are determined based on the model parameter update strategy.
[0096] S209. Construct a pre-defined objective function for the generative adversarial network (GAN), and iteratively train the GAN based on the pre-defined objective function, loss function, target noise parameters, and updated model parameters.
[0097] S210. In response to the request for data anonymization of financial data, a trained generative adversarial network is used to generate anonymized data corresponding to the financial data in order to respond to the data anonymization request.
[0098] Example 3
[0099] Figure 3 This is a structural block diagram of a data desensitization device provided in Embodiment 3 of the present invention. This embodiment is applicable to situations where comprehensive feature processing and model iteration of financial data are performed based on generative adversarial networks to generate desensitized data that conforms to the distribution of real financial data. The data desensitization device provided in this embodiment can execute the data desensitization method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method. This data desensitization device can be implemented in hardware and / or software and configured in an electronic device with data desensitization function, such as... Figure 3 As shown, the data anonymization device may specifically include:
[0100] The vector determination module 301 is used to respond to the data desensitization request for financial data, determine the pre-stored financial data table, and perform normalization feature processing on the financial data table to obtain the normalized feature vector corresponding to the financial data table.
[0101] The function determination module 302 is used to determine the values of the generated discrete variables corresponding to each target sample in the financial data table, and to determine the loss function of the preset generative adversarial network based on the cross-entropy loss function according to the values of the generated discrete variables.
[0102] The data desensitization module 303 is used to iteratively train the preset generative adversarial network based on the preset objective function, loss function, preset noise parameter update strategy and model parameter update strategy, and use the trained generative adversarial network to generate desensitized data corresponding to financial data in order to respond to data desensitization requests.
[0103] The technical solution of this invention, in response to a data anonymization request for financial data, firstly determines a pre-stored financial data table and performs normalized feature processing on the financial data table to obtain a normalized feature vector corresponding to the financial data table; secondly, it determines the values of the generated discrete variables corresponding to each target sample in the financial data table, and based on the values of the generated discrete variables, determines the loss function of a preset generative adversarial network (GAN) based on the cross-entropy loss function; thirdly, iteratively trains the preset GAN based on its objective function, loss function, preset noise parameter update strategy, and model parameter update strategy, and uses the trained GAN to generate anonymized data corresponding to the financial data, thus responding to the data anonymization request. By performing comprehensive feature processing and model iteration on tabular and highly correlated financial data, anonymized data that conforms to the distribution of real financial data can be generated, ensuring that personal privacy, trade secrets, or legally protected information are not leaked when the data is shared, tested, or analyzed.
[0104] Furthermore, the vector determination module 301 is specifically used for:
[0105] Identify at least two modalities associated with the financial data table, and for each target sample in the financial data table, assign the target sample to the modality with the highest probability to obtain the modality normalization feature corresponding to the target sample;
[0106] Identify at least two target features of the target sample recorded in the financial data table, and based on a preset coding strategy, determine the coding features corresponding to the discrete and continuous features in the target features respectively;
[0107] Based on the coding features and modal normalization features corresponding to each target sample in the financial data table, the normalized feature vector corresponding to the financial data table is determined.
[0108] Furthermore, the function determination module 302 is specifically used for:
[0109] For each target sample, a corresponding generated discrete variable is selected with equal probability, and probabilistic modeling is performed on all values of the generated discrete variable to obtain the probability of each value of the generated discrete variable.
[0110] Based on the probability of each value of the generated discrete variable, the value of the generated discrete variable corresponding to the target sample is determined.
[0111] Furthermore, the function determination module 302 is also used for:
[0112] Based on the values of the discrete variables generated for each target sample, the final condition variables are determined to construct the loss function of the discriminator in the generative adversarial network;
[0113] Based on the values of the generated discrete variables, the values of the real discrete variables, and the preset weight coefficients, a cross-entropy loss is constructed between the features of the real discrete variables and the features of the generated discrete variables. Based on the cross-entropy loss, the final condition variable, and the ground motion distance, a loss function for the generator in the generative adversarial network is constructed.
[0114] Furthermore, the data anonymization module 303 may include:
[0115] The determination unit is used to determine the target noise parameters corresponding to the current training round based on the preset noise parameter update strategy during the iterative training of the preset generative adversarial network, and to determine the updated model parameters corresponding to the current training round based on the model parameter update strategy.
[0116] The training unit is used to construct the objective function of the pre-defined generative adversarial network (GAN) and iteratively train the GAN based on the objective function, loss function, target noise parameters, and updated model parameters.
[0117] Furthermore, the training unit is specifically used for:
[0118] Based on the final condition variable, the preset first activation function, the preset second activation function, the normalization layer, and the fully connected layer, the hidden layer structure of the generator is constructed, and noise is added to the output of the generator based on the reparameterization function. Combined with the hidden layer structure, the objective function of the generator is constructed.
[0119] The objective function of the discriminator is constructed based on the final condition variable, the preset neuron inactivation strategy, the preset third activation function, and the fully connected layer.
[0120] Example 4
[0121] Figure 4 This is a schematic diagram of the structure of the electronic device provided in Embodiment 4 of the present invention. Figure 4A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0122] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0123] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0124] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as data desensitization methods.
[0125] In some embodiments, the data anonymization method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the data anonymization method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the data anonymization method by any other suitable means (e.g., by means of firmware).
[0126] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0127] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0128] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0129] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0130] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0131] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0132] In one embodiment, the present invention further includes a computer program product, which includes a computer program that, when executed by a processor, implements the data desensitization method of any embodiment of the present invention.
[0133] In the implementation of a computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages as well as conventional procedural programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0134] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0135] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A data anonymization method, characterized in that, include: In response to the request for data anonymization of financial data, a pre-stored financial data table is determined, and the financial data table is subjected to normalized feature processing to obtain the normalized feature vector corresponding to the financial data table. Determine the values of the generated discrete variables corresponding to each target sample in the financial data table, and based on the values of the generated discrete variables, determine the loss function of the preset generative adversarial network based on the cross-entropy loss function; Based on the preset objective function, loss function, preset noise parameter update strategy, and model parameter update strategy of the generative adversarial network, the preset generative adversarial network is iteratively trained, and the trained generative adversarial network is used to generate de-identified data corresponding to financial data in order to respond to data de-identification requests.
2. The method according to claim 1, characterized in that, Normalize the features of the financial data table to obtain the normalized feature vector corresponding to the financial data table, including: Identify at least two modalities associated with the financial data table, and for each target sample in the financial data table, assign the target sample to the modality with the highest probability to obtain the modality normalization feature corresponding to the target sample; Identify at least two target features of the target sample recorded in the financial data table, and based on a preset coding strategy, determine the coding features corresponding to the discrete and continuous features in the target features respectively; Based on the coding features and modal normalization features corresponding to each target sample in the financial data table, the normalized feature vector corresponding to the financial data table is determined.
3. The method according to claim 1, characterized in that, Determine the values of the generated discrete variables corresponding to each target sample in the financial data table, including: For each target sample, a corresponding generated discrete variable is selected with equal probability, and probabilistic modeling is performed on all values of the generated discrete variable to obtain the probability of each value of the generated discrete variable. Based on the probability of each value of the generated discrete variable, the value of the generated discrete variable corresponding to the target sample is determined.
4. The method according to claim 1, characterized in that, Based on the values of the generated discrete variables, and using the cross-entropy loss function, the loss function of the pre-defined generative adversarial network is determined, including: Based on the values of the discrete variables generated for each target sample, the final condition variables are determined to construct the loss function of the discriminator in the generative adversarial network; Based on the values of the generated discrete variables, the values of the real discrete variables, and the preset weight coefficients, a cross-entropy loss is constructed between the features of the real discrete variables and the features of the generated discrete variables. Based on the cross-entropy loss, the final condition variable, and the ground motion distance, a loss function for the generator in the generative adversarial network is constructed.
5. The method according to claim 1, characterized in that, Based on the pre-defined objective function, loss function, pre-defined noise parameter update strategy, and model parameter update strategy of the generative adversarial network, iterative training of the pre-defined generative adversarial network is performed, including: During the iterative training of the preset generative adversarial network, the target noise parameters corresponding to the current training round are determined based on the preset noise parameter update strategy, and the updated model parameters corresponding to the current training round are determined based on the model parameter update strategy. Construct a pre-defined objective function for the generative adversarial network (GAN), and iteratively train the GAN based on the pre-defined objective function, loss function, target noise parameters, and updated model parameters.
6. The method according to claim 5, characterized in that, in, The pre-defined generative adversarial network (GAN) includes a generator and a discriminator. Correspondingly, the objective function for constructing the pre-defined GAN includes: Based on the final condition variable, the preset first activation function, the preset second activation function, the normalization layer, and the fully connected layer, the hidden layer structure of the generator is constructed, and noise is added to the output of the generator based on the reparameterization function. Combined with the hidden layer structure, the objective function of the generator is constructed. The objective function of the discriminator is constructed based on the final condition variable, the preset neuron inactivation strategy, the preset third activation function, and the fully connected layer.
7. A data anonymization device, characterized in that, include: The vector determination module is used to respond to the data anonymization request for financial data, determine the pre-stored financial data table, and perform normalization feature processing on the financial data table to obtain the normalized feature vector corresponding to the financial data table. The function determination module is used to determine the values of the generated discrete variables corresponding to each target sample in the financial data table, and based on the values of the generated discrete variables, determine the loss function of the preset generative adversarial network based on the cross-entropy loss function. The data anonymization module is used to iteratively train a preset generative adversarial network (GAN) based on its objective function, loss function, preset noise parameter update strategy, and model parameter update strategy. It then uses the trained GAN to generate anonymized data corresponding to the financial data in response to data anonymization requests.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that is executed by the at least one processor to enable the at least one processor to perform the data desensitization method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the data desensitization method according to any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the data desensitization method according to any one of claims 1-6.