An industrial system public data protection method and system

By generating and adding pseudo-data against the adversarial autoencoder algorithm, the risk of inference and leakage of public data in industrial systems is solved, and the dual guarantee of data security and availability is achieved.

CN114154183BActive Publication Date: 2025-06-20XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111470390.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2025-06-20
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

The prior art is difficult to effectively protect public data from industrial systems, especially when facing the risks of data leakage and secondary leakage, traditional encryption and desensitization methods cannot avoid the inference of sensitive information implicitly in the data.

Method used

Adversarial autoencoder algorithm is used to train a group of adversarial networks through generators and discriminators to generate "pseudo-data" with the same format as the real data but different contents, and add these pseudo-data to the data to increase noise and prevent key parameters from being inferred.

Benefits of technology

It realizes hierarchical protection of public data in industrial systems, ensures the security and availability of data, avoids the risk of data being inferred and leaked, and does not affect the normal use of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114154183B_ABST
    Figure CN114154183B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for protecting public data of an industrial system, including: sorting and selecting public data of the industrial system; training a pseudo-data generator based on the original public data set; training a data discriminator based on the original public data set and a generative adversarial network; designing a hierarchical protection method for public data for objects with different security levels based on a noise addition strategy, and implementing equal distribution of public data for objects with different security levels and subsequent content restoration for high-security-level objects based on this method; the present invention is simple to implement and has a low computational complexity, effectively reducing the computational resource overhead during the later use of the model through the pre-training of the model. The present invention realizes that while public data is disclosed, it can prevent untrusted third parties from obtaining higher-security-level information through means such as inference inversion, so that the key information hidden in the disclosed data is not leaked, and through the differential hierarchical protection method, it ensures the complete acquisition of the real data content by the trusted party.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of artificial intelligence and data security, and particularly relates to a method and a system for protecting public data of an industrial system. Background Art

[0002] Modern industrial systems are large-scale and complex cyber-physical systems that contain a large amount of data information, such as systematic control parameters, state estimation model parameters, industrial system real-time measurement data, user-level terminal data, etc. Taking the power system as an example, a large amount of power system-related data needs to be made public to different social groups. On the other hand, due to the cyber-physical integration characteristics of industrial systems, some key data of industrial systems are at risk of leakage or secondary leakage when other relevant data are made public.

[0003] Traditional data protection means usually include data encryption and data desensitization. Encryption schemes are generally used to protect secret data for specific personnel, and their characteristics of high confidentiality level and strong pertinence are not applicable to public data. Therefore, at present, desensitization methods are usually used to protect public data, that is, by deleting or modifying some sensitive information in the data (such as privacy information such as personal information and enterprise operations) or adding some data noise (such as face blurring) to achieve the separation of sensitive information. However, these methods can only prevent the direct utilization of data sensitive information by the external environment, and cannot avoid some indirect and secondary technical means to obtain the sensitive information hidden in the data. For example, for the network parameters in industrial system data, they are generally considered as key data that should not be made public, and once leaked, it may cause serious consequences. However, there are already some technical means to infer network parameters from industrial system public data. At this time, the traditional desensitization method will face two failure problems: 1. If the original data is deleted, modified or blindly added with noise, it will cause data loss and lead to serious errors in some functions of the industrial system itself, such as state estimation and power flow calculation; 2. If only the system parameters are hidden without processing the public data, it will face the risk of being inferred.

[0004] Therefore, new data protection methods are needed, which can have a certain anti-inference and anti-leakage ability while ensuring that the normal use of data is not affected. Summary of the Invention

[0005] To overcome the above-mentioned drawbacks of the prior art, the object of the present invention is to provide a method and system for protecting public data in an industrial system. By combining the adversarial autoencoder algorithm, a generator and a discriminator in a set of adversarial networks are trained using the original public data. The generator generates "pseudo-data" with the same content format and similar content as the real data, and the discriminator is used by the trusted party to "eliminate" the pseudo-data. For low-level objects, a dataset with "pseudo-data" noise added is publicly released, so that the key parameters of the industrial system are protected from the risk of being inferred. For high-level trusted parties, the discriminator is given the task of denoising the publicly released dataset with noise added, ensuring the normal utilization of the relevant data of the industrial system for specific purposes, and achieving the purpose of hierarchical protection of public data. Its advantages are as follows: Compared with the desensitization method, its noise addition strategy is rule-based (generated by the generation network); when applied to industrial system data, the similarity between the "noise" and the real data can be improved by adding prior information, and this degree is controllable. The present invention realizes a controllable and reversible data noise addition method through a generation rule based on deep learning training, which not only ensures the security of the public data in the industrial system, but also ensures the availability of the public data. It has the advantages of easy data acquisition, simple model training, wide user coverage, easy later deployment, and low computational consumption during system operation, making this application have obvious advantages compared with traditional methods and systems.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] An industrial system public data protection method, characterized by comprising:

[0008] Step 1, obtaining multiple groups of historical data {x} of the public data of the industrial system to be protected in the format of multiple groups of column vectors;

[0009] Step 2, determining the generation mechanism model of the public data selected in Step 1 and the corresponding implicit key parameters based on the prior knowledge related to the industrial system;

[0010] Step 3, constructing a generation network G for generating pseudo-data of the public data selected in Step 1 based on an autoencoder, and determining the activation function of the encoder layer and the setting of some network structure parameters in the generation network G according to the mathematical characteristics of the generation mechanism model and the implicit key parameters obtained in Step 2. Then, input the dataset {x} obtained in Step 1 to train the autoencoder until the decoder layer outputs the reconstruction result of the original data and meets the training termination condition. Specifically:

[0011] Step 3.1, determining a suitable activation function f of the autoencoder network according to the generation mechanism model a ;

[0012] Step 3.2, set the autoencoder network structure parameters {α} based on the mathematical properties of the implicit key parameters, and set the network training parameters {β}, and select an appropriate network loss function f l ;

[0013] Step 3.3, input the original data set {x1, x2,..., x m} into the autoencoder network G for training until the total error value of the samples is less than the set threshold or the maximum number of training times is reached. At the same time, adjust the network training parameters {β} to obtain a better training result, where m is the number of groups of original data;

[0014] Step 4, select an appropriate probability distribution p(z′) based on the prior knowledge obtained in Step 2 to generate the pseudo-code z′, and then based on the adversarial network training framework, use the real code z obtained from the encoding layer of the network G generated by the real data input in Step 3 as the negative sample and the pseudo-code z′ as the positive sample to train the discriminator network D, and synchronously train the generator network G to make the probability distribution of the real code z closer to the predefined distribution p(z′). Specifically:

[0015] Step 4.1, select an appropriate pseudo-code probability distribution p(z′) based on the generation mechanism of the public data;

[0016] Step 4.2, generate a sufficient set of pseudo-code vectors {z′} by sampling on the probability distribution p(z′);

[0017] Step 4.3, construct the discriminator network D, set the network structure parameters {α D} and training parameters {β D}, use the set of pseudo-code vectors {z′} as the positive sample, and use the set of real data codes {z} output by the encoding layer of the generator network G trained in Step 3 as the negative sample to train the discriminator network D until the total error value is less than the set threshold or the maximum number of training times is reached;

[0018] Step 4.4, while training the discriminator network D in Step 4.3, synchronously train the network parameters of the generator network G based on the loss function f GD related to the training error of the discriminator network to make the output of the encoder layer of the generator network closer to the real data code of the pseudo-code;

[0019] Step 5, after the training of the generator network G and the discriminator network D in Steps 3 and 4 is completed, use the generator network G to generate a sufficient amount of pseudo-data set {x′}, and then add the pseudo-data to the real data set at the mixing ratio γ to obtain the publicly available data set after noise addition. Specifically:

[0020] Step 5.1, resample on the probability distribution p(z′) to generate a sufficient amount of pseudo-code {z″};

[0021] Step 5.2, input the pseudo-coding set {z″} into the decoding layer of the trained generation network G to output the pseudo-dataset {x′};

[0022] Step 5.3, assume the mixing ratio is γ, then add γm groups of pseudo-data to the real dataset to obtain the publicly available dataset after noise addition processing;

[0023] Step 6, publicly disclose different classified datasets to corresponding objects according to the security classification strategy. Specifically:

[0024] Step 6.1, directly disclose the publicly available dataset after noise addition processing in Step 5 to low-classified objects;

[0025] Step 6.2, deploy the network model of the discriminant network D trained in Step 4 at the data receiving end of high-classified objects, and publicly disclose the publicly available dataset after noise addition processing;

[0026] Step 6.3, use the discriminator D to perform noise removal processing on the publicly available dataset, so as to obtain the complete real dataset for high-classified objects.

[0027] Furthermore, the publicly disclosed data of the industrial system to be protected in the present invention refers to the low-classified data to be publicly disclosed in the industrial system, such as the data related to power generation, power consumption, power transmission, and electricity price in the power system, etc. Specifically, the data to be protected refers to that this type of data is restricted by a part of mathematical models, and attackers may perform calculations such as inference and inversion based on this part of the models or data, so as to obtain some non-public high-classified data. Multiple groups of historical data {x} refer to multiple data sections of this publicly disclosed data taken within a period of time, and each group of data is stored in the form of an n-dimensional column vector (n is determined by the nature of this type of data itself), so as to form the original dataset for model training.

[0028] Furthermore, the data generation mechanism model mentioned in Step 2 refers to a mathematical model established according to the physical model for generating this type of publicly disclosed data and the corresponding physical constraints, which determines the structure and mathematical properties of the data and can be used as the prior knowledge for constructing the pseudo-data generator; the corresponding implicit key parameters refer to some high-classified industrial system parameters in certain mathematical models that are directly related to this type of data and may be illegally obtained by calculation methods such as inference, inversion estimation, etc.

[0029] Furthermore, the autoencoder mentioned in Step 3 is an artificial neural network algorithm for learning the encoded representation of input data. Through a deep learning network with an input-hidden layer-output structure, the autoencoder can realize the replication output of the input data, and at the same time, the encoding output by the hidden layer can be regarded as a low-dimensional and efficient representation of the input data, so that similar data of the original data can be generated through more encodings. Therefore, the autoencoder is also often used as a generation model.

[0030] Further, the network structure parameters mentioned in step 3.2 refer to the construction scheme of the artificial neural network, specifically including the setting of the depth of the network hierarchy (mainly referring to the number of hidden layers), the number of neuron nodes in each layer (in the present invention, the neural network is mainly used for the generation model, so the number of nodes in the input layer and the output layer is the same as the dimension of the original data; the hidden layer is set according to experience, generally less than the number of nodes in the input layer, playing an encoding effect), the setting of the neuron activation function, etc. The network training parameters refer to a series of setting values in the training process of the artificial neural network, such as the initial weight of the nodes, the batch size, the number of iterations of the data set (epoch), the loss function, the allowable error (iteration terminates when the condition is met), etc.

[0031] Further, the adversarial network training framework mentioned in step 4 refers to the Generative Adversarial Network (GAN), an artificial neural network training framework for training a generative network that can generate pseudo-samples extremely similar to real samples and a discriminative network that can distinguish between original samples and pseudo-samples. Through the adversarial training of the generative network and the discriminative network, the generative network tries to generate pseudo-samples that can deceive the discriminative network as much as possible, while the discriminative network tries to distinguish between real samples and pseudo-samples as much as possible.

[0032] The present invention also provides an industrial system public data protection system, including four modules: a data acquisition module, a model training module, a noise mixing module, and a noise screening module, which can realize the hierarchical protection of public data and realize the data disclosure of objects with different confidentiality levels through noise addition and noise removal strategies. It is characterized in that it includes:

[0033] The data acquisition module obtains multiple groups of historical data {x} of the industrial system public data to be protected from the industrial system data management center, and generates a sufficient amount of pseudo-codes from a pre-set probability distribution p(z′);

[0034] The model training module constructs a generative network G based on the autoencoder, inputs multiple groups of historical data {x} to train the network to output the reconstruction result of the original data; then constructs a neural network D as the discriminative network, and based on the pseudo-codes generated from p(z′) and the real codes generated by the generative network G, uses the adversarial network framework to co-train the generative network G and the discriminative network D, so that the real codes output by the encoding layer of the generative network G approximately follow the probability distribution p(z′) while the discriminative network D can realize the discrimination between real codes and pseudo-codes;

[0035] The noise mixing module uses a pre-set probability distribution p(z′) to generate a sufficient amount of pseudo-codes, and then generates pseudo-data through the decoding layer of the generative network G, and mixes it with the original data at a proportional coefficient γ to generate a noisy public data set;

[0036] The noise filtering module uses the discriminant network D to discriminate the noisy dataset, filter out the false data, and thus obtain a complete real dataset for high-level objects.

[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0038] 1. In the process of establishing the data noise addition model, the computational consumption is mainly concentrated in the process of training the generation network and the discriminant network in the early stage. In the actual later use process, the computational resources consumed for data noise addition and denoising are much lower than the traditional methods.

[0039] 2. The system model structure after training is simple, with a small size, and it is relatively easy to realize lightweight application-level deployment.

[0040] 3. Through one-time training to generate a model, it can complete the hierarchical protection of public data, and realize controllable noise addition and denoising restoration of data. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a block diagram of the method for protecting public data of industrial systems according to the present invention.

[0042] Figure 2 is a framework diagram of the system for protecting public data of industrial systems according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The following will describe the embodiments of the present invention in detail with reference to the drawings and embodiments.

[0044] The method for protecting public data of industrial systems in the present invention specifically includes a data acquisition process, a model training process, a noise mixing process, and a noise filtering process, which can realize hierarchical protection of public data and realize data disclosure for objects of different confidentiality levels through noise addition and denoising strategies. Figure 1 is a block diagram of the method for protecting public data of industrial systems according to the present invention. The system in the present invention is described in the form of a block diagram. Figure 2 is a framework diagram of the system for protecting public data of industrial systems according to the present invention.

[0045] Data Acquisition Process

[0046] Table 1 is an example of node injection power data of a 30-node power system. Table 2 is an example of 20 groups of pseudo-coded data sampled from a preset 10-dimensional standard normal distribution p(z′).

[0047] Table 1

[0048]

[0049] Table 2

[0050]

[0051]

[0052] The specific process of data acquisition is as follows:

[0053] (1) Obtain the publicly available industrial system data to be disclosed through the original data sources of the industrial system (such as the data centers of the relevant departments of power generation, power sales, and power consumption in the power system in this example). In this example, the power system node injection power data is used as an example to demonstrate the subsequent implementation process. This data is a power grid measurement data for real-time monitoring and estimation calculations and is usually relatively easy to obtain;

[0054] (2) Determine the distribution function of the multi-dimensional probability distribution p(z′) according to the mechanism model of the publicly available industrial system data. The dimension should be consistent with the coding layer in the generation network G;

[0055] (3) Generate a sufficient amount of pseudo-codes z′ based on p(z′).

[0056] Model training process

[0057] Build a generation network G based on the autoencoder, input multiple groups of historical data to train the network to output the reconstruction results of the original data; then build a neural network D as the discriminant network. Based on the pseudo-codes generated by p(z′) and the real codes generated by the generation network G, use the adversarial network framework to co-train the network G and the network D, so that the real codes output by the coding layer of the network G approximately follow the distribution p(z′) while the network D can distinguish between real codes and pseudo-codes. The specific process is as follows:

[0058] (1) Build an autoencoder network G based on the autoencoder to generate pseudo-data of this type of publicly available data. Determine the appropriate activation function f according to the generation mechanism model a , and set the structure parameters {α} and network training parameters {β} of the autoencoder network based on the mathematical characteristics of the implicit key parameters, and select an appropriate loss function f l ;

[0059] (2) Input the original data set {x1, x2,..., x m} into the autoencoder network G for training until the total error value of the samples is less than the set threshold or the maximum number of training times is reached. At the same time, adjust the network training parameters {β} to obtain better training results;

[0060] (3) Build a discriminant network D, set the structure parameters {α D} and training parameters {β D}. Use a partial set of pseudo-code vectors {z′} as positive samples and the coding set {z} of the real data output by the coding layer of the generation network G obtained in step 3 as negative samples to train the discriminant network D until the total error value is less than the set threshold or the maximum number of training times is reached;

[0061] (4) While training the discriminant network D, based on the loss function f related to the training error of the discriminant network GD Synchronously train the network parameters of the generator network G to make the output of the generator network encoder layer closer to the true data encoding of the pseudo-code.

[0062] Noise mixing process

[0063] Generate a sufficient amount of pseudo-codes with a preset probability distribution p(z′), then generate pseudo-data through the decoder layer of the generator network G, and mix it with the original data at a proportionality coefficient γ to generate a noisy public dataset. The specific implementation process is as follows:

[0064] (1) Obtain another part of the dataset {z″} from the sufficient amount of pseudo-codes generated by upsampling the probability distribution p(z′);

[0065] (2) Input the pseudo-code set {z″} into the decoder layer of the trained generator network G to output the pseudo-dataset {x′};

[0066] (3) Let the mixing ratio be γ, then add γm groups of pseudo-data to the real dataset (m is the number of groups of the real dataset) to obtain the noisy public dataset.

[0067] Table 3 shows an example of the mixed data with a proportionality coefficient of 0.5, and the serial numbers 11 - 20 are all pseudo-data.

[0068] Table 3

[0069]

[0070] It can be seen that the numerical distribution of the mixed pseudo-data and real data is approximate, and there is no feasibility of manual screening. At the same time, without other reliable homologous industrial system-related data, data recipients cannot eliminate the noisy data through technical methods such as state estimation and error identification.

[0071] Noise screening process

[0072] Use the discriminant network D to discriminate the noisy dataset and screen out the pseudo-data, so as to obtain a complete real dataset for high-level classified objects. The specific implementation process is as follows:

[0073] (1) Deploy the network model of the discriminant network D trained during the model training process at the data receiving end of the high-level classified object and give it the noisy dataset;

[0074] (2) Use the discriminator D to denoise the public dataset to obtain a complete real dataset for high-level classified objects.

[0075] Table 4 is an example of a mixed dataset obtained by shuffling the dataset serial numbers in Table 3 and then performing denoising discrimination using a discriminant network.

[0076] Table 4

[0077]

[0078] Among them, the columns in regular font are the correctly discriminated real data, the columns in bold font (color cannot be used for distinction, it is recommended to use italics for distinction) are the correctly discriminated pseudo data, and the columns in italics are the pseudo data misjudged as real data. In the test experiment of this example, the relative error between the denoised data and the real data results is less than 0.15%, that is, among every thousand groups of publicly available injection power measurement data, the number of groups of real data lost due to misjudgment or pseudo data misjudged as real data does not exceed two groups. This error is negligible in actual application scenarios.

Claims

1. An industrial system public data protection method, characterized in that, It includes the following steps: Step 1: Obtain multiple groups of historical data {x} of the industrial system public data to be protected in the format of multiple groups of column vectors; Step 2: Based on the prior knowledge related to the industrial system, determine the generation mechanism model of the public data selected in Step 1 and the corresponding implicit key parameters; Step 3: Based on the autoencoder, construct a generation network G for generating pseudo-data of the public data selected in Step 1. According to the mathematical characteristics of the generation mechanism model and the implicit key parameters obtained in Step 2, determine the activation function of the encoder layer in the generation network G and the setting of some network structure parameters. Then input the data set {x} obtained in Step 1 to train the autoencoder until the decoder layer outputs the reconstruction result of the original data and meets the training termination condition; Step 4: Select the probability distribution p(z') to generate the pseudo-code z' according to the prior knowledge obtained in Step 2. Then, based on the adversarial network training framework, use the real code z obtained by inputting the real data into the encoding layer of the generation network G in Step 3 as the negative sample and the pseudo-code z' as the positive sample to train the discriminant network D, and synchronously train the generation network G to make the probability distribution of the real code z closer to the predefined distribution p(z'); Step 5: After the training of the generation network G and the discriminant network D is completed, use the generation network G to generate a sufficient amount of pseudo-data set {x'}. Then add the pseudo-data to the real data set at the mixing ratio γ to obtain the publicly available data set after noise addition processing; Step 6: Disclose the data sets of different security levels to the corresponding objects according to the security classification strategy, as follows: Step 6.1: Directly disclose the publicly available data set after noise addition processing in Step 5 to the low-security-level objects; Step 6.2: Deploy the network model of the discriminant network D trained in Step 4 at the data receiving end of the high-security-level objects, and disclose the publicly available data set after noise addition processing; Step 6.3: Use the discriminant network D to perform denoising processing on the publicly available data set to obtain the complete real data set for the high-security-level objects.

2. The industrial system public data protection method according to claim 1, characterized in that, In Step 1, the industrial system public data to be protected refers to the low-security-level data to be publicly disclosed in the industrial system. This type of data is restricted by a part of the mathematical model, and there is a possibility that attackers can obtain some non-disclosable high-security-level data based on this part of the model or data.

3. The industrial system public data protection method according to claim 1, characterized in that, In Step 1, the multiple groups of historical data {x} refer to multiple groups of data sections of the public data within a certain period of time. Each group of data is stored in the form of an n-dimensional column vector, thus forming the original data set for model training.

4. The industrial system public data protection method according to claim 1, characterized in that, In Step 2, the data generation mechanism model refers to the mathematical model established according to the physical model generating the public data and the corresponding physical constraints, which determines the structure and mathematical properties of the data and can be used as the prior knowledge for constructing the pseudo-data generator; the corresponding implicit key parameters refer to some high-security-level industrial system parameters in the mathematical model that have a direct relationship with the data and may be illegally obtained.

5. The industrial system public data protection method according to claim 1, characterized in that, Step 3 specifically includes: Step 3.1, determine the activation function f of the autoencoder network according to the generation mechanism model a ; Step 3.2, set the autoencoder network structure parameters {α} based on the mathematical characteristics of the implicit key parameters, and set the network training parameters {β}, and select the network loss function f l ; Step 3.3, input the original data set {x} = {x1, x2,..., x m} into the generation network G for training until the total error value of the samples is less than the set threshold or the maximum number of training times is reached, and at the same time adjust the network training parameters {β} to obtain a better training result, where m is the number of groups of the original data.

6. The industrial system public data protection method according to claim 1, characterized in that, Step 4 specifically includes: Step 4.1: Based on the generation mechanism of the public data, select the pseudo-code probability distribution p(z'); Step 4.2: Generate a sufficient amount of pseudo-code vector set {z'} by sampling on the probability distribution p(z'); Step 4.3, construct the discriminant network D, and set the initial structural parameters {α D} and training parameters {β D} of the network. Use the pseudo-coding vector set {z'} as the positive samples, and use the coding set {z} of the real data output by the coding layer of the generative network G obtained in Step 3 as the negative samples to train the discriminant network D until the total error value is less than the set threshold or the maximum number of training times is reached; Step 4.4, while training the discriminant network D in Step 4.3, synchronously train the network parameters of the generator network G based on the loss function f related to the training error of the discriminant network, so that the output of the encoder layer of the generator network is closer to the true data encoding of the pseudo-code. GD ​ 7. The industrial system public data protection method according to claim 1, characterized in that, The specific steps of step 5 include: Step 5.1, resample on the probability distribution p(z') to generate a sufficient amount of pseudo-codes {z''}; Step 5.2, input the pseudo-code set {z''} into the decoding layer of the trained generation network G to output a pseudo-dataset {x'}; Step 5.3, let the mixing ratio be γ, then add γm groups of pseudo-data to the real dataset to obtain a publicly available dataset after noise addition processing, where m is the number of groups of original data.

8. The industrial system public data protection method according to claim 7, characterized in that, The mixing ratio γ determines the ratio of pseudo-data to real data in the noisy dataset. This value takes a decimal within the range of (0, 1), and the larger the value, the stronger the ability of the dataset to resist inference and inversion attacks, but the more difficult it is to restore the real dataset.

9. An industrial system public data protection system, characterized in that, It includes: A data acquisition module, which acquires multiple groups of historical data {x} of the publicly available industrial system data to be protected in the format of multiple column vectors from the industrial system data management center; determines the generation mechanism model of the selected publicly available data and the corresponding implicit key parameters based on the prior knowledge related to the industrial system, and generates a sufficient amount of pseudo-codes z' from the pre-set probability distribution p(z') according to the acquired prior knowledge; A model training module, which constructs a generation network G for generating pseudo-data of the publicly available data selected by the data acquisition module based on an autoencoder, and determines the activation function of the encoder layer and the setting of some network structure parameters in the generation network G according to the mathematical characteristics of the acquired generation mechanism model and implicit key parameters, and then inputs multiple groups of historical data {x} to train the network to output the reconstruction result of the original data until the reconstruction result of the original data is output by the decoding layer and meets the training termination condition; Then construct a neural network as a discriminative network D, use the pseudo-codes z' generated by p(z') as positive samples, and the real codes z generated by inputting real data into the encoding layer of the generation network G as negative samples, and use the adversarial network framework to co-train the generation network G and the discriminative network D, so that the real codes z output by the encoding layer of the generation network G approximately follow the probability distribution p(z') while the discriminative network D can realize the discrimination between the real codes z and the pseudo-codes z'; A noise mixing module, after the training of the generation network G and the discriminative network D is completed, uses the generation network G to generate a sufficient amount of pseudo-datasets {x'}, and then adds the pseudo-data to the real dataset at the mixing ratio γ to obtain a publicly available dataset after noise addition processing; A noise screening module, uses the discriminative network D to discriminate the noisy dataset and screen out the pseudo-data, so as to obtain a complete real dataset for high-level objects.

Citation Information

Patent Citations

  • Deeply differential privacy protection method based on generative adversarial network

    CN107368752A

  • Service providing method based on internet big data

    CN110765337A