A method, apparatus, device, medium and product for detecting data poisoning attacks
By preprocessing and reconstructing error calculations on the initial data set of the recommendation system, identifying and replacing abnormal data, the problem of data poisoning attacks is solved, and the accuracy and user experience of the recommendation system are improved.
Patent Information
- Application Number
- CN202411650746.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-19
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-11-19
AI Technical Summary
Recommendation systems are vulnerable to data poisoning attacks, resulting in a decrease in recommendation accuracy and fairness and affecting user experience.
By preprocessing and reconstructing the initial data set, reconstruction errors are calculated, abnormal data is identified and replaced, and sample data is generated using the Dirichlet distribution to fill the abnormal locations to ensure the accuracy of the data set.
It improves the accuracy of the recommendation information of the recommendation system and improves the user experience.
Smart Images

Figure CN119598146B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of artificial intelligence technology. More specifically, embodiments of the present invention relate to a method, apparatus, device, medium and product for detecting data poisoning attacks. Background Art
[0002] This section aims to provide background or context for the embodiments of the present invention recited in the claims. The description herein is not admitted to be prior art merely because it is included in this section.
[0003] In modern information society, recommendation systems play a crucial role and are widely used in many fields such as e-commerce and social media, providing users with convenient information acquisition channels. It can accurately recommend items of interest based on users' historical interaction data, greatly enhancing the user experience. However, recommendation systems are extremely vulnerable to the threat of data poisoning attacks. Attackers can inject false users or false data into the system to cleverly manipulate the learning process and recommendation decisions of the recommendation system for the purpose of promoting specific items or interfering with normal recommendations. Such attack behaviors will seriously damage the accuracy and fairness of the recommendation system and also bring bad shopping and information acquisition experiences to users. Summary of the Invention
[0004] In this context, embodiments of the present invention are expected to provide a method, apparatus, device, medium and product for detecting data poisoning attacks.
[0005] In the first aspect of the embodiments of the present invention, a method for detecting data poisoning attacks is provided, including:
[0006] Preprocessing the initial data in the obtained initial dataset to obtain a preprocessed dataset;
[0007] Respectively reconstructing each preprocessed data in the preprocessed dataset to obtain the reconstructed data of each preprocessed data;
[0008] Based on each preprocessed data and the reconstructed data of each preprocessed data, calculating the reconstruction error of each preprocessed data;
[0009] Determining the preprocessed data with a reconstruction error greater than a preset error threshold in the preprocessed dataset as abnormal data.
[0010] In an embodiment of this embodiment, the preprocessing the initial data in the obtained initial dataset to obtain a preprocessed dataset includes:
[0011] Performing redundancy processing on the initial data in the obtained initial dataset to obtain a target dataset; wherein, the target dataset includes a plurality of target data;
[0012] Normalize the target data in the target dataset to obtain a preprocessed dataset.
[0013] In one embodiment of this implementation manner, reconstructing one target preprocessed data in the preprocessed dataset to obtain the reconstructed data of the target preprocessed data includes:
[0014] Obtain the encoding weight matrix and encoding bias vector of a pre-constructed encoder, and the decoding weight matrix and decoding bias vector of a pre-constructed decoder;
[0015] Encode one target preprocessed data in the preprocessed dataset based on the encoding weight matrix and the encoding bias vector to obtain encoded data;
[0016] Reconstruct the encoded data based on the decoding weight matrix and the decoding bias vector to obtain the reconstructed data of the target preprocessed data.
[0017] In one embodiment of this implementation manner, the calculation formula for the reconstruction error of the target preprocessed data is:
[0018]
[0019] where L i represents the reconstruction error of the target preprocessed data, N represents the number of preprocessed data in the preprocessed dataset, x i represents the target preprocessed data, represents the reconstructed data of the target preprocessed data.
[0020] In one embodiment of this implementation manner, after determining the preprocessed data with a reconstruction error greater than the preset error threshold in the preprocessed dataset as abnormal data, the method further includes:
[0021] Determine the neighborhood set of the abnormal data from the preprocessed dataset;
[0022] Determine the set of embedding vectors corresponding to the neighborhood set;
[0023] Calculate the average embedding vector of the set of embedding vectors;
[0024] Calculate based on the preset parameter for the average embedding vector to obtain Dirichlet distribution data;
[0025] Generate sample data based on the Dirichlet distribution data;
[0026] Replace the abnormal data in the preprocessed dataset with the sample data to obtain a normal dataset.
[0027] In one embodiment of this embodiment, determining the neighborhood set of the abnormal data from the preprocessed data set includes:
[0028] Convert the format of the preprocessed data set to obtain a tree-shaped data structure data set;
[0029] Based on the tree-shaped data structure data set, calculate the current distance between each piece of preprocessed data and the abnormal data;
[0030] Determine the target preprocessed data in the preprocessed data whose current distance is less than the preset distance threshold;
[0031] Construct the neighborhood set of the abnormal data based on the target preprocessed data.
[0032] In the second aspect of the embodiments of the present invention, a data poisoning attack detection device is provided, and the device includes:
[0033] A preprocessing unit, configured to preprocess the initial data in the obtained initial data set to obtain a preprocessed data set;
[0034] A reconstruction unit, configured to reconstruct each piece of preprocessed data in the preprocessed data set respectively to obtain the reconstructed data of each piece of preprocessed data;
[0035] A calculation unit, configured to calculate the reconstruction error of each piece of preprocessed data based on each piece of preprocessed data and the reconstructed data of each piece of preprocessed data;
[0036] A determination unit, configured to determine the preprocessed data in the preprocessed data set whose reconstruction error is greater than the preset error threshold as abnormal data.
[0037] In the third aspect of the embodiments of the present invention, a computing device is provided, and the computing device includes: at least one processor, a memory, and an input / output unit; wherein, the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the method according to any one of the first aspect.
[0038] In the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, which includes instructions that, when running on a computer, cause the computer to execute the method according to any one of the first aspect.
[0039] In the fifth aspect of the embodiments of the present invention, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, it implements the method according to any one of the first aspect.
[0040] The data poisoning attack detection method, device, equipment, medium and product according to the embodiments of the present invention can normalize data, so that data with different features are on the same scale. Furthermore, each preprocessed data in the preprocessed data set can be reconstructed, and the reconstruction error between the preprocessed data and the reconstructed data can be calculated. Thus, the preprocessed data with a large reconstruction error can be determined as abnormal data. In this way, the abnormal data in the input initial data set can be accurately determined, ensuring that the system for recommending to users based on the initial data set can work more accurately, improving the accuracy of the recommendation information generated by the recommendation system, and further enhancing the user experience of the recommendation system. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown by way of example and not limitation, wherein:
[0042] Figure 1 is a flowchart of the data poisoning attack detection method provided by an embodiment of the present invention;
[0043] Figure 2 is a schematic structural diagram of the data poisoning attack detection device provided by an embodiment of the present invention;
[0044] Figure 3 schematically shows a structural diagram of a medium according to an embodiment of the present invention;
[0045] Figure 4 schematically shows a structural diagram of a computing device according to an embodiment of the present invention.
[0046] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and then implement the present invention, and do not limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to convey the scope of the present disclosure fully to those skilled in the art.
[0048] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, equipment, method or computer program product. Therefore, the present disclosure can be specifically implemented in the following forms: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0049] According to an embodiment of the present invention, a method, device, equipment, medium and product for detecting data poisoning attacks are provided.
[0050] It should be noted that the number of any element in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.
[0051] Next, with reference to several representative embodiments of the present invention, the principles and spirit of the present invention will be explained in detail.
[0052] Exemplary method
[0053] Next, with reference to Figure 1 , Figure 1 is a schematic flowchart of a method for detecting data poisoning attacks provided in an embodiment of the present invention. It should be noted that the embodiments of the present invention can be applied to any applicable scenario.
[0054] Figure 1 The process of the method for detecting data poisoning attacks provided in an embodiment of the present invention shown in includes:
[0055] Step S101, preprocess the initial data in the obtained initial dataset to obtain a preprocessed dataset.
[0056] In an embodiment of the present invention, the initial data in the initial dataset may be the behavior data of users obtained by a recommendation system, and the recommendation system can infer the content that users like to watch based on the obtained behavior data of users. It can be seen that the behavior data of users (i.e., the initial data) obtained by the recommendation system is crucial for the recommendation content output by the recommendation system to users.
[0057] In an embodiment of the present invention, preprocessing the initial data may include cleaning redundant data from the initial data and normalizing the initial data.
[0058] As an optional embodiment, the manner in which step S101 preprocesses the initial data in the obtained initial dataset to obtain a preprocessed dataset may include:
[0059] Perform redundancy processing on the initial data in the obtained initial dataset to obtain a target dataset; wherein, the target dataset includes multiple target data;
[0060] Perform normalization processing on the target data in the target dataset to obtain a preprocessed dataset.
[0061] In an embodiment of the present invention, the formula for normalizing the target data may be:
[0062]
[0063] Among them, X is the initial data, and X min is the smallest initial data in the initial dataset, and X max is the largest initial data in the initial dataset, and X normalized is the obtained preprocessed data.
[0064] In this way, the initial data is mapped to the range of [0, 1], and then the dimensional difference of the data can be eliminated.
[0065] Step S102: Reconstruct each preprocessed data in the preprocessed dataset respectively to obtain the reconstructed data of each preprocessed data.
[0066] In the embodiment of the present invention, the preprocessed data can be reconstructed by a pre-constructed autoencoder, and the autoencoder can include an encoder and a decoder. The autoencoder needs to be trained in advance to obtain a better data reconstruction effect.
[0067] In the embodiment of the present invention, before training the autoencoder, a training dataset needs to be obtained. The training data in the training dataset is obtained from real-world datasets related to the recommendation system, such as the FilmTrust dataset, ML-100K (MovieLens-100K), and ML-1M (MovieLens-1M), and these datasets can be obtained on relevant open data platforms.
[0068] After that, the obtained training dataset can be preprocessed. First, filter the training data of cold-start users (with less than 15 ratings) in the training dataset that seriously affect the recommendation system, and at the same time clean the redundant training data. Secondly, use the common min-max normalization method to process the training dataset, map the training dataset to a unified range (such as [0, 1]) to obtain the target training dataset.
[0069] After obtaining the target training dataset, the target training dataset can be simulatedly attacked to obtain the attacked target training dataset. Multiple attack methods can be used to simulate the attack on the target training dataset, such as AUSH, Average, Random, PGA, TNA, Infmix, etc. attacks. To simulate the attack situations that the recommendation system may suffer in the actual scenario. Assume that the target training dataset is D. In the simulated attack, a certain attack scale p is used to modify some data values in the target training dataset to generate the attacked target training dataset X.
[0070] At this time, the autoencoder can be trained based on the attacked target training dataset X.
[0071] The encoder transforms the target training dataset X = {x1, x2, x3, ...} ∈ R n For each training data x i Mapped to the hidden layer, representing h i ∈R m (where m <n),这一过程通过线性变换实现:
[0072] h i =f(x i )=σ(Wx i +b)
[0073] Where W∈R m×n is the encoding weight matrix of the encoder, b∈R m is the encoder’s encoding bias vector, and σ is the activation function (i.e., Sigmoid function).
[0074] The decoder reconstructs the hidden layer representation {h1,h2,h3,…} and obtains the output
[0075]
[0076] Where W'∈R n×m is the decoding weight matrix of the decoder, b'∈R n is the decoding bias vector of the decoder and σ' is the corresponding activation function.
[0077] During the training process of the autoencoder, the mean square error (MSE) is used as the loss function to measure the reconstruction effect:
[0078]
[0079] Among them, x i is the i-th sample of the original data, is the i-th sample of the reconstructed data, and N is the number of samples.
[0080] The parameters W, b, W', b' are continuously adjusted through optimization algorithms (such as stochastic gradient descent algorithm) to minimize the loss function, thereby obtaining the optimal autoencoder model.
[0081] As an optional implementation manner, step S102 reconstructs a target preprocessed data in the preprocessed data set, and a method of obtaining the reconstructed data of the target preprocessed data may include:
[0082] Obtaining an encoding weight matrix and an encoding bias vector of a pre-built encoder, and a decoding weight matrix and a decoding bias vector of a pre-built decoder;
[0083] Encoding a target preprocessed data in the preprocessed dataset based on the encoding weight matrix and the encoding bias vector to obtain encoded data;
[0084] Reconstructing the encoded data based on the decoding weight matrix and the decoding bias vector to obtain a reconstructed data of the target preprocessed data.
[0085] Step S103: Calculating a reconstruction error of each preprocessed data based on each preprocessed data and the reconstructed data of each preprocessed data.
[0086] In the embodiment of the present invention, the calculation formula for the reconstruction error of the target preprocessed data is:
[0087]
[0088] where L i represents the reconstruction error of the target preprocessed data, N represents the number of preprocessed data in the preprocessed dataset, and x i represents the target preprocessed data, represents the reconstructed data of the target preprocessed data.
[0089] Step S104: Determining the preprocessed data with a reconstruction error greater than a preset error threshold in the preprocessed dataset as abnormal data.
[0090] In the embodiment of the present invention, the preset error threshold can be determined by using the percentile p to set the threshold, specifically:
[0091] Calculating the reconstruction error of each target training data in all existing target training datasets Suppose there are N reconstruction errors of target training data, which are respectively denoted as L1, L2,..., L N .
[0092] Sorting these reconstruction errors from smallest to largest to obtain L (1) , L (2) ,..., L (N) .
[0093] The calculation formula for the position of the percentile p is: index = p × N;
[0094] where, if index is an integer, then the preset error threshold T of the p-th percentile is the average value of L (index) and L (index+1) ; if index is not an integer, rounding up to then the preset error threshold T of the p-th percentile is
[0095] When the reconstruction error of a piece of preprocessed data is greater than a preset error threshold, the data is considered abnormal; otherwise, it is considered normal.
[0096] As an alternative implementation, after step S104, the following steps may further be included:
[0097] Determine the neighborhood set of the abnormal data from the preprocessed data set;
[0098] Determine the set of embedding vectors corresponding to the neighborhood set;
[0099] Calculate the average embedding vector of the set of embedding vectors;
[0100] Based on preset parameters, calculate the average embedding vector to obtain Dirichlet distribution data;
[0101] Generate sample data based on the Dirichlet distribution data;
[0102] Use the sample data to replace the abnormal data in the preprocessed data set to obtain a normal data set.
[0103] As an alternative implementation, the method for determining the neighborhood set of the abnormal data from the preprocessed data set may include:
[0104] Convert the format of the preprocessed data set to obtain a tree-structured data set;
[0105] Based on the tree-structured data set, calculate the current distance between each piece of preprocessed data and the abnormal data;
[0106] Determine the target preprocessed data in the preprocessed data whose current distance is less than a preset distance threshold;
[0107] Construct the neighborhood set of the abnormal data based on the target preprocessed data.
[0108] In the embodiments of the present invention, for a given preprocessed data set X = x1, x2, x3, …, x n , the preprocessed data set X is constructed into a tree-structured data set of a KDTree tree structure for fast distance query.
[0109] For the detected abnormal data x i , query other data points that are closer to it through the tree-structured data set.
[0110] Specifically, by setting a preset distance threshold ∈, find the data point x that satisfies d(x i , x j ) ≤ ∈ j, these points form the neighborhood set \(N(x\) i i ) of \(x\).)
[0111] Among them, \(d(x\) i , \(x\) j ) represents the distance between \(x\) i and \(x\) j , and a metric standard such as the Euclidean distance can be used to calculate it.
[0112] For two points \(x=(x_1,x_2,\cdots,x\) d ) and \(y=(y_1,y_2,\cdots,y\) d ) in a \(d -\)dimensional space, the Euclidean distance is defined as:
[0113]
[0114] For the neighborhood set \(N(x\) i ) of the detected abnormal data \(x\), the corresponding set of embedding vectors is \(E = \{e\) i | \(x\) j \in N(x\) j ) \}, where \(e\) i is the embedding vector of the data point \(x\) j . j
[0115] Calculate the convex hull of this set of embedding vectors. For the abnormal data \(x\) i , if the convex hull vertices exist, calculate its average embedding vector
[0116]
[0117] Among them, \(V\) is the set of convex hull vertices, and ensure that the elements in the average embedding vector are all greater than a minimum value (e.g., \(1e - 10\)) to avoid numerical problems.
[0118] If calculating the convex hull fails, it may be because the data points do not meet the conditions for the existence of the convex hull. In this case, directly use the embedding vectors of all neighborhood sets as the basis to calculate the average embedding vector.
[0119] Add a preset parameter \(\alpha\) (usually a positive number) to the average embedding vector as the parameter of the Dirichlet distribution data.
[0120] The Dirichlet distribution data is a multivariate probability distribution. For the parameter vector \(\alpha=(\alpha_1,\alpha_2,\cdots,\alpha\) n ), its probability density function is:
[0121]
[0122] Among them, \(x\) i \geq0,\) B(α) is a normalization constant, and Γ(·) is the gamma function.
[0123] For each dimension l = 1, 2, …, d, we generate a random variable Y l according to the gamma distribution Gamma(a l ). The probability density function of the gamma distribution is:
[0124]
[0125] After generating the gamma-distributed random variables Y1, Y2, …, Y d corresponding to all dimensions, we calculate each dimension of the new sample vector x new :
[0126]
[0127] In this way, we obtain a new sample vector x new = (x new,1 , x new,2 , …, x new,d ), and this sample data can be used to fill in the positions of the abnormal data.
[0128] The present invention can accurately identify the abnormal data in the input initial data set, ensuring that the system for making recommendations to users based on the initial data set can work more accurately, improving the accuracy of the recommendation information generated by the recommendation system, and thus also enhancing the user experience of the recommendation system.
[0129] Exemplary device
[0130] After introducing the method of the exemplary embodiment of the present invention, next, with reference to Figure 2 we will describe a data poisoning attack detection device of the exemplary embodiment of the present invention, which includes:
[0131] A preprocessing unit 201, configured to preprocess the initial data in the obtained initial data set to obtain a preprocessed data set;
[0132] A reconstruction unit 202, configured to reconstruct each preprocessed data in the preprocessed data set respectively to obtain the reconstructed data of each preprocessed data;
[0133] A calculation unit 203, configured to calculate the reconstruction error of each preprocessed data based on each preprocessed data and the reconstructed data of each preprocessed data;
[0134] A determination unit 204, configured to determine preprocessing data in the preprocessing data set with a reconstruction error greater than a preset error threshold as abnormal data.
[0135] As an alternative implementation manner, the preprocessing unit 201 preprocesses the initial data in the obtained initial data set to obtain the preprocessing data set in the following specific manner:
[0136] Perform redundancy processing on the initial data in the obtained initial data set to obtain a target data set; wherein, the target data set includes a plurality of target data;
[0137] Perform normalization processing on the target data in the target data set to obtain the preprocessing data set.
[0138] As an alternative implementation manner, the reconstruction unit 202 reconstructs a target preprocessing data in the preprocessing data set to obtain the reconstruction data of the target preprocessing data in the following specific manner:
[0139] Obtain the encoding weight matrix and encoding bias vector of a pre-constructed encoder, and the decoding weight matrix and decoding bias vector of a pre-constructed decoder;
[0140] Encode a target preprocessing data in the preprocessing data set based on the encoding weight matrix and the encoding bias vector to obtain encoded data;
[0141] Reconstruct the encoded data based on the decoding weight matrix and the decoding bias vector to obtain the reconstruction data of the target preprocessing data.
[0142] In the embodiments of the present invention, the calculation formula for the reconstruction error of the target preprocessing data is:
[0143]
[0144] Wherein, L i represents the reconstruction error of the target preprocessing data, N represents the number of preprocessing data in the preprocessing data set, x i represents the target preprocessing data, represents the reconstruction data of the target preprocessing data.
[0145] As an alternative implementation manner, the determination unit 204 is further configured to:
[0146] After determining preprocessing data in the preprocessing data set with a reconstruction error greater than a preset error threshold as abnormal data, determine a neighborhood set of the abnormal data from the preprocessing data set;
[0147] Determine an embedding vector set corresponding to the neighborhood set;
[0148] Calculate the average embedding vector of the set of embedding vectors;
[0149] Based on preset parameters, calculate the average embedding vector to obtain Dirichlet distribution data;
[0150] Generate sample data based on the Dirichlet distribution data;
[0151] Use the sample data to replace the abnormal data in the preprocessed data set to obtain a normal data set.
[0152] As an alternative implementation, the manner in which the determination unit 204 determines the neighborhood set of the abnormal data from the preprocessed data set may specifically be:
[0153] Convert the format of the preprocessed data set to obtain a tree-shaped data structure data set;
[0154] Based on the tree-shaped data structure data set, calculate the current distance between each piece of the preprocessed data and the abnormal data;
[0155] Determine the target preprocessed data in the preprocessed data whose current distance is less than a preset distance threshold;
[0156] Construct the neighborhood set of the abnormal data based on the target preprocessed data.
[0157] The present invention can accurately determine the abnormal data in the input initial data set, ensuring that the system for recommending to users based on the initial data set can work more accurately, improving the accuracy of the recommendation information generated by the recommendation system, and thus also improving the user experience of the recommendation system.
[0158] Exemplary medium
[0159] After introducing the methods and devices of the exemplary embodiments of the present invention, next, refer to Figure 3 Describe the computer-readable storage medium of the exemplary embodiments of the present invention. Please refer to Figure 3, which shows that the computer-readable storage medium is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will implement the steps recorded in the above method embodiments. For example, preprocess the initial data in the obtained initial data set to obtain a preprocessed data set; reconstruct each preprocessed data in the preprocessed data set respectively to obtain the reconstructed data of each preprocessed data; calculate the reconstruction error of each preprocessed data based on each preprocessed data and the reconstructed data of each preprocessed data; determine the preprocessed data with a reconstruction error greater than a preset error threshold in the preprocessed data set as abnormal data; the specific implementation manners of each step will not be repeated here.
[0160] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated one by one here.
[0161] Exemplary computing device
[0162] After introducing the methods, devices, and media of the exemplary embodiments of the present invention, next, refer to Figure 4 a computing device for detecting data poisoning attacks according to the exemplary embodiments of the present invention.
[0163] Figure 4 The block diagram of an exemplary computing device 40 suitable for implementing the embodiments of the present invention is shown. The computing device 40 may be a computer system or a server. Figure 4 The shown computing device 40 is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0164] As Figure 4 shown, the components of the computing device 40 may include, but are not limited to: one or more processors or processing units 401, a system memory 402, and a bus 403 connecting different system components (including the system memory 402 and the processing unit 401).
[0165] The computing device 40 typically includes a variety of computer system-readable media. These media can be any available media accessible by the computing device 40, including volatile and non-volatile media, removable and non-removable media.
[0166] System memory 402 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 4021 and / or cache memory 4022. Computing device 40 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, ROM 4023 may be used to read and write to non-removable, non-volatile magnetic media ( Figure 4 not shown in the figure, commonly referred to as a "hard disk drive"). Although not shown in Figure 4 the figure, a disk drive for reading and writing to removable non-volatile disks (such as a "floppy disk"), and an optical disk drive for reading and writing to removable non-volatile optical disks (such as a CD-ROM, DVD-ROM or other optical media) may be provided. In these cases, each drive may be connected to bus 403 through one or more data media interfaces. System memory 402 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the embodiments of the present invention.
[0167] A program / utilities 4025 having a set (at least one) of program modules 4024 may be stored, for example, in system memory 402, and such program modules 4024 include but are not limited to: an operating system, one or more application programs, other program modules, and program data, and an implementation of a network environment may be included in each or some combination of these examples. Program modules 4024 generally perform the functions and / or methods in the embodiments described in the present invention.
[0168] Computing device 40 may also communicate with one or more external devices 404 (such as a keyboard, pointing device, display, etc.). Such communication may be through an input / output (I / O) interface 405. Also, computing device 40 may further communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN) and / or a public network, such as the Internet) through a network adapter 406. As Figure 4 shown, network adapter 406 communicates with other modules (such as processing unit 401, etc.) of computing device 40 through bus 403. It should be understood that although Figure 4 not shown in the figure, other hardware and / or software modules may be used in conjunction with computing device 40.
[0169] The processing unit 401 executes various functional applications and data processing by running programs stored in the system memory 402. For example, it preprocesses the initial data in the obtained initial dataset to obtain a preprocessed dataset; reconstructs each preprocessed data in the preprocessed dataset to obtain the reconstructed data of each preprocessed data; calculates the reconstruction error of each preprocessed data based on each preprocessed data and the reconstructed data of each preprocessed data; and determines the preprocessed data with a reconstruction error greater than the preset error threshold in the preprocessed dataset as abnormal data. The specific implementation manners of each step will not be repeated here. It should be noted that although several units / modules or sub-units / sub-modules of the data poisoning attack detection device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present invention, the features and functions of the two or more units / modules described above can be embodied in one unit / modules. Conversely, the features and functions of one unit / modules described above can be further divided and embodied by multiple units / modules.
[0170] In the description of the present invention, it should be noted that the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0171] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, device, and unit can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0172] In several embodiments provided by the present invention, it should be understood that the disclosed system, device, and method can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some communication interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.
[0173] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0174] In addition, in each embodiment of the present invention, each functional unit may be integrated into a processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit.
[0175] If the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0176] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, and are not intended to limit them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or make equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
[0177] In addition, although the operations of the method of the present invention are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.
[0178] In an exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
Claims
1. A method for detecting data poisoning attacks, characterized in that The method includes: Preprocessing the initial data in the obtained initial dataset to obtain a preprocessed dataset; Respectively reconstructing each preprocessed data in the preprocessed dataset to obtain the reconstructed data of each preprocessed data; Based on each preprocessed data and the reconstructed data of each preprocessed data, calculating the reconstruction error of each preprocessed data; Determining the preprocessed data with a reconstruction error greater than a preset error threshold in the preprocessed dataset as abnormal data; And, after determining the preprocessed data with a reconstruction error greater than a preset error threshold in the preprocessed dataset as abnormal data, the method further includes: Determining the neighborhood set of the abnormal data from the preprocessed dataset; Determining the corresponding embedding vector set of the neighborhood set; Calculating the average embedding vector of the embedding vector set; Calculating based on the preset parameters for the average embedding vector to obtain Dirichlet distribution data; Generating sample data based on the Dirichlet distribution data; Using the sample data to replace the abnormal data in the preprocessed dataset to obtain a normal dataset; Wherein, determining the neighborhood set of the abnormal data from the preprocessed dataset includes: Converting the format of the preprocessed dataset to obtain a tree - shaped data structure dataset; Based on the tree - shaped data structure dataset, calculating the current distance between each preprocessed data and the abnormal data; Determining the target preprocessed data with a current distance less than a preset distance threshold in the preprocessed data; Constructing the neighborhood set of the abnormal data based on the target preprocessed data; And, calculating the average embedding vector of the embedding vector set includes: Calculating the convex hull of the embedding vector set; For the abnormal data, if the convex hull vertices exist, calculating the average embedding vector of the embedding vector set, and the calculation formula of the average embedding vector is: Among them, represents the average embedding vector, V is the set of convex hull vertices, and all elements in the average embedding vector are greater than the minimum value 1e-10; If calculating the convex hull fails, then using the embedding vectors of all neighborhood sets as the basis to calculate the average embedding vector.
2. The data poisoning attack detection method according to claim 1, wherein, The preprocessing the initial data in the obtained initial dataset to obtain a preprocessed dataset includes: Performing redundancy processing on the initial data in the obtained initial dataset to obtain a target dataset; wherein, the target dataset includes multiple target data; Performing normalization processing on the target data in the target dataset to obtain a preprocessed dataset.
3. The data poisoning attack detection method according to claim 1, characterized in that The reconstructing a target preprocessed data in the preprocessed dataset to obtain the reconstructed data of the target preprocessed data includes: Obtaining the encoding weight matrix and encoding bias vector of a pre - constructed encoder, and the decoding weight matrix and decoding bias vector of a pre - constructed decoder; Encoding a target preprocessed data in the preprocessed dataset based on the encoding weight matrix and the encoding bias vector to obtain encoded data; Reconstructing the encoded data based on the decoding weight matrix and the decoding bias vector to obtain the reconstructed data of the target preprocessed data.
4. The data poisoning attack detection method according to claim 3, characterized in that, The calculation formula of the reconstruction error of the target preprocessed data is: Among them, L i represents the reconstruction error of the target preprocessed data, N represents the number of preprocessed data in the preprocessed dataset, and x i represents the target preprocessed data, represents the reconstructed data of the target preprocessed data.
5. A data poisoning attack detection device, characterized in that, The device includes: A preprocessing unit for preprocessing the initial data in the acquired initial dataset to obtain a preprocessed dataset; A reconstruction unit for reconstructing each preprocessed data in the preprocessed dataset respectively to obtain the reconstructed data of each preprocessed data; A calculation unit for calculating the reconstruction error of each preprocessed data based on each preprocessed data and the reconstructed data of each preprocessed data; A determination unit for determining the preprocessed data with a reconstruction error greater than a preset error threshold in the preprocessed dataset as abnormal data; And, the determination unit is further configured to: After determining the preprocessed data with a reconstruction error greater than a preset error threshold in the preprocessed dataset as abnormal data, determine a neighborhood set of the abnormal data from the preprocessed dataset; Determine the corresponding embedding vector set of the neighborhood set; Calculate the average embedding vector of the embedding vector set; Calculate based on a preset parameter the average embedding vector to obtain Dirichlet distribution data; Generate sample data based on the Dirichlet distribution data; Use the sample data to replace the abnormal data in the preprocessed dataset to obtain a normal dataset; Wherein, the manner in which the determination unit determines the neighborhood set of the abnormal data from the preprocessed dataset is specifically: Convert the format of the preprocessed dataset to obtain a tree-structured dataset; Based on the tree-structured dataset, calculate the current distance between each preprocessed data and the abnormal data; Determine the target preprocessed data with a current distance less than a preset distance threshold in the preprocessed data; Construct the neighborhood set of the abnormal data based on the target preprocessed data; And, the manner in which the determination unit calculates the average embedding vector of the embedding vector set is specifically: calculate the convex hull of the embedding vector set; For the abnormal data, if the convex hull vertices exist, calculate the average embedding vector of the embedding vector set, and the calculation formula of the average embedding vector is: Among them, represents the average embedding vector, V is the set of vertices of the convex hull, and all elements in the average embedding vector are greater than the minimum value of 1e-10; If calculating the convex hull fails, then use the embedding vectors of all neighborhood sets as a basis to calculate the average embedding vector.
6. A computing device, the computing device comprising: At least one processor, a memory, and an input / output unit; Wherein, the memory is used for storing a computer program, and the processor is used for calling the computer program stored in the memory to execute the method according to any one of claims 1 to 4.
7. A computer-readable storage medium, which includes instructions that, when running on a computer, cause the computer to execute the method according to any one of claims 1 to 4.
8. A computer program product comprising a computer program, characterized in that, The computer program, when executed by the processor, implements the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Power system network attack detection method and system considering category imbalance
CN116827603A