A neural network training, image recognition method and device

By calculating and removing the opposite part of the anti-sample backpass gradient from the training sample, and updating the neural network parameters, the problem of low accuracy in the identification model output in the prior art is solved, and effective defense of the anti-sample and high accuracy output of the recognition model is achieved.

CN114358272BActive Publication Date: 2025-05-06HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011033839.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-27
Publication Date
2025-05-06
Estimated Expiration
2040-09-27

AI Technical Summary

Technical Problem

During the training process, existing recognition models have low output accuracy due to sample problems, especially the existence of anti-samples will have adverse effects on the training process.

Method used

By obtaining the training sample and its adversarial samples, the backward transmission gradient of both are calculated, and the part of the backward transmission gradient of the adversarial sample that is opposite to the backward transmission gradient of the training sample is removed, the target backward transmission gradient is obtained, and the neural network parameters are updated until the training is completed.

Benefits of technology

This method can weaken the negative effect of adversarial samples on training samples, reduce the adverse effects of adversarial samples on training processes, improve the output accuracy of recognition models, and be able to defend against attacks on adversarial samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114358272B_ABST
    Figure CN114358272B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a neural network training and image recognition method and device, the method comprising: inputting training samples and adversarial samples into the neural network to obtain output data of the neural network; calculating the back propagation gradient of the initial sample and the back propagation gradient of the adversarial sample based on the output data; removing the part of the back propagation gradient of the adversarial sample that is opposite to the back propagation gradient of the initial sample to obtain the target back propagation gradient; the removal process can weaken the negative effect of the adversarial sample on the initial sample, or reduce the performance damage of the adversarial sample to the initial sample, thereby reducing the adverse effects of the adversarial sample on the training process; then, the back propagation gradient of the initial sample and the target back propagation gradient are back-propagated to update the parameters of the neural network; the recognition model output obtained after the iterative update is high in accuracy and can defend against the attack of the adversarial sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of machine learning technology, and in particular to a neural network training and image recognition method and device. Background Art

[0002] Some recognition models can be obtained by training neural networks. The training process generally includes: obtaining some sample data, iteratively updating the parameters in the neural network based on these sample data, and obtaining the recognition model after the update is completed. For example, by training the neural network based on sample images, an image recognition model can be obtained, such as a face recognition model or a license plate recognition model, etc. The face recognition model can determine the identity of a person by recognizing a face image, and the license plate recognition model can determine the license plate number by recognizing a license plate image. Alternatively, the neural network can be trained based on a sample video to obtain a video recognition model, which can be used to recognize information such as faces and license plates in the video, or to track targets. For another example, by training the neural network based on sample corpus, a semantic recognition model can be obtained, and the semantic recognition model can determine the meaning of a sentence by recognizing corpus data. For another example, by training the neural network based on sample speech, a speech recognition model can be obtained, and the speech recognition model can determine the meaning of a sentence by recognizing speech data.

[0003] However, when using recognition models for recognition, such as the above-mentioned face recognition, license plate recognition, tracking target recognition, semantic recognition, speech recognition, etc., the recognition model may give incorrect output due to problems with the samples during the training process, and the output accuracy of the recognition model is low. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a neural network training, image recognition method and device to improve the output accuracy of the recognition model.

[0005] To achieve the above objectives, the present application provides a neural network training method, including:

[0006] Obtaining training samples and adversarial samples of the training samples, wherein the training samples include any one of sample images, sample videos, sample corpora, and sample voices;

[0007] Inputting the training sample and the adversarial sample into a neural network to obtain output data of the neural network;

[0008] Based on the output data, respectively calculate the back propagation gradient of the training sample and the back propagation gradient of the adversarial sample;

[0009] Removing the part of the back-propagated gradient of the adversarial sample that is opposite to the back-propagated gradient of the training sample to obtain a target back-propagated gradient;

[0010] The parameters of the neural network are updated by back-propagating the back-propagated gradient of the training sample and the target back-propagated gradient, and then returning to the step of inputting the training sample and the adversarial sample into the neural network until the training is completed.

[0011] Optionally, removing the portion of the back propagation gradient of the adversarial sample that is opposite to the back propagation gradient of the training sample to obtain the target back propagation gradient includes:

[0012] Among the back-propagated gradients of the adversarial sample, a back-propagated gradient whose similarity with the back-propagated gradient of the training sample does not satisfy a preset similarity condition is selected as the back-propagated gradient to be processed;

[0013] The part of the back-propagated gradient to be processed that is opposite to the back-propagated gradient of the training sample is removed to obtain the target back-propagated gradient.

[0014] Optionally, selecting, from the back-propagation gradients of the adversarial sample, a back-propagation gradient whose similarity with the back-propagation gradient of the training sample does not satisfy a preset similarity condition as the back-propagation gradient to be processed, includes:

[0015] For each adversarial sample's back-propagation gradient, calculate the similarity between the back-propagation gradient of the adversarial sample and the back-propagation gradient of the training sample;

[0016] Determining whether the similarity satisfies a preset similarity condition;

[0017] If not, the back propagation gradient of the adversarial sample is determined as the candidate processing back propagation gradient;

[0018] The back propagation gradient to be processed is selected from the determined candidate back propagation gradients according to the set selection probability in the current parameter update process; the selection probability increases as the number of parameter updates of the neural network increases.

[0019] Optionally, the method further includes:

[0020] If the similarity satisfies a preset similarity condition, the back propagation gradient of the adversarial sample is determined as the target back propagation gradient;

[0021] And / or, determining a candidate processed back-propagation gradient that is not selected as the back-propagation gradient to be processed as a target back-propagation gradient.

[0022] Optionally, the preset similarity condition is: an angle between the back-propagation gradient of the adversarial sample and the back-propagation gradient of the training sample is not greater than 90 degrees;

[0023] The removing the part of the back-propagated gradient to be processed that is opposite to the back-propagated gradient of the training sample to obtain the target back-propagated gradient includes:

[0024] Calculate the projection of the back-propagated gradient to be processed in a target direction to obtain a projected gradient, wherein the target direction is the direction of the back-propagated gradient of the training sample;

[0025] In the back-propagated gradient to be processed, the projected gradient is removed to obtain the target back-propagated gradient.

[0026] Optionally, before removing the part of the back propagation gradient of the adversarial sample that is opposite to the back propagation gradient of the training sample to obtain the target back propagation gradient, the method further includes:

[0027] Based on the overall back propagation gradient of the training samples in the previous parameter update process and the back propagation gradient of all training samples in the current parameter update process, the overall back propagation gradient of the training samples in the current parameter update process is calculated;

[0028] The removing the part of the back propagation gradient of the adversarial sample that is opposite to the back propagation gradient of the training sample to obtain the target back propagation gradient includes:

[0029] The back-propagated gradient of the adversarial sample is removed in the opposite direction to the overall back-propagated gradient of the training sample in the current parameter update process to obtain the target back-propagated gradient.

[0030] Optionally, the calculating the overall back propagation gradient of the training samples in the current parameter updating process based on the overall back propagation gradient of the training samples in the previous parameter updating process and the back propagation gradient of all training samples in the current parameter updating process includes:

[0031] Calculate the product of the overall back propagation gradient of the training sample in the last parameter update process and the hyperparameter in the moving average algorithm as the first value;

[0032] Calculate the sum of the back-propagated gradients of all training samples in this parameter update process as the second value;

[0033] Calculate the sum of the first value and the second value as a third value;

[0034] Calculate the influence coefficient of the overall back propagation gradient of the training sample in the previous parameter updating process on the overall back propagation gradient of the training sample in the current parameter updating process as the fourth value;

[0035] Calculate the product of the fourth value and the hyperparameter in the moving average algorithm as a fifth value;

[0036] Calculating the sum of the fifth value and a preset parameter as a sixth value;

[0037] The ratio of the third value to the sixth value is calculated as the overall back propagation gradient of the training sample in this parameter updating process.

[0038] Optionally, if the training sample is a sample image, the method further includes:

[0039] Obtain an image to be recognized;

[0040] Inputting the image to be recognized into the image recognition model trained based on the neural network to obtain an image recognition result;

[0041] If the training sample is a sample video, the method further includes:

[0042] Get the video to be identified;

[0043] Inputting the video to be recognized into a video recognition model trained based on the neural network to obtain a video recognition result;

[0044] If the training sample is a sample corpus, the method further includes:

[0045] Obtain the corpus to be recognized;

[0046] Inputting the corpus to be recognized into a semantic recognition model obtained by training the neural network to obtain a semantic recognition result;

[0047] If the training sample is a sample speech, the method further includes:

[0048] Get the speech to be recognized;

[0049] The speech to be recognized is input into a speech recognition model trained based on the neural network to obtain a speech recognition result.

[0050] To achieve the above object, the present application also provides an image recognition method, including:

[0051] Obtain an image to be recognized;

[0052] Inputting the image to be recognized into a pre-trained image recognition model to obtain an image recognition result;

[0053] The process of training the image recognition model includes:

[0054] Obtaining a sample image and an adversarial sample of the sample image;

[0055] Inputting the sample image and the adversarial sample into a neural network to obtain output data of the neural network;

[0056] Based on the output data, respectively calculate the back-propagation gradient of the sample image and the back-propagation gradient of the adversarial sample;

[0057] Removing the part of the back-propagated gradient of the adversarial sample that is opposite to the back-propagated gradient of the sample image to obtain a target back-propagated gradient;

[0058] The back propagation gradient of the sample image and the target back propagation gradient are back-propagated to update the parameters of the neural network, and then the step of inputting the sample image and the adversarial sample into the neural network is returned to execute until the training is completed, thereby obtaining an image recognition model.

[0059] To achieve the above objectives, the present application also provides a neural network training device, including:

[0060] A first acquisition module, used to acquire training samples and adversarial samples of the training samples, wherein the training samples include any one of sample images, sample videos, sample corpora, and sample voices;

[0061] A first input module, used to input the training sample and the adversarial sample into the neural network to obtain output data of the neural network;

[0062] A first calculation module, used for respectively calculating the back propagation gradient of the training sample and the back propagation gradient of the adversarial sample based on the output data;

[0063] A removal module, used to remove the part of the back-propagation gradient of the adversarial sample that is opposite to the back-propagation gradient of the training sample, to obtain a target back-propagation gradient;

[0064] An update module is used to update the parameters of the neural network by back-propagating the back-propagation gradient of the training sample and the target back-propagation gradient, and then re-trigger the first input module until the training is completed.

[0065] Optionally, the removal module includes:

[0066] A selection submodule, configured to select, from the back-propagation gradients of the adversarial sample, a back-propagation gradient whose similarity with the back-propagation gradient of the training sample does not satisfy a preset similarity condition as the back-propagation gradient to be processed;

[0067] The removal submodule is used to remove the part of the back-propagation gradient to be processed that is opposite to the back-propagation gradient of the training sample to obtain the target back-propagation gradient.

[0068] Optionally, the selection submodule is specifically used for:

[0069] For each adversarial sample's back-propagation gradient, calculate the similarity between the back-propagation gradient of the adversarial sample and the back-propagation gradient of the training sample;

[0070] Determining whether the similarity satisfies a preset similarity condition;

[0071] If not, the back propagation gradient of the adversarial sample is determined as the candidate processing back propagation gradient;

[0072] The back propagation gradient to be processed is selected from the determined candidate back propagation gradients according to the set selection probability in the current parameter update process; the selection probability increases as the number of parameter updates of the neural network increases.

[0073] Optionally, the device further includes: a first determining module and / or a second determining module, wherein:

[0074] A first determination module, configured to determine the back propagation gradient of the adversarial sample as a target back propagation gradient when the similarity satisfies a preset similarity condition;

[0075] The second determination module is used to determine the candidate processed back-propagation gradient that has not been selected as the back-propagation gradient to be processed as the target back-propagation gradient.

[0076] Optionally, the preset similarity condition is: an angle between the back-propagation gradient of the adversarial sample and the back-propagation gradient of the training sample is not greater than 90 degrees;

[0077] The removal submodule is specifically used for:

[0078] Calculate the projection of the back-propagated gradient to be processed in a target direction to obtain a projected gradient, wherein the target direction is the direction of the back-propagated gradient of the training sample;

[0079] In the back-propagated gradient to be processed, the projected gradient is removed to obtain the target back-propagated gradient.

[0080] Optionally, the device further comprises:

[0081] A second calculation module is used to calculate the overall back propagation gradient of the training samples in the current parameter update process based on the overall back propagation gradient of the training samples in the previous parameter update process and the back propagation gradient of all training samples in the current parameter update process;

[0082] The removal module is specifically used to remove the part of the back-propagation gradient of the adversarial sample that is opposite to the overall back-propagation gradient of the training sample in the current parameter updating process to obtain the target back-propagation gradient.

[0083] Optionally, the second computing module is specifically configured to:

[0084] Calculate the product of the overall back propagation gradient of the training sample in the last parameter update process and the hyperparameter in the moving average algorithm as the first value;

[0085] Calculate the sum of the back-propagated gradients of all training samples in this parameter update process as the second value;

[0086] Calculate the sum of the first value and the second value as a third value;

[0087] Calculate the influence coefficient of the overall back propagation gradient of the training sample in the previous parameter updating process on the overall back propagation gradient of the training sample in the current parameter updating process as the fourth value;

[0088] Calculate the product of the fourth value and the hyperparameter in the moving average algorithm as a fifth value;

[0089] Calculating the sum of the fifth value and a preset parameter as a sixth value;

[0090] The ratio of the third value to the sixth value is calculated as the overall back propagation gradient of the training sample in this parameter updating process.

[0091] Optionally, the device further comprises:

[0092] A first recognition module is used to obtain an image to be recognized when the training sample is a sample image and an image recognition model is obtained through training; the image to be recognized is input into the image recognition model to obtain an image recognition result;

[0093] The second recognition module is used to obtain a video to be recognized when the training sample is a sample video and a video recognition model is obtained through training; the video to be recognized is input into the video recognition model to obtain a video recognition result;

[0094] The third recognition module is used to obtain the corpus to be recognized when the training sample is a sample corpus and the semantic recognition model is obtained through training; input the corpus to be recognized into the semantic recognition model to obtain a semantic recognition result;

[0095] The fourth recognition module is used to obtain the speech to be recognized when the above-mentioned training sample is a sample speech and the speech recognition model is trained; input the speech to be recognized into the speech recognition model to obtain the speech recognition result.

[0096] To achieve the above object, the present application also provides an image recognition device, including:

[0097] A second acquisition module is used to acquire an image to be identified;

[0098] A second input module is used to input the image to be recognized into a pre-trained image recognition model to obtain an image recognition result;

[0099] The device also includes a training module, which is used to train and obtain the image recognition model;

[0100] The training module comprises:

[0101] An acquisition submodule, used to acquire a sample image and an adversarial sample of the sample image;

[0102] An input submodule, used for inputting the sample image and the adversarial sample into a neural network to obtain output data of the neural network;

[0103] A calculation submodule, used for respectively calculating the back-propagation gradient of the sample image and the back-propagation gradient of the adversarial sample based on the output data;

[0104] A removal submodule, used to remove the part of the back-propagated gradient of the adversarial sample that is opposite to the back-propagated gradient of the sample image, to obtain a target back-propagated gradient;

[0105] The update submodule is used to update the parameters of the neural network by back-propagating the back-propagation gradient of the sample image and the target back-propagation gradient, and then return to trigger the input submodule until the training is completed to obtain an image recognition model.

[0106] To achieve the above object, an embodiment of the present application further provides an electronic device, including a processor and a memory;

[0107] Memory, used to store computer programs;

[0108] The processor is used to implement any of the above-mentioned neural network training or any of the image recognition methods when executing the program stored in the memory.

[0109] Using the embodiment shown in the present application, the training samples and adversarial samples are input into the neural network to obtain the output data of the neural network; the back propagation gradient of the training samples and the back propagation gradient of the adversarial samples are calculated based on the output data; the part of the back propagation gradient of the adversarial sample that is opposite to the back propagation gradient of the training sample is removed to obtain the target back propagation gradient; this removal process can weaken the negative effect of the adversarial sample on the training sample, or reduce the performance damage of the adversarial sample on the training sample, thereby reducing the adverse effects of the adversarial sample on the training process; then, the back propagation gradient of the training sample and the target back propagation gradient are back propagated to update the parameters of the neural network; the recognition model output obtained after the iterative update is completed has a high accuracy rate and can defend against the attack of the adversarial sample. It can be seen that this scheme can not only weaken the negative effect of the adversarial sample on the training sample, but also maintain the defense capability of the recognition model against the adversarial sample, thereby improving the output accuracy rate of the recognition model.

[0110] Of course, implementing any product or method of the present application does not necessarily require achieving all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0111] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0112] Figure 1 A schematic diagram of a first flow chart of a neural network training method provided in an embodiment of the present application;

[0113] Figure 2 A vector decomposition schematic diagram provided for an embodiment of the present application;

[0114] Figure 3 A second flow chart of the neural network training method provided in the embodiment of the present application;

[0115] Figure 4 A flowchart of an image recognition method provided in an embodiment of the present application;

[0116] Figure 5 A third flow chart of the neural network training method provided in the embodiment of the present application;

[0117] Figure 6 A schematic diagram of the structure of a neural network training device provided in an embodiment of the present application;

[0118] Figure 7 A schematic diagram of the structure of an image recognition device provided in an embodiment of the present application;

[0119] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0120] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0121] In some related solutions, when using recognition models to identify adversarial examples, the recognition models usually give incorrect outputs. Adversarial examples contain adversarial perturbations, which can be perturbations that are imperceptible to the human eye. They do not affect human judgment but will cause the recognition model to give an incorrect output with high confidence. For example, when using a face recognition model to identify adversarial examples of face image type, the identity of the person cannot be accurately determined; when using a license plate recognition model to identify adversarial examples of license plate image type, the license plate number cannot be accurately determined; when using a semantic recognition model to identify adversarial examples of corpus type, the meaning of the sentence cannot be accurately determined; when using a speech recognition model to identify adversarial examples of speech type, the meaning of the sentence cannot be accurately determined.

[0122] In order to improve the output accuracy of the recognition model, some related solutions generate adversarial samples of training samples during the training of neural networks, and train the neural networks based on the training samples and adversarial samples. However, in this solution, the existence of adversarial samples will have an adverse effect on the training process.

[0123] In order to achieve the above-mentioned purpose, the embodiment of the present application provides a neural network training and image recognition method and device, which can be applied to various electronic devices without specific limitation. The neural network training method is first described in detail below.

[0124] Figure 1 A first flow chart of a neural network training method provided in an embodiment of the present application includes:

[0125] S101: Obtain training samples and adversarial samples of the training samples, where the training samples include any one of sample images, sample videos, sample corpora, and sample voices.

[0126] If the training sample is a sample image, the embodiment of the present application can be used to train an image recognition model; further subdivided, if the sample image contains a face area, the embodiment of the present application can be used to train a face recognition model; if the sample image contains a license plate area, the embodiment of the present application can be used to train a license plate recognition model, and so on.

[0127] If the training sample is a sample video, the embodiment of the present application can be used to train a video recognition model; further subdivided, if the sample video contains a face area, the embodiment of the present application can be used to train a face recognition model; if the sample video contains a license plate area, the embodiment of the present application can be used to train a license plate recognition model; if the sample video contains a target to be tracked, the embodiment of the present application can be used to train a target tracking model, and so on.

[0128] If the training sample is a sample corpus, the semantic recognition model can be trained by applying the embodiments of the present application. Corpus is language material, which can be understood as data including texts such as sentences and words. If the training sample is a sample speech, the speech recognition model can be trained by applying the embodiments of the present application.

[0129] The training samples mentioned in the embodiments of the present application can be understood as normal samples without adding adversarial disturbances, or non-adversarial samples. For example, after obtaining the training samples, adversarial samples can be generated based on the training samples, and the generation method is not limited.

[0130] For example, suppose the training sample is x = [x1, x2, ..., x n ], the adversarial sample generated based on the training sample is Here it is assumed that the training samples correspond one to one with the adversarial samples.

[0131] S102: Input the training samples and adversarial samples into the neural network to obtain output data of the neural network.

[0132] The input and output in S102 can be understood as forward propagation.

[0133] S103: Based on the output data, the back propagation gradient of the training sample and the back propagation gradient of the adversarial sample are calculated respectively.

[0134] The back propagation gradient of the sample can be solved based on the back propagation algorithm, and the specific solution process is not limited.

[0135] S104: Remove the part of the back-propagated gradient of the adversarial sample that is opposite to the back-propagated gradient of the training sample to obtain the target back-propagated gradient.

[0136] The applicant discovered that the existence of adversarial samples can have an adverse effect on the training process because the existence of adversarial samples can have a negative effect on the training samples, or in other words, the adversarial samples damage the performance of the training samples. Based on this discovery, in the embodiments of the present application, removing the part of the back-propagation gradient of the adversarial sample that is opposite to the back-propagation gradient of the training sample can weaken the negative effect of the adversarial sample on the training sample, or reduce the performance damage to the training sample.

[0137] For example, refer to Figure 2 As shown, the back propagation gradient of the adversarial sample is represented as a vector Represent the back-propagated gradient of the training sample as a vector Will Perform orthogonal decomposition to obtain vector and and In the opposite direction, remove In get That is, it represents the target back-propagation gradient.

[0138] In one embodiment, S104 may include: among the back propagation gradients of the adversarial sample, selecting a back propagation gradient whose similarity with the back propagation gradient of the training sample does not meet a preset similarity condition as the back propagation gradient to be processed; removing the part of the back propagation gradient to be processed that is in the opposite direction to the back propagation gradient of the training sample to obtain a target back propagation gradient.

[0139] If the similarity between the back-propagated gradient of the adversarial sample and the back-propagated gradient of the training sample meets the preset conditions, it means that the adversarial sample will not have a negative effect on the training sample. In this case, the back-propagated gradient of the adversarial sample will not be removed.

[0140] It can be seen that in this embodiment, firstly, only the back-propagated gradients of some adversarial samples are removed, which reduces the amount of calculation and improves the processing efficiency compared to removing the back-propagated gradients of all adversarial samples. Secondly, only the back-propagated gradients of adversarial samples that have a negative effect on training samples are removed, and the removal efficiency is higher.

[0141] In one case, for each back propagation gradient of an adversarial sample, the similarity between the back propagation gradient of the adversarial sample and the back propagation gradient of the training sample can be calculated; and it can be determined whether the similarity satisfies a preset similarity condition; if not, the back propagation gradient of the adversarial sample is determined as the back propagation gradient to be processed.

[0142] Alternatively, in another case, for each back propagation gradient of an adversarial sample, the similarity between the back propagation gradient of the adversarial sample and the back propagation gradient of the training sample can be calculated; a determination can be made as to whether the similarity satisfies a preset similarity condition; if not, the back propagation gradient of the adversarial sample is determined as a candidate for processing back propagation gradient; and a back propagation gradient to be processed is selected from the determined candidate for processing back propagation gradients.

[0143] In this case, a part of the back-propagated gradients of the adversarial samples that have a negative effect on the training samples is selected as the back-propagated gradients to be processed, that is, only the back-propagated gradients of this part of the adversarial samples are removed.

[0144] Some of the back-propagated gradients to be processed may be selected from the determined candidate back-propagated gradients, for example, randomly selected, or randomly selected according to a set probability, or selected in a specified order, etc. The specific selection method is not limited.

[0145] In one embodiment, selecting the back propagation gradient to be processed from the determined candidate back propagation gradients may include: selecting the back propagation gradient to be processed from the determined candidate back propagation gradients according to the set selection probability in the current parameter update process; the selection probability increases as the number of parameter updates of the neural network increases.

[0146] The training process of the neural network can be understood as an iterative update process of parameters. One execution of S102-S106 is one iteration, or one parameter update. When the conditions for the termination of the iteration are met, for example, when the preset number of iterations is reached, or when the loss convergence condition is met, the training is completed, and the neural network after the training is completed is the recognition model. The conditions for the termination of the iteration are not limited.

[0147] In this implementation, different selection probabilities can be set for each parameter update process. For example, the following formula can be used to set the selection probability in each parameter update process:

[0148]

[0149] Among them, p represents the selection probability, γ represents the preset parameter, T represents the total number of parameter updates during the neural network training process, t represents the parameter update this time is the tth parameter update, and exp represents the exponential function with the natural constant e as the base. As t increases, p also increases.

[0150] In this implementation, at the beginning of neural network training, a small number of back propagated gradients of adversarial samples are selected for removal processing; as the training progresses, the back propagated gradient directions of training samples become more and more accurate, and more and more back propagated gradients of adversarial samples are selected for removal processing; thus, on the one hand, compared with removing the back propagated gradients of a large number of adversarial samples throughout the training process, this implementation reduces the amount of calculation; on the other hand, compared with removing the back propagated gradients of a small number of adversarial samples throughout the training process, the training accuracy is improved; it can be seen that this implementation reasonably plans the amount of calculation and improves the training accuracy.

[0151] In one implementation, if the similarity between the back propagation gradient of the adversarial sample and the back propagation gradient of the training sample meets a preset similarity condition, the back propagation gradient of the adversarial sample can be directly determined as the target back propagation gradient and directly participate in the subsequent operation of S105. In this way, the richness of the sample can be improved.

[0152] In one implementation, the candidate processed back-propagation gradient that is not selected as the back-propagation gradient to be processed can be determined as the target back-propagation gradient, that is, the candidate processed back-propagation gradient that is not selected as the back-propagation gradient to be processed directly participates in the subsequent operation of S105. In this way, the richness of the samples can be improved.

[0153] Alternatively, in another implementation, the candidate processed back-propagation gradient that is not selected as the back-propagation gradient to be processed may not be determined as the target back-propagation gradient, that is, the candidate processed back-propagation gradient that is not selected as the back-propagation gradient to be processed may not participate in subsequent operations in this parameter update process. In this way, the negative effect of adversarial samples on training samples can be reduced.

[0154] In the above implementation, it is necessary to calculate the similarity between the back propagation gradient of the adversarial sample and the back propagation gradient of the training sample. For example, there are many ways to calculate the similarity, such as calculating cosine similarity, calculating distance similarity, etc., which are not specifically limited. Correspondingly, the preset similarity condition can be an angle condition, a distance condition, etc., and the specific similarity condition can be set according to the actual situation.

[0155] In one implementation, the angle between the back propagation gradient of the adversarial sample and the back propagation gradient of the training sample can be calculated; and it is determined whether the angle is greater than 90 degrees; if it is, indicating that the preset similarity condition is not met, then the back propagation gradient of the adversarial sample can be determined as the back propagation gradient to be processed, or the back propagation gradient of the adversarial sample can be determined as a candidate processing back propagation gradient, and then the back propagation gradient to be processed is selected from the candidate processing back propagation gradients.

[0156] In this implementation, S104 may include: calculating the projection of the back-propagation gradient to be processed in the target direction to obtain the projected gradient, where the target direction is the direction of the back-propagation gradient of the training sample; in the back-propagation gradient to be processed, removing the projected gradient to obtain the target back-propagation gradient.

[0157] For example, refer to Figure 2 As shown, the back propagation gradient of the adversarial sample is represented as a vector Represent the back-propagated gradient of the training sample as a vector Will Perform orthogonal decomposition to obtain vector and and In the opposite direction, remove In get That is, it represents the target back-propagation gradient.

[0158] In the above content, the "back propagation gradient of the training sample" in "removing the part of the back propagation gradient of the adversarial sample that is opposite to the back propagation gradient of the training sample" can be: the sum of the back propagation gradients of some training samples, or it can also be the sum of the back propagation gradients of all training samples. Alternatively, in one embodiment, the back propagation gradient of the training sample can be: the overall back propagation gradient of the training sample in this parameter update process. In this embodiment, the overall back propagation gradient of the training sample in this parameter update process can be calculated in the following manner:

[0159] Based on the overall back propagation gradient of the training samples in the previous parameter update process and the back propagation gradient of all training samples in the current parameter update process, the overall back propagation gradient of the training samples in the current parameter update process is calculated.

[0160] In this implementation, S104 includes: removing the part of the back propagation gradient of the adversarial sample that is opposite to the overall back propagation gradient of the training sample in the current parameter update process, and obtaining the target back propagation gradient. In other contents of the embodiments of the present application, if the previous parameter update process is not emphasized separately, the data in the current parameter update process is processed.

[0161] In this implementation, a moving average algorithm may be used to calculate the overall back propagation gradient of the training samples in this parameter update process. The calculation process may include:

[0162] Calculate the product of the overall back propagation gradient of the training sample in the last parameter update process and the hyperparameter in the moving average algorithm as the first value;

[0163] Calculate the sum of the back-propagated gradients of all training samples in this parameter update process as the second value;

[0164] Calculate the sum of the first value and the second value as a third value;

[0165] Calculate the influence coefficient of the overall back propagation gradient of the training sample in the previous parameter updating process on the overall back propagation gradient of the training sample in the current parameter updating process as the fourth value;

[0166] Calculate the product of the fourth value and the hyperparameter in the moving average algorithm as a fifth value;

[0167] Calculating the sum of the fifth value and a preset parameter as a sixth value;

[0168] The ratio of the third value to the sixth value is calculated as the overall back propagation gradient of the training sample in this parameter updating process.

[0169] The above calculation process can be expressed as the following formula:

[0170]

[0171] Among them, G t represents the overall back-propagation gradient of the training sample during the tth parameter update process, G t-1 represents the overall back propagation gradient of the training sample during the t-1th parameter update process, α represents the hyperparameter in the moving average algorithm, α can be a preset value, and the specific value is not limited; i represents the identifier of the training sample, n represents the number of training samples, represents the back-propagation gradient of the i-th training sample during the t-th parameter update process, represents the sum of the back-propagated gradients of all training samples during the t-th parameter update process, N t-1 N represents the influence coefficient of the overall back propagation gradient of the training sample in the t-1th parameter update process on the overall back propagation gradient of the training sample in the tth parameter update process. t =αN t-1 +1. The 1 in the denominator of the above formula represents the above preset parameter.

[0172] The tth parameter updating process can be understood as this parameter updating process.

[0173] Alternatively, the average of the back propagation gradients of all training samples in the last parameter update process and the back propagation gradients of all training samples in this parameter update process can be calculated as the above-mentioned "back propagation gradient of training samples". Alternatively, the weighted average of the back propagation gradients of all training samples in the last parameter update process and the back propagation gradients of all training samples in this parameter update process can be calculated as the above-mentioned "back propagation gradient of training samples". I will not list them one by one.

[0174] Similarly, the target direction in the above content can be: the direction of the sum of the back-propagation gradients of some training samples, or can also be the direction of the sum of the back-propagation gradients of all training samples. Or, in one embodiment, the target direction can be: the direction of the overall back-propagation gradient of the training samples in this parameter update process. They are not listed one by one.

[0175] In the above content, the "back propagation gradient of the training sample" in "calculating the similarity between the back propagation gradient of the adversarial sample and the back propagation gradient of the training sample" can be: the sum of the back propagation gradients of some training samples, or it can also be the sum of the back propagation gradients of all training samples. Alternatively, in one embodiment, the back propagation gradient of the training sample can be: the overall back propagation gradient of the training sample during this parameter update process. In this embodiment, the overall back propagation gradient of the training sample can be calculated in the above manner, which will not be repeated here. In this embodiment, the similarity between the back propagation gradient of the adversarial sample and the overall back propagation gradient of the training sample during this parameter update process is calculated.

[0176] In one of the above implementations, the angle between the back propagation gradient of the adversarial sample and the back propagation gradient of the training sample is calculated. If the back propagation gradient of the training sample is the overall back propagation gradient of the training sample, the following formula can be used to calculate the angle between the back propagation gradient of the adversarial sample and the overall back propagation gradient of the training sample:

[0177]

[0178] in, represents the angle between the overall back-propagation gradient of the training sample and the back-propagation gradient of the i-th adversarial sample, G t represents the overall back-propagation gradient of the training sample during the t-th parameter update process, i represents the identifier of the adversarial sample, Represents adversarial examples The back-propagation gradient of L represents the loss value, represents the gradient operator, Indicates: The adversarial sample The output data input to the neural network is relative to the true label y i The loss value.

[0179] if Less than or equal to 90 degrees, then it is wrong Remove the If the back propagation gradient of the adversarial sample is greater than 90 degrees, and the back propagation gradient of the adversarial sample is determined as the back propagation gradient to be processed, the following formula is used to calculate To remove:

[0180]

[0181] in, represents the target back-propagation gradient, represents the projected gradient.

[0182] As described above, in one embodiment, if the similarity between the back-propagation gradient of the adversarial sample and the back-propagation gradient of the training sample meets the preset similarity condition, the back-propagation gradient of the adversarial sample is not removed, and the back-propagation gradient of the adversarial sample can be directly determined as the target back-propagation gradient. In this case, the target back-propagation gradient can be calculated by the following formula:

[0183]

[0184] in, Represents adversarial examples The back-propagation gradient of represents the projected gradient, It represents the angle between the overall back-propagation gradient of the training sample and the back-propagation gradient of the i-th adversarial sample.

[0185] The neural network includes multiple layers. In S104, the back propagation gradients corresponding to some layers in the neural network can be removed, or the back propagation gradients corresponding to all layers in the neural network can be removed, and the specific layers are not limited.

[0186] S105: Update the parameters of the neural network by back-propagating the back-propagated gradient of the training sample and the target back-propagated gradient.

[0187] S106: Determine whether the training is completed, if not, return to execute S102, if yes, execute S107.

[0188] S107: Obtain a recognition model.

[0189] As mentioned above, assume that the training sample is x = [x1, x2, ..., x n ], the adversarial sample generated based on the training sample is The training process can be understood as the process of making the parameters θ in the neural network gradually satisfy the following formula:

[0190]

[0191] Among them, θ represents the parameters in the neural network, i represents the sample identifier, and n represents the number of samples. The samples mentioned here refer to both training samples and adversarial samples, and the two correspond one to one; L(x i |y i ) means: the training sample x i The output data input to the neural network is relative to the true label y i The loss value, Indicates: The adversarial sample The output data input to the neural network is relative to the true label y i The loss value.

[0192] The t+1th parameter update can be expressed as:

[0193]

[0194] Among them, θ t+1 represents the parameters in the neural network obtained after the t+1th parameter update, θ t represents the parameters in the neural network obtained after the tth parameter update, Represents the training sample x i The back-propagation gradient of represents the above target back-propagation gradient, β represents the learning rate, β can be a fixed value or a function that changes with t: the smaller t is, the larger β is, and when t increases, β becomes smaller. That is to say, in the early stage of neural network training, the learning rate is large, and as the training progresses, the learning rate gradually decreases.

[0195] As described above, the training process of the neural network can be understood as an iterative update process of parameters. Executing S102-S106 once is an iteration, or in other words, the parameters are updated once. When the conditions for termination of the iteration are met, for example, when the preset number of iterations is reached, or when the loss convergence condition is met, the training is completed, and the neural network after the training is completed is the recognition model. The conditions for termination of the iteration are not limited.

[0196] If the training sample is a sample image, the image recognition model can be trained by applying the embodiment of the present application, and the image recognition model can be used for image recognition. The embodiment of the present application can further provide an image recognition method, including: obtaining an image to be recognized; inputting the image to be recognized into the image recognition model to obtain an image recognition result, such as a face recognition result, a license plate recognition result, etc., without specific limitation. The image recognition model can defend against attacks from adversarial samples, and the accuracy of the image recognition result obtained is relatively high.

[0197] If the training sample is a sample video, the video recognition model can be trained by applying the embodiment of the present application, and the video recognition model can be used for video recognition. The embodiment of the present application can further provide a video recognition method, including: obtaining a video to be recognized; inputting the video to be recognized into the video recognition model to obtain a video recognition result, such as a face recognition result, a license plate recognition result, a target tracking result, etc., without specific limitation. The video recognition model can defend against the attack of adversarial samples, and the accuracy of the obtained video recognition result is relatively high.

[0198] If the training sample is a sample corpus, a semantic recognition model can be trained by applying the embodiment of the present application, and the semantic recognition model can be used for corpus recognition. The embodiment of the present application can further provide a corpus recognition method, including: obtaining a corpus to be recognized; inputting the corpus to be recognized into the semantic recognition model to obtain a semantic recognition result. The semantic recognition model can defend against attacks from adversarial samples, and the accuracy of the obtained semantic recognition result is relatively high.

[0199] If the training sample is a sample speech, the speech recognition model can be trained by applying the embodiment of the present application, and the speech recognition model can be used for speech recognition. The embodiment of the present application can further provide a speech recognition method, including: obtaining a speech to be recognized; inputting the speech to be recognized into the speech recognition model to obtain a speech recognition result. The speech recognition model can defend against attacks from adversarial samples, and the accuracy of the speech recognition result obtained is relatively high.

[0200] Using the embodiments shown in the present application, training samples and adversarial samples are input into a neural network to obtain output data of the neural network; the back propagation gradient of the training sample and the back propagation gradient of the adversarial sample are respectively calculated based on the output data; the part of the back propagation gradient of the adversarial sample that is opposite to the back propagation gradient of the training sample is removed to obtain the target back propagation gradient; this removal process can weaken the negative effect of the adversarial sample on the training sample, or reduce the performance damage of the adversarial sample to the training sample, thereby reducing the adverse effects of the adversarial sample on the training process; then, the back propagation gradient of the training sample and the target back propagation gradient are back-propagated to update the parameters of the neural network; the recognition model output obtained after the iterative update is completed has a high accuracy rate and can defend against the attack of the adversarial sample.

[0201] It can be seen that this scheme can not only weaken the negative impact of adversarial samples on training samples, but also maintain the recognition model's defense capability against adversarial samples, thereby improving the output accuracy of the recognition model.

[0202] Figure 3 A second flow chart of the neural network training method provided in the embodiment of the present application includes:

[0203] S301: Obtain training samples and adversarial samples of the training samples, where the training samples include any one of sample images, sample videos, sample corpora, and sample voices.

[0204] If the training sample is a sample image, the embodiment of the present application can be used to train an image recognition model; further subdivided, if the sample image contains a face area, the embodiment of the present application can be used to train a face recognition model; if the sample image contains a license plate area, the embodiment of the present application can be used to train a license plate recognition model, and so on.

[0205] If the training sample is a sample video, the embodiment of the present application can be used to train a video recognition model; further subdivided, if the sample video contains a face area, the embodiment of the present application can be used to train a face recognition model; if the sample video contains a license plate area, the embodiment of the present application can be used to train a license plate recognition model; if the sample video contains a target to be tracked, the embodiment of the present application can be used to train a target tracking model, and so on.

[0206] If the training sample is a sample corpus, the semantic recognition model can be trained by applying the embodiments of the present application. Corpus is language material, which can be understood as data including texts such as sentences and words. If the training sample is a sample speech, the speech recognition model can be trained by applying the embodiments of the present application.

[0207] The training samples mentioned in this embodiment can be understood as normal samples without adding adversarial disturbances, or non-adversarial samples. For example, after obtaining the training samples, adversarial samples can be generated based on the training samples, and the generation method is not limited.

[0208] For example, suppose the training sample is x = [x1, x2, ..., x n ], the adversarial sample generated based on the training sample is Here it is assumed that the training samples correspond one to one with the adversarial samples.

[0209] S302: Input the training samples and adversarial samples into the neural network to obtain output data of the neural network.

[0210] S303: Based on the output data, the back propagation gradient of the training sample and the back propagation gradient of the adversarial sample are calculated respectively.

[0211] S304: Calculate the overall back propagation gradient of the training samples in the current parameter updating process based on the overall back propagation gradient of the training samples in the previous parameter updating process and the back propagation gradient of all training samples in the current parameter updating process.

[0212] The training process of the neural network can be understood as an iterative update process of parameters. One execution of S302-S313 is one iteration, or one parameter update. When the conditions for the termination of the iteration are met, for example, when the preset number of iterations is reached, or when the loss convergence condition is met, the training is completed, and the neural network after the training is completed is the recognition model. The conditions for the termination of the iteration are not limited.

[0213] In this implementation, a moving average algorithm may be used to calculate the overall back propagation gradient of the training sample in this parameter update process. The calculation process may include:

[0214] Calculate the product of the overall back propagation gradient of the training sample in the last parameter update process and the hyperparameter in the moving average algorithm as the first value;

[0215] Calculate the sum of the back-propagated gradients of all training samples in this parameter update process as the second value;

[0216] Calculate the sum of the first value and the second value as a third value;

[0217] Calculate the influence coefficient of the overall back propagation gradient of the training sample in the previous parameter updating process on the overall back propagation gradient of the training sample in the current parameter updating process as the fourth value;

[0218] Calculate the product of the fourth value and the hyperparameter in the moving average algorithm as a fifth value;

[0219] Calculating the sum of the fifth value and a preset parameter as a sixth value;

[0220] The ratio of the third value to the sixth value is calculated as the overall back propagation gradient of the training sample in this parameter updating process.

[0221] The above calculation process can be expressed as the following formula:

[0222]

[0223] Among them, G t represents the overall back-propagation gradient of the training sample during the tth parameter update process, G t-1 represents the overall back propagation gradient of the training sample during the t-1th parameter update process, α represents the hyperparameter in the moving average algorithm, α can be a preset value, and the specific value is not limited; i represents the identifier of the training sample, n represents the number of training samples, represents the back-propagation gradient of the i-th training sample during the t-th parameter update process, represents the sum of the back-propagated gradients of all training samples during the t-th parameter update process, N t-1 N represents the influence coefficient of the overall back propagation gradient of the training sample in the t-1th parameter update process on the overall back propagation gradient of the training sample in the tth parameter update process. t =αN t-1 +1. The 1 in the denominator of the above formula represents the above preset parameter.

[0224] The tth parameter updating process can be understood as this parameter updating process.

[0225] S305: For each back propagation gradient of the adversarial sample, calculate the angle between the back propagation gradient of the adversarial sample and the overall back propagation gradient of the training sample.

[0226] In other contents of the embodiments of the present application, unless the previous parameter update process is emphasized separately, the data in the current parameter update process is processed.

[0227] In one implementation, the following formula may be used to calculate the angle between the back propagation gradient of the adversarial sample and the overall back propagation gradient of the training sample:

[0228]

[0229] in, represents the angle between the overall back-propagation gradient of the training sample and the back-propagation gradient of the i-th adversarial sample, G t represents the overall back-propagation gradient of the training sample during the t-th parameter update process, i represents the identifier of the adversarial sample, Represents adversarial examples The back-propagation gradient of L represents the loss value, represents the gradient operator, Indicates: The adversarial sample The output data input to the neural network is relative to the true label y i The loss value.

[0230] S306: Determine whether the angle is greater than 90 degrees; if not, execute S307; if greater, execute S308.

[0231] S307: Determine the back propagation gradient of the adversarial sample as the target back propagation gradient.

[0232] S308: Determine the back propagation gradient of the adversarial sample as a candidate processing back propagation gradient.

[0233] S309: Select a back-propagated gradient to be processed from the determined candidate back-propagated gradients.

[0234] In one case, all candidate processed back-propagation gradients may be determined as the back-propagation gradients to be processed. Alternatively, in another case, some back-propagation gradients to be processed may be selected from the determined candidate processed back-propagation gradients, such as by random selection, or by random selection according to a set probability, or by selection according to a specified order, etc., and the specific selection method is not limited.

[0235] In one implementation, a selection probability can be set in this parameter update process; the selection probability increases as the number of parameter updates of the neural network increases; and the back-propagation gradient to be processed is selected from the determined candidate back-propagation gradients using the selection probability.

[0236] The following formula can be used to set the selection probability during each parameter update process:

[0237]

[0238] Among them, p represents the selection probability, γ represents the preset parameter, T represents the total number of parameter updates during the neural network training process, t represents the parameter update this time is the tth parameter update, and exp represents the exponential function with the natural constant e as the base. As t increases, p also increases.

[0239] In this implementation, at the beginning of neural network training, a small number of back propagated gradients of adversarial samples are selected for removal processing; as the training progresses, the back propagated gradient directions of training samples become more and more accurate, and more and more back propagated gradients of adversarial samples are selected for removal processing; thus, on the one hand, compared with removing the back propagated gradients of a large number of adversarial samples throughout the training process, this implementation reduces the amount of calculation; on the other hand, compared with removing the back propagated gradients of a small number of adversarial samples throughout the training process, the training accuracy is improved; it can be seen that this implementation reasonably plans the amount of calculation and improves the training accuracy.

[0240] S310: Calculate the projection of the back-propagated gradient to be processed in the target direction to obtain the projected gradient, where the target direction is the direction of the overall back-propagated gradient of the training sample.

[0241] For example,

[0242] Among them, G t represents the overall back-propagation gradient of the training sample during the t-th parameter update process, i represents the identifier of the adversarial sample, Represents adversarial examples The back-propagation gradient of L represents the loss value, represents the gradient operator, Indicates: The adversarial sample The output data input to the neural network is relative to the true label y i The loss value.

[0243] S311: Remove the projected gradient from the back-propagated gradient to be processed to obtain the target back-propagated gradient.

[0244] Figure 3 In the embodiment shown, the target back propagation gradient can be calculated using the following formula:

[0245]

[0246] in, Represents adversarial examples The back-propagation gradient of represents the projected gradient, It represents the angle between the overall back-propagation gradient of the training sample and the back-propagation gradient of the i-th adversarial sample.

[0247] A neural network consists of multiple layers. Figure 3 In the illustrated embodiment, the back propagation gradients corresponding to some levels in the neural network may be removed, or the back propagation gradients corresponding to all levels in the neural network may be removed, without limitation to the specific levels.

[0248] S312: Update the parameters of the neural network by back-propagating the back-propagation gradient of the training sample and the target back-propagation gradient.

[0249] S313: Determine whether the training is completed, if not, return to execute S302, if yes, execute S314.

[0250] S314: Obtain a recognition model.

[0251] As mentioned above, assume that the training sample is x = [x1, x2, ..., x n ], the adversarial sample generated based on the training sample is The training process can be understood as the process of making the parameters θ in the neural network gradually satisfy the following formula:

[0252]

[0253] Among them, θ represents the parameters in the neural network, i represents the sample identifier, and n represents the number of samples. The samples mentioned here refer to both training samples and adversarial samples, and the two correspond one to one; L(x i |y i ) means: the training sample x i The output data input to the neural network is relative to the true label y i The loss value, Indicates: The adversarial sample The output data input to the neural network is relative to the true label y i The loss value.

[0254] The t+1th parameter update can be expressed as:

[0255]

[0256] Among them, θ t+1 represents the parameters in the neural network obtained after the t+1th parameter update, θ t represents the parameters in the neural network obtained after the tth parameter update, Represents the training sample x i The back-propagation gradient of represents the above target back-propagation gradient, β represents the learning rate, β can be a fixed value or a function that changes with t: the smaller t is, the larger β is, and when t increases, β becomes smaller. That is to say, in the early stage of neural network training, the learning rate is large, and as the training progresses, the learning rate gradually decreases.

[0257] As described above, the training process of the neural network can be understood as an iterative update process of parameters. Executing S302-S313 once is an iteration, or in other words, the parameters are updated once. When the conditions for iteration termination are met, for example, when the preset number of iterations is reached, or when the loss convergence conditions are met, the training is completed. The neural network after training is the recognition model.

[0258] If the training sample is a sample image, the image recognition model can be trained by applying the embodiment of the present application, and the image recognition model can be used for image recognition. The embodiment of the present application can further provide an image recognition method, including: obtaining an image to be recognized; inputting the image to be recognized into the image recognition model to obtain an image recognition result, such as a face recognition result, a license plate recognition result, etc., without specific limitation. The image recognition model can defend against attacks from adversarial samples, and the accuracy of the image recognition result obtained is relatively high.

[0259] If the training sample is a sample video, the video recognition model can be trained by applying the embodiment of the present application, and the video recognition model can be used for video recognition. The embodiment of the present application can further provide a video recognition method, including: obtaining a video to be recognized; inputting the video to be recognized into the video recognition model to obtain a video recognition result, such as a face recognition result, a license plate recognition result, a target tracking result, etc., without specific limitation. The video recognition model can defend against the attack of adversarial samples, and the accuracy of the obtained video recognition result is relatively high.

[0260] If the training sample is a sample corpus, a semantic recognition model can be trained by applying the embodiment of the present application, and the semantic recognition model can be used for corpus recognition. The embodiment of the present application can further provide a corpus recognition method, including: obtaining a corpus to be recognized; inputting the corpus to be recognized into the semantic recognition model to obtain a semantic recognition result. The semantic recognition model can defend against attacks from adversarial samples, and the accuracy of the obtained semantic recognition result is relatively high.

[0261] If the training sample is a sample speech, the speech recognition model can be trained by applying the embodiment of the present application, and the speech recognition model can be used for speech recognition. The embodiment of the present application can further provide a speech recognition method, including: obtaining a speech to be recognized; inputting the speech to be recognized into the speech recognition model to obtain a speech recognition result. The speech recognition model can defend against attacks from adversarial samples, and the accuracy of the speech recognition result obtained is relatively high.

[0262] Using the embodiments shown in the present application, training samples and adversarial samples are input into a neural network to obtain output data of the neural network; the back propagation gradient of the training sample and the back propagation gradient of the adversarial sample are respectively calculated based on the output data; the part of the back propagation gradient of the adversarial sample that is opposite to the back propagation gradient of the training sample is removed to obtain the target back propagation gradient; this removal process can weaken the negative effect of the adversarial sample on the training sample, or reduce the performance damage of the adversarial sample to the training sample, thereby reducing the adverse effects of the adversarial sample on the training process; then, the back propagation gradient of the training sample and the target back propagation gradient are back-propagated to update the parameters of the neural network; the recognition model output obtained after the iterative update is completed has a high accuracy rate and can defend against the attack of the adversarial sample.

[0263] It can be seen that this scheme can not only weaken the negative impact of adversarial samples on training samples, but also maintain the recognition model's defense capability against adversarial samples.

[0264] The present embodiment also provides an image recognition method, such as Figure 4 As shown, including:

[0265] S401: Acquire an image to be recognized.

[0266] S402: Input the image to be recognized into a pre-trained image recognition model to obtain an image recognition result.

[0267] The process of training the image recognition model is as follows: Figure 5 As shown, including:

[0268] S501: Obtain a sample image and an adversarial sample of the sample image.

[0269] S502: Input the sample image and the adversarial sample into the neural network to obtain output data of the neural network.

[0270] S503: Based on the output data, the back-propagation gradient of the sample image and the back-propagation gradient of the adversarial sample are respectively calculated.

[0271] S504: Remove the part of the back-propagated gradient of the adversarial sample that is opposite to the back-propagated gradient of the sample image to obtain the target back-propagated gradient.

[0272] S505: Update the parameters of the neural network by back-propagating the back-propagated gradient of the sample image and the target back-propagated gradient.

[0273] S506: Determine whether the training is completed, if not, return to execute S502, if yes, execute S507.

[0274] S507: Obtain an image recognition model.

[0275] The process of training the neural network refers to the above Figure 1-Figure 3 The embodiments shown will not be described in detail here.

[0276] To further subdivide, if the sample image contains a face area, the face recognition model can be trained by applying the embodiment of the present application. In this case, the image to be recognized can include a face area, and the image recognition result can be a face recognition result.

[0277] If the sample image includes a license plate area, the license plate recognition model can be trained by applying the embodiment of the present application. In this case, the image to be recognized can include the license plate area, and the image recognition result can be a license plate recognition result.

[0278] Using the embodiment shown in the present application, the initial image and the adversarial sample are input into the neural network to obtain the output data of the neural network; the back propagation gradient of the initial image and the back propagation gradient of the adversarial sample are calculated based on the output data; the part of the back propagation gradient of the adversarial sample that is opposite to the back propagation gradient of the initial image is removed to obtain the target back propagation gradient; this removal process can weaken the negative effect of the adversarial sample on the initial image, or reduce the performance damage of the adversarial sample on the initial image, thereby reducing the adverse effect of the adversarial sample on the training process; then, the back propagation gradient of the initial image and the target back propagation gradient are back propagated to update the parameters of the neural network; the image recognition model obtained after the iterative update is completed has a high output accuracy and can defend against the attack of the adversarial sample. It can be seen that this scheme can not only weaken the negative effect of the adversarial sample on the initial image, but also maintain the image recognition model's defense capability against the adversarial sample, thereby improving the output accuracy of the recognition model.

[0279] Corresponding to the above-mentioned neural network training method embodiment, the present application embodiment also provides a neural network training device, such as Figure 6 As shown, including:

[0280] A first acquisition module 601 is used to acquire training samples and adversarial samples of the training samples, wherein the training samples include any one of sample images, sample videos, sample corpora, and sample voices;

[0281] A first input module 602, used to input the training sample and the adversarial sample into the neural network to obtain output data of the neural network;

[0282] A first calculation module 603, used to calculate the back propagation gradient of the training sample and the back propagation gradient of the adversarial sample based on the output data;

[0283] A removal module 604 is used to remove the part of the back-propagation gradient of the adversarial sample that is opposite to the back-propagation gradient of the training sample to obtain a target back-propagation gradient;

[0284] The updating module 605 is used to update the parameters of the neural network by back-propagating the back-propagation gradient of the training sample and the target back-propagation gradient, and then re-trigger the first input module until the training is completed.

[0285] In one implementation, the removal module 604 includes: a selection submodule and a removal submodule (not shown in the figure), wherein:

[0286] A selection submodule, configured to select, from the back-propagation gradients of the adversarial sample, a back-propagation gradient whose similarity with the back-propagation gradient of the training sample does not satisfy a preset similarity condition as the back-propagation gradient to be processed;

[0287] The removal submodule is used to remove the part of the back-propagation gradient to be processed that is opposite to the back-propagation gradient of the training sample to obtain the target back-propagation gradient.

[0288] In one implementation, the selection submodule is specifically used to:

[0289] For each adversarial sample's back-propagation gradient, calculate the similarity between the back-propagation gradient of the adversarial sample and the back-propagation gradient of the training sample;

[0290] Determining whether the similarity satisfies a preset similarity condition;

[0291] If not, the back propagation gradient of the adversarial sample is determined as the candidate processing back propagation gradient;

[0292] The back propagation gradient to be processed is selected from the determined candidate back propagation gradients according to the set selection probability in the current parameter update process; the selection probability increases as the number of parameter updates of the neural network increases.

[0293] In one implementation, the device further includes: a first determining module and / or a second determining module (not shown in the figure), wherein:

[0294] A first determination module, configured to determine the back propagation gradient of the adversarial sample as a target back propagation gradient when the similarity satisfies a preset similarity condition;

[0295] The second determination module is used to determine the candidate processed back-propagation gradient that has not been selected as the back-propagation gradient to be processed as the target back-propagation gradient.

[0296] In one embodiment, the preset similarity condition is: the angle between the back propagation gradient of the adversarial sample and the back propagation gradient of the training sample is no greater than 90 degrees;

[0297] The removal submodule is specifically used for:

[0298] Calculate the projection of the back-propagated gradient to be processed in a target direction to obtain a projected gradient, wherein the target direction is the direction of the back-propagated gradient of the training sample;

[0299] In the back-propagated gradient to be processed, the projected gradient is removed to obtain the target back-propagated gradient.

[0300] In one embodiment, the device further comprises:

[0301] A second calculation module (not shown in the figure) is used to calculate the overall back propagation gradient of the training samples in the current parameter update process based on the overall back propagation gradient of the training samples in the previous parameter update process and the back propagation gradient of all training samples in the current parameter update process;

[0302] The removal module 604 is specifically used to remove the part of the back propagation gradient of the adversarial sample that is opposite to the overall back propagation gradient of the training sample in the current parameter update process to obtain the target back propagation gradient.

[0303] In one implementation, the second calculation module is specifically configured to:

[0304] Calculate the product of the overall back propagation gradient of the training sample in the last parameter update process and the hyperparameter in the moving average algorithm as the first value;

[0305] Calculate the sum of the back-propagated gradients of all training samples in this parameter update process as the second value;

[0306] Calculate the sum of the first value and the second value as a third value;

[0307] Calculate the influence coefficient of the overall back propagation gradient of the training sample in the previous parameter updating process on the overall back propagation gradient of the training sample in the current parameter updating process as the fourth value;

[0308] Calculate the product of the fourth value and the hyperparameter in the moving average algorithm as a fifth value;

[0309] Calculating the sum of the fifth value and a preset parameter as a sixth value;

[0310] The ratio of the third value to the sixth value is calculated as the overall back propagation gradient of the training sample in this parameter updating process.

[0311] In one embodiment, the device further includes: a first identification module, a third identification module and a fourth identification module (not shown in the figure), wherein:

[0312] A first recognition module is used to obtain an image to be recognized when the training sample is a sample image and an image recognition model is obtained through training; the image to be recognized is input into the image recognition model to obtain an image recognition result;

[0313] The second recognition module is used to obtain a video to be recognized when the training sample is a sample video and a video recognition model is obtained through training; the video to be recognized is input into the video recognition model to obtain a video recognition result;

[0314] The third recognition module is used to obtain the corpus to be recognized when the training sample is a sample corpus and the semantic recognition model is obtained through training; input the corpus to be recognized into the semantic recognition model to obtain a semantic recognition result;

[0315] The fourth recognition module is used to obtain the speech to be recognized when the above-mentioned training sample is a sample speech and the speech recognition model is trained; input the speech to be recognized into the speech recognition model to obtain the speech recognition result.

[0316] Using the embodiment shown in the present application, the training samples and adversarial samples are input into the neural network to obtain the output data of the neural network; the back propagation gradient of the training samples and the back propagation gradient of the adversarial samples are calculated based on the output data; the part of the back propagation gradient of the adversarial samples that is opposite to the back propagation gradient of the training samples is removed to obtain the target back propagation gradient; this removal process can weaken the negative effect of the adversarial samples on the training samples, or reduce the performance damage of the adversarial samples on the training samples, thereby reducing the adverse effects of the adversarial samples on the training process; then, the back propagation gradient of the training samples and the target back propagation gradient are back propagated to update the parameters of the neural network; the recognition model output obtained after the iterative update is completed has a high accuracy rate and can defend against the attack of the adversarial samples. It can be seen that this scheme can not only weaken the negative effect of the adversarial samples on the training samples, but also maintain the defense capability of the recognition model against the adversarial samples.

[0317] Corresponding to the above-mentioned image recognition method embodiment, the present application embodiment also provides an image recognition device, such as Figure 7 As shown, including:

[0318] The second acquisition module 701 is used to acquire the image to be recognized;

[0319] The second input module 702 is used to input the image to be recognized into a pre-trained image recognition model to obtain an image recognition result;

[0320] The device also includes a training module 703, which is used to train and obtain the image recognition model;

[0321] The training module 703 includes:

[0322] An acquisition submodule 7031 is used to acquire a sample image and an adversarial sample of the sample image;

[0323] An input submodule 7032, used to input the sample image and the adversarial sample into the neural network to obtain output data of the neural network;

[0324] A calculation submodule 7033, used to calculate the back propagation gradient of the sample image and the back propagation gradient of the adversarial sample based on the output data;

[0325] A removal submodule 7034 is used to remove the part of the back-propagated gradient of the adversarial sample that is opposite to the back-propagated gradient of the sample image, so as to obtain a target back-propagated gradient;

[0326] The updating submodule 7035 is used to update the parameters of the neural network by back-propagating the back-propagation gradient of the sample image and the target back-propagation gradient, and then return to trigger the input submodule until the training is completed to obtain an image recognition model.

[0327] Using the embodiment shown in the present application, the initial image and the adversarial sample are input into the neural network to obtain the output data of the neural network; the back propagation gradient of the initial image and the back propagation gradient of the adversarial sample are calculated based on the output data; the part of the back propagation gradient of the adversarial sample that is opposite to the back propagation gradient of the initial image is removed to obtain the target back propagation gradient; this removal process can weaken the negative effect of the adversarial sample on the initial image, or reduce the performance damage of the adversarial sample on the initial image, thereby reducing the adverse effect of the adversarial sample on the training process; then, the back propagation gradient of the initial image and the target back propagation gradient are back propagated to update the parameters of the neural network; the image recognition model obtained after the iterative update has a high output accuracy and can defend against the attack of the adversarial sample. It can be seen that this scheme can not only weaken the negative effect of the adversarial sample on the initial image, but also maintain the image recognition model's defense capability against the adversarial sample.

[0328] The present application also provides an electronic device, such as Figure 8 As shown, it includes a processor 801 and a memory 802,

[0329] Memory 802, used for storing computer programs;

[0330] The processor 801 is used to implement any of the above-mentioned neural network training or any of the image recognition methods when executing the program stored in the memory 802.

[0331] The memory mentioned in the above electronic device may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the above processor.

[0332] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0333] In another embodiment provided in the present application, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, any one of the above-mentioned neural network training or any one of the image recognition methods is implemented.

[0334] In another embodiment provided by the present application, a computer program product including instructions is also provided, which, when executed on a computer, enables the computer to execute any one of the above-mentioned neural network training or any one of the image recognition methods.

[0335] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive Solid State Disk (SSD)), etc.

[0336] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0337] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, equipment embodiment, computer-readable storage medium embodiment, and computer program product embodiment, since they are basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0338] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are included in the protection scope of the present application.

Claims

1. A neural network training method, characterized in that: include: Obtaining training samples and adversarial samples of the training samples, wherein the training samples include any one of sample images, sample videos, sample corpora, and sample voices; Inputting the training sample and the adversarial sample into a neural network to obtain output data of the neural network; Based on the output data, respectively calculate the back propagation gradient of the training sample and the back propagation gradient of the adversarial sample; Removing the part of the back-propagated gradient of the adversarial sample that is opposite to the back-propagated gradient of the training sample to obtain a target back-propagated gradient; The parameters of the neural network are updated by back-propagating the back-propagated gradient of the training sample and the target back-propagated gradient, and then returning to the step of inputting the training sample and the adversarial sample into the neural network until the training is completed.

2. The method according to claim 1, characterized in that The removing the part of the back propagation gradient of the adversarial sample that is opposite to the back propagation gradient of the training sample to obtain the target back propagation gradient includes: Among the back-propagated gradients of the adversarial sample, a back-propagated gradient whose similarity with the back-propagated gradient of the training sample does not satisfy a preset similarity condition is selected as the back-propagated gradient to be processed; The part of the back-propagated gradient to be processed that is opposite to the back-propagated gradient of the training sample is removed to obtain the target back-propagated gradient.

3. The method according to claim 2, characterized in that The step of selecting, from the back-propagation gradients of the adversarial sample, a back-propagation gradient whose similarity with the back-propagation gradient of the training sample does not satisfy a preset similarity condition as the back-propagation gradient to be processed includes: For each adversarial sample's back-propagation gradient, calculate the similarity between the back-propagation gradient of the adversarial sample and the back-propagation gradient of the training sample; Determining whether the similarity satisfies a preset similarity condition; If not, the back propagation gradient of the adversarial sample is determined as the candidate processing back propagation gradient; The back propagation gradient to be processed is selected from the determined candidate back propagation gradients according to the set selection probability in the current parameter update process; the selection probability increases as the number of parameter updates of the neural network increases.

4. The method according to claim 3, characterized in that: The method further comprises: If the similarity satisfies a preset similarity condition, the back propagation gradient of the adversarial sample is determined as the target back propagation gradient; And / or, determining a candidate processed back-propagation gradient that is not selected as the back-propagation gradient to be processed as a target back-propagation gradient.

5. The method according to claim 2 or 3, characterized in that: The preset similarity condition is: the angle between the back-propagation gradient of the adversarial sample and the back-propagation gradient of the training sample is no greater than 90 degrees; The removing the part of the back-propagated gradient to be processed that is opposite to the back-propagated gradient of the training sample to obtain the target back-propagated gradient includes: Calculate the projection of the back-propagated gradient to be processed in a target direction to obtain a projected gradient, wherein the target direction is the direction of the back-propagated gradient of the training sample; In the back-propagated gradient to be processed, the projected gradient is removed to obtain the target back-propagated gradient.

6. The method according to claim 1, characterized in that Before removing the part of the back propagation gradient of the adversarial sample that is opposite to the back propagation gradient of the training sample to obtain the target back propagation gradient, the method further includes: Based on the overall back propagation gradient of the training samples in the previous parameter update process and the back propagation gradient of all training samples in the current parameter update process, the overall back propagation gradient of the training samples in the current parameter update process is calculated; The removing the part of the back propagation gradient of the adversarial sample that is opposite to the back propagation gradient of the training sample to obtain the target back propagation gradient includes: The back-propagated gradient of the adversarial sample is removed in the opposite direction to the overall back-propagated gradient of the training sample in the current parameter update process to obtain the target back-propagated gradient.

7. The method according to claim 6, characterized in that The overall back propagation gradient of the training samples in the last parameter updating process and the back propagation gradient of all training samples in the current parameter updating process are calculated, including: Calculate the product of the overall back propagation gradient of the training sample in the last parameter update process and the hyperparameter in the moving average algorithm as the first value; Calculate the sum of the back-propagated gradients of all training samples in this parameter update process as the second value; Calculate the sum of the first value and the second value as a third value; Calculate the influence coefficient of the overall back propagation gradient of the training sample in the previous parameter updating process on the overall back propagation gradient of the training sample in the current parameter updating process as the fourth value; Calculate the product of the fourth value and the hyperparameter in the moving average algorithm as a fifth value; Calculating the sum of the fifth value and a preset parameter as a sixth value; The ratio of the third value to the sixth value is calculated as the overall back propagation gradient of the training sample in this parameter updating process.

8. The method according to claim 1, characterized in that: If the training sample is a sample image, the method further includes: Obtain an image to be recognized; Inputting the image to be recognized into the image recognition model trained based on the neural network to obtain an image recognition result; If the training sample is a sample video, the method further includes: Get the video to be identified; Inputting the video to be recognized into a video recognition model trained based on the neural network to obtain a video recognition result; If the training sample is a sample corpus, the method further includes: Obtain the corpus to be recognized; Inputting the corpus to be recognized into a semantic recognition model obtained by training the neural network to obtain a semantic recognition result; If the training sample is a sample speech, the method further includes: Get the speech to be recognized; The speech to be recognized is input into a speech recognition model trained based on the neural network to obtain a speech recognition result.

9. An image recognition method, characterized in that: include: Obtain an image to be recognized; Inputting the image to be recognized into a pre-trained image recognition model to obtain an image recognition result; The process of training the image recognition model includes: Obtaining a sample image and an adversarial sample of the sample image; Inputting the sample image and the adversarial sample into a neural network to obtain output data of the neural network; Based on the output data, respectively calculate the back-propagation gradient of the sample image and the back-propagation gradient of the adversarial sample; Removing the part of the back-propagated gradient of the adversarial sample that is opposite to the back-propagated gradient of the sample image to obtain a target back-propagated gradient; The back propagation gradient of the sample image and the target back propagation gradient are back-propagated to update the parameters of the neural network, and then the step of inputting the sample image and the adversarial sample into the neural network is returned to execute until the training is completed, thereby obtaining an image recognition model.

10. A neural network training device, characterized in that: include: A first acquisition module, used to acquire training samples and adversarial samples of the training samples, wherein the training samples include any one of sample images, sample videos, sample corpora, and sample voices; A first input module, used to input the training sample and the adversarial sample into the neural network to obtain output data of the neural network; A first calculation module, used for respectively calculating the back propagation gradient of the training sample and the back propagation gradient of the adversarial sample based on the output data; A removal module, used to remove the part of the back-propagation gradient of the adversarial sample that is opposite to the back-propagation gradient of the training sample, to obtain a target back-propagation gradient; An update module is used to update the parameters of the neural network by back-propagating the back-propagation gradient of the training sample and the target back-propagation gradient, and then re-trigger the first input module until the training is completed.

11. The device according to claim 10, characterized in that The removal module comprises: A selection submodule, configured to select, from the back-propagation gradients of the adversarial sample, a back-propagation gradient whose similarity with the back-propagation gradient of the training sample does not satisfy a preset similarity condition as the back-propagation gradient to be processed; The removal submodule is used to remove the part of the back-propagation gradient to be processed that is opposite to the back-propagation gradient of the training sample to obtain the target back-propagation gradient.

12. An image recognition device, characterized in that: include: A second acquisition module is used to acquire an image to be recognized; A second input module is used to input the image to be recognized into a pre-trained image recognition model to obtain an image recognition result; The device also includes a training module, which is used to train and obtain the image recognition model; The training module comprises: An acquisition submodule, used to acquire a sample image and an adversarial sample of the sample image; An input submodule, used for inputting the sample image and the adversarial sample into a neural network to obtain output data of the neural network; A calculation submodule, used for respectively calculating the back-propagation gradient of the sample image and the back-propagation gradient of the adversarial sample based on the output data; A removal submodule, used to remove the part of the back-propagated gradient of the adversarial sample that is opposite to the back-propagated gradient of the sample image, to obtain a target back-propagated gradient; The update submodule is used to update the parameters of the neural network by back-propagating the back-propagation gradient of the sample image and the target back-propagation gradient, and then return to trigger the input submodule until the training is completed to obtain an image recognition model.

Citation Information

Patent Citations

  • Convolutional neural network training method, image recognition method and corresponding devices

    CN110647992A

  • Large-batch adversarial sample generation method and system

    CN111275123A