Neural network search method and device

By using the pre-trained model to generate adversarial samples, determine the influence of neural network parameters and sort and prune it, the problems of long training time and high search cost in the deep learning model in the existing technology are solved, and the search and adversarial training of highly robust neural network structures are realized, which improves the security of the model.

CN114897165BActive Publication Date: 2025-05-23JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210628804.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-06
Publication Date
2025-05-23
Estimated Expiration
2042-06-06

AI Technical Summary

Technical Problem

When improving the robustness of deep learning models, the existing technology faces problems such as too long training time, slow model convergence speed, high training difficulty and high search cost, which leads to the inability to effectively solve the security problems of neural networks in many scenarios.

Method used

By processing the training samples using the pre-trained model, perturbation information is generated and superimposed with the training samples to generate adversarial samples. Then, the adversarial samples are used to determine the influence of parameters in the to-process neural network on the expression ability value and convergence speed value, and sort and prune them to obtain the target neural network and conduct adversarial training.

Benefits of technology

This method can effectively search for neural network structures with high robustness, reduce training time and search cost, and improve the security of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114897165B_ABST
    Figure CN114897165B_ABST
Patent Text Reader

Abstract

The present disclosure provides a neural network search method and device, which relates to the field of trusted artificial intelligence. The neural network search method includes: using a pre-trained model to process a training sample to generate disturbance information; superimposing the training sample and the disturbance information to generate an adversarial sample; determining the influence of each parameter in all parameters in the neural network to be processed on the expression ability value and the convergence speed value of the neural network to be processed; according to a first sorting rule, all parameters are sorted according to the influence of the expression ability value to obtain a first sorting result; according to a second sorting rule, all parameters are sorted according to the influence of the convergence speed value to obtain a second sorting result; according to the sequence number of each parameter in the first sorting result and the sequence number in the second sorting result, the pruning evaluation effect is determined; a preset number of parameters with the worst pruning evaluation effect are deleted to obtain a target neural network; and adversarial training is performed using the target neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of trusted artificial intelligence, in particular to the field of automatic machine learning, and in particular to a neural network search method and device. Background Art

[0002] At present, deep learning has achieved remarkable results in the fields of face recognition, autonomous driving, image classification and retrieval, recommendation systems, etc., but with it comes concerns about the security of deep learning algorithms. Studies have shown that images or special markers (i.e. adversarial samples) generated by malicious algorithms can cause deep learning systems to make wrong judgments, such as making autonomous driving systems misjudge road signs and face payment systems misidentify payers, which will pose a huge threat to fields with high security requirements. In order to solve this problem, in related technologies, one way is to focus on the training method of the model and improve the robustness of the model through different adversarial training methods. Another way is to search for robust neural network structures through neural network search technology. Summary of the invention

[0003] The inventors note that the disadvantages of the first method are that the training time is too long, the model convergence speed is slow, the training difficulty is high, and many scenarios are limited by the parameter scale, and the robustness that can be achieved is limited. The disadvantage of the second method is that the search cost is very high, which makes the search cost completely unaffordable in many scenarios.

[0004] Based on this, the present disclosure provides a neural network search solution that can effectively search for a neural network structure with high robustness, thereby better solving the security problem of the neural network.

[0005] According to a first aspect of an embodiment of the present disclosure, a neural network search method is provided, comprising: processing a training sample using a pre-trained model to generate disturbance information; superimposing the training sample and the disturbance information to generate an adversarial sample; using the adversarial sample, determining the influence of each parameter of all parameters in the neural network to be processed on the expression ability value and the convergence speed value of the neural network to be processed; sorting all the parameters according to the influence of the expression ability value according to a first sorting rule to obtain a first sorting result; sorting all the parameters according to the influence of the convergence speed value according to a second sorting rule to obtain a second sorting result, wherein the first sorting rule is opposite to the second sorting rule; determining the pruning evaluation effect of each parameter according to the serial number of each parameter in the first sorting result and the serial number in the second sorting result; deleting a preset number of parameters with the worst pruning evaluation effect to obtain a target neural network; and performing adversarial training using the target neural network.

[0006] In some embodiments, using the adversarial sample, determining the influence of each parameter of all the parameters in the neural network to be processed on the expression ability value and the convergence speed value of the neural network to be processed includes: inputting the adversarial sample into the neural network to be processed, calculating the expression ability value R0 and the convergence speed value H0 of the neural network to be processed; based on the expression ability value R0, calculating the influence of each parameter of all the parameters in the neural network to be processed on the expression ability value of the neural network to be processed; based on the convergence speed value H0, calculating the influence of each parameter on the convergence speed value of the neural network to be processed.

[0007] In some embodiments, based on the expressive power value R0, calculating the influence of each parameter of all the parameters in the neural network to be processed on the expressive power value of the neural network to be processed includes: deleting the jth parameter in the neural network to be processed to obtain an intermediate neural network, where 1≤j≤N, N is the total number of parameters; inputting the adversarial sample into the intermediate neural network to calculate the expressive power value Rj of the intermediate neural network; and determining the influence of the jth parameter on the expressive power value of the neural network to be processed based on the expressive power value R0 and the expressive power value Rj.

[0008] In some embodiments, the influence of the jth parameter on the expressiveness value of the neural network to be processed is the absolute value of the difference between the expressiveness value R0 and the expressiveness value Rj.

[0009] In some embodiments, the expressivity value is a linear region value.

[0010] In some embodiments, based on the convergence rate value H0, calculating the influence of each parameter on the convergence rate value of the neural network to be processed includes: deleting the jth parameter in the neural network to be processed to obtain an intermediate neural network, where 1≤j≤N, N is the total number of parameters; inputting the adversarial sample into the intermediate neural network to calculate the convergence rate value Hj of the intermediate neural network; and determining the influence of the jth parameter on the convergence rate value of the neural network to be processed based on the convergence rate value H0 and the convergence rate value Hj.

[0011] In some embodiments, the influence of the jth parameter on the convergence speed value of the neural network to be processed is the absolute value of the difference between the convergence speed value H0 and the convergence speed value Hj.

[0012] In some embodiments, the convergence rate value is a neural tangent kernel NTK value.

[0013] In some embodiments, the NTK value is

[0014]

[0015] in, is the gradient of the target loss function with respect to the parameter w.

[0016] In some embodiments, processing training samples using a pre-trained model includes: processing the training samples using a preset loss function to generate disturbance information; wherein the loss function is associated with the predicted probability value of the pre-trained model on the correct classification label of the training sample and the predicted probability value of the pre-trained model on the incorrect classification label of the training sample.

[0017] In some embodiments, the loss function for:

[0018]

[0019] Among them, x v is the training sample, v is the generated disturbance information, is the predicted probability value of the pre-trained model on the correct classification label of the training sample, is the predicted probability value of the pre-trained model on the i-th misclassified label of the training sample, -κ is the confidence value, and max is the maximum value function.

[0020] In some embodiments, the first sorting rule is to sort all the parameters in order from small to large according to the influence of the expression ability value; the second sorting rule is to sort all the parameters in order from large to small according to the influence of the convergence speed value.

[0021] In some embodiments, the pruning evaluation effect of each parameter is represented by a weighted sum of the sequence number of each parameter in the first sorting result and the sequence number of each parameter in the second sorting result.

[0022] In some embodiments, deleting a preset number of parameters with the worst pruning evaluation effects includes: deleting a preset number of parameters with the largest weighted sum.

[0023] In some embodiments, the weight of the sequence number in the first sorting result is the same as the weight of the sequence number in the second sorting result.

[0024] In some embodiments, the weight of the sequence number in the first sorting result and the weight of the sequence number in the second sorting result are 1.

[0025] According to a second aspect of an embodiment of the present disclosure, a neural network search device is provided, comprising: a first processing module, configured to process a training sample using a pre-trained model to generate disturbance information, and superimpose the training sample and the disturbance information to generate an adversarial sample; a second processing module, configured to determine, using the adversarial sample, the influence of each parameter of all parameters in the neural network to be processed on the expression ability value and the convergence speed value of the neural network to be processed; a third processing module, configured to sort all the parameters according to the influence of the expression ability value according to a first sorting rule to obtain a first sorting result, and sort all the parameters according to the influence of the convergence speed value according to a second sorting rule to obtain a second sorting result, wherein the first sorting rule is opposite to the second sorting rule, and the pruning evaluation effect of each parameter is determined according to the serial number of each parameter in the first sorting result and the serial number in the second sorting result; a fourth processing module, configured to delete a preset number of parameters with the worst pruning evaluation effect to obtain a target neural network, and use the target neural network for adversarial training.

[0026] According to a third aspect of an embodiment of the present disclosure, a neural network search device is provided, comprising: a memory configured to store instructions; a processor coupled to the memory, the processor being configured to execute a method as described in any of the above embodiments based on the instructions stored in the memory.

[0027] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and when the instructions are executed by a processor, the method involved in any of the above embodiments is implemented.

[0028] Other features and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0030] Figure 1 A flowchart of a neural network search method according to an embodiment of the present disclosure;

[0031] Figure 2 A flowchart of a neural network search method according to another embodiment of the present disclosure;

[0032] Figure 3 A schematic diagram of the structure of a neural network search device according to an embodiment of the present disclosure;

[0033] Figure 4 This is a schematic diagram of the structure of a neural network search device according to another embodiment of the present invention. DETAILED DESCRIPTION

[0034] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all of the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is by no means intended to limit the present disclosure and its application or use. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0035] Unless specifically stated otherwise, the relative arrangement of components and steps, the numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present disclosure.

[0036] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0037] Technologies, methods, and apparatus known to ordinary technicians in the relevant field may not be discussed in detail, but where appropriate, such technologies, methods, and apparatus should be considered part of the authorization specification.

[0038] In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limiting. Therefore, other examples of the exemplary embodiments may have different values.

[0039] It should be noted that like reference numerals and letters refer to similar items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0040] Figure 1 The flowchart of the neural network search method of one embodiment of the present disclosure is shown in FIG. In some embodiments, the following neural network search method is performed by a neural network search device.

[0041] In step 101, a training sample is processed using a pre-trained model to generate disturbance information.

[0042] In some embodiments, the training samples are processed using a preset loss function to generate disturbance information. The loss function is associated with the predicted probability value of the pre-trained model on the correct classification label of the training samples and the predicted probability value of the pre-trained model on the incorrect classification label of the training samples.

[0043] For example, the loss function It is the following formula (1).

[0044]

[0045] Among them, x v is the training sample, v is the generated disturbance information, is the predicted probability value of the pre-trained model on the correct classification label of the training sample, is the predicted probability value of the pre-trained model on the i-th misclassified label of the training sample, -k is the confidence value, and max is the maximum value function.

[0046] It should be noted that the goal of the loss function is to minimize the predicted probability value on the target class to make the model make an incorrect judgment. During the iteration process, the gradient of the loss function with respect to the sample is calculated each time, and multiple rounds of superposition are performed and cropped to a certain range, so that the disturbance of the image superposition will not cause too much interference to the original image.

[0047] In step 102, the training sample and the perturbation information are superimposed to generate an adversarial sample.

[0048] In step 103, the adversarial sample is used to determine the influence of each parameter of all parameters in the neural network to be processed on the expression ability value and the convergence speed value of the neural network to be processed.

[0049] like Figure 2 As shown, the above step 103 may include:

[0050] In step 201, the adversarial sample is input into the neural network to be processed, and the expression ability value R0 and the convergence speed value H0 of the neural network to be processed are calculated.

[0051] In step 202, based on the expression capability value R0, the influence of each parameter of all parameters in the neural network to be processed on the expression capability value of the neural network to be processed is calculated.

[0052] For example, the jth parameter in the neural network to be processed is deleted to obtain an intermediate neural network, where 1≤j≤N, and N is the total number of parameters.

[0053] Next, the adversarial sample is input into the intermediate neural network, and the expressive power value Rj of the intermediate neural network is calculated.

[0054] Then, the influence of the jth parameter on the expression value of the neural network to be processed is determined according to the expression value R0 and the expression value Rj. For example, the influence of the jth parameter on the expression value of the neural network to be processed is the absolute value of the difference between the expression value R0 and the expression value Rj.

[0055] In some embodiments, the expression ability value of the neural network is the linear region value of the neural network. For example, the linear region value is defined as follows:

[0056]

[0057] Among them, z(x 0 ; θ) represents the pre-activation function, for example, the pre-activation function is the ReLu function. P(z) represents two modes of the pre-activation function. For example, the value can be set to {1, -1}, then the number of linear regions is the total number of non-empty sets, which represents the expressive power of the neural network.

[0058] In step 203, the influence of each parameter on the convergence speed value of the neural network to be processed is calculated according to the convergence speed value H0.

[0059] For example, the jth parameter in the neural network to be processed is deleted to obtain an intermediate neural network, where 1≤j≤N, and N is the total number of parameters.

[0060] Next, the adversarial sample is input into the intermediate neural network, and the convergence speed value Hj of the intermediate neural network is calculated.

[0061] Then, according to the convergence speed value H0 and the convergence speed value Hj, the influence of the jth parameter on the convergence speed value of the neural network to be processed is determined. For example, the influence of the jth parameter on the convergence speed value of the neural network to be processed is the absolute value of the difference between the convergence speed value H0 and the convergence speed value Hj.

[0062] In some embodiments, the convergence speed value of the neural network is a NTK (Neural Tangent Kernel) value.

[0063] For example, the NTK value is

[0064]

[0065] in, is the gradient of the target loss function with respect to the parameter w.

[0066] return Figure 1 In step 104, all parameters are sorted according to the influence of the expression ability values ​​according to the first sorting rule to obtain a first sorting result.

[0067] In some embodiments, the first sorting rule is to sort all parameters in order from small to large according to the influence of the expression ability value.

[0068] In other words, the smaller the influence of the expressiveness value, the higher the ranking of the corresponding parameter.

[0069] In step 105, all parameters are sorted according to the influence of the convergence speed values ​​according to a second sorting rule to obtain a second sorting result, wherein the first sorting rule is opposite to the second sorting rule.

[0070] In some embodiments, the second sorting rule is to sort all parameters in descending order according to the influence of the convergence speed value.

[0071] In other words, the greater the influence of the convergence speed value, the higher the ranking of the corresponding parameter.

[0072] In step 106, the pruning evaluation effect of each parameter is determined according to the sequence number of each parameter in the first sorting result and the sequence number of each parameter in the second sorting result.

[0073] In some embodiments, the pruning evaluation effect of each parameter is represented by a weighted sum of the sequence number of each parameter in the first sorting result and the sequence number of each parameter in the second sorting result.

[0074] For example, the weight of the sequence number in the first sorting result and the weight of the sequence number in the second sorting result may be the same or different. In some embodiments, the weight of the sequence number in the first sorting result is 1.

[0075] For example, the influence of parameter a on the expressiveness value of the neural network to be processed ranks fifth in the first sorting result, and the influence of parameter a on the convergence speed value of the neural network to be processed ranks second in the second sorting result, then the weighted sum of parameter a is 5+2=7.

[0076] In step 107, a preset number of parameters with the worst pruning evaluation effects are deleted to obtain a target neural network.

[0077] For example, when the first sorting rule is to sort all parameters in order of the influence of the expression ability value from small to large, and the second sorting rule is to sort all parameters in order of the influence of the convergence speed value from large to small, the preset number of parameters with the largest weighted sum are deleted. In other words, in this sorting method, the larger the weighted sum of the parameter, the worse the pruning evaluation effect of the parameter.

[0078] For example, there are 20 parameters in total. When the first sorting rule is to sort all parameters in order of influence from small to large according to the expression ability value, and the second sorting rule is to sort all parameters in order of influence from large to small according to the convergence speed value, the influence of parameter b on the expression ability value of the neural network to be processed is ranked 19th in the first sorting result, and the influence of parameter b on the convergence speed value of the neural network to be processed is ranked 20th in the second sorting result, then the weighted sum of parameter b is 19+20=39. If the weighted sum of parameter b is the largest among the 20 parameters, it indicates that the pruning evaluation effect of parameter b is the worst, in which case parameter b is deleted.

[0079] It should be noted here that the preset number can be flexibly set according to the usage scenario. For example, in the mobile phone usage scenario, the preset number is large, that is, by deleting more parameters, the operating efficiency of the target neural network is improved. In the server usage scenario, the preset number is small, that is, by deleting smaller parameters, the calculation accuracy of the target neural network is improved.

[0080] It should also be noted that if different sorting rules are used, that is, all parameters are sorted in order from large to small according to the influence of the expression ability value to obtain the first sorting result, and all parameters are sorted in order from small to large according to the influence of the convergence speed value to obtain the second sorting result, then after determining the weighted sum of each parameter according to the serial number of each parameter in the first order result and the serial number in the second order result, the preset number of parameters with the smallest weighted sum are deleted to obtain the target neural network. In other words, under this sorting method, the smaller the weighted sum of the parameter, the worse the pruning evaluation effect of the parameter.

[0081] For example, there are 20 parameters in total. If all parameters are sorted in descending order of the influence of the expression ability value to obtain the first sorting result, and all parameters are sorted in descending order of the influence of the convergence speed value to obtain the second sorting result, the influence of parameter b on the expression ability value of the neural network to be processed is ranked second in the first sorting result, and the influence of parameter b on the convergence speed value of the neural network to be processed is ranked first in the second sorting result, then the weighted sum of parameter b is 2+1=3. If the weighted sum of parameter b is the smallest among the 20 parameters, it indicates that the pruning evaluation effect of parameter b is the worst, and in this case, parameter b is deleted.

[0082] In step 108, adversarial training is performed using the target neural network.

[0083] In the neural network search method provided in the above embodiment of the present disclosure, according to the influence of each parameter in the neural network to be processed on the expression ability value and the convergence speed value of the neural network to be processed, the parameters to be deleted are selected and deleted (i.e., pruning), so that a neural network structure with high robustness can be searched, thereby better solving the security problem of the neural network.

[0084] Figure 3 FIG. 1 is a schematic diagram of the structure of a neural network search device according to an embodiment of the present disclosure. Figure 3 As shown, the neural network search device includes a first processing module 31, a second processing module 32, a third processing module 33 and a fourth processing module 34.

[0085] The first processing module 31 is configured to process the training sample using the pre-training model to generate disturbance information, and superimpose the training sample and the disturbance information to generate an adversarial sample.

[0086] In some embodiments, the first processing module 31 processes the training sample using a preset loss function to generate disturbance information. The loss function is associated with the predicted probability value of the pre-trained model on the correct classification label of the training sample and the predicted probability value of the pre-trained model on the wrong classification label of the training sample.

[0087] For example, the loss function As shown in the above formula (1).

[0088] The second processing module 32 is configured to use the adversarial sample to determine the influence of each parameter of all parameters in the neural network to be processed on the expression ability value and the convergence speed value of the neural network to be processed.

[0089] In some embodiments, the second processing module 32 inputs the adversarial sample into the neural network to be processed, and calculates the expression ability value R0 and the convergence speed value H0 of the neural network to be processed.

[0090] Next, the second processing module 32 calculates the influence of each parameter of all parameters in the neural network to be processed on the expression ability value of the neural network to be processed according to the expression ability value R0.

[0091] For example, the jth parameter in the neural network to be processed is deleted to obtain an intermediate neural network, where 1≤j≤N, and N is the total number of parameters. The adversarial sample is input into the intermediate neural network, and the expression ability value Rj of the intermediate neural network is calculated.

[0092] Then, the influence of the jth parameter on the expression value of the neural network to be processed is determined according to the expression value R0 and the expression value Rj. For example, the influence of the jth parameter on the expression value of the neural network to be processed is the absolute value of the difference between the expression value R0 and the expression value Rj.

[0093] In some embodiments, the expression ability value of the neural network is the linear region value of the neural network. For example, the definition of the linear region value is as shown in the above formula (2).

[0094] Next, the second processing module 32 calculates the influence of each parameter on the convergence speed value of the neural network to be processed according to the convergence speed value H0.

[0095] For example, the jth parameter in the neural network to be processed is deleted to obtain an intermediate neural network, where 1≤j≤N, and N is the total number of parameters. The adversarial sample is input into the intermediate neural network, and the convergence speed value Hj of the intermediate neural network is calculated.

[0096] Then, according to the convergence speed value H0 and the convergence speed value Hj, the influence of the jth parameter on the convergence speed value of the neural network to be processed is determined. For example, the influence of the jth parameter on the convergence speed value of the neural network to be processed is the absolute value of the difference between the convergence speed value H0 and the convergence speed value Hj.

[0097] In some embodiments, the convergence rate value of the neural network is the NTK value.

[0098] For example, the NTK value is as shown in the above formula (3).

[0099] The third processing module 33 is configured to sort all the parameters according to the influence of the expression ability value according to the first sorting rule to obtain a first sorting result, and sort all the parameters according to the influence of the convergence speed value according to the second sorting rule to obtain a second sorting result, wherein the first sorting rule is opposite to the second sorting rule, and the pruning evaluation effect of each parameter is determined according to the serial number of each parameter in the first sorting result and the serial number in the second sorting result.

[0100] In some embodiments, the first sorting rule is to sort all parameters in order from small to large according to the influence of the expression ability value, and the second sorting rule is to sort all parameters in order from large to small according to the influence of the convergence speed value.

[0101] It should be noted that the smaller the influence of the expression ability value, the higher the ranking of the corresponding parameter. The greater the influence of the convergence speed value, the higher the ranking of the corresponding parameter.

[0102] In some embodiments, the pruning evaluation effect of each parameter is represented by a weighted sum of the sequence number of each parameter in the first sorting result and the sequence number of each parameter in the second sorting result.

[0103] For example, the weight of the sequence number in the first sorting result and the weight of the sequence number in the second sorting result may be the same or different. In some embodiments, the weight of the sequence number in the first sorting result is 1.

[0104] For example, the influence of parameter a on the expressiveness value of the neural network to be processed ranks fifth in the first sorting result, and the influence of parameter a on the convergence speed value of the neural network to be processed ranks second in the second sorting result, then the weighted sum of parameter a is 5+2=7.

[0105] The fourth processing module 34 is configured to delete a preset number of parameters with the worst pruning evaluation effects to obtain a target neural network, and use the target neural network to perform adversarial training.

[0106] For example, when the first sorting rule is to sort all parameters in order of the influence of the expression ability value from small to large, and the second sorting rule is to sort all parameters in order of the influence of the convergence speed value from large to small, the preset number of parameters with the largest weighted sum are deleted. In other words, in this sorting method, the larger the weighted sum of the parameter, the worse the pruning evaluation effect of the parameter.

[0107] It should also be noted that if different sorting rules are used, that is, all parameters are sorted in order from large to small according to the influence of the expression ability value to obtain the first sorting result, and all parameters are sorted in order from small to large according to the influence of the convergence speed value to obtain the second sorting result, then after determining the weighted sum of each parameter according to the serial number of each parameter in the first order result and the serial number in the second order result, the preset number of parameters with the smallest weighted sum are deleted to obtain the target neural network. In other words, under this sorting method, the smaller the weighted sum of the parameter, the worse the pruning evaluation effect of the parameter.

[0108] Figure 4 FIG. 1 is a schematic diagram of the structure of a neural network search device according to another embodiment of the present disclosure. Figure 4 As shown, the neural network search device includes a memory 41 and a processor 42.

[0109] The memory 41 is used to store instructions. The processor 42 is coupled to the memory 41. The processor 42 is configured to execute the instructions stored in the memory to implement the following. Figure 1-2 The method of any one of the embodiments.

[0110] like Figure 4As shown, the neural network search device also includes a communication interface 43 for information exchange with other devices. At the same time, the neural network search device also includes a bus 44, through which the processor 42, the communication interface 43, and the memory 41 communicate with each other.

[0111] The memory 41 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory. The memory 41 may also be a memory array. The memory 41 may also be divided into blocks, and the blocks may be combined into virtual volumes according to certain rules.

[0112] In addition, the processor 42 may be a central processing unit (CPU), or may be an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present disclosure.

[0113] The present disclosure also relates to a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, which are executed by a processor to implement the following Figure 1 The method of any one of the embodiments.

[0114] In some embodiments, the functional unit module described above can be implemented as a general-purpose processor, a programmable logic controller (PLC), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components or any appropriate combination thereof for performing the functions described in the present disclosure.

[0115] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0116] The description of the present disclosure is given for the purpose of illustration and description, and is not intended to be exhaustive or to limit the present disclosure to the disclosed form. Many modifications and variations will be apparent to those of ordinary skill in the art. The embodiments are selected and described in order to better illustrate the principles and practical applications of the present disclosure, and to enable those of ordinary skill in the art to understand the present disclosure and thereby design various embodiments with various modifications suitable for specific uses.

Claims

1. A neural network search method, applied to face recognition or image classification retrieval, include: Using the pre-trained model to process the training samples to generate perturbation information; Superimposing the training sample and the disturbance information to generate an adversarial sample; Using the adversarial sample, determine the influence of each parameter of all parameters in the neural network to be processed on the expression ability value and the convergence speed value of the neural network to be processed; According to the first sorting rule, all the parameters are sorted according to the influence of the expression ability value to obtain a first sorting result; According to a second sorting rule, sorting all the parameters according to the influence of the convergence speed value to obtain a second sorting result, wherein the first sorting rule is opposite to the second sorting rule; Determine a pruning evaluation effect of each parameter according to a sequence number of each parameter in the first sorting result and a sequence number of each parameter in the second sorting result; Deleting a preset number of parameters with the worst pruning evaluation effects to obtain a target neural network; Using the target neural network to perform adversarial training; Wherein, using the adversarial sample, determining the influence of each parameter of all parameters in the neural network to be processed on the expression ability value and the convergence speed value of the neural network to be processed includes: Input the adversarial sample into the neural network to be processed, and calculate the expression ability value R0 and the convergence speed value H0 of the neural network to be processed, wherein the expression ability value is the linear region value of the neural network to be processed, and the convergence speed value is the neural tangent kernel NTK value; According to the expression ability value R0, calculating the influence of each parameter of all parameters in the neural network to be processed on the expression ability value of the neural network to be processed; According to the convergence speed value H0, the influence of each parameter on the convergence speed value of the neural network to be processed is calculated.

2. The method according to claim 1, in, According to the expression capability value R0, calculating the influence of each parameter of all parameters in the neural network to be processed on the expression capability value of the neural network to be processed includes: Deleting the jth parameter in the neural network to be processed to obtain an intermediate neural network, where 1≤j≤N, N is the total number of parameters; Input the adversarial sample into the intermediate neural network, and calculate the expression ability value Rj of the intermediate neural network; According to the expression ability value R0 and the expression ability value Rj, the influence of the jth parameter on the expression ability value of the neural network to be processed is determined.

3. The method according to claim 2, in, The influence of the jth parameter on the expression ability value of the neural network to be processed is the absolute value of the difference between the expression ability value R0 and the expression ability value Rj.

4. The method according to claim 1, in, According to the convergence speed value H0, calculating the influence of each parameter on the convergence speed value of the neural network to be processed includes: Deleting the jth parameter in the neural network to be processed to obtain an intermediate neural network, where 1≤j≤N, N is the total number of parameters; Input the adversarial sample into the intermediate neural network, and calculate the convergence speed value Hj of the intermediate neural network; According to the convergence speed value H0 and the convergence speed value Hj, the influence of the jth parameter on the convergence speed value of the neural network to be processed is determined.

5. The method according to claim 4, in, The influence of the jth parameter on the convergence speed value of the neural network to be processed is the absolute value of the difference between the convergence speed value H0 and the convergence speed value Hj.

6. The method according to claim 5, in, The NTK value is in, is the gradient of the target loss function with respect to the parameter w.

7. The method according to claim 1, in, Using the pre-trained model to process the training samples includes: The training samples are processed using a preset loss function to generate disturbance information; The loss function is associated with the predicted probability value of the pre-trained model on the correct classification label of the training sample and the predicted probability value of the pre-trained model on the incorrect classification label of the training sample.

8. The method according to claim 7, in, The loss function for: in, is the training sample, The disturbance information generated is is the predicted probability value of the pre-trained model on the correct classification label of the training sample, is the predicted probability value of the pre-trained model on the i-th misclassified label of the training sample, is the confidence value, and max is the maximum value function.

9. The method according to any one of claims 1 to 8, in, The first sorting rule is to sort all the parameters in ascending order according to the influence of the expression ability value; The second sorting rule is to sort all the parameters in descending order according to the influence of the convergence speed value.

10. The method according to claim 9, in, The pruning evaluation effect of each parameter is represented by a weighted sum of the sequence number of each parameter in the first sorting result and the sequence number of each parameter in the second sorting result.

11. The method according to claim 10, in, The deleting of the preset number of parameters with the worst pruning evaluation effect comprises: Delete a preset number of parameters with the largest weighted sum.

12. The method according to claim 10, in, The weight of the sequence number in the first sorting result is the same as the weight of the sequence number in the second sorting result.

13. The method according to claim 12, in, The weight of the sequence number in the first sorting result and the weight of the sequence number in the second sorting result are 1.

14. A neural network search device, applied to face recognition or image classification retrieval, include: A first processing module is configured to process the training sample using the pre-training model to generate disturbance information, and superimpose the training sample and the disturbance information to generate an adversarial sample; The second processing module is configured to use the adversarial sample to determine the influence of each parameter of all parameters in the neural network to be processed on the expression ability value and the convergence speed value of the neural network to be processed, wherein the adversarial sample is input into the neural network to be processed, and the expression ability value R0 and the convergence speed value H0 of the neural network to be processed are calculated, wherein the expression ability value is the linear region value of the neural network to be processed, and the convergence speed value is the neural tangent kernel NTK value, and according to the expression ability value R0, the influence of each parameter of all parameters in the neural network to be processed on the expression ability value of the neural network to be processed is calculated, and according to the convergence speed value H0, the influence of each parameter on the convergence speed value of the neural network to be processed is calculated; A third processing module is configured to sort all the parameters according to the influence of the expression ability value according to the first sorting rule to obtain a first sorting result, and sort all the parameters according to the influence of the convergence speed value according to the second sorting rule to obtain a second sorting result, wherein the first sorting rule is opposite to the second sorting rule, and the pruning evaluation effect of each parameter is determined according to the sequence number of each parameter in the first sorting result and the sequence number in the second sorting result; The fourth processing module is configured to delete a preset number of parameters with the worst pruning evaluation effects to obtain a target neural network, and use the target neural network to perform adversarial training.

15. A neural network search device, include: a memory configured to store instructions; A processor is coupled to the memory, and the processor is configured to execute the method according to any one of claims 1 to 13 based on instructions stored in the memory.

16. A computer-readable storage medium, in, The computer-readable storage medium stores computer instructions, and when the instructions are executed by a processor, the method according to any one of claims 1 to 13 is implemented.

Citation Information

Patent Citations

  • Feature filtering defense method for deep reinforcement learning model

    CN111600851A

  • Adaptive high-precision compression method and system for convolutional neural network model

    CN113011570A