A method for constructing the critical path of a deep network based on iterative optimization of gate parameters
By adding gate parameters to the deep learning convolutional network model and building a critical path model, the problem of huge structure and poor interpretability of deep learning networks is solved, and structure simplification and recognition accuracy are improved.
Patent Information
- Application Number
- CN202211461783.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-22
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-11-22
AI Technical Summary
The network structure of the deep learning convolutional network model is huge, resulting in high storage space and computing resources consumption, and the evolution law of information transmission is unclear, resulting in poor interpretability.
The deep network critical path construction method based on gate parameters is adopted. By adding gate parameters to the deep convolution neural network model, the importance of each convolution kernel is calculated, the convolution kernel with high contribution is selected, the critical path model of the deep network is constructed, and only the gate parameters are updated and optimized.
The deep network structure is simplified, the interpretability of the mechanism of the deep learning network structure is improved, the covariate offset is reduced, and the recognition accuracy of the network model is improved.
Smart Images

Figure CN115759200B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a method for constructing a critical path of a deep network based on iterative optimization of gate parameters. Background Art
[0002] Although deep learning technology has achieved many good results, there are still the following problems: 1. The network structure of deep learning convolutional network models is huge, resulting in high storage space and computing resource consumption, making it difficult to implement on various hardware platforms. For currently mainstream networks, such as VGG16, the number of parameters is more than 130 million, occupying more than 500 MB of space, and more than 30 billion floating-point operations are required to complete an image recognition task. 2. The structure of deep neural networks is complex and huge, and the law of information transmission and evolution in the network is not clear, resulting in poor interpretability of intelligent algorithms based on deep learning, unclear mechanism of intelligent decision-making, and users cannot clearly know the internal decision-making process and basis of deep learning convolutional network models, bringing potential risks to practical applications. Therefore, in recent years, research on the transparency, interpretability, and credibility of deep learning convolutional network models has gradually increased.
[0003] Aiming at the problem of the complex structure and information evolution law of deep learning networks, this application studies a method for constructing a critical path of a deep learning network, analyzes the key neuron nodes and paths of the deep network, quantifies the contribution degree of neurons and simplifies the deep network, thereby improving the interpretability of the action mechanism of the deep learning network structure. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method for constructing a critical path of a deep network based on iterative optimization of gate parameters, which has a simple structure and reasonable design, calculates the importance of each convolution kernel based on gate parameters, selects convolution kernels with large contribution degrees to obtain a critical path model of the deep network, and when repairing the constructed critical path model of the deep network, only updates and optimizes the gate parameters, does not update other parameters of the deep network model, and does not change the parameters and structure of the deep network model; the repair reduces covariate shift and has good use effects.
[0005] To solve the above technical problem, the technical solution adopted by the present invention is: A method for constructing a critical path of a deep network based on iterative optimization of gate parameters, characterized by comprising the following steps:
[0006] Step 1, obtain a data set: collect multiple images to obtain a data set W=(X, Y), where X represents input data and Y represents data labels, and divide the data set W into a training set, a validation set, and a test set according to a ratio;
[0007] Step 2. Construct a deep convolutional neural network model: Define the objective function of the deep convolutional neural network model, use the training set as the input of the deep convolutional neural network model, and solve the optimal parameters of the deep convolutional neural network model to complete the training of the deep convolutional neural network model;
[0008] Step 3. Add gate parameters to the trained deep convolutional neural network model:
[0009] Step 301. Add gate parameters φ i,j , φ i,j to the convolutional kernels of each layer of the trained deep convolutional neural network model, where the gate parameter φ i,j is (1, C i,j , 1, 1), and C i,j represents the number of image channels of the gate parameter φ i,j . 1 ≤ i ≤ n, where n represents the number of network layers, and 1 ≤ j ≤ m, where m represents the number of convolutional kernels in the i-th layer;
[0010] Step 302. Use the training set to retrain the deep convolutional neural network model with the added gate parameter φ i,j , and use the validation set to verify the deep convolutional neural network model with the added gate parameter φ i,j to obtain the retrained deep convolutional neural network model and the gate parameter φ i,j ;
[0011] Step 4. Construct a critical path model of the deep network based on the gate parameter:
[0012] Step 401. Calculate the importance of each convolutional kernel based on the gate parameter: The computer calculates the importance Θ(φ i,j ) of the j-th convolutional kernel in the i-th layer of the retrained deep convolutional neural network model according to the formula Θ(φ i,j ) = KL(ΔL(φ Ω ||ΔL i,j )(0)), where KL(·) represents the KL divergence calculation, ΔL(φ i,j ) represents the difference between the predicted value and the true value of the retrained deep convolutional neural network model in Step 3, and ΔL Ω (0) represents the difference between the predicted value of the deep convolutional neural network model completed in training in Step 2 and the true value Y. Ω represents the network model parameters other than φ i,j ;
[0013] Step 402. Set the gate parameters with low convolutional kernel importance to zero using a threshold: If Θ(φ i,j ) < Θ(φ), then set the gate parameter φ i,jAssign it to 0, where Θ(φ) represents the optimization threshold, and obtain the critical path model of the deep network;
[0014] Step 403: Repair the constructed critical path model of the deep network: Input the training set into the critical path model of the deep network for iterative training, update and optimize the gate parameters of the critical path model of the deep network through the gradient descent algorithm. The zeroed gate parameters remain zero and do not participate in the update. During the iterative training of the gate parameters, the loss function of the critical path model of the deep network is L'=(1 - α)L1 + αL2, where α represents the weight, X k represents the k-th data in the training set, Y k represents X k corresponding sample label, 1 ≤ k ≤ N, N represents the number of samples in the training set, Δ represents the label smoothing factor, p(X k , θ - ) represents the predicted probability of the training sample X k , θ - represents the parameters of the critical path model;
[0015] Step 404: Use the validation set to verify the critical path model of the deep network that has completed the update and optimization in Step 403, and obtain the repaired critical path model of the deep network;
[0016] Step 405: Input the test set into the repaired critical path model of the deep network to obtain the classification result of the test set.
[0017] In the above method for constructing the critical path of a deep network based on iterative optimization of gate parameters, it is characterized in that: in Step 1, the image is a radar image. First, extract the foreground image from the collected radar image, then perform Gaussian blur processing, and then extract the feature vector to obtain the data set W=(X, Y).
[0018] In the above method for constructing the critical path of a deep network based on iterative optimization of gate parameters, it is characterized in that: in Step 302, add the gate parameter φ i,j to the loss function of the network model where L(X, Y:θ) represents the cross-entropy loss function, θ represents the parameters of the network model, represents the sparse constraint of the gate parameter, and λ represents the weight of the sparse constraint.
[0019] In the above method for constructing the critical path of a deep network based on iterative optimization of gate parameters, it is characterized in that: in Step 401, L represents the loss function, δL represents the increment of the loss function L, and δφ i,j represents the increment of the gate parameter φ i,j ; It represents the sample prediction result, and Y represents the sample label.
[0020] The present invention has the following advantages compared with the prior art:
[0021] 1. The structure of the present invention is simple, reasonably designed, and convenient to implement and operate.
[0022] 2. The present invention adds a gate parameter to the trained deep convolutional neural network model, calculates the importance of each convolutional kernel based on the gate parameter, obtains convolutional kernel nodes with large contribution and strong information interaction, can improve the interpretability of the action mechanism of the deep network structure, and selecting convolutional kernels with large contribution can obtain the key path model of the deep network, playing a role in simplifying the deep network.
[0023] 3. When repairing the key path model of the constructed deep network, the present invention only updates and optimizes the gate parameter of the key path model of the deep network through the gradient descent algorithm. The zeroed gate parameter remains zero and does not participate in the update, and does not update other parameters of the deep network model, without changing the parameters and structure of the deep network model.
[0024] 4. When repairing and training the network model, the present invention constructs a new loss function by superimposing two loss functions. The loss function L1 adds the constraint condition of the gate parameter φ i,j The loss function L2 introduces a label smoothing factor to smooth the label pair, reducing the noise impact in the label, improving the accuracy of the label pair data, further reducing the covariate shift, and improving the recognition accuracy of the network model.
[0025] In summary, the present invention has a simple structure and reasonable design. It calculates the importance of each convolutional kernel based on the gate parameter, selects convolutional kernels with large contribution to obtain the key path model of the deep network. When repairing the constructed key path model of the deep network, only the gate parameter is updated and optimized, and other parameters of the deep network model are not updated, without changing the parameters and structure of the deep network model; the repair reduces the covariate shift and has good use effects.
[0026] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. Brief Description of the Drawings
[0027] Figure 1 It is the circuit principle block diagram of the present invention. Detailed Embodiments
[0028] Next, the method of the present invention will be further described in detail in combination with the drawings and embodiments of the present invention.
[0029] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0030] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0032] For ease of description, spatial relative terms such as "above", "on top of", "on the upper surface", "above", etc. can be used herein to describe the spatial positional relationship between a device or feature shown in the figure and other devices or features. It should be understood that the spatial relative terms are intended to include different orientations in use or operation in addition to the orientation of the device described in the figure. For example, if the device in the figure is inverted, the device described as "above other devices or structures" or "on top of other devices or structures" will then be positioned "below other devices or structures" or "beneath other devices or structures". Thus, the exemplary term "above" can include both the orientation of "above" and "below". The device can also be positioned in other different ways (rotated 90 degrees or in other orientations), and corresponding interpretations are made for the spatial relative descriptions used herein.
[0033] As Figure 1 shown, the present invention includes the following steps:
[0034] Step 1: Obtain a data set: Collect multiple images to obtain a data set W = (X, Y), where X represents the input data and Y represents the data label. Divide the data set W into a training set, a validation set, and a test set according to a ratio.
[0035] In actual use, the image is a radar image collected according to the application scenario, or the MSTAR dataset is directly used. The MSTAR dataset has been widely used in the research of radar image target recognition.
[0036] For the collected radar image, first extract the foreground image, then perform Gaussian blur processing, and then extract the feature vector. The complexity of the obtained feature vector becomes lower, removing the offset and unimportant features, reducing the complexity of the input information, improving the distribution uniformity rate of the samples to covariates, and improving the model effect.
[0037] The feature vector is saved or stored on an electronic or computer-readable storage medium and / or transmitted to another program or application for further processing or use.
[0038] Step 2: Construct a deep convolutional neural network model: Define the objective function of the deep convolutional neural network model, use the training set as the input of the deep convolutional neural network model, and solve the optimal parameters of the deep convolutional neural network model to complete the training of the deep convolutional neural network model.
[0039] It should be noted that the convolutional neural network model can be selected from the vgg model, resnet-34 model, resnet-50 or resnet-56 model. The deep learning neural network model, deep neural network model, neural network model, network model, and deep network in the following text all refer to the deep convolutional neural network model.
[0040] Step 3: Add gate parameters to the trained deep convolutional neural network model:
[0041] Step 301: Add the gate parameter φ to the convolutional kernels of each layer of the trained deep convolutional neural network model i,j , φ i,j represents the gate parameter added to the j-th convolutional kernel of the i-th convolutional layer. The gate parameter φ i,j is (1, C i , 1, 1), C i represents the number of image channels of the gate parameter φ i,j , where 1 ≤ i ≤ n, n represents the number of network layers, and 1 ≤ j ≤ m, m represents the number of convolutional kernels in the i-th layer.
[0042] The gate parameter is a gradient-updatable variable. The role of the convolutional layer is to extract the features of the image. The input data of the convolutional layer is called the input feature map, and the output data is called the output feature map. The setting of the gate parameter multiplies the output of the convolutional layer by a gate parameter as the input of the next convolutional layer, which is used to identify the attributes of the convolutional kernel and changes the structure of the standard neural network model. The parameter of the gate parameter φ i,j is (1, C i,j, 1, 1), C i,j represents the number of image channels of the i-th gate parameter. In actual use, if it is for the global simplification of the neural network model, the gate parameter φ is added after each convolutional layer i,j , and C i,j ≠1. If it is for the partial simplification of the neural network model, the C of the convolutional layer that does not need to be simplified i,j =1.
[0043] Step 302. Use the training set to retrain the deep convolutional neural network model with the added gate parameter φ i,j . In actual use, the loss function of the network model with the added gate parameter φ i,j is where L(X, Y:θ) represents the cross-entropy loss function, θ represents the parameters of the network model, represents the sparse constraint of the gate parameter, and λ represents the weight of the sparse constraint.
[0044] After training is completed, fix the convolutional kernel parameters, and use the validation set to verify the deep convolutional neural network model with the added gate parameter φ i,j to obtain the retrained deep convolutional neural network model and the gate parameter φ i,j .
[0045] In actual use, use the training set to perform sparse training on the deep convolutional neural network model with the added gate parameter φ i,j . The sparse constraint reduces the dependence between each convolutional kernel, making them more independent and facilitating the further interpretation of the contribution degree of the convolutional kernel. During training, only the gate parameter φ i,j is updated, and each time a subset of the training data is used to train the neural network model.
[0046] Step Four. Construct a key path model of the deep network based on the gate parameter:
[0047] Step 401. Calculate the importance of each convolutional kernel based on the gate parameter: The computer calculates the importance Θ(φ i,j ) of the j-th convolutional kernel in the i-th layer of the retrained deep convolutional neural network model according to the formula Θ(φ i,j ) = KL(ΔL(φ Ω ||ΔL i,j i,j ), where KL(·) represents the KL divergence calculation, ΔL(φ) represents the difference between the predicted value and the true value of the retrained deep convolutional neural network model in Step 3, and ΔL Ω (0) represents the difference between the predicted value of the deep convolutional neural network model after training in Step 2 and the true value Y, and Ω represents except for φ i,jNetwork model parameters other than
[0048] In a deep convolutional neural network, the role of the convolutional kernel is to extract the features of the input radar data. For a convolutional kernel responsible for extracting a certain feature, if it detects its corresponding feature from the upper-layer input, the node is activated, allowing useful feature information to propagate to the next layer. If this node believes that the input does not contain its corresponding feature, the node will be masked, blocking the propagation of useless feature information to the next layer. Therefore, calculating the contribution degree of each convolutional kernel based on the gate parameter, and obtaining convolutional kernel nodes with large contribution degree and strong information interaction is beneficial to realizing the interpretability of the deep network and constructing the key path model.
[0049] Based on the gate parameter φ i,j Calculate the contribution degree of each convolutional kernel in each layer of the deep network. The contribution degree is used as the basis for judging whether the convolutional kernel is a key node. In actual use, L represents the loss function, δL represents the increment of the loss function L, and δφ i,j represents the increment of the gate parameter φ i,j , represents the sample prediction result, and Y represents the sample label.
[0050] Step 402, set the gate parameters with low importance of the convolutional kernel to zero: If Θ(φ i,j ) < Θ(φ), then assign the gate parameter φ i,j to 0. Θ(φ) represents the optimization threshold, and the key path model is obtained. In actual use, set the optimization threshold Θ(φ) to screen the importance of the convolutional kernel. The convolutional kernel with a contribution degree lower than the preset optimization threshold Θ(φ) can be regarded as a "non-critical" node. Therefore, assign its corresponding gate parameter φ i,j to 0. Assigning the gate parameter to 0 is equivalent to deleting the corresponding convolutional kernel. The remaining convolutional kernels are the convolutional kernels that play a key role in processing image information, and the remaining convolutional kernels constitute the key path model of the deep network. The key path model can be used to analyze the evolution of information on the key path of the deep network, thus completing the interpretable elaboration of the deep network.
[0051] Step 403, repair the constructed key path model of the deep network: Input the training set into the key path model of the deep network for iterative training, and update and optimize the gate parameters of the key path model of the deep network through the gradient descent algorithm. The zeroed gate parameters remain zero and do not participate in the update. During the iterative training of the gate parameters, the loss function of the key path model of the deep network is L'=(1 - α)L1 + αL2, where α represents the weight, X k represents the k-th data in the training set, Y k represents Xk The corresponding sample label, 1 ≤ k ≤ N, where N represents the number of samples in the training set, Δ represents the label smoothing factor, and p(X k , θ - ) represents the predicted probability of the training sample X k , and θ - represents the parameters of the critical path model.
[0052] In actual use, the loss value of iterative training is backpropagated, and the gradient of the weights of the critical path model for the interpretability of the deep network in the backpropagation process is extracted. According to the extracted gradient, only the gate parameters of the critical path model of the deep network are updated and optimized through the gradient descent algorithm, and other parameters of the neural network model are not updated.
[0053] It should be noted that pruning the convolutional kernels of the original deep network model is likely to cause information loss, resulting in covariate shift in the optimized neural network model and deteriorating the model performance. Therefore, during the repair training, the optimization of the gate parameter φ i,j is considered, so a constraint condition of the gate parameter φ i,j is added to the loss function L1.
[0054] The loss function L2 can measure the difference between the predicted probability and the true probability of X k . In the network model, the difference between the true probability distribution and the predicted probability distribution is manifested as loss. The smaller the loss, the better the prediction classification effect of the network model. Δ = 0.05. Smoothing the labels reduces the noise impact in the labels, improves the accuracy of the label pairs for the data, further reduces the covariate shift, and improves the recognition accuracy of the network model.
[0055] Step 404: Use the validation set to verify the deep network critical path model that has completed the update and optimization in step 403 to obtain the repaired deep network critical path model;
[0056] Step 405: Input the test set into the repaired deep network critical path model to obtain the classification result of the test set.
[0057] The effectiveness of the present invention can be further verified by the following simulation experiments:
[0058] 1. Experimental conditions and methods
[0059] The hardware platform is: Inter(R) Core(TM) i5-10600K CPU@4.10GHZ,
[0060] 16.0GB RAM;
[0061] The software platform is: Pytorch 1.10;
[0062] 2. Simulation Content and Results
[0063] In this embodiment, the deep network selects the VGG16 neural network model. The VGG16 neural network model can be used for image processing, including the processing of radar image information. Image processing is used in various technical applications. As an example, it can be used for the localization of specific types or kinds of objects in images, such as in computer vision applications for image and video analysis. This embodiment does not make specific limitations.
[0064] Table 1 is a comparison table of the number of convolutional kernel nodes and classification accuracy of each layer of the original deep network and the key path model of the deep network when the VGG16 model is selected for the deep network.
[0065] Table 1 Comparison Table of the Original Deep Network and the Key Path Model of the Deep Network
[0066]
[0067]
[0068] For example, in the first layer, the original number of convolutional kernels is 64. Through contribution degree screening, only 51 convolutional kernels that play a role in the network performance remain. It can be seen from the results that the key path model of the deep network only retains 1452 convolutional kernels of the original neural network model, that is, only 34.37% of the number of nodes is retained. After simplification and repair, the recognition accuracy rate of the key path model of the deep network on the data set is still 92.46%. The key path model is very close to the original deep network model in terms of performance, indicating the reliability of the key path model of the deep network and verifying the effectiveness of the construction method of this application.
[0069] Moreover, from the changes in the number of convolutional kernel nodes of each layer, it can be seen that whether it is the key path model of the deep network or the original deep network model, the number of convolutional kernel nodes in the first layer is relatively large, which is convenient for extracting sufficient information from the input samples for classification; the number of convolutional kernel nodes in the last layer is also relatively large, which is to provide sufficient information for the final classification. In contrast, the number of convolutional kernel nodes in the middle layer is relatively small, indicating that in the process of feature transfer from input to output, the contribution degree of the middle layer nodes to information transfer and extraction is relatively low, and the neural network requires fewer high-level and abstract features for task discrimination, so there are a large number of redundancies in the middle layer nodes. Using the key path can explain the evolution and transfer process of signals in the network model, thereby improving the interpretability of the action mechanism of the deep learning network structure.
[0070] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.
[0071] As described above, these are only embodiments of the present invention and do not impose any limitations on the present invention. Any simple modifications, changes, and equivalent structural changes made to the above embodiments based on the technical essence of the present invention still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A method for constructing a critical path of a deep network based on iterative optimization of gate parameters, characterized in that: Including the following steps: Step 1, obtaining a data set: collecting multiple images to obtain a data set W = (X, Y), where X represents input data and Y represents data labels, and dividing the data set W into a training set, a validation set, and a test set according to a ratio; Step 2, constructing a deep convolutional neural network model: defining an objective function of the deep convolutional neural network model, using the training set as the input of the deep convolutional neural network model, and solving the optimal parameters of the deep convolutional neural network model to complete the training of the deep convolutional neural network model; Step 3, adding a gate parameter to the trained deep convolutional neural network model: Step 301: Add the gate parameter φ to the convolutional kernels of each layer in the trained deep convolutional neural network model i,j , where φ i,j represents the gate parameter added to the j-th convolutional kernel in the i-th convolutional layer. The gate parameter φ i,j is (1, C i,j , 1, 1), where C i,j represents the number of image channels of the gate parameter φ i,j . 1 ≤ i ≤ n, where n represents the number of network layers, and 1 ≤ j ≤ m, where m represents the number of convolutional kernels in the i-th layer; Step 302: Retrain the deep convolutional neural network model with the gate parameter φ added using the training set, and verify the deep convolutional neural network model with the gate parameter φ added using the validation set to obtain the retrained deep convolutional neural network model and the gate parameter φ i,j Step 302: Retrain the deep convolutional neural network model with the gate parameter φ added using the training set, and verify the deep convolutional neural network model with the gate parameter φ added using the validation set to obtain the retrained deep convolutional neural network model and the gate parameter φ i,j Step 302: Retrain the deep convolutional neural network model with the gate parameter φ added using the training set, and verify the deep convolutional neural network model with the gate parameter φ added using the validation set to obtain the retrained deep convolutional neural network model and the gate parameter φ i,j ; Step 4, constructing a critical path model of the deep network based on the gate parameter: Step 401. Calculate the importance of each convolution kernel based on the gate parameter: The computer calculates the importance Θ(φ i,j ) of the j-th convolution kernel in the i-th layer of the retrained deep convolutional neural network model according to the formula Θ(φ i,j ) = KL(ΔL(φ Ω ||ΔL i,j )(0)), where KL(·) represents the KL divergence calculation, ΔL(φ i,j ) represents the difference between the predicted value and the true value of the retrained deep convolutional neural network model in Step 3, and ΔL Ω (0) represents the difference between the predicted value of the deep convolutional neural network model completed in training in Step 2 and the true value Y. Ω represents the network model parameters other than φ i,j ; Step 402, set the gate parameters with low importance of the convolution kernel to zero using a threshold: If Θ(φ i,j ) < Θ(φ), then assign the gate parameter φ i,j to 0. Θ(φ) represents the optimization threshold, and the critical path model of the deep network is obtained; Step 403: Repair the constructed deep network critical path model: Input the training set into the deep network critical path model for iterative training, update and optimize the gate parameters of the deep network critical path model through the gradient descent algorithm. The zeroed gate parameters remain zero and do not participate in the update. During the iterative training of the gate parameters, the loss function of the deep network critical path model is L' = (1 - α)L1 + αL2, where α represents the weight, X k represents the k-th data in the training set, Y k represents X k corresponding sample label, 1 ≤ k ≤ N, N represents the number of samples in the training set, Δ represents the label smoothing factor, p(X k , θ - ) represents the predicted probability of the training sample X k , θ - represents the parameters of the critical path model; Step 404, validating the updated and optimized deep network critical path model in step 403 using the validation set to obtain a repaired deep network critical path model; Step 405, inputting the test set into the repaired deep network critical path model to obtain the classification result of the test set.
2. A method for constructing a critical path of a deep network based on iterative optimization of gate parameters according to claim 1, characterized in that: In step 1, the image is a radar image. First, the foreground image is extracted from the collected radar image, then Gaussian blur processing is performed, and then feature vectors are extracted to obtain the data set W = (X, Y).
3. A method for constructing a critical path of a deep network based on iterative optimization of gate parameters according to claim 1, characterized in that: In step 302, the gate parameter φ is added. i,j The loss function of the network model where L(X, Y:θ) represents the cross-entropy loss function, θ represents the parameters of the network model, represents the sparse constraint of the gate parameter, and λ represents the weight of the sparse constraint.
4. A method for constructing a critical path of a deep network based on iterative optimization of gate parameters according to claim 1, characterized in that: In step 401, L represents the loss function, δL represents the increment of the loss function L, and δφ i,j represents the increment of the gate parameter φ i,j , represents the sample prediction result, and Y represents the sample label.
Citation Information
Patent Citations
Cutting method of deep convolutional neural network model based on grey correlation analysis
CN110647990A
Deep network model compression method based on discrete coefficient
CN115131646A