Training Method of Network Model, Electronic Device and Computer Readable Storage Medium

By using the differential network architecture in the training of network models, the feature information and training parameters of the front cascade unit are obtained and the current unit is trained, which solves the problem of excessive training time of network model in the existing technology, and achieves the effect of improving search efficiency and accuracy.

CN113988288BActive Publication Date: 2025-05-30SHANGHAI JINSHENG COMM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111183804.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-11
Publication Date
2025-05-30
Estimated Expiration
2041-10-11

AI Technical Summary

Technical Problem

The existing network model training methods take huge time to search network architectures, resulting in too long training time.

Method used

Using a differential network architecture, the current unit is trained by obtaining the feature information output from the first two cascaded units and the training parameters of the previous same type of unit to improve search efficiency and shorten the search time.

Benefits of technology

By quickly learning the information of the previous same type of unit, the accuracy of the network model is improved and the training time of the network model is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113988288B_ABST
    Figure CN113988288B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of network models, and discloses a training method, an electronic device, and a computer-readable storage medium for a network model. The architecture of the network model is a differential network architecture, and the differential network architecture includes a basic unit and a compression unit. The method includes: obtaining the feature information output by the previous two cascaded units and the training parameters of the previous unit of the same type; wherein, the feature information is obtained by training the basic unit and / or the compression unit in the network model based on training samples; inputting the training parameters and the feature information into the current unit, so that the current unit outputs the feature information to the next unit to train the network model. By the above method, the search efficiency of each unit can be improved, the search time can be shortened, and thus the training time of the network model can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of network models, and particularly to a training method for network models, an electronic device, and a computer-readable storage medium. Background Art

[0002] In recent years, the rise of deep learning neural networks has gradually enabled people to discover the great potential of neural networks in a large number of complex tasks such as image recognition and speech processing. The data processing ability of neural networks is extremely powerful, exceeding many traditional computing methods and becoming the mainstream computing method in many fields. The performance of neural networks is closely related to the parameters and weights of the network. The vast majority of related neural network model architectures are designed manually. During the network design process, a large amount of research and experiments are required to try and explore the effects of different network model structures.

[0003] Therefore, the demand for automated search design of network structures has emerged. This is a method in which a computer designs different search spaces according to a pre-established search strategy and automatically searches for the optimal network structure in a specific search space. Summary of the Invention

[0004] The main technical problem to be solved by this application is to provide a training method for network models, an electronic device, and a computer-readable storage medium, which can improve the search efficiency of each unit, shorten the search time, and thus reduce the training time of the network model.

[0005] To solve the above problems, a technical solution adopted by this application is to provide a training method for a network model. The architecture of the network model is a differential network architecture, and the differential network architecture includes a basic unit and a compression unit. The method includes: obtaining the feature information output by the first two cascaded units and the training parameters of the previous unit of the same type; wherein, the feature information is obtained by training the basic unit and / or the compression unit in the network model based on training samples; inputting the training parameters and the feature information into the current unit so that the current unit outputs the feature information to the next unit to train the network model.

[0006] To solve the above problems, another technical solution adopted by this application is to provide an electronic device, which includes a processor and a memory coupled to the processor. A computer program is stored in the memory, and the processor is used to execute the computer program to implement the method provided by the above technical solution.

[0007] To solve the above problems, another technical solution adopted by this application is to provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method provided by the above technical solution.

[0008] The beneficial effects of the present application are as follows: Different from the prior art, a training method for a network model provided by the present application uses the feature information output by the previous two cascaded units and the training parameters of the previous unit of the same type during the training of the current unit to train the current unit, so that the current unit can quickly learn the information of the previous unit of the same type, thereby improving the accuracy of the network model. Moreover, training based on the training parameters of the previous unit of the same type can improve the search efficiency of each unit, shorten the search time, and thus reduce the training time of the network model. Description of the Drawings

[0009] Figure 1 is a schematic flowchart of an embodiment of the training method for the network model provided by the present application;

[0010] Figure 2 is a schematic structural diagram of the differential network architecture provided by the present application;

[0011] Figure 3 is a schematic structural diagram of a single unit in the differential network architecture provided by the present application;

[0012] Figure 4 is a schematic flowchart of another embodiment of the training method for the network model provided by the present application;

[0013] Figure 5 is a schematic flowchart of an embodiment of step 41 provided by the present application;

[0014] Figure 6 is a schematic flowchart of an embodiment of step 412 provided by the present application;

[0015] Figure 7 is a schematic flowchart of an embodiment of step 42 provided by the present application;

[0016] Figure 8 is a schematic flowchart of an embodiment of step 422 provided by the present application;

[0017] Figure 9 is a schematic diagram of an application scenario of the training method for the network model provided by the present application;

[0018] Figure 10 is a schematic flowchart of another embodiment of the training method for the network model provided by the present application;

[0019] Figure 11 is a schematic flowchart of another embodiment of the training method for the network model provided by the present application;

[0020] Figure 12 is a schematic structural diagram of an embodiment of the electronic device provided by the present application;

[0021] Figure 13 It is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by the present application. Detailed implementation manners

[0022] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. It can be understood that the specific embodiments described herein are only used to explain the present application, rather than limiting the present application. In addition, it should be noted that for the convenience of description, only the parts related to the present application are shown in the drawings, rather than all the structures. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0023] Referring to "embodiment" herein means that the specific features, structures or characteristics described in conjunction with the embodiment may be included in at least one embodiment of the present application. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.

[0024] The related mainstream automatic network architecture search methods include the automatic search method based on reinforcement learning, the search method based on evolutionary algorithms, the differentiable search method, etc. The earliest network architecture search method was designed based on the reinforcement learning method. According to the output feedback of the network, the network structure was searched and optimized. This method would consume a huge search time: NASNet based on reinforcement learning required 2000 days of search time on the CIFAR-10 dataset. Many subsequent network architecture search methods, such as the network architecture search method based on evolutionary algorithms and the network architecture search method based on differentiation, made corresponding changes in the search strategy and method to adapt to different task requirements and improved the effect of the model. However, they still required a large amount of search time. For example, AmoebaNet based on evolutionary algorithms required 3150 days of search time on the CIFAR-10 dataset, and DARTS (Differentiable Architecture Search) designed based on the differential method required 1.5 - 4 days of search time on the CIFAR-10 dataset.

[0025] Based on this, the present application proposes the following technical solutions to solve the problem that the search time is too long when training a network model by the network architecture search method.

[0026] Refer to Figure 1 , Figure 1It is a schematic flowchart of an embodiment of the training method of the network model provided by this application. The architecture of the network model is a differential network architecture, and the differential network architecture includes a basic unit and a compression unit. The method includes:

[0027] Step 11: Obtain the feature information output by the first two cascaded units and the training parameters of the previous same-type unit.

[0028] Among them, the feature information is obtained by training the basic unit and / or compression unit in the network model based on the training samples.

[0029] Among them, the training samples can be an image-based training set. The true labels of the image content are marked in the images of the training set.

[0030] Step 12: Input the training parameters and the feature information into the current unit, so that the current unit outputs the feature information to the next unit to train the network model.

[0031] Since the previous same-type unit has obtained more training parameters through training, such as numerous prediction labels, node connection relationships, and weights. When the current unit is trained, having these training parameters is equivalent to having a larger label library and information library. Therefore, not only can the training time be reduced, but also the training accuracy can be improved.

[0032] In some embodiments, the training parameters can be extracted from the previous same-type unit by using the self-distillation method. For example, when the previous same-type unit outputs the feature information, the training parameters such as the prediction labels, node connection relationships, and weights of the previous same-type unit are extracted by using the self-distillation method.

[0033] Refer to Figure 2 for an explanation of the network architecture search for the differential network architecture:

[0034] In this embodiment, the differential network architecture is divided into 6 basic units and 2 compression units. The 6 basic units and 2 compression units are cascaded in the order of 2 basic units and 1 compression unit. As Figure 2 shown, if the 6 basic units and 2 compression units are numbered from 1 to 8, then the units numbered 1 and 2 are basic units, the unit numbered 3 is a compression unit, the units numbered 4 and 5 are basic units, the unit numbered 6 is a compression unit, and the units numbered 7 and 8 are basic units.

[0035] Each basic unit and each compression unit have two inputs. For details, refer to Figure 3 , Figure 3 which is the schematic diagram of the basic structure of a single unit. Among them, the two inputs of the basic unit numbered 1 are the same training sample. That is, when inputting the training sample, the training sample can be copied to form two training samples.

[0036] Two inputs of the basic unit numbered 2, one of which is a training sample and the other is the feature information output by the basic unit numbered 1. In some embodiments, the training sample may be an image, and the feature information output by each unit is a feature map.

[0037] Two inputs of the compression unit numbered 3, one of which is the feature information output by the basic unit numbered 1 and the other is the feature information output by the basic unit numbered 2.

[0038] Two inputs of the basic unit numbered 4, one of which is the feature information output by the basic unit numbered 2 and the other is the feature information output by the compression unit numbered 3.

[0039] Two inputs of the basic unit numbered 5, one of which is the feature information output by the compression unit numbered 3 and the other is the feature information output by the basic unit numbered 4.

[0040] Two inputs of the compression unit numbered 6, one of which is the feature information output by the basic unit numbered 4 and the other is the feature information output by the basic unit numbered 5.

[0041] Two inputs of the basic unit numbered 7, one of which is the feature information output by the basic unit numbered 5 and the other is the feature information output by the compression unit numbered 6.

[0042] Two inputs of the basic unit numbered 8, one of which is the feature information output by the compression unit numbered 6 and the other is the feature information output by the basic unit numbered 7.

[0043] Among them, each unit includes 7 nodes. As Figure 3 shown, it includes two input nodes (Input 1 and Input 2), 4 intermediate nodes and one output node. The input of each node is the output of all its predecessor nodes. That is to say, if it is Node 1, the input of Node 1 is the output of Node 0. If it is Node 3, the input of Node 3 is the output of Node 0, Node 1 and Node 2. The two input nodes are connected to all intermediate nodes.

[0044] During the training process (i.e., the search process), the appropriate edges are selected to connect nodes to nodes. Here, the edges also include the corresponding operations between nodes. Each unit can be regarded as a directed acyclic graph with 6 nodes. The edges between the nodes represent possible operations, such as 3×3sep convolution. Initially, the specific operations are unknown. The available operations are 8 types, such as no operation, 3×3 max pooling, 3×3 average pooling, residual connection, 3×3sep convolution, 5×5sep convolution, 3×3 dilated convolution, and 5×5 dilated convolution.

[0045] The search space is continuously relaxed, and each edge is regarded as a mixture of all sub-operations (softmax weight superposition).

[0046] Then, joint optimization is performed to update the edge hyperparameters (i.e., the architecture search task) and the architecture-independent network parameters on the sub-operation mixture probabilities.

[0047] After optimization, directly select the sub-operation with the highest probability.

[0048] Among them, through the above method, using the network model with a differential network architecture, the basic unit and the compression unit with the best internal node connection relationship and weights are trained. Then, these units are connected to form a large network, and the hyperparameter layers can control how many units are connected. For example, layers = 20 means that 20 cells (units) are connected in sequence.

[0049] Among them, the sizes of the input and output feature maps of the basic unit are the same.

[0050] The size of the feature map output by the compression unit is reduced by half compared to the size of the input feature map.

[0051] In some embodiments, the structures of the basic unit and the compression unit are the same, but the operations are different. Among them, for the convolutional network, the inputs of the two input nodes are the outputs of the previous two layers (layers) of units. For the recurrent network, the input is the input of the current layer and the state of the previous layer.

[0052] The input of each intermediate node is obtained by summing the outputs of its predecessor nodes through the corresponding relationship of the edges. The predecessor node refers to the node connected to the input end of each intermediate node.

[0053] The output of the output node is obtained by combining the outputs of each intermediate node using the concat function.

[0054] Edges represent operations (such as 3×3 convolutions). During the process of converging to obtain a structure, all edges between two nodes will exist and participate in training, and finally, a weighted average is taken. This weight is what the network model of the differential network architecture needs to train. The desired result is the edge with the best effect, so its weight is the largest.

[0055] In some embodiments, a group of edges can be processed by softmax normalization. It can be known that each operation corresponds to a weight value, that is, the above-mentioned training parameter. We call these weight values a weight matrix. The larger the weight value, the more important the operation represented in this group of edges.

[0056] It is hoped to converge to a weight matrix in the end. Among the edges in this matrix, the larger the weight value, the better the effect after being retained.

[0057] During the training process, through the previously defined search space, the weight matrix is optimized by gradient descent. At this time, a weight matrix is trained so that the edges with large weights are retained. Therefore, after the differential network structure converges, a process of generating the final unit is still required.

[0058] For each intermediate node, at most retain the connection relationships with the two strongest predecessor nodes; for the edges between two nodes, only retain the edge with the largest weight. Assume that a node has three predecessors. The edge with the highest weight between the predecessor node and the current node represents the strength of that predecessor node, and select the two strongest predecessor nodes. Only retain the edge with the largest weight between nodes.

[0059] In this embodiment, when using the differential network architecture to search for the network model architecture, during the training of the current unit, by using the feature information output by the previous two cascaded units and the training parameters of the previous unit of the same type, the current unit is trained so that the current unit can quickly learn the information of the previous unit of the same type, thereby improving the accuracy of the network model. And training based on the training parameters of the previous unit of the same type can improve the search efficiency of the current unit, shorten the search time, and thus reduce the training time of the network model.

[0060] When training the model in the above manner, the following technical solutions can also be adopted. Specifically, refer to Figure 4 , Figure 4 is a schematic flowchart of another embodiment of the training method of the network model provided by this application. The method includes:

[0061] Step 41: When the current unit is a basic unit, determine the first loss value between the basic unit and the previous basic unit.

[0062] After training the current basic unit, the loss value can be calculated to determine whether the basic unit is close to the previous basic unit, so as to determine whether the current basic unit has learned the knowledge of the previous basic unit. In this way, in the subsequent iterative training process, the basic unit can quickly approach the previous basic unit.

[0063] After training the current basic unit using the training parameters and feature information, it is necessary to determine the similarity between the current basic unit and the previous basic unit, from which it can be determined whether the current basic unit approaches the previous basic unit.

[0064] In some embodiments, referring to Figure 5 , step 41 may be the following process:

[0065] Step 411: Obtain the first feature information output by the previous basic unit and the second feature information output by the basic unit.

[0066] Step 412: Determine the first loss value based on the first feature information and the second feature information.

[0067] In some embodiments, the feature information is a feature map, that is, the first feature information is the first feature map and the second feature information is the second feature map. The number of channels between the first feature maps and the second feature maps of some basic units is different, that is, the feature maps are not equivalent. Combining Figure 2 Explanation: For example, for the feature maps output by the basic unit numbered 2 and the basic unit numbered 4, since the size of the feature map output by the compression unit numbered 3 is halved, the size of the feature map output by the basic unit numbered 4 corresponds to the size of the feature map output by the compression unit numbered 3.

[0068] For example, the present application proposes the following solution. Referring to Figure 6 , step 412 may be the following process:

[0069] Step 4121: Upsample the second feature information.

[0070] The upsampling can adopt an interpolation algorithm, such as the nearest neighbor interpolation method, the bilinear interpolation method, the high-order interpolation method, the edge-based image interpolation algorithm, and the region-based image interpolation algorithm. An appropriate algorithm can be selected according to actual needs for upsampling.

[0071] Step 4122: Determine the first loss value using the first feature information and the upsampled second feature information.

[0072] The way to determine the loss value can be to use the softmax function to normalize the first feature information and the upsampled second feature information respectively to obtain the probability of each classification result in the first feature information and the second feature information.

[0073] Compare the probabilities of each classification result in the first feature information and the second feature information to obtain a first loss value.

[0074] Step 42: When the current unit is a compression unit, determine a second loss value between the compression unit and the previous compression unit.

[0075] After training the current compression unit, it is possible to determine whether the compression unit is close to the previous compression unit by calculating the loss value, so as to determine whether the current compression unit has learned the knowledge of the previous compression unit. In this way, during the subsequent iterative training process, the compression unit can quickly approach the previous compression unit.

[0076] In some embodiments, referring to Figure 7 , step 42 may be the following process:

[0077] Step 421: Obtain third feature information output by the previous compression unit and fourth feature information output by the compression unit.

[0078] Step 422: Determine the second loss value based on the third feature information and the fourth feature information.

[0079] In some embodiments, the size of the feature map output by the compression unit is halved. Therefore, the number of channels between the third feature information and the fourth feature information is different, that is, the feature information is not equivalent. The present application proposes the following solution. Referring to Figure 8 , step 422 may be the following process:

[0080] Step 4221: Upsample the fourth feature information.

[0081] The upsampling can adopt an interpolation algorithm, such as nearest neighbor interpolation, bilinear interpolation, high-order interpolation, edge-based image interpolation algorithm, and region-based image interpolation algorithm. An appropriate algorithm can be selected according to actual needs for upsampling.

[0082] Step 4222: Determine the second loss value using the third feature information and the upsampled fourth feature information.

[0083] The way to determine the loss value can be to use the softmax function to normalize the third feature information and the upsampled fourth feature information respectively to obtain the probability of each classification result in the third feature information and the fourth feature information.

[0084] Compare the probabilities of each classification result in the third feature information and the fourth feature information to obtain the second loss value.

[0085] Step 43: Train the network model based on the first loss value and the second loss value.

[0086] In the iterative training of a network model, by making the current unit quickly approach the previous unit of the same type, the current unit can quickly learn the information of the previous unit of the same type, thereby improving the accuracy of the network model. Moreover, training based on the training parameters of the previous unit of the same type can improve the search efficiency of the current unit, shorten the search time, and thus improve the training time of the network model.

[0087] It can be understood that the terms "first", "second", etc. in this application are used to distinguish different objects, rather than to describe a specific order.

[0088] In an application scenario, in combination with Figure 9 it is described as follows:

[0089] When training the 6 basic units and 2 compression units of the differential network architecture, the self-distillation method is also adopted to output the training parameters of the previous unit of the same type (i.e., Figure 9 the distillation information in

[0090] to the current unit, so that the current unit can be trained based on the training parameters.

[0091]

[0092] Among them, A m and A m+1 respectively represent the output feature maps of the two previous and subsequent units of the same type. Φ(·) represents the softmax function, B(·) represents the upsampling method, represents the distillation learning function. Among them, Among them, A mi represents the i-th channel of the feature map A m , ρ represents the exponential parameter, and C m represents the total number of channels. Among them, ρ can be set to 2. When ρ = 2, the training of the network model is more effective.

[0093] The total loss in each iterative training of the network model can be expressed as:

[0094]

[0095] Among them, y represents the predicted label, represents the true label, and γ is the distillation coefficient, which can be set according to actual needs.

[0096] In Figure 9In the application, a training technique of self-distillation is introduced, which effectively improves the final accuracy of the supernetwork after search. Through the self-distillation technique, more information can be provided for the training of the current unit under limited resource constraints, improving the training degree of each unit, so that the performance of the trained subnetwork is closer to the real performance, thereby improving the consistency between the evaluation value of the subnetwork in the supernetwork and its real performance. It can be understood that the network model of the untrained differential network architecture is the supernetwork. After training, each unit after selecting the optimal node connection method can be used as a subnetwork.

[0097] Furthermore, since the deep network can learn the information of the shallow network more quickly, most of the training time is saved, enabling the supernetwork to achieve the same or better performance as before in a relatively short time, improving the search efficiency and shortening the search time.

[0098] Refer to Figure 10 , Figure 10 FIG. is a schematic flowchart of another embodiment of the training method of the network model provided by this application. The method includes:

[0099] Step 101: Obtain training samples.

[0100] Step 102: Input the training samples into the network model to train the network model.

[0101] Step 103: Obtain the third loss value of the network model.

[0102] Step 104: Determine whether the third loss value meets the set threshold.

[0103] When the third loss value meets the set threshold, execute Step 105.

[0104] In other embodiments, during the training process of the network model, the accuracy of the network model after this training can also be obtained. When the accuracy of the network model meets the preset accuracy, execute Step 105.

[0105] Step 105: Obtain the feature information output by the first two cascaded units and the training parameters of the previous unit of the same type.

[0106] Step 106: Input the training parameters and feature information into the current unit so that the current unit outputs feature information to the next unit to train the network model.

[0107] Step 105 and Step 106 have the same or similar technical solutions as any of the above embodiments and will not be elaborated here.

[0108] In the subsequent iterative training process, the model is trained in the manner of Step 105 and Step 106 until the network model converges.

[0109] In this embodiment, by performing conventional training on the network model in the initial stage and only executing step 105 when the loss value meets the set threshold, it can effectively prevent the network model from being overly biased due to directly using the training parameters and feature information to train the current unit, thus avoiding the problem of the network model's accuracy decline.

[0110] In the training of the network model, a corresponding way to end the training needs to be set. For example, the total number of iterations can be set, and the training ends after the total number of iterations is reached.

[0111] In this application, the following method is used to determine whether to end the training. Refer to Figure 11 , Figure 11 is a schematic flowchart of another embodiment of the training method of the network model provided by this application. The method includes:

[0112] Step 111: Determine whether the training of the network model meets the convergence condition.

[0113] If it meets the condition, execute step 112; if it does not meet the condition, execute step 114 and continue to train according to the solution of any of the above embodiments.

[0114] In the training of the network model, usually the set total number of iterations exceeds several hundred times. However, by using the training method of the network model provided by this application, the training speed can be effectively improved. Therefore, the training requirements may be met quickly, and there is no need to keep iterating. Therefore, step 111 needs to be executed after each iteration training.

[0115] Step 112: End the training of the network model and obtain the corresponding node connection relationships and weights in the basic unit and the compression unit.

[0116] Step 113: Construct a network model using the node connection relationships and weights that meet the preset requirements.

[0117] For example, for the intermediate nodes of each unit, at most retain the connection relationships with the two strongest predecessor nodes; for the edges between two nodes, only retain the edge with the largest weight. Assume a node has three predecessor nodes. Represent the strength of that predecessor node by the edge with the highest weight between the predecessor node and the current node, and select the two strongest predecessor nodes. Only retain the edge with the largest weight between the nodes.

[0118] Step 114: Continue training.

[0119] Refer to Figure 12 , Figure 12It is a schematic structural diagram of an embodiment of an electronic device provided by this application. The electronic device 120 includes a processor 121 and a memory 122 coupled to the processor 121. A computer program is stored in the memory 122, and the processor 121 is configured to execute the computer program to implement the following method:

[0120] Obtain the feature information output by the first two cascaded units and the training parameters of the previous unit of the same type; wherein, the feature information is obtained by training basic units and / or compression units in the network model based on training samples; input the training parameters and the feature information into the current unit, so that the current unit outputs the feature information to the next unit to train the network model.

[0121] Optionally, the processor 121 is further configured to execute the computer program to implement the method of any of the above embodiments. For details, refer to the corresponding methods above and will not be elaborated here.

[0122] Refer to Figure 13 , Figure 13 It is a schematic structural diagram of an embodiment of a computer-readable storage medium provided by this application. The computer-readable storage medium 130 stores a computer program 131. When the computer program 131 is executed by a processor, the following method is implemented:

[0123] Obtain the feature information output by the first two cascaded units and the training parameters of the previous unit of the same type; wherein, the feature information is obtained by training basic units and / or compression units in the network model based on training samples; input the training parameters and the feature information into the current unit, so that the current unit outputs the feature information to the next unit to train the network model.

[0124] Optionally, when the computer program 131 is executed by a processor, it implements the method of any of the above embodiments. For details, refer to the corresponding methods above and will not be elaborated here.

[0125] In summary, in this application, when training the current unit, the feature information output by the previous cascaded unit and the training parameters of the previous unit of the same type are used to train the current unit, so that the current unit can quickly learn the information of the previous unit of the same type, thereby improving the accuracy of the network model. Moreover, training based on the training parameters of the previous unit of the same type can improve the search efficiency of the current unit, shorten the search time, and thus reduce the training time of the network model, thereby increasing the feasibility of the network architecture search technology in industrial applications.

[0126] Secondly, since the differential-based network architecture search method alternately updates the network architecture parameters and neural network weights, the training process is unstable and the network performance fluctuates greatly. In addition, the differential network architecture search technology evaluates the performance of the sub-network based on the super-network. Therefore, the consistency of the super-network with a low degree of training is poor, and the performance of the sub-network obtained by the search is also poor.

[0127] Moreover, the search process of the network architecture search method based on differentiation usually only includes dozens of rounds, while it usually takes hundreds of rounds to fully train the subnetwork. Therefore, in the search process, due to the limitations of search time and resources, the training of the supernet and the different subnetworks in the supernet is very insufficient. Therefore, in the evaluation process based on the supernetwork, there must be a large gap between the performance obtained by the subnetwork and the actual performance. Through any of the above technical solutions proposed in this application, the training degree of the supernetwork can be improved under limited resources, and the performance consistency problem between the supernetwork and the subnetwork can be effectively improved, so as to more effectively search for a better subnetwork structure.

[0128] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only illustrative, for example, the division of the modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0129] If the integrated units in the above other embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), disk or optical disk and other media that can store program codes.

[0130] The above description is only an implementation method of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A training method for a network model, characterized in that, the architecture of the network model is a differential network architecture, and the differential network architecture includes a basic unit and a compression unit. The method includes: Obtaining the feature information output by the first two cascaded units and the training parameters of the previous same-type unit; wherein, the feature information is obtained by training the basic unit and / or the compression unit in the network model based on training samples; Inputting the training parameters and the feature information into the current unit, so that the current unit outputs feature information to the next unit to train the network model; wherein, the training samples are based on an image training set, and the true labels of the image content are marked in the images of the training set.

2. The method according to claim 1, characterized in that, the training parameters are extracted from the previous same-type unit by means of self-distillation.

3. The method according to claim 1, characterized in that, the method further includes: When the current unit is the basic unit, determining a first loss value between the basic unit and the previous basic unit; When the current unit is the compression unit, determining a second loss value between the compression unit and the previous compression unit; Training the network model based on the first loss value and the second loss value.

4. The method according to claim 3, characterized in that, the determining of the first loss value between the basic unit and the previous basic unit includes: Obtaining first feature information output by the previous basic unit and second feature information output by the basic unit; Determining the first loss value based on the first feature information and the second feature information.

5. The method according to claim 4, characterized in that, the determining of the first loss value based on the first feature information and the second feature information includes: Performing upsampling on the second feature information; Determining the first loss value by using the first feature information and the upsampled second feature information.

6. The method according to claim 3, characterized in that, the determining of the second loss value between the compression unit and the previous compression unit includes: Obtaining third feature information output by the previous compression unit and fourth feature information output by the compression unit; Determining the second loss value based on the third feature information and the fourth feature information.

7. The method according to claim 6, characterized in that, the determining of the second loss value based on the third feature information and the fourth feature information includes: Performing upsampling on the fourth feature information; Determining the second loss value by using the third feature information and the upsampled fourth feature information.

8. The method according to claim 1, characterized in that, before obtaining the feature information output by the first two cascaded units and the training parameters of the previous same-type unit, it includes: Obtaining the training samples; Inputting the training samples into the network model to train the network model; Obtaining a third loss value of the network model; Determine whether the third loss value meets the set threshold; If so, perform the steps of obtaining the feature information output by the first two cascaded units and the training parameters of the previous unit of the same type.

9. The method according to claim 1, wherein, the method further includes: determine whether the training of the network model meets the convergence condition; If so, end the training of the network model, and obtain the corresponding node connection relationship and weights in the basic unit and the compression unit; Construct the network model by using the node connection relationship and the weights that meet the preset requirements.

10. An electronic device, wherein, the electronic device includes a processor and a memory coupled to the processor, the memory stores a computer program, and the processor is configured to execute the computer program to implement the method according to any one of claims 1-9.

11. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1-9 is implemented.

Citation Information

Patent Citations

  • Neural network structure search method and device, electronic equipment and storage medium

    CN111612134A

  • Neural network structure searching method, image processing method and device

    CN112445823A