A neural network training method
By extending the convolutional layer in deep neural networks and fusing the extended convolution kernel module, the problem of low computing efficiency and inability to run in real time in the edge-end devices is solved, and performance improvement and real-time guarantee are achieved.
Patent Information
- Application Number
- CN202110663008.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-15
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-06-15
AI Technical Summary
Deep learning models are inefficient in computing in edge-end devices and cannot run in real time, mainly due to the limited computing power of micro GPUs, which requires quantitative compression and careful design of networks to meet the recognition accuracy and real-time requirements.
By extending the convolution layer of the original deep neural network, a convolution layer with the same number of channels but a small size of the convolution kernel is next to it, improving the network complexity, and fusing the extended convolution kernel module into the original network after training is completed, ensuring that the calculation amount does not increase.
Without increasing the amount of computing, the performance of deep neural networks has been improved, the accuracy rate of 1.4% is improved, and the real-time operation capability is ensured.
Smart Images

Figure CN113313237B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of neural network training, and in particular, to a neural network training method. Background Art
[0002] With the continuous update and development of computer vision technology, object detection technology plays an important role in multiple fields such as intelligent transportation, image retrieval, and face recognition. Deep learning, which has been developing increasingly popular in recent years, serves as a more efficient tool to assist our research and discovery in the field of object detection.
[0003] Currently, deep learning has greatly surpassed traditional vision algorithms in the field of object detection. Under big data, deep learning can autonomously learn effective features, and the learned features far exceed the algorithm features designed by hand in terms of quantity and performance.
[0004] In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art: Although deep learning has performance advantages over traditional vision algorithms, the disadvantages of deep learning are also obvious. The models using deep learning contain a large number of parameters, which brings a significant increase in the computational amount during training, reduces the computational efficiency of the model, and at the same time, the large amount of computation increases the error rate of model calculation and reduces the accuracy of model calculation. Moreover, the huge computational amount of training deep models results in that deep learning cannot run in real time on edge devices such as robots. The computing power of such edge micro GPUs is limited, and the network needs to be carefully designed and quantized and compressed to meet the real-time requirements under the premise of recognition accuracy. Summary of the Invention
[0005] The embodiments of the present invention provide a neural network training method to solve the problem that the limited computing power of the micro CPUs carried by edge devices such as robots in the prior art leads to the inability of deep learning to run in real time, or it is difficult to meet the accuracy rate after the quantization and compression of deep learning.
[0006] In a first aspect, the embodiments of the present invention provide a neural network training method, including:
[0007] Collect a sample to be analyzed as a training sample, and perform a first training on the original deep neural network based on the training sample;
[0008] Expand the convolutional layer of the deep neural network to obtain an expanded deep neural network;
[0009] Perform a second training on the expanded deep neural network based on the same configuration parameters in the first training, fuse the convolutional layer of the expanded deep neural network after the second training, and perform batch normalization processing.
[0010] Preferably, the deep neural network at least includes a first convolutional layer, and the first convolutional layer is connected to a first batch normalization layer and an activation layer. The first convolutional layer, the first batch normalization layer, and the activation layer form a convolutional block.
[0011] To expand the convolutional layer of the deep neural network, specifically including:
[0012] Connect a second convolutional layer beside the first convolutional layer. The first convolutional layer is connected to the first batch normalization layer, and the second convolutional layer is connected to the second batch normalization layer. The first batch normalization layer and the second batch normalization layer are connected to the activation layer to obtain an expanded deep neural network; the number of channels of the first convolutional layer is the same as that of the second convolutional layer, and the size of the convolutional kernel of the first convolutional layer is larger than that of the second convolutional layer.
[0013] Preferably, the parameters of the first convolutional layer and the first batch normalization layer inherit the parameters of the convolutional block of the original deep neural network, and the parameters of the second convolutional layer and the second batch normalization layer are randomly initialized.
[0014] Preferably, fuse the convolutional layers of the expanded deep neural network after the second training and perform batch normalization processing, specifically including:
[0015] Determine the input-output relationship of the convolutional layer:
[0016] Out conv =W conv *In + B conv
[0017] Where, In is the input; W conv is the convolutional kernel parameter, with a size of [Cout, Cin, K, K], and K is the size of the convolutional kernel; B conv is the bias of the convolutional layer;
[0018] Determine the input-output relationship of the batch normalization layer:
[0019]
[0020]
[0021]
[0022] In bn is the input of the batch normalization layer, and μ, σ, γ, β are the mean, variance, scaling coefficient, and translation coefficient obtained by training the current training samples respectively; μ n 、σ n 、γ n 、β nThe mean, variance, scaling factor, and translation factor are obtained by training for the nth training sample;
[0023] Let In bn = Out conv , B conv = 0, and the fused expression is obtained as:
[0024] Out = W bn * W conv * In + b bn = W f * In + b f
[0025] In the formula, In is the input, and W f and b f are the parameters of the fused convolutional layer;
[0026] Batch normalization processing is performed.
[0027] Preferably, it further includes:
[0028] Determine the calculation formulas of the first convolutional layer and the second convolutional layer after fusion, determine the expansion method of the second convolutional layer based on the calculation formulas of the first convolutional layer and the second convolutional layer, so as to expand the second convolutional layer to the same size as the first convolutional layer, and determine the structure parameters of the fused convolutional block obtained after the fusion of the first convolutional layer and the second convolutional layer.
[0029] In a second aspect, an embodiment of the present invention provides a neural network training system, including:
[0030] An initial training module, which collects a sample to be analyzed as a training sample and performs the first training on the original deep neural network based on the training sample;
[0031] An expansion module, which expands the convolutional layer of the deep neural network to obtain an expanded deep neural network;
[0032] An enhanced training module, which performs the second training on the expanded deep neural network based on the same configuration parameters in the first training, fuses the convolutional layers of the expanded deep neural network after the second training, and performs batch normalization processing.
[0033] Preferably, the deep neural network includes at least a first convolutional layer, and the first convolutional layer is connected to a first batch normalization layer and an activation layer, and the first convolutional layer, the first batch normalization layer, and the activation layer form a convolutional block;
[0034] The expansion of the convolutional layer of the deep neural network specifically includes:
[0035] A second convolutional layer is connected in parallel to the first convolutional layer. The first convolutional layer is connected to a first batch normalization layer, and the second convolutional layer is connected to a second batch normalization layer. The first batch normalization layer and the second batch normalization layer are connected to an activation layer to obtain an extended deep neural network. The number of channels of the first convolutional layer is the same as that of the second convolutional layer, and the size of the convolutional kernel of the first convolutional layer is larger than that of the second convolutional layer.
[0036] Preferably, the parameters of the first convolutional layer and the first batch normalization layer inherit the convolutional block parameters of the original deep neural network, and the parameters of the second convolutional layer and the second batch normalization layer are randomly initialized.
[0037] In a third aspect, an embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the neural network training method described in the first aspect embodiment of the present invention are implemented.
[0038] In a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the neural network training method described in the first aspect embodiment of the present invention are implemented.
[0039] A neural network training method and system provided by an embodiment of the present invention improve the performance of the network without increasing the computing amount of the deployed network. By connecting a convolutional layer with the same number of channels and a smaller convolutional kernel size in parallel to the convolutional layer of the original deep neural network, the complexity of the deep neural network is improved. After training, the convolutional module with the parallel convolutional kernel is fused into the original network to achieve the purpose of not increasing the computing amount. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0041] Figure 1 It is a flowchart of the neural network training method provided by an embodiment of the present invention;
[0042] Figure 2 It is a schematic diagram of the convolutional block structure provided by an embodiment of the present invention;
[0043] Figure 3 It is a schematic diagram of the extended convolutional block structure provided by an embodiment of the present invention;
[0044] Figure 4 Schematic diagram of the optimized structure of the convolution block provided by the embodiment of the present invention;
[0045] Figure 5 Schematic diagram of the convolution block structure of the fusion-expanded deep neural network after expanding the bypass branch provided by the embodiment of the present invention
[0046] Figure 6 Schematic diagram of the entity structure provided by the embodiment of the present invention. Detailed implementation manners
[0047] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0048] As used herein, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase does not necessarily refer to the same embodiment at every occurrence in the specification, nor is it an independent or alternative embodiment mutually exclusive of other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0049] With the development of artificial intelligence technology and the popularization of deep neural networks, the computing power demand for devices has increased sharply. Edge devices such as robots will mount or integrate micro graphics processing units (GPUs). Generally, the computing power of such edge micro GPUs is limited, and the network needs to be carefully designed and quantized and compressed to meet the real-time requirements under the premise of recognition accuracy.
[0050] Therefore, the embodiments of the present invention provide a neural network training method and system, which can improve the performance of the deep neural network in edge devices such as robots without increasing the computational amount of the deployed network. By connecting a convolution layer with the same number of channels and a smaller convolution kernel size beside the convolution layer of the original deep neural network, the complexity of the deep neural network is improved. After training, the convolution module with the bypassed convolution kernel is fused into the original network to achieve the purpose of not increasing the computational amount. The following will be described and introduced through multiple embodiments.
[0051] Figure 1This embodiment of the present invention provides a neural network training method, which can be applied to mount or integrate a deep neural network in a micro-graphics processing unit (GPU) of edge devices such as robots, so as to improve the performance of the deep neural network in edge devices such as robots. The method includes:
[0052] Collect samples to be analyzed as training samples, and perform the first training on the original deep neural network based on the training samples. It can be understood that, as an application scenario of this embodiment of the present invention, the micro-graphics processing unit (GPU) integrated in the robot can perform image detection using the original deep neural network, and the training samples in this embodiment are the sample images to be analyzed.
[0053] Expand the convolutional layer of the deep neural network to obtain an expanded deep neural network;
[0054] Perform the second training on the expanded deep neural network based on the same configuration parameters in the first training, fuse the convolutional layer of the expanded deep neural network after the second training, and perform batch normalization processing.
[0055] Specifically, as Figure 2 shown in, the convolutional layer is a basic unit of the visual deep neural network, and has attributes such as a convolutional window (Kernel), a convolutional stride (Stride), the number of input and output channels (Cin, Cout), etc. It is usually combined with a batch normalization layer (BatchNormalization) and an activation layer (Relu) to form a convolutional block. The deep convolutional neural network is usually stacked by the above convolutional blocks.
[0056] In this embodiment, the network width is expanded by adding bypass branches, the network is retrained, and then the bypass branches are fused to ensure that the network computation amount remains unchanged.
[0057] Based on the above embodiment, as a preferred implementation manner, the deep neural network includes at least a first convolutional layer, the first convolutional layer is connected to a first batch normalization layer and an activation layer, and the first convolutional layer, the first batch normalization layer, and the activation layer form a convolutional block;
[0058] The expansion of the convolutional layer of the deep neural network specifically includes:
[0059] Connect a second convolutional layer beside the first convolutional layer, the first convolutional layer is connected to the first batch normalization layer, the second convolutional layer is connected to the second batch normalization layer, and the first batch normalization layer and the second batch normalization layer are connected to the activation layer to obtain an expanded deep neural network; the number of channels of the first convolutional layer is the same as that of the second convolutional layer, and the convolutional kernel size of the first convolutional layer is larger than that of the second convolutional layer.
[0060] Specifically, as shown in Figure 3 , in the original deep neural network, a convolution block with a convolution kernel size of 3*3 is selected, and a convolution layer with a convolution kernel size of 1*1 is connected in parallel. The extended convolution block is defined as shown in Figure 3 .
[0061] Based on the above embodiments, as a preferred embodiment, the parameters of the first convolution layer and the first batch normalization layer inherit the convolution block parameters of the original deep neural network, and the parameters of the second convolution layer and the second batch normalization layer are randomly initialized.
[0062] According to the above method of the embodiments of the present invention, all convolution blocks with a convolution kernel size of 3*3 in the original deep neural network are replaced with the above extended convolution blocks. Among them, the left branch (the first convolution layer) in the extended convolution block, that is, the convolution layer branch with a convolution kernel of 3*3 and the parameters of its corresponding batch normalization layer inherit the convolution block parameters of the original network; the parameters of the right branch in the extended convolution block are randomly initialized.
[0063] Adopt the same configuration parameters and loss function as those for training the original deep neural network, retrain the extended deep neural network, and after training is completed, save the parameter values of the extended deep neural network.
[0064] Based on the above embodiments, as a preferred embodiment, fuse the convolution layers of the extended deep neural network after the second training and perform batch normalization processing, specifically including:
[0065] Determine the input-output relationship of the convolution layer:
[0066] Out conv =W conv *In + B conv (1)
[0067] where, In is the input; W conv is the convolution kernel parameter, with a size of [Cout, Cin, K, K], and K is the size of the convolution kernel; B conv is the bias of the convolution layer;
[0068] According to the definition of the batch normalization layer, the input-output relationship of the batch normalization layer is defined as follows:
[0069]
[0070] The above formula can be rewritten as:
[0071]
[0072]
[0073]
[0074] In bn is the input of the batch normalization layer, and μ, σ, γ, and β are the mean, variance, scaling coefficient, and translation coefficient obtained from the training of the current training sample; μ n , σ n , γ n , β n are the mean, variance, scaling coefficient, and translation coefficient obtained from the training of the nth training sample;
[0075] Let In bn = Out conv , B conv = 0, and the fused expression is obtained as:
[0076] Out = W bn * W conv * In + b bn = W f * In + b f (6)
[0077] In the formula, In is the input, and W f and b f are the parameters of the fused convolutional layer;
[0078] Perform batch normalization processing.
[0079] According to the above formula, the extended convolutional blocks in the trained extended deep neural network can be fused and optimized as shown in Figure 4 as follows.
[0080] Based on the above embodiments, as a preferred implementation manner, it further includes:
[0081] Determine the calculation formulas of the first convolutional layer and the second convolutional layer after fusion, determine the expansion method of the second convolutional layer based on the calculation formulas of the first convolutional layer and the second convolutional layer, so as to expand the second convolutional layer to the same size as the first convolutional layer, and determine the structural parameters of the fused convolutional block obtained after the fusion of the first convolutional layer and the second convolutional layer.
[0082] Specifically, when fusing the extended bypass branches of the extended deep neural network, according to the fused result, the calculation formulas of the two branches are defined as follows:
[0083]
[0084] Since the sizes of W 3*3 and W 1*1 are [Cout, Cin, 3, 3] and [Cout, Cin, 1, 1] respectively, it is necessary to expand W1*1 into W 3*3The same size, and the expansion method is as follows:
[0085]
[0086] Therefore, the Out formula can be rewritten as:
[0087] Out = (W 1*1→3*3 + W 3*3 ) * In + (b 3*3 + b 1*1 ) = W * In + b
[0088] The 1×1 convolutional layer branch of the bypass is fused into the 3×3 convolutional layer of the main branch. The structure of the fused convolutional block is as Figure 5 shown. Figure 5 FIG. is a schematic diagram of the convolutional block structure after the expansion bypass branch of the fusion extended deep neural network provided by the embodiment of the present invention.
[0089] The calculation speed and accuracy of the fused extended neural network and the original deep neural network are shown in the following table.
[0090] Original network Extended fusion network Computing speed 42ms 42ms Accuracy 91.7% 93.1%
[0091] It can be clearly seen from the tabular data that the method proposed in this patent does not increase the calculation time, but brings a 1.4% improvement in accuracy.
[0092] The embodiment of the present invention also provides a neural network training system, based on the neural network training method in the above embodiments, including:
[0093] An initial training module, which collects samples to be analyzed as training samples and performs the first training on the original deep neural network based on the training samples;
[0094] An expansion module, which expands the convolutional layer of the deep neural network to obtain an extended deep neural network;
[0095] An enhanced training module, which performs the second training on the extended deep neural network based on the same configuration parameters in the first training, fuses the convolutional layer of the extended deep neural network after the second training, and performs batch normalization processing.
[0096] Preferably, the deep neural network includes at least a first convolutional layer, the first convolutional layer is connected with a first batch normalization layer and an activation layer, and the first convolutional layer, the first batch normalization layer and the activation layer form a convolutional block;
[0097] Expanding the convolutional layer of the deep neural network specifically includes:
[0098] A second convolutional layer is connected in parallel to the first convolutional layer. The first convolutional layer is connected to a first batch normalization layer, and the second convolutional layer is connected to a second batch normalization layer. The first batch normalization layer and the second batch normalization layer are connected to an activation layer to obtain an extended deep neural network. The number of channels of the first convolutional layer is the same as that of the second convolutional layer, and the kernel size of the first convolutional layer is larger than that of the second convolutional layer.
[0099] Preferably, the parameters of the first convolutional layer and the first batch normalization layer inherit the convolutional block parameters of the original deep neural network, and the parameters of the second convolutional layer and the second batch normalization layer are randomly initialized.
[0100] Based on the same concept, an embodiment of the present invention further provides a schematic diagram of an entity structure, as Figure 6 shown. The server may include: a processor 601, a communication interface 602, a memory 603, and a communication bus 604. Among them, the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604. The processor 601 can call the logical instructions in the memory 603 to execute the steps of the neural network training method described in the above embodiments. For example, it includes:
[0101] Collect a sample to be analyzed as a training sample, and perform a first training on the original deep neural network based on the training sample;
[0102] Expand the convolutional layer of the deep neural network to obtain an extended deep neural network;
[0103] Perform a second training on the extended deep neural network based on the same configuration parameters in the first training, fuse the convolutional layers of the extended deep neural network after the second training, and perform batch normalization processing.
[0104] In addition, when the logic instructions in the above-mentioned memory 603 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0105] Based on the same concept, an embodiment of the present invention further provides a non-transitory computer-readable storage medium. This computer-readable storage medium stores a computer program, and this computer program includes at least one segment of code. This at least one segment of code can be executed by a main control device to control the main control device to implement the steps of the neural network training method as described in the above-mentioned various embodiments. For example, it includes:
[0106] Collect a sample to be analyzed as a training sample, and perform a first training on the original deep neural network based on the training sample;
[0107] Expand the convolutional layer of the deep neural network to obtain an expanded deep neural network;
[0108] Perform a second training on the expanded deep neural network based on the same configuration parameters in the first training, fuse the convolutional layer of the expanded deep neural network after the second training, and perform batch normalization processing.
[0109] Based on the same technical concept, an embodiment of the present application further provides a computer program. When this computer program is executed by a main control device, it is used to implement the above-mentioned method embodiments.
[0110] The program can be stored in whole or in part on a storage medium packaged together with the processor, or can be stored in part or in whole on a memory not packaged together with the processor.
[0111] Based on the same technical concept, an embodiment of the present application further provides a processor. This processor is used to implement the above-mentioned method embodiments. The above-mentioned processor can be a chip.
[0112] In summary, the neural network training method and system provided by the embodiments of the present invention improve the performance of the network without increasing the computing amount of the deployed network. By connecting a convolutional layer with the same number of channels and a smaller convolutional kernel size beside the convolutional layer of the original deep neural network, the complexity of the deep neural network is improved. After training, the convolutional module with the connected convolutional kernel is fused into the original network to achieve the purpose of not increasing the computing amount.
[0113] The various embodiments of the present invention can be combined arbitrarily to achieve different technical effects.
[0114] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in accordance with the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state disk).
[0115] Those of ordinary skill in the art can understand all or part of the processes in the methods of the above embodiments. The processes can be completed by relevant hardware instructed by a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage medium includes: various media such as ROM or random access memory RAM, magnetic disk, or optical disc that can store program codes.
[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A neural network training method, applied to a robot for mounting or integrating a deep neural network in a micro graphics computing unit for image detection, characterized in that, Including: Collecting a sample to be analyzed as a training sample, and performing a first training on the original deep neural network based on the training sample; Expanding the convolutional layer of the original deep neural network to obtain an expanded deep neural network; Performing a second training on the expanded deep neural network based on the same configuration parameters in the first training, fusing the convolutional layers of the expanded deep neural network after the second training, and performing batch normalization processing; The original deep neural network includes at least a first convolutional layer, the first convolutional layer is connected to a first batch normalization layer and an activation layer, and the first convolutional layer, the first batch normalization layer and the activation layer form a convolutional block; Expanding the convolutional layer of the original deep neural network specifically includes: Connecting a second convolutional layer beside the first convolutional layer, the first convolutional layer is connected to a first batch normalization layer, the second convolutional layer is connected to a second batch normalization layer, and the first batch normalization layer and the second batch normalization layer are connected to an activation layer to obtain an expanded deep neural network; The number of channels of the first convolutional layer is the same as that of the second convolutional layer, and the convolution kernel size of the first convolutional layer is larger than that of the second convolutional layer; The parameters of the first convolutional layer and the first batch normalization layer inherit the parameters of the convolutional block of the original deep neural network, and the parameters of the second convolutional layer and the second batch normalization layer are randomly initialized.
2. The neural network training method according to claim 1, characterized in that, Fusing the convolutional layers of the expanded deep neural network after the second training and performing batch normalization processing specifically includes: Determining the input-output relationship of the convolutional layer: Out conv = W conv *In + B conv where, In is the input; W conv is the convolutional kernel parameter, with a size of [Cout, Cin, K, K], where K is the size of the convolutional kernel; B conv is the bias of the convolutional layer; Determining the input-output relationship of the batch normalization layer: In bn is the input to the batch normalization layer, where μ, σ, γ, and β are the mean, variance, scaling factor, and translation factor obtained from training the current training sample; μ n , σ n , γ n , and β n are the mean, variance, scaling factor, and translation factor obtained from training the nth training sample; Let In bn = Out conv , B conv = 0, and the fused expression is obtained as: Out = W bn *W conv *In + b bn = W f *In + b f where, In is the input, W f and b f are the parameters of the fused convolutional layer; Performing batch normalization processing.
3. The neural network training method according to claim 2, characterized in that, Also including: Determining the calculation formula of the first convolutional layer and the second convolutional layer after fusion, determining the expansion method of the second convolutional layer based on the calculation formulas of the first convolutional layer and the second convolutional layer to expand the second convolutional layer to the same size as the first convolutional layer, and determining the structural parameters of the fused convolutional block obtained by fusing the first convolutional layer and the second convolutional layer.
4. A neural network training system, applied to a robot for mounting or integrating a deep neural network in a micro graphics computing unit for image detection, characterized in that, Including: An initial training module that collects a sample to be analyzed as a training sample and performs a first training on the original deep neural network based on the training sample; An expansion module that expands the convolutional layer of the original deep neural network to obtain an expanded deep neural network; An enhanced training module that performs a second training on the expanded deep neural network based on the same configuration parameters in the first training, fuses the convolutional layers of the expanded deep neural network after the second training, and performs batch normalization processing; The original deep neural network includes at least a first convolutional layer, the first convolutional layer is connected to a first batch normalization layer and an activation layer, and the first convolutional layer, the first batch normalization layer and the activation layer form a convolutional block; Expanding the convolutional layer of the original deep neural network specifically includes: Connecting a second convolutional layer beside the first convolutional layer, the first convolutional layer is connected to a first batch normalization layer, the second convolutional layer is connected to a second batch normalization layer, and the first batch normalization layer and the second batch normalization layer are connected to an activation layer to obtain an expanded deep neural network; The number of channels of the first convolutional layer is the same as that of the second convolutional layer, and the kernel size of the first convolutional layer is larger than that of the second convolutional layer; The parameters of the first convolutional layer and the first batch normalization layer inherit the convolutional block parameters of the original deep neural network, and the parameters of the second convolutional layer and the second batch normalization layer are randomly initialized.
5. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that,When the processor executes the program, the steps of the neural network training method according to any one of claims 1 to 3 are implemented.
6. A non-transitory computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, the steps of the neural network training method according to any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Feature fusion coefficient learnable image semantic segmentation method
CN107766794A
Convolutional neural network training method based on LRU
CN110197261A