A target recognition method and system suitable for low computing power devices
By introducing bypass ideas and R-ECA modules into the target recognition network of low-computing power equipment, the problems of gradient vanishing and gradient explosion are solved, and the recognition accuracy and stability are improved. They are suitable for smart watches and low-computing power vehicle-mounted autonomous driving systems.
Patent Information
- Application Number
- CN202310115691.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-10
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2043-02-10
AI Technical Summary
The target recognition network of low-computing power equipment is prone to gradient vanishing and gradient explosion problems in deep separable convolutional networks, resulting in insufficient recognition accuracy and stability.
The idea of bypass is introduced in the deep separable convolution network, and the weights of each channel are obtained through 1-dimensional convolution operation and Sigmoid activation function, and feature map processing is combined with the R-ECA module to improve the accuracy and stability of the network.
It effectively improves the accuracy and stability of the target recognition network of low-computing equipment, and is suitable for smart watches and low-computing vehicle-mounted autonomous driving systems.
Smart Images

Figure CN116206127B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image recognition technology and relates to a target recognition method and system, and specifically relates to a target recognition method and system with low computing power, high precision and high stability. Background Art
[0002] In the era of big data, machine learning has begun to be applied to all walks of life. At the same time, low-computing power devices have become popular, and machine learning models suitable for low-computing power devices have also attracted people's attention.
[0003] Compared to traditional convolutions, depthwise separable convolutions have fewer parameters and are well-suited for low-computing devices. However, object recognition networks constructed using these convolutions can suffer from vanishing and exploding gradients as network depth increases, a pressing issue for low-computing devices. Summary of the Invention
[0004] To address these issues, the present invention proposes a low-computing-power, high-precision, and high-stability target recognition method and system. This approach introduces a "detour" approach into a network composed of depthwise separable convolutions. A one-dimensional convolution with a kernel size of k is then performed, and a Sigmoid activation function is applied to obtain the weights w for each channel. These weights are then multiplied by the corresponding elements of the original input feature map to obtain the final output feature map. This operation effectively improves the accuracy and stability of the network and is suitable for low-computing-power devices.
[0005] The technical solution adopted by the method of the present invention is: a target recognition method suitable for low-computing power devices, which inputs the collected original target image into the target recognition network and outputs a high-precision and high-stability feature map;
[0006] The target recognition network includes a 3×3 conv basic unit, a dp basic unit, an R-ECA module, two dp basic units, one R-ECA module, two dp basic units, one R-ECA module, five dp basic units, one R-ECA module, two dp basic units, and one R-ECA module, which are connected in sequence; wherein conv represents a traditional convolutional layer and dp represents a depth-separable convolutional layer; the outputs of the conv basic unit and the dp basic unit are processed by the BN layer and the ReLU nonlinear activation function before being input into the next unit;
[0007] The R-ECA module includes a bypass mechanism submodule and a cross-channel interaction submodule. The workflow of the bypass mechanism is as follows:
[0008]
[0009] Among them, W lis a 1×1 convolution operation, W 1d is a one-dimensional convolution, W dp Represents the depthwise separable convolution operation; the bypass mechanism submodule includes n depthwise separable convolution modules and a dimension-upgrading module. The convolution kernel size of the depthwise convolution in the depthwise separable convolution module is 3×3 and the step size is 1; the convolution kernel size of the dimension-upgrading module is 1×1 and the initial step size is 1; let the input be x l , for x l Perform the following operations: First, the branch's dimension-upgrading module first l By increasing the dimension, we can get H(x l ), while the main line is x l Perform a series of depth-separable convolution operations to obtain F(x l ), and finally H(x l ) and F(x l ) are added; the output of the bypass mechanism is obtained, where the output represents the output of the R-ECA module;
[0010] The workflow of the cross-channel interaction submodule is as follows:
[0011]
[0012] Among them, x l is the input, W l is a 1×1 convolution operation, W 1d is a one-dimensional convolution, W GAP is the global pooling operation, W dp Depthwise separable convolution operation; the main line part is for input x l Using the convolution operation, we get F(x l ), F(x l ) Through global average pooling, the h and w dimensions are both changed to 1, where h and w represent the length and width of the feature map respectively, and only the channel dimension is retained. Then, through 1D convolution, the channels of each layer interact with the channels of the adjacent layers and share weights. Sigmoid is used for processing. The input x l Multiply it by the processed feature map weight to get F′(x l ); At the same time, the branch line inputs x l Use the dimension-raising module to increase the dimension and get the output H(x l ) and save it, so that H(x l ) dimension and F′(x l ) dimension is the same, and a detour is completed at this time; H(x l ) and F′(x l ) and output x after adding l+1 .
[0013] The technical solution adopted by the system of the present invention is: a target recognition system suitable for low-computing power devices, comprising:
[0014] one or more processors;
[0015] A storage device is used to store one or more programs, which, when executed by the one or more processors, enable the one or more processors to implement the target recognition method applicable to low-computing power devices.
[0016] This invention is suitable for low-computing devices, such as smartwatches and some low-computing vehicle-mounted autonomous driving systems. These low-computing devices require not only a small target recognition system but also strong recognition stability for safe and stable operation. Large-scale deep learning networks are not well suited to low-computing devices, and this invention appropriately addresses this shortcoming. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a structural diagram of a DP basic unit according to an embodiment of the present invention;
[0018] Figure 2 This is a structural diagram of the bypass mechanism submodule in the R-ECA module according to an embodiment of the present invention;
[0019] Figure 3 This is a structural diagram of the cross-channel interaction sub-module in the R-ECA module of an embodiment of the present invention. DETAILED DESCRIPTION
[0020] In order to facilitate ordinary technicians in this field to understand and implement the present invention, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the implementation examples described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0021] In the era of big data, machine learning is becoming increasingly prevalent. Machine learning models suitable for low-computing devices are also attracting attention. Depthwise separable convolutions have fewer parameters than traditional convolutions, making them ideal for low-computing devices. However, object recognition networks based on these convolutions can suffer from vanishing and exploding gradients as network depth increases. Therefore, introducing a "detour" approach within networks composed of depthwise separable convolutions can effectively address this issue. A 1D convolution operation with a kernel size of k is then performed, followed by a sigmoid activation function to obtain the weights w for each channel. These weights are then multiplied by the corresponding elements of the original input feature map to produce the final output feature map. This operation effectively improves network accuracy and stability, making it suitable for low-computing devices.
[0022] The present invention provides a target recognition method suitable for low-computing-power devices, which inputs the collected original target image into the target recognition network and outputs a feature map with high precision and high stability.
[0023] The target recognition network of this embodiment includes a 3×3 conv basic unit, 1 dp basic unit, 1 R-ECA module, 2 dp basic units, 1 R-ECA module, 2 dp basic units, 1 R-ECA module, 5 dp basic units, 1 R-ECA module, 2 dp basic units, and 1 R-ECA module, which are connected sequentially; conv represents a traditional convolutional layer, and dp represents a depth-wise separable convolutional layer; the outputs of the conv basic unit and the dp basic unit are processed by the BN layer and the ReLU nonlinear activation function before being input into the next unit.
[0024] Please see Figure 1 The DP basic unit of this embodiment includes a depth convolution layer and a point convolution layer. The input is C0 channels. After passing through C0 depth convolution layers with a convolution kernel size of 3×3 and a channel of 1, a feature map with the number of channels C0 is obtained. Then, after passing through C1 point convolution layers with a convolution kernel size of 1×1 and the number of channels n, a feature map with a channel size of C1 is obtained. Among them, the convolution kernel size of the depth convolution layer is all 3×3, and the step size is all 1; the convolution kernel size of the point convolution layer is all 1×1, and the step size is all 1. The number of channels of the initial input data is 64.
[0025] The R-ECA module of this embodiment includes a bypass mechanism submodule and a cross-channel interaction submodule; see Figure 2 , the workflow of the bypass mechanism is:
[0026]
[0027] Among them, W l is a 1×1 convolution operation, W 1d is a one-dimensional convolution, W dp Represents the depthwise separable convolution operation; the bypass mechanism submodule includes n depthwise separable convolution modules and a dimension-upgrading module. The convolution kernel size of the depthwise convolution in the depthwise separable convolution module is 3×3 and the step size is 1; the convolution kernel size of the dimension-upgrading module is 1×1 and the initial step size is 1; let the input be x l , for x l Perform the following operations: First, the branch's dimension-upgrading module first l By increasing the dimension, we can get H(x l ), while the main line is x l Perform a series of depth-separable convolution operations to obtain F(x l ), and finally H(x l ) and F(x l) are added; the output of the bypass mechanism is obtained, where the output represents the output of the R-ECA module;
[0028] Please see Figure 3 The workflow of the cross-channel interaction submodule in this embodiment is as follows:
[0029]
[0030] Among them, x l is the input, W l is a 1×1 convolution operation, W 1d is a one-dimensional convolution, W GAP is the global pooling operation, W dp Depthwise separable convolution operation; at the same time, the main line part is on the input x l Using the convolution operation, we get F(x l ), F(x l ) Through global average pooling, the h and w dimensions are both changed to 1, where h and w represent the length and width of the feature map respectively, and only the channel dimension is retained. Then, through 1D convolution, the channels of each layer interact with the channels of the adjacent layers and share weights. Sigmoid is used for processing. The input x l Multiply it by the processed feature map weight, then the weight will be added to the feature map, and we get F′(x l ); At the same time, the branch line inputs x l Use the dimension-raising module to increase the dimension and get the output H(x l ) and save it, so that H(x l ) dimension and F′(x l ) dimension is the same, and a detour is completed at this time; H(x l ) and F′(x l ) and output x after adding l+1 .
[0031] This embodiment proposes a module to improve the accuracy and stability of the target recognition network, named R-ECA module. This embodiment uses depthwise separable convolution to build a backbone network, and then embeds the R-ECA module into the backbone network.
[0032] This embodiment uses depthwise separable convolution to build the backbone network and combines it with standard convolution, because depthwise separable convolution has less computational complexity than standard convolution and is more suitable for low-computing devices. The details are as follows:
[0033] The calculation formula for the standard convolution operation amount is:
[0034] FLOP s =(2×C0×K 2 -1)H×W×C1;
[0035] Parameter calculation formula:
[0036] K 2 ×C0×C1;
[0037] Where C0 is the number of input channels, K is the convolution kernel size, H, W are the sizes of the input feature maps, and C1 is the size of the output channel.
[0038] Assume that the input feature dimension is D F ×D F ×M, where M is the number of channels. The parameters of the convolution kernel are
[0039] D K ×D K ×1×M;
[0040] The convolution kernel dimension and input dimension are both M. During convolution, each channel of the feature map corresponds to only one convolution kernel, and the feature dimension after output depth convolution is as follows:
[0041] FLOP S =M×D F ×D F ×D K ;
[0042] Then input the features after deep convolution, the dimension is D F ×D F ×M convolution, the parameters are:
[0043] 1×1×M×N;
[0044] The output dimension is
[0045] D F ×D F ×N;
[0046] During the convolution process, a standard 1×1 convolution is performed on each feature. The details are as follows:
[0047] FLOP s =N×D F ×D F ×M;
[0048] Adding the parameters of formula (5) and formula (7) gives
[0049] D K ×D K ×M+M×N;
[0050] Formula (9) represents the number of parameters of depth-wise separable convolution.
[0051] The ratio of the parameters is:
[0052]
[0053] The R-ECA module of this embodiment includes two parts: setting the "bypass" mechanism and obtaining the weights of each channel; the details are as follows:
[0054] (1) Bypass mechanism:
[0055] This embodiment designs a "bypass" method. In the target recognition network, there are many bypass branches that directly connect the input to the subsequent layers, so that the subsequent layers can directly learn these inputs. This is a "bypass".
[0056] This solution designs five bypasses to connect the input directly to the subsequent layers, allowing the subsequent layers to directly learn (obtain) these inputs. Each bypass saves the input of the feature map that has not yet been upgraded and increases the dimension to the same as the last output dimension after the feature increase. It is then added to the last output after the feature map feature increase as the next input. In this way, as the network depth and feature map dimension increase, setting up multiple bypasses can effectively alleviate the problems of gradient vanishing and gradient explosion. Traditional convolutional layers or fully connected layers will more or less have problems such as information loss and loss when transmitting information. The "bypass" solution solves this problem to some extent. By directly bypassing the input information to the output, the integrity of the information is protected. The entire network only needs to learn the difference between the input and output, simplifying the learning objectives and difficulty.
[0057] (2) Obtain the weight of each channel:
[0058] The output of the bypass mechanism is used as the input feature map for global average pooling. A one-dimensional convolution operation with a kernel size of k (usually k = 5) is performed, and the weight w of each channel is obtained through the Sigmoid activation function. The formula is as follows:
[0059] w=σ(C1D k (y));
[0060] Among them, σ represents the sigmoid function, C1D k Represents a 1-dimensional convolution operation with a convolution kernel size of k; multiply the weights by the corresponding elements of the original input feature map to obtain the final output feature map.
[0061] The R-ECA mechanism of this embodiment mainly includes a detour mechanism and cross-channel interaction. Figure 2 As shown in the figure, the hierarchical structure of the bypass mechanism has been given above. For cross-channel interaction, it mainly includes a global pooling layer, adaptive convolution and a sigmoid function. Among them, the convolution kernel size of the adaptive convolution layer can be adjusted according to the input data. Because it involves multiple classifications, the sigmoid function is specifically a softmax function. Figure 3As shown, the output X of the previous R-ECA mechanism l As the input this time, after the depth-wise separable convolution, it is input into the global pooling layer, and then after the adaptive convolution and sigmoid function processing, the weight of each channel is obtained, and the weight is combined with X l Multiply to get the result, and X l Also processed by the bypass mechanism, the processing result is added to the appeal result to get X l+1 , as the next input.
[0062] The target recognition network in this embodiment is a trained target recognition network. Training data uses images of diseased apple leaves. Data augmentation is performed on several diseased apple leaves, primarily using rotation, cropping, and noise addition. During training, the bitchsize is set to 64, and the epoch is set to 1000. Each network is trained using the following method: The initial learning rate of the network is set to 0.03, and the batch_size (the number of samples selected for a single training run; its size influences model optimization and speed and is adjusted based on network parameters and GPU memory size) is set to 64. The loss is then trained for 400 epochs (ranging from 2.2256 to 0.8195) to obtain the initial loss value. The learning rate was then lowered to 0.003, and the loss was trained for 400 epochs within a range of 0.8142 to 0.6962. Finally, the learning rate was raised to 0.0003, and the loss was trained for 200 epochs within a narrow range of 0.6928 to 0.6909, yielding the final training set loss. After multiple experimental tests, the test loss in the first 400 epochs dropped from 2.23 to 0.82, a significant decrease of 63%. The test loss in the second 400 epochs dropped from 0.82 to 0.67, a moderate decrease of 18%. Finally, the test loss in the final 200 epochs dropped from 0.67 to 0.66, a smaller decrease of 1.4%.
[0063] It should be understood that the above description of the preferred embodiment is relatively detailed and cannot be regarded as limiting the scope of protection of the patent of the present invention. Under the guidance of the present invention, ordinary technicians in this field can also make substitutions or modifications without departing from the scope of protection of the claims of the present invention, which all fall within the scope of protection of the present invention. The scope of protection requested by the present invention shall be based on the attached claims.
Claims
1. A target recognition method suitable for low-computing-power devices, characterized by: Input the collected original target image into the target recognition network and output a high-precision and high-stability feature map; The target recognition network includes a 3×3 conv basic unit, a dp basic unit, an R-ECA module, two dp basic units, one R-ECA module, two dp basic units, one R-ECA module, five dp basic units, one R-ECA module, two dp basic units, and one R-ECA module, which are connected in sequence; wherein conv represents a traditional convolutional layer and dp represents a depth-separable convolutional layer; the outputs of the conv basic unit and the dp basic unit are processed by the BN layer and the ReLU nonlinear activation function before being input into the next unit; The R-ECA module includes a bypass mechanism submodule and a cross-channel interaction submodule. The workflow of the bypass mechanism is as follows: W l x l ⊕(W dp (W dp ...(W dp x l )...)); Among them, W l is a 1×1 convolution operation, W 1d is a one-dimensional convolution, W dp Represents the depthwise separable convolution operation; the bypass mechanism submodule includes n depthwise separable convolution modules and a dimension-upgrading module. The convolution kernel size of the depthwise convolution in the depthwise separable convolution module is 3×3 and the step size is 1; the convolution kernel size of the dimension-upgrading module is 1×1 and the initial step size is 1; let the input be x l , for x l Perform the following operations: First, the branch's dimension-upgrading module first l By increasing the dimension, we can get H(x l ), while the main line is x l Perform a series of depth-separable convolution operations to obtain F(x l ), and finally H(x l ) and F(x l ) are added; the output of the bypass mechanism is obtained, where the output represents the output of the R-ECA module; The workflow of the cross-channel interaction submodule is as follows: Among them, x l is the input, W l is a 1×1 convolution operation, W 1d is a one-dimensional convolution, W GAP is the global pooling operation, W dp Depthwise separable convolution operation; the main line part is for input x l Using the convolution operation, we get F(x l ), F(x l ) Through global average pooling, the h and w dimensions are both changed to 1, where h and w represent the length and width of the feature map respectively, and only the channel dimension is retained. Then, through 1D convolution, the channels of each layer interact with the channels of the adjacent layers and share weights. Sigmoid is used for processing. The input x l Multiply it by the processed feature map weight to get F′(x l ); At the same time, the branch line inputs x l Use the dimension-raising module to increase the dimension and get the output H(x l ) and save it, so that H(x l ) dimension and F′(x l ) dimension is the same, and a detour is completed at this time; H(x l ) and F′(x l ) and output x after adding l+1 .
2. The target recognition method applicable to low computing power devices according to claim 1, characterized in that: The DP basic unit includes a depth convolution layer and a point convolution layer. The input is C0 channels. After passing through C0 depth convolution layers with a convolution kernel size of 3×3 and a channel of 1, a feature map with the number of channels C0 is obtained. Then, after passing through C1 point convolution layers with a convolution kernel size of 1×1 and a channel number of n, a feature map with a channel size of C1 is obtained. Among them, the convolution kernel size of the depth convolution layer is all 3×3, and the step size is all 1; the convolution kernel size of the point convolution layer is all 1×1, and the step size is all 1. The number of channels of the initial input data is 64.
3. The target recognition method applicable to low computing power devices according to claim 1 or 2, characterized in that: The target recognition network is a trained target recognition network; The number of epochs in training is set to 1000. During training, the original data is preprocessed to increase the data volume and enrich the data differences. The initial learning rate of the network is adjusted to 0.03, the batch_size of the training sample is adjusted to 64, and the loss value is allowed to decrease in the range of (2.2256-0.8195) for 400 epochs to obtain the initial loss value. The learning rate is then lowered to 0.003, and the loss is allowed to float in the range of (0.8142-0.6962) for 400 epochs. Finally, the learning rate is adjusted to 0.000 3, and the loss is allowed to float in the range of (0.6928-0.6909) for 200 epochs to obtain the final training set loss value.
4. A target recognition system suitable for low computing power devices, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the target recognition method suitable for low-computing power devices as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Neural network model compression method based on splicing convolution
CN111882053A
SAR image automatic target identification method based on depth separable convolutional neural network
CN113177465A