Lightweight target detection method and system based on convolutional neural network
By building a lightweight object detection model and optimization of centroid pruning algorithm, the problems of large computing volume and high memory usage on edge devices are solved, and efficient real-time object detection is achieved to meet the practical application needs of resource-constrained devices.
Patent Information
- Application Number
- CN202510715240.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The existing object detection methods have large computing volume and high memory usage on edge devices, making it difficult to meet the balance between computing complexity and detection accuracy, and cannot meet the real-time requirements.
A lightweight object detection model is built, and feature extraction and fusion is used using the FF-ELAN module, the EConv module and the MFA module, and combined with the optimization model of the centroid pruning algorithm based on the centroid to reduce the computational complexity and memory footprint.
Realize efficient and real-time object detection on edge devices, reduce computing complexity and memory usage, while maintaining high detection accuracy to meet practical application needs.
Smart Images

Figure CN120258049A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object detection, and particularly to a lightweight object detection method and system based on a convolutional neural network. Background Art
[0002] Object detection, as a key technology in the field of computer vision, has been widely applied in many scenarios such as robot navigation, intelligent video surveillance, and industrial inspection. It uses a vision model to replace manual detection, effectively reducing the consumption of human capital, and has become a hot research direction in both theoretical research and practical applications in recent years. With the vigorous development of deep learning technology, object detection algorithms based on deep convolutional neural networks have gradually occupied the mainstream position. Algorithms such as Faster R-CNN, YOLO, and SSD have achieved remarkable results in detection accuracy with their end-to-end learning methods.
[0003] However, when these advanced object detection algorithms are actually applied to edge computing devices, many problems are exposed. Edge computing devices (such as embedded systems, mobile terminals, etc.) usually have limited computing power, scarce memory resources, and limited power consumption. Most existing detection algorithms rely on complex network structures and have disadvantages such as large computational complexity, large number of parameters, and frequent memory access, bringing huge computational pressure and memory burden to edge devices and making it difficult to be directly deployed. Although researchers have proposed various lightweight object detection methods, such as MoblieNet - SSD, and tried to reduce computational complexity and memory occupancy through means such as model compression and network structure optimization, in the actual application process, they still face a series of severe challenges: 1. The trade-off problem between computational complexity and accuracy: Although existing lightweight methods have optimized the computational amount, when dealing with complex scenarios or high-resolution images, the detection accuracy will significantly decrease. Taking YOLOv8 as an example, while pursuing high accuracy, it still has a large demand for memory and computational resources and is difficult to operate efficiently on resource-constrained edge devices, and cannot well balance the relationship between computational complexity and detection accuracy.
[0004] 2. The memory access bottleneck restricts performance: The storage bandwidth and memory capacity of edge devices are inherently insufficient. Existing object detection methods need to frequently access memory when dealing with high-resolution images, which makes memory access a key bottleneck affecting performance, seriously reducing the inference speed and thus affecting the real-time performance and accuracy of object detection.
[0005] 3. Difficulty in meeting strict real-time requirements: In scenarios with extremely high real-time requirements such as autonomous driving and security monitoring, existing lightweight methods still cannot meet the actual needs of low latency and high throughput. For example, when the YOLO series of algorithms process high-resolution images, the inference speed is slow, unable to meet the requirements of real-time monitoring of high-bitrate videos, and it is difficult to ensure the efficient and stable operation of the system.
[0006] In summary, there is an urgent need for a new object detection method that can significantly reduce computational complexity and memory occupancy while ensuring detection accuracy, so as to adapt to resource-constrained edge computing devices and meet the strict requirements of real-time and efficiency in actual application scenarios. Summary of the Invention
[0007] To this end, the embodiments of the present invention provide a lightweight object detection method and system based on a convolutional neural network to solve the problems of large computational volume and high memory occupancy of existing object detection methods on edge devices.
[0008] To solve the above problems, the embodiments of the present invention provide a lightweight object detection method based on a convolutional neural network, the method comprising: Construct a lightweight object detection model based on a convolutional neural network, the model consisting of an FF-ELAN module, an EConv module, an MFA module, and a detection head, wherein the FF-ELAN module is used to extract shallow features and deep features of the input image simultaneously, the EConv module is used to replace traditional conventional convolutions with a lightweight convolution structure to reduce the computational complexity of the model and maintain the feature expression ability; the MFA module is used to achieve multi-level feature reuse and cross-feature fusion, and the detection head is used to predict the target category and location; The overall process of the model is: the input image sequentially passes through multiple FF-ELAN modules for feature extraction, and the features output by each FF-ELAN module are processed by the EConv module and then input to the MFA module for feature fusion; finally, the fused features are input to the detection head to complete the object detection task; Optimize the model using a centroid-based pruning algorithm to obtain an optimized model; Deploy the optimized model on an edge device; The edge device uses the optimized model to perform object detection on the collected image data and outputs the detection results.
[0009] Preferably, the FF-ELAN module adopts a dual-branch structure, including a left branch and a right branch; the left branch processes the input feature map using pointwise convolution to retain shallow features; the right branch uses a multi-layer connection combination of partial convolution to extract deep features, and each layer of partial convolution only calculates on partial channels of the input feature map.
[0010] Preferably, in the right branch, the output features of each layer of partial convolution are used as the input of the next layer of partial convolution, gradually enhancing the feature expression ability; the output features of the left branch and the right branch are combined in the channel dimension to generate combined features, and the combined features are fused through pointwise convolution to obtain features with rich semantic information.
[0011] Preferably, the MFA module constructs an EConv module through pointwise convolution and depthwise separable convolution to replace the traditional conventional convolution operation; the EConv module decomposes the conventional convolution into a depthwise separable convolution and a pointwise convolution. The depthwise separable convolution performs spatial convolution on each input channel separately, and the pointwise convolution adjusts the output of the depthwise separable convolution in the channel dimension.
[0012] Preferably, the calculation process of the EConv module is as follows: first, perform spatial convolution on the input feature map through a depthwise separable convolution, and then perform channel dimension reduction or expansion on the output of the depthwise separable convolution through a pointwise convolution to reduce the amount of calculation and memory occupancy and achieve feature fusion.
[0013] Preferably, the centroid-based pruning algorithm includes the following steps: For each convolutional kernel in the model, calculate its centroid in the spatial dimension. The centroid is the average value of the convolutional kernel in the spatial dimension, and the calculation formula for the centroid is as follows: ; In the formula, is the centroid of the convolutional kernel, represents the th convolutional kernel, is the number of convolutional kernels in this layer; Calculate the Euclidean distance between each convolutional kernel and the centroid. The Euclidean distance calculation formula is: ; In the formula, is the size of the convolutional kernel, and respectively represent the values of the convolutional kernel and the centroid at the position ( ); Set a distance threshold, and according to the distance calculation result, delete redundant convolutional kernels whose distance from the centroid is less than the distance threshold.
[0014] Preferably, the distance threshold is adaptively determined according to the distribution characteristics of each layer of convolutional kernels, and the calculation formula is: ; In the formula, is the distance threshold, is the average value of the distances between all convolutional kernels and the centroid, is the standard deviation, is the adjustment coefficient.
[0015] Preferably, the detection head includes a classification branch and a regression branch. The classification branch is used to predict the class probability of the target, and the regression branch is used to predict the position and size of the target; the detection head receives the feature map processed by the FF-ELAN module and the MFA module, and performs target class and position detection.
[0016] The embodiment of the present invention also provides a lightweight target detection system based on a convolutional neural network. This system is used to implement the above-mentioned lightweight target detection method based on a convolutional neural network, and specifically includes: A model construction module that constructs a lightweight target detection model based on a convolutional neural network. The model consists of an FF-ELAN module, an EConv module, an MFA module, and a detection head. The FF-ELAN module is used to simultaneously extract shallow features and deep features of the input image. The EConv module is used to replace traditional conventional convolutions with a lightweight convolution structure to reduce the model calculation complexity and maintain the feature expression ability; the MFA module is used to achieve multi-level feature reuse and cross-feature fusion, and the detection head is used to predict the target class and position; The overall process of the model is: the input image passes through multiple FF-ELAN modules in sequence for feature extraction. The features output by each FF-ELAN module are processed by the EConv module and then input to the MFA module for feature fusion; finally, the fused features are input to the detection head to complete the target detection task; A model optimization module, which is used to optimize the model using a centroid-based pruning algorithm to obtain an optimized model; A model deployment module, which is used to deploy the optimized model on edge devices; A target detection module, which is used for the edge device to perform target detection on the collected image data using the optimized model and output the detection result.
[0017] An embodiment of the present invention also provides a computer storage medium, which stores a computer software product. The computer software product includes a number of instructions for causing a computer device to execute the above-mentioned lightweight object detection method based on a convolutional neural network.
[0018] As can be seen from the above technical solutions, the present invention application has the following beneficial effects: (1) Low computational complexity: The present invention constructs a lightweight object detection model composed of FF-ELAN modules, EConv modules, MFA modules, etc., and uses a lightweight convolutional structure to replace traditional conventional convolutions, reducing the computational complexity of the model and the computational overhead when running on edge devices. At the same time, the centroid-based pruning algorithm deletes redundant convolutional kernels, further reducing the computational amount, enabling the model to run quickly under the limited computing power of edge devices.
[0019] (2) Low memory occupancy: On the one hand, the lightweight convolutional structure and feature fusion method adopted in the model design reduce the number of parameters and the amount of memory access; on the other hand, the centroid-based pruning algorithm optimizes the model, removes redundant convolutional kernels, reduces the model scale, and thus reduces memory occupancy, adapting to the characteristics of limited memory resources of edge devices.
[0020] (3) Good detection effect: The dual-branch structure of the FF-ELAN module can extract shallow and deep features simultaneously, the MFA module realizes multi-level feature reuse and cross-feature fusion, and the detection head accurately predicts the target category and location. Combining the optimized model, while reducing the computational amount and memory occupancy, it can still maintain a high object detection accuracy and meet the actual application requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly describe the drawings required to be used in the embodiments. By referring to the drawings, the features and advantages of the present invention will be more clearly understood. The drawings are schematic and should not be construed as limiting the present invention in any way. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them: Figure 1 It is a flowchart of a lightweight object detection method based on a convolutional neural network provided by the present invention; Figure 2 It is a schematic diagram of the overall architecture of a lightweight object detection model based on a convolutional neural network constructed by the present invention; Figure 3 It is a schematic diagram of a partial convolutional calculation process of the present invention; Figure 4It is a block diagram of a lightweight object detection system based on a convolutional neural network provided by the present invention. Specific Embodiments
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0023] Embodiment 1:
[0024] To solve the problems of large computational complexity and high memory occupancy of existing object detection methods on edge devices, as Figure 1 shown, an embodiment of the present invention proposes a lightweight object detection method based on a convolutional neural network. The method includes: S1: Construct a lightweight object detection model based on a convolutional neural network; S2: Optimize the model using a centroid-based pruning algorithm to obtain an optimized model; S3: Deploy the optimized model on an edge device; S4: The edge device uses the optimized model to perform object detection on the collected image data and outputs the detection result.
[0025] As can be seen from the above technical solutions, the present invention proposes a lightweight object detection method based on a convolutional neural network. By constructing a lightweight object detection model based on a convolutional neural network and using the unique designs of the FF-ELAN module, EConv module, MFA module, and detection head, the computational complexity of the model is effectively reduced and the feature extraction and fusion ability is improved, initially solving the problem of large computational complexity on edge devices; the centroid-based pruning algorithm further optimizes the model, deletes redundant convolutional kernels, reduces the computational amount and memory occupancy, and makes the model more adaptable to the resource limitations of edge devices; the optimized model is deployed on the edge device to ensure that the model can run stably in the actual edge scenario; finally, the edge device uses the model to perform object detection on the collected image data and outputs the result, realizing efficient and real-time object detection, successfully solving the problems of large computational complexity and high memory occupancy of existing object detection methods on edge devices, and meeting the actual application requirements.
[0026] In step S1, a lightweight object detection model based on a convolutional neural network is constructed, as Figure 2As shown in the figure, the model consists of an FF-ELAN module, an EConv module, an MFA module, and a detection head. The FF-ELAN module is used to extract shallow features and deep features of the input image simultaneously. The EConv module is used to replace the traditional conventional convolution with a lightweight convolution structure to reduce the computational complexity of the model and maintain the feature expression ability. The MFA module is used to achieve multi-level feature reuse and cross-feature fusion, and the detection head is used to predict the target category and location. The overall process of the model is as follows: The input image passes through multiple FF-ELAN modules in sequence for feature extraction. The features output by each FF-ELAN module are processed by the EConv module and then input to the MFA module for feature fusion. Finally, the fused features are input to the detection head to complete the target detection task.
[0027] Specifically, the FF-ELAN (Focused Feature Extraction Layer Aggregation Network) module: adopts a dual-branch structure, including a left branch and a right branch. The left branch uses pointwise convolution (1x1 convolution) to process the input feature map and retains shallow features. The right branch extracts deep features through the multi-layer connection combination of partial convolution. Each layer of partial convolution only calculates some channels of the input feature map, reducing the computational amount and memory access. The output of each layer of partial convolution is used as the input of the next layer, gradually enhancing the feature expression ability. Finally, the outputs of the left and right branches are combined in the channel dimension and then fused through pointwise convolution to obtain features with rich semantic information.
[0028] Conventional convolution: performs global feature extraction on the input feature map to capture the overall information of the image. The formula is as follows: ; In the formula, is the convolution output feature map, is the weight of the convolution kernel, is the input feature map, is the bias term.
[0029] Pointwise convolution: reduces the dimensionality and increases the dimensionality of the feature map in the channel dimension through 1x1 convolution, reducing the computational amount. The formula is as follows: ; In the formula, is the pointwise convolution output feature map, is the pointwise convolution kernel weight, is the bias term.
[0030] Partial Convolution: Reduces the number of convolutions on channels, reduces the computation for irrelevant regions, and pays more attention to extracting features from the central region. The computational cost of partial convolution is only of that of conventional convolution, significantly reducing the computational complexity and memory access. As Figure 3 shown, the specific calculation process is as follows: ; Among them, is the output feature map of partial convolution, and are the height and width of the feature map, is the convolution kernel size, is the number of input channels.
[0031] Through the dual-branch structure, the FF-ELAN module can extract shallow and deep features simultaneously, enhancing the diversity of features. The left branch retains the spatial detail information of the shallow layer, and the right branch gradually extracts the deep semantic information through partial convolution. Finally, combined features with rich semantic information are generated through feature fusion. The use of partial convolution significantly reduces the computational amount, and the introduction of pointwise convolution further reduces the computational overhead of the channel dimension. The dual-branch structure improves the computational efficiency of the module through parallel computing, which is especially suitable for resource-constrained edge devices.
[0032] Further, in the structural design of the MFA module, the present invention proposes a feature fusion bottleneck block (Bottleneck Block) aimed at efficiently fusing multi-level features while significantly reducing the computational complexity. This bottleneck block (layer) constructs an efficient EConv module through the combination of pointwise convolution and depthwise separable convolution to replace the traditional conventional convolution operation. This design not only reduces the computational amount and memory occupancy but also improves the efficiency of feature fusion, which is especially suitable for resource-constrained edge devices.
[0033] Specifically, the MFA module consists of three 1x1 convolutions (pointwise convolutions) and a bottleneck layer. The input data first enters the first 1x1 convolution for channel transformation. At the same time, the input data of another branch first passes through the second 1x1 convolution and then enters the bottleneck layer for feature processing. The third 1x1 convolution also operates on the input data. Finally, the outputs of these three branches converge at a node to achieve feature fusion, thereby completing multi-level feature reuse and cross-feature fusion. The EConv module consists of three parts. The top and bottom are 1x1 convolutions (pointwise convolutions), and the middle is DSC (depthwise separable convolution operation). The input data first passes through the 1x1 convolution at the top for preliminary channel adjustment, then enters the DSC in the middle for spatial convolution and other operations, and finally is further processed by the 1x1 convolution at the bottom. Moreover, the output of the 1x1 convolution at the top is also fused with the output of the 1x1 convolution at the bottom through a branch to achieve the purpose of reducing the computational complexity and maintaining the feature expression ability. The bottleneck layer consists of two EConv modules and two 1x1 convolutions. The input data first enters the first EConv module for processing, and its output then enters the second EConv module. At the same time, the input data also directly enters the second 1x1 convolution through a branch. The output of the first EConv module and the output of the second 1x1 convolution converge at a node and finally are processed by the first 1x1 convolution to achieve feature fusion and information integration.
[0034] Furthermore, the EConv module decomposes the conventional convolution into depthwise separable convolution and pointwise convolution, significantly reducing the computational complexity. The depthwise separable convolution performs spatial convolution on each input channel separately, while the pointwise convolution adjusts the output of the depthwise separable convolution in the channel dimension. The computational amount of the depthwise separable convolution is significantly lower than that of the conventional convolution.
[0035] The computational process of the EConv module is as follows: First, spatial convolution is performed on the input feature map through the depthwise separable convolution, and its computational amount is: ; In the formula, is the output feature map of the depthwise separable convolution, and are the height and width of the feature map, is the convolution kernel size, is the number of input channels.
[0036] Then, the output of the depthwise separable convolution is dimensionally reduced or increased in the channel dimension through the pointwise convolution, and the computational amount is:
[0037] In the formula, Number of output channels.
[0038] The total computational volume of the Econv module is:
[0039] In the formula, is the output feature map of the Econv module. Compared with conventional convolution, the computational volume of the EConv module is reduced by about 50%.
[0040] Furthermore, the Detection Head includes a classification branch and a regression branch. The classification branch is used to predict the class probability of the target, and the regression branch is used to predict the location and size of the target; the Detection Head receives the feature map processed by the FF-ELAN module and the MFA module to detect the target class and location.
[0041] Through the combination of depthwise separable convolution and pointwise convolution, the EConv module of the present invention significantly reduces the computational volume and memory occupancy. The bottleneck layer (block) enhances the diversity and expressive power of features through multi-level feature reuse and cross-feature fusion. Pointwise convolution realizes information interaction between channels, while depthwise separable convolution retains rich spatial information, and finally generates a more discriminative feature map through feature fusion. The design of the EConv module has a highly modular characteristic and can be flexibly embedded into different network structures to improve the overall performance of the model.
[0042] In step S2, the above model is optimized using a centroid-based pruning algorithm to obtain an optimized model.
[0043] The present invention proposes a centroid-based pruning (CDP) algorithm for the post-processing stage of the model, aiming to significantly reduce the computational volume and memory occupancy of the model by calculating the centroid of the convolution kernel and deleting redundant convolution kernels close to the centroid. The CDP algorithm optimizes the model structure by systematically identifying and removing redundant convolution kernels, while maintaining high detection accuracy, and is particularly suitable for resource-constrained edge devices.
[0044] The core idea of the CDP algorithm is to calculate the centroid of the convolution kernel and delete redundant convolution kernels close to the centroid, thereby reducing the number of model parameters and computational volume. The specific steps are as follows: Centroid calculation: For the convolution kernel of each layer in the model, calculate its centroid in the spatial dimension. The centroid is the average value of the convolution kernel in the spatial dimension, which reflects the overall distribution characteristics of the convolution kernel of this layer. The formula for calculating the centroid is as follows: ; In the formula, is the centroid of the convolutional kernel, represents the th convolutional kernel, and
[0045] Distance metric: Calculate the Euclidean distance between each convolutional kernel and the centroid, where the Euclidean distance is calculated as follows: ; In the formula, is the size of the convolutional kernel, and respectively represent the values of the convolutional kernel and the centroid at the position ([[]] ).
[0046] Pruning strategy: Set a distance threshold, and according to the distance calculation result, delete the redundant convolutional kernels whose distance from the centroid is less than the distance threshold. Specifically, the distance threshold is adaptively determined according to the distribution characteristics of the convolutional kernels in each layer, and the calculation formula is: ; In the formula, is the distance threshold, is the average value of the distances between all convolutional kernels and the centroid, is the standard deviation, is the adjustment coefficient.
[0047] By calculating the centroid of the convolutional kernel, redundant convolutional kernels can be quickly identified and deleted, significantly reducing the number of model parameters and the computational complexity. Since the convolutional kernels with a large distance from the centroid are retained, these convolutional kernels usually contain more unique information. Therefore, the pruned model can maintain a high detection accuracy. The pruning strategy is adaptively adjusted according to the distribution characteristics of the convolutional kernels in each layer, avoiding the accuracy loss caused by a fixed pruning rate in traditional pruning methods.
[0048] The CDP algorithm is particularly suitable for resource-constrained edge devices, and can reduce the model's computational complexity and memory occupancy while maintaining a high detection accuracy. For example, in scenarios with high real-time requirements such as autonomous driving and security monitoring, the CDP algorithm can significantly improve the inference speed of the model and meet the requirements of low latency and high throughput.
[0049] In step S3, the optimized model is deployed on the edge device.
[0050] In step S4, the edge device uses the optimized model to perform object detection on the collected image data and outputs the detection result.
[0051] Embodiment 2:
[0052] As Figure 4As shown in the figure, the present invention provides a lightweight object detection system based on a convolutional neural network. This system is used to implement the lightweight object detection method based on a convolutional neural network in the first embodiment above, and specifically includes: A model construction module 100 constructs a lightweight object detection model based on a convolutional neural network. The model consists of an FF-ELAN module, an EConv module, an MFA module, and a detection head. Among them, the FF-ELAN module is used to extract shallow features and deep features of the input image simultaneously. The EConv module is used to replace traditional conventional convolutions with a lightweight convolution structure to reduce the computational complexity of the model and maintain the feature expression ability. The MFA module is used to achieve multi-level feature reuse and cross-feature fusion. The detection head is used to predict the target category and location. The overall process of the model is as follows: The input image passes through multiple FF-ELAN modules in sequence for feature extraction. The features output by each FF-ELAN module are processed by the EConv module and then input to the MFA module for feature fusion. Finally, the fused features are input to the detection head to complete the object detection task. A model optimization module 200 is used to optimize the model using a centroid-based pruning algorithm to obtain an optimized model. A model deployment module 300 is used to deploy the optimized model on edge devices. An object detection module 400 is used for edge devices to perform object detection on the collected image data using the optimized model and output the detection results.
[0053] A lightweight object detection system based on a convolutional neural network in this embodiment is used to implement the aforementioned lightweight object detection method based on a convolutional neural network. Therefore, the specific implementation manners in the lightweight object detection system based on a convolutional neural network can be seen in the embodiment part of the aforementioned lightweight object detection method based on a convolutional neural network. For example, the model construction module 100, the model optimization module 200, the model deployment module 300, and the object detection module 400 are respectively used to implement steps S1, S2, S3, and S4 in the aforementioned lightweight object detection method based on a convolutional neural network. Therefore, its specific implementation manners can refer to the descriptions of the corresponding individual part embodiments. To avoid redundancy, they will not be elaborated here.
[0054] Embodiment Three:
[0055] The embodiment of the present invention provides a computer storage medium. The computer storage medium stores a computer software product. The computer software product includes several instructions for causing a computer device to execute the aforementioned lightweight object detection method based on a convolutional neural network.
[0056] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0057] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0058] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means, and the instruction means implements the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0059] Obviously, the above embodiments are only examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or variations derived therefrom are still within the protection scope of the present invention.
Claims
1. A lightweight object detection method based on convolutional neural network, characterized in that Including: Construct a lightweight object detection model based on a convolutional neural network. The model consists of an FF-ELAN module, an EConv module, an MFA module, and a detection head. Among them, the FF-ELAN module is used to extract shallow features and deep features of the input image simultaneously. The EConv module is used to replace the traditional conventional convolution with a lightweight convolution structure to reduce the computational complexity of the model and maintain the feature expression ability. The MFA module is used to achieve multi-level feature reuse and cross-feature fusion. The detection head is used to predict the target category and location. The overall process of the model is as follows: The input image passes through multiple FF-ELAN modules in sequence for feature extraction. After the features output by each FF-ELAN module are processed by the EConv module, they are input into the MFA module for feature fusion. Finally, the fused features are input into the detection head to complete the object detection task. Among them, the calculation process of the EConv module is as follows: First, perform spatial convolution on the input feature map through depthwise separable convolution, and then perform dimensionality reduction or dimensionality increase on the output of the depthwise separable convolution in the channel dimension through pointwise convolution to reduce the computational amount and memory occupancy and achieve feature fusion. Optimize the model using a centroid-based pruning algorithm to obtain an optimized model. Deploy the optimized model on an edge device. The edge device uses the optimized model to perform object detection on the collected image data and outputs the detection results.
2. The lightweight object detection method based on a convolutional neural network according to claim 1, wherein The FF-ELAN module adopts a dual-branch structure, including a left branch and a right branch. The left branch uses pointwise convolution to process the input feature map to retain shallow features. The right branch uses a multi-layer connection combination of partial convolution to extract deep features. Each layer of partial convolution only calculates part of the channels of the input feature map.
3. The lightweight object detection method based on a convolutional neural network according to claim 2, wherein In the right branch, the output features of each layer of partial convolution are used as the input of the next layer of partial convolution, gradually enhancing the feature expression ability. The output features of the left branch and the right branch are combined in the channel dimension to generate combined features, and the combined features are fused through pointwise convolution to obtain features with rich semantic information.
4. The lightweight object detection method based on a convolutional neural network according to claim 1, characterized in that The MFA module constructs an EConv module through pointwise convolution and depthwise separable convolution to replace the traditional conventional convolution operation. The EConv module decomposes the conventional convolution into depthwise separable convolution and pointwise convolution. The depthwise separable convolution performs spatial convolution on each input channel separately, and the pointwise convolution adjusts the channel dimension of the output of the depthwise separable convolution.
5. The lightweight object detection method based on convolutional neural network according to claim 1, characterized in that The centroid-based pruning algorithm includes the following steps: For each convolutional kernel in the model, calculate its centroid in the spatial dimension. The centroid is the average value of the convolutional kernel in the spatial dimension. The formula for the centroid is as follows: ; wherein, is the centroid of the convolution kernel, represents the th convolution kernel, is the number of convolution kernels of this layer; Calculate the Euclidean distance between each convolutional kernel and the centroid, where the Euclidean distance The calculation formula is: ; In the formula, is the size of the convolution kernel, and respectively represent the values of the convolution kernel and the centroid at the position ([[]] ); Set a distance threshold. According to the distance calculation result, delete the redundant convolutional kernels whose distance from the centroid is less than the distance threshold.
6. The lightweight object detection method based on convolutional neural network according to claim 5, characterized in that The distance threshold is adaptively determined according to the distribution characteristics of each layer of convolutional kernels. The calculation formula is: ; Wherein, is the distance threshold,[ is the average value of the distances between all convolution kernels and the centroid,[ is the standard deviation,[ is the adjustment coefficient.[ 7. The lightweight object detection method based on a convolutional neural network according to claim 1, characterized in that The detection head includes a classification branch and a regression branch. The classification branch is used to predict the class probability of the target, and the regression branch is used to predict the position and size of the target. The detection head receives the feature map processed by the FF-ELAN module and the MFA module to detect the target class and position.
8. A lightweight object detection system based on a convolutional neural network, characterized in that, The system is used to implement the lightweight object detection method based on convolutional neural network according to any one of claims 1 to 7, and specifically includes: A model construction module that constructs a lightweight object detection model based on a convolutional neural network. The model consists of an FF-ELAN module, an EConv module, an MFA module, and a detection head. The FF-ELAN module is used to extract shallow features and deep features of the input image simultaneously. The EConv module is used to replace the traditional conventional convolution with a lightweight convolution structure to reduce the computational complexity of the model and maintain the feature expression ability. The MFA module is used to achieve multi-level feature reuse and cross-feature fusion, and the detection head is used to predict the target class and position. The overall process of the model is as follows: The input image is sequentially subjected to feature extraction by multiple FF-ELAN modules. The features output by each FF-ELAN module are processed by the EConv module and then input to the MFA module for feature fusion. Finally, the fused features are input to the detection head to complete the object detection task. A model optimization module that is used to optimize the model using a centroid-based pruning algorithm to obtain an optimized model. A model deployment module that is used to deploy the optimized model on an edge device. An object detection module that is used for the edge device to perform object detection on the collected image data using the optimized model and output the detection results.
9. A computer storage medium, characterized in that, The computer storage medium stores a computer software product, and the computer software product includes several instructions for causing a computer device to execute the lightweight object detection method based on convolutional neural network according to any one of claims 1 to 7.
Citation Information
Patent Citations
Lightweight double-branch convolutional neural network for image target detection and detection method thereof
CN114648684A
Lightweight target detection method, and lightweight method and device of target detection model
CN119942137A
Cited By
Mask detection method fusing machine vision and edge calculation
CN120599430A