Target detection algorithm lightweight method and system for embedded low-computing-power platform

By applying self-supervised pre-training, distillation learning, pruning and quantization technologies on an embedded low-computing platform, the object detection model is lightweighted, and the problems of limited computing resources, large memory consumption and high power consumption in the existing technology are solved, and efficient and real-time object detection is achieved.

CN120147772APending Publication Date: 2025-06-13NEWLAND YOUMAIJIE (GUANGDONG PROVINCE) TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510053597.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing object detection algorithm based on deep learning has limited computing resources, large memory consumption and high power consumption on embedded low-computing platforms, resulting in poor real-time performance and short battery life.

Method used

Through self-supervised pre-training, distillation learning, structured pruning and quantization, the object detection model is lightweighted, replaced the backbone network of the Yolov8-x model as MobileOne, trained the student model using the detection head cross-distillation method, and reduced the computational volume and storage requirements of the model through pruning and quantization.

Benefits of technology

It significantly reduces the model's dependence on computing resources and storage, improves inference speed and detection accuracy, and meets the real-time requirements of embedded devices when computing resources are limited.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147772A_ABST
    Figure CN120147772A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection algorithm lightweight method and system for an embedded low-computing-power platform, and the method comprises the steps: carrying out the image collection of a task object, and obtaining a data set with diversified samples; selecting a Yov8-x model as a teacher model, performing MIM self-supervision pre-training on a backbone network of the Yov8-x model, and performing fine tuning on the Yov8-x model; selecting MobileOne to carry out ConvModule module replacement as a student model, and training the student model through a detection head cross distillation method; and a structured pruning method is used to prune the trunk network of the student model after distillation learning training, the student model after pruning is subjected to lightweight processing, and the student model after lightweight processing is used for designating embedded platform deployment. According to the method, the problems of limited computing resources, overlarge memory consumption and high power consumption of a target detection algorithm based on deep learning are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a method and system for lightweighting target detection algorithms for embedded low-computing-power platforms. Background Art

[0002] In the field of embedded applications, target detection is a key computer vision task, which involves identifying and locating specific objects in images or videos. Traditional target detection algorithms, such as those based on HOG (Histogram of Oriented Gradients) or SIFT (Scale-Invariant Feature Transform), perform well in some cases, but usually require high computing resources and memory and have poor robustness. In recent years, with the rise of deep learning, methods such as convolutional neural networks (CNNs) have made significant progress in target detection tasks, mainly in the following aspects:

[0003] ① The rise of single-stage detectors: Single-stage object detectors, such as Yolo and SSD, have achieved great success in the field of target detection. They no longer rely on complex two-stage detection processes and simplify the target detection problem into an end-to-end regression task. These algorithms achieve faster inference speeds, making real-time target detection possible.

[0004] ② Further optimization of backbone networks: Continuous optimization and innovation in the backbone network part have led to higher detection accuracy. Many new backbone network structures improve the performance of target detection by introducing special attention mechanisms and residual structures.

[0005] ③ Continuous optimization of feature fusion modules: Feature fusion modules are usually used for the fusion of multi-scale features. Many algorithms innovatively propose modules such as FPN, PAN, SPP, and GD to improve the feature fusion effect and the target scale adaptability of target detection algorithms.

[0006] ④ Continuous optimization of detection head modules: Innovative designs of detection heads such as decoupled detection heads, Anchor-free detection heads, and rotation angle detection heads continuously improve the accuracy of detection algorithms.

[0007] However, existing deep learning-based target detection algorithms still have deficiencies, mainly including:

[0008] ①Limited computing resources: Advanced object detection algorithms typically use deep convolutional neural networks, which contain a large number of layers and parameters. For each input image, these models need to perform a large number of convolutional, pooling, and fully connected operations, resulting in a high computational load. In addition, the residual structure in the network architecture requires a large number of memory access operations, seriously affecting efficiency. Embedded computing hardware, especially low-power NPUs or CPUs, often struggles to cope with this high computational requirement, leading to high latency and poor real-time performance.

[0009] ②Excessive memory consumption: Modern object detection models usually require a large amount of memory to store model parameters and intermediate feature maps. The sizes of these parameters and feature maps typically far exceed the memory capacity of embedded devices. This not only limits the scale of the model but also reduces the flexibility of the model because large models cannot be loaded when the memory capacity is limited.

[0010] ③High power consumption: High computational requirements and memory consumption lead to high power consumption problems. This poses a threat to the battery life of embedded devices, especially for applications such as mobile devices and drones that need to run for a long time. High power consumption may also cause the device to overheat, reducing its reliability and service life.

[0011] Chinese Patent No. CN117173395A discloses a YOLOv8 partial convolution network object detection method, including the following steps: Select the MS COCO 2017 dataset; construct a feature extraction network based on YOLOv8, replace the backbone network of the YOLOv8 model with the partial convolution network FasterNet, and perform feature extraction on the initial target image; add a BiFormer attention module to the c2f module outside the backbone network of the YOLOv8 model to purify the extracted features; determine the region of interest, and apply token-to-token attention to capture key information in the input tensor; replace the CIoU loss function with the Wise-IoU loss function to complete the construction of the initial target detection model; use the new dataset to train the initial target detection model to obtain the final target detection model; use the final target detection model to perform object detection on the image to be detected and evaluate the performance of the model. Although this invention may have certain advantages in purifying features and improving accuracy by introducing innovative modules such as the FasterNet convolution network, BiFormer attention module, and Wise-IoU loss function in the YOLOv8 model, it also introduces a relatively high computational overhead and increases the "black box" nature of the model, affecting the credibility and debuggability of the model. Summary of the Invention

[0012] To solve the problems existing in the above-mentioned prior art, the present invention provides a method and system for lightweighting a target detection algorithm for an embedded low-computation platform, which lightweight the target detection algorithm to make it more suitable for embedded systems with limited resources.

[0013] The technical solution of the present invention is as follows:

[0014] On the one hand, the present invention provides a method for lightweighting a target detection algorithm for an embedded low-computation platform, including the following steps:

[0015] Collect images of task objects to obtain a dataset with diverse samples, and use a random partitioning method to partition and output the dataset as a training set, a validation set, and a test set.

[0016] Select the Yolov8-x model as the teacher model, perform MIM self-supervised pre-training on the backbone network of the Yolov8-x model, and fine-tune the Yolov8-x model. After the fine-tuning is completed, output the teacher model weights for subsequent distillation learning of the student model.

[0017] Select MobileOne to replace the ConvModule module in the backbone network of the Yolov8-x model as the student model, and train the student model through the detection head cross-distillation method. After the training is completed, output the student model weights for subsequent lightweight deployment.

[0018] Use the structured pruning method to prune the backbone network of the student model that has completed distillation learning training, and perform lightweight processing on the pruned student model. The lightweight processed student model is used for deployment on a specified embedded platform.

[0019] Preferably, the MIM self-supervised pre-training of the backbone network of the Yolov8-x model is specifically as follows:

[0020] Design a corresponding image decoder according to the backbone network of the Yolov8-x model and the downsampling ratio of its corresponding feature layer, and combine it with the backbone network of the Yolov8-x model as an encoder to form an AutoEncoder network.

[0021] Use the MIM self-supervised learning method to perform random masking of the input training images with a preset area ratio, and output the training images with masks to guide the decoder of the AutoEncoder network to learn the correlation between different local semantics.

[0022] Pre-train the AutoEncoder network. During the feature extraction process of training, the encoder uses sparse convolution to skip the masked parts and only extracts features from the non-masked regions for image reconstruction of the masked parts. During the process of reconstructing the image in training, the decoder uses transposed convolution to upsample the output features of the encoder to obtain an output image with the same resolution as the input image.

[0023] Use the L1 loss to calculate the reconstruction error between the input image and the output image in the masked region.

[0024] Update the parameters of the AutoEncoder network through the SGD optimizer to minimize the reconstruction error, and finally obtain the pre-trained weights of the backbone network of the teacher model.

[0025] Preferably, the fine-tuning of the Yolov8-x model is specifically as follows:

[0026] Add the pre-trained Encoder weights to the original backbone network of the Yolov8-x model, freeze the weights of the first N convolutional layers of the backbone network, retain the general features of the shallow layers to improve the generalization of the subsequent model, and fine-tune the Yolov8-x model based on the training set.

[0027] Preferably, the selection of MobileOne to replace the ConvModule module in the backbone network of the Yolov8-x model as the student model is specifically as follows:

[0028] Select MobileOne to replace the ConvModule module in the backbone network of the Yolov8-x model. Take the 1 / 8, 1 / 16, and 1 / 32 times downsampled feature layers of MobileOne as the input of the Neck and denote them as C3, C4, and C5 respectively. The number of channels is set to 128, 256, and 512 respectively. Keep the remaining structures of the Yolov8-x model, including the Neck network and the Head network, unchanged as the student model.

[0029] Preferably, the training of the student model by the detection head cross-distillation method is specifically as follows:

[0030] Freeze all the weights of the teacher model.

[0031] Extract C3, C4, and C5 of the student model during the training process and send them to the corresponding detection heads of the teacher model, and use a 1x1 convolution module to align the number of channels for image prediction and inference.

[0032] Calculate the similarity error between the image prediction results of the teacher model and the image prediction results of the student model through the L2 loss, and weight and merge them with the Ground-Truth error of the student model according to the preset weight ratio to obtain the final training loss.

[0033] Update the parameters of the student model using the SGD optimizer to train the student model by minimizing the training loss.

[0034] Preferably, use a structured pruning method to prune the backbone network of the student model that has completed distillation learning training, specifically:

[0035] According to the activation values of the BN layers in the backbone network of the student model, delete the convolutional kernels with activation values less than the preset threshold and the subsequent convolutional kernels corresponding to them.

[0036] Fine-tune the pruned student model using the SGD optimizer and output the pruned student model.

[0037] Preferably, the lightweight processing of the pruned student model specifically includes:

[0038] Reconstruct the parameters of the student model, merge the residual 3x3 convolution weights, residual 1x1 convolution weights, and BN layer weights in the residual convolution module to obtain a straight-through 3x3 convolution backbone network.

[0039] Perform PTQ quantization on the student model, use M typical scenario images for quantization calibration, and quantize the student model weights from FP32 to INT8.

[0040] On the other hand, the present invention provides a lightweight system for object detection algorithms for embedded low-computing power platforms, including a data acquisition module, a teacher model pre-training and fine-tuning module, a student model distillation training module, and a student model pruning and lightweight module.

[0041] The data acquisition module is used to collect images of task objects to obtain a dataset with diverse samples, and use a random partitioning method to partition and output the dataset as a training set, a validation set, and a test set.

[0042] The teacher model pre-training and fine-tuning module is used to select the Yolov8-x model as the teacher model, perform MIM self-supervised pre-training on the backbone network of the Yolov8-x model, and fine-tune the Yolov8-x model. After fine-tuning, output the teacher model weights for subsequent distillation learning of the student model.

[0043] The student model distillation training module is used to select MobileOne to replace the ConvModule module in the backbone network of the Yolov8-x model as the student model, and train the student model through the detection head cross-distillation method. After training, output the student model weights for subsequent lightweight deployment.

[0044] Student model pruning and lightweight module, which is used to prune the backbone network of the student model that has completed distillation learning training using a structured pruning method, and lightweight the pruned student model. The lightweighted student model is used for deployment on a specified embedded platform.

[0045] In another aspect, the present invention also provides an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the target detection algorithm lightweighting method for an embedded low-computing-power platform as described in any embodiment of the present invention.

[0046] In another aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the target detection algorithm lightweighting method for an embedded low-computing-power platform as described in any embodiment of the present invention.

[0047] Compared with the prior art, the present invention has the following technical effects:

[0048] By combining multiple technologies such as self-supervised pre-training, distillation learning, pruning, and quantization, the present invention significantly improves the performance and efficiency of the target detection model. Among them, self-supervised pre-training enhances the feature representation ability of the teacher model, and fine-tuning further optimizes the performance of the model in the target detection task. Through the detection head cross-distillation method, the knowledge of the teacher model is efficiently transferred to the student model, enabling it to achieve high detection accuracy while maintaining a small scale. And lightweight means such as structured pruning and quantization are adopted to significantly reduce the computational amount and storage requirements of the model, so that it can be efficiently deployed on an embedded low-computing-power platform. Overall, the method proposed by the present invention can significantly reduce the dependence of the model on computing resources and storage while ensuring high accuracy, improve the inference speed and reduce energy consumption, meeting the real-time requirements of embedded devices with limited computing resources. Description of the Drawings

[0049] Figure 1 is the overall flowchart of the target detection algorithm lightweighting method for an embedded low-computing-power platform described in the present invention;

[0050] Figure 2 is the overall structure diagram of a Yolov8 model that can be selected as the teacher model;

[0051] Figure 3 is the structural schematic diagram of the MobileOne module used by the student model to replace the ConvModule module in the teacher model;

[0052] Figure 4 is the model detection effect diagram after the final training of the application example is completed. Specific implementation manner

[0053] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will combine specific embodiments of the present application and refer to the accompanying drawings to clearly and completely describe the technical solutions of the present invention.

[0054] Embodiment 1

[0055] This embodiment provides a method for lightweighting a target detection algorithm for an embedded low-computing-power platform, trains a teacher model using self-supervised pre-training + fine-tuning, trains a student model using detection head cross-distillation, and finally uses lightweight means such as model pruning + parameter reconstruction + quantization to complete the deployment on the embedded side. Refer to Figure 1 as shown, and specifically includes the following steps:

[0056] Collect images of task objects to obtain a dataset with diverse samples, and use a random partitioning method to partition and output the dataset as a training set, a validation set, and a test set. The partitioning ratio of the dataset is not limited, and various factors such as the size of the data volume or task requirements can be considered for dataset partitioning. In this embodiment, the dataset is preferably partitioned into a training set, a validation set, and a test set in a ratio of 6:2:2. The following models will use these three parts of the dataset for corresponding training, parameter tuning, and evaluation. Specifically: the training set is used for training all models, the validation set is used for performance verification during the training process of all models, and the test set is used for performance testing after all models are trained.

[0057] As Figure 2 shown, select the Yolov8-x model as the teacher model, perform MIM self-supervised pre-training on the backbone network of the Yolov8-x model, and fine-tune the Yolov8-x model. After fine-tuning, output the teacher model weights for subsequent distillation learning of the student model.

[0058] As Figure 3 shown, select MobileOne to replace the ConvModule module in the backbone network of the Yolov8-x model as the student model, and train the student model using the detection head cross-distillation method. After training, output the student model weights for subsequent lightweight deployment.

[0059] Use the structured pruning method to prune the backbone network of the student model that has completed distillation learning training, and perform lightweight processing on the pruned student model. The lightweight processed student model is used for deployment on a specified embedded platform.

[0060] As a preferred implementation manner of this embodiment, the specific operation of performing MIM self-supervised pre-training on the backbone network of the Yolov8-x model is as follows:

[0061] According to the backbone network of the Yolov8-x model and the downsampling ratio of its corresponding feature layers, a corresponding image decoder, namely Decoder, is designed. Together with the backbone network of the Yolov8-x model as Encoder, they form an AutoEncoder network, and the AutoEncoder network includes Encoder and Decoder.

[0062] Using the MIM self-supervised learning method, randomly mask the input training images with a preset area ratio, and output the training images with masks, which are used to guide the decoder of the AutoEncoder network to learn the correlation between different local semantics. The preset area ratio is set according to the actual task requirements and is not limited. In this embodiment, the area ratio of the random mask is preferably 75%.

[0063] Pre-train the AutoEncoder network. During the feature extraction process of training, Encoder uses sparse convolution to skip the masked parts and only extracts features from the unmasked areas for image reconstruction of the masked parts; during the process of reconstructing the image in training, Decoder uses transposed convolution to upsample the output features of Encoder to obtain an output image with the same resolution as the input image.

[0064] Use the L1 loss to calculate the reconstruction error between the input image and the output image in the masked area.

[0065] Update the parameters of the AutoEncoder network through the SGD optimizer to minimize the reconstruction error, and finally obtain the pre-trained weights of the backbone network of the teacher model.

[0066] As a preferred implementation of this embodiment, the fine-tuning of the Yolov8-x model is specifically as follows:

[0067] Add the pre-trained Encoder weights to the original backbone network of the Yolov8-x model, freeze the weights of the first N convolutional layers of the backbone network, retain the general features of the shallow layers to improve the generalization of the subsequent model, and fine-tune the Yolov8-x model based on the training set.

[0068] As a preferred implementation of this embodiment, the specific method of selecting MobileOne to replace the ConvModule module in the backbone network of the Yolov8-x model as the student model is as follows:

[0069] Select MobileOne to replace the ConvModule module in the backbone network of the Yolov8-x model. Take the 1 / 8, 1 / 16, and 1 / 32 downsampled feature layers of MobileOne as the inputs of the Neck and denote them as C3, C4, and C5 respectively. The number of channels is set to 128, 256, and 512 respectively. Keep the remaining structures of the Yolov8-x model, including the Neck network and the Head network, unchanged as the student model.

[0070] As a preferred implementation manner of this embodiment, training the student model by the detection head cross-distillation method is specifically as follows:

[0071] Freeze all the weights of the teacher model.

[0072] Respectively extract C3, C4, and C5 of the student model during the training process and send them to the corresponding detection heads of the teacher model, and use a 1x1 convolution module to align the number of channels for image prediction inference. Since the network architectures, depths, and numbers of convolution kernels of the student model and the teacher model are different, the feature maps at different levels may have different numbers of channels. In this embodiment, a 1x1 convolution module is used to adjust the number of channels output by the student model without changing the spatial dimensions (width and height) of the feature maps, so that it matches the number of channels of the corresponding layer of the teacher model, ensuring that the output dimensions of both are the same, which is convenient for calculating and comparing errors.

[0073] Calculate the similarity error between the image prediction results of the teacher model and the student model through the L2 loss, and weighted merge it with the Ground-Truth error of the student model (that is, the loss between the prediction result and the true label when the student model performs task prediction. In this embodiment, the loss is calculated through Box-Loss, Cls-Loss, and Dfl-Loss) according to a preset weight ratio to obtain the final training loss. The specific weight ratio is adjusted according to the actual training situation and is not limited here. In this embodiment, the preferred weight ratio of the L2 loss to the Ground-Truth error is 1:9.

[0074] Update the parameters of the student model through the SGD optimizer to train the student model by minimizing the training loss.

[0075] As a preferred implementation manner of this embodiment, using the structured pruning method to prune the backbone network of the student model that has completed the distillation learning training is specifically as follows:

[0076] According to the activation values of the BN layers in the backbone network of the student model, delete the convolution kernels with activation values less than the preset threshold and the subsequent convolution kernels corresponding to them.

[0077] Fine-tune and train the pruned student model using the SGD optimizer, and output the pruned student model.

[0078] As a preferred implementation of this embodiment, the lightweight processing of the pruned student model specifically includes:

[0079] Perform parameter reconstruction on the student model, merge the residual 3x3 convolution weights, residual 1x1 convolution weights, and BN layer weights in the residual convolution module to obtain a straight-through 3x3 convolution backbone network.

[0080] Perform PTQ quantization on the student model, use M typical scenario images for quantization calibration, and quantize the student model weights from FP32 to INT8.

[0081] If the embedded platform supports it, continue to use perceptual quantization to fine-tune the INT8 student model to maintain its accuracy to the greatest extent.

[0082] To verify the effectiveness and superiority of the method provided in this embodiment, the following provides some specific cases:

[0083] During the training of the teacher model, in the training of the AutoEncoder network, the learning rate of the SGD optimizer is set to 0.0002, and the weights of the first 5 convolutional layers of the backbone network are frozen during the fine-tuning of the teacher model, so as to use the strategy of self-supervised pre-training and fine-tuning based on sparse convolution MIM for training. As shown in Table 1, compared with directly training through Ground-Truth, the mean average precision of the teacher model is improved when calculating at the threshold of IoU = 0.5.

[0084] Table 1 Comparison table of teacher model mAP@0.5 training indicators

[0085] Dataset Direct training / mAP@0.5 Training in this embodiment / mAP@0.5 Bottle cap detection (3000 images) 99.5 99.7 New energy battery detection (5000 images) 98.9 99.6 Electronic device detection (80000 images) 96.5 99.0

[0086] During the training of the student model, the learning rate of the SGD optimizer is set to 0.0001, and the detection head cross-distillation method is used to train the student model. As shown in Table 2, compared with directly training through Ground-Truth, the accuracy of the student model is improved.

[0087] Table 2 Comparison table of student model mAP@0.5 training indicators

[0088] Dataset Direct training / mAP@0.5 Training in this embodiment / mAP@0.5 Bottle cap detection (3000 images) 97.7 98.9 New energy battery detection (5000 images) 96.5 98.1 Electronic device detection (80000 images) 91.2 94.6

[0089] Perform pruning and lightweight processing on the student model that has completed distillation learning training. The learning rate of the SGD optimizer is set to 0.0001, and train for 10 Epochs. As shown in Table 3, appropriate lightweight operations can be selected to balance the running efficiency and performance of the model.

[0090] Table 3 Comparison table of model performance and running efficiency

[0091]

[0092] As Figure 4 shown, the student model obtained through the above training, parameter tuning and evaluation can effectively perform object detection and recognition.

[0093] Embodiment 2

[0094] Correspondingly, this embodiment provides a lightweight system for object detection algorithms for embedded low-computation platforms. The system is used to implement the lightweight method for object detection algorithms for embedded low-computation platforms as described in any embodiment of the present invention, and includes a data acquisition module, a teacher model pre-training and fine-tuning module, a student model distillation training module, and a student model pruning and lightweight module.

[0095] The data acquisition module is used to collect images of task objects to obtain a dataset with diverse samples, and use a random partitioning method to partition and output the dataset as a training set, a validation set, and a test set.

[0096] The teacher model pre-training and fine-tuning module is used to select the Yolov8-x model as the teacher model, perform MIM self-supervised pre-training on the backbone network of the Yolov8-x model, and fine-tune the Yolov8-x model. After fine-tuning, the weights of the teacher model are output for subsequent distillation learning of the student model.

[0097] The student model distillation training module is used to select MobileOne to replace the ConvModule module in the backbone network of the Yolov8-x model as the student model, and train the student model through the detection head cross-distillation method. After training, the weights of the student model are output for subsequent lightweight deployment.

[0098] The student model pruning and lightweight module is used to use a structured pruning method to prune the backbone network of the student model that has completed distillation learning training, and perform lightweight processing on the pruned student model. The lightweight processed student model is used for deployment on a specified embedded platform.

[0099] Embodiment 3

[0100] This embodiment provides an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the lightweight method for object detection algorithms for embedded low-computation platforms as described in any embodiment of the present invention.

[0101] Embodiment 4

[0102] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the target detection algorithm lightweight method for an embedded low-computing-power platform as described in any embodiment of the present invention.

[0103] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent the case where A exists alone, A and B exist simultaneously, or B exists alone. Where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c may represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c may be single or multiple.

[0104] Those of ordinary skill in the art can realize that the units and algorithm steps described in the embodiments disclosed herein can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0105] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0106] In several embodiments provided by the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (hereinafter referred to as ROM), random access memories (hereinafter referred to as RAM), magnetic disks, or optical discs that can store program codes.

[0107] The above are only embodiments of the present invention, and thus do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.

Claims

1. A lightweight method for target detection algorithm for embedded low-computing power platforms, characterized in that: The following steps are involved: Capture images of the task object to obtain a dataset with diverse samples. Use random partitioning to divide the dataset and output it into training set, validation set, and test set. Select the Yolov8-x model as the teacher model, perform MIM self-supervised pre-training on the backbone network of the Yolov8-x model, and fine-tune the Yolov8-x model. After fine-tuning, output the teacher model weights for subsequent distillation learning of the student model. Select MobileOne to replace the ConvModule module in the Yolov8-x model backbone network as the student model, train the student model through the detection head cross distillation method, and output the student model weights after training for subsequent lightweight deployment; A structured pruning method is used to prune the backbone network of the student model that has completed distillation learning training, and the pruned student model is lightweight. The lightweight student model is used for deployment on a specified embedded platform.

2. The lightweight method for target detection algorithm for embedded low computing power platform according to claim 1 is characterized in that: The MIM self-supervised pre-training of the backbone network of the Yolov8-x model is as follows: According to the backbone network of the Yolov8-x model and the downsampling ratio of its corresponding feature layer, the corresponding image decoder is designed, and the backbone network of the Yolov8-x model is used as an encoder to form the AutoEncoder network. Using the MIM self-supervised learning method, the input training image is randomly masked with a preset area ratio, and the output is a training image with a mask, which is used to guide the decoder of the AutoEncoder network to learn the correlation between different local semantics; The AutoEncoder network is pre-trained. During the feature extraction process of the training, the encoder uses sparse convolution to skip the masked part and only extracts features from the non-masked area for image reconstruction of the masked part. During the reconstructed image training process, the decoder uses reverse convolution to upsample the output features of the encoder to obtain an output image with the same resolution as the input image. Use L1 loss to calculate the reconstruction error between the input image and the output image in the mask area; The parameters of the AutoEncoder network are updated through the SGD optimizer to minimize the reconstruction error, and finally the pre-trained weights of the backbone network of the teacher model are obtained.

3. The lightweight method for target detection algorithm for embedded low computing power platform according to claim 2 is characterized in that: Fine-tuning the Yolov8-x model is as follows: The pre-trained Encoder weights are added to the original backbone network of the Yolov8-x model, the weights of the first N convolutional layers of the backbone network are frozen, the common features of the shallow layers are retained to improve the generalization of subsequent models, and the Yolov8-x model is fine-tuned based on the training set.

4. The lightweight method for target detection algorithm for embedded low computing power platform according to claim 1 is characterized in that: Select MobileOne to replace the ConvModule module in the Yolov8-x model backbone network as the student model: Select MobileOne to replace the ConvModule module in the backbone network of the Yolov8-x model, take the 1 / 8, 1 / 16 and 1 / 32 times downsampling feature layers of MobileOne as the input of Neck and record them as C3, C4, C5 respectively, and set the number of channels to 128, 256, and 512 respectively. Keep the rest of the structure of the Yolov8-x model including the Neck network and the Head network unchanged as the student model.

5. The lightweight method for target detection algorithm for embedded low computing power platform according to claim 4 is characterized in that: The student model is trained by the detection head cross distillation method as follows: Freeze all weights of the teacher model; Extract C3, C4, and C5 of the student model during training and send them to the corresponding detection head of the teacher model. Use a 1x1 convolution module to align the number of channels and perform image prediction inference. The image prediction results of the teacher model and the image prediction results of the student model are used to calculate the similarity error through L2 loss, and then weighted and combined with the Ground-Truth error of the student model according to the preset weight ratio to obtain the final training loss; The parameters of the student model are updated through the SGD optimizer to minimize the training loss for student model training.

6. The lightweight method for target detection algorithm for embedded low computing power platform according to claim 1, characterized in that: Using the structured pruning method, the backbone network of the student model that has completed distillation learning training is pruned as follows: According to the activation value of the BN layer in the backbone network of the student model, the convolution kernels with activation values ​​less than the preset threshold and the corresponding subsequent convolution kernels are deleted; The pruned student model is fine-tuned using the SGD optimizer to output the pruned student model.

7. The lightweight method for target detection algorithm for embedded low computing power platform according to claim 1, characterized in that: The lightweight processing of the pruned student model specifically includes: The parameters of the student model are reconstructed, and the residual 3x3 convolution weights and residual 1x1 convolution weights in the residual convolution module are merged with the BN layer weights to obtain a straight 3x3 convolution backbone network. Perform PTQ quantization on the student model, use M typical scene images for quantization correction, and quantize the student model weights from FP32 to INT8.

8. A lightweight target detection algorithm system for embedded low-computing power platforms, characterized in that: The system is used to implement the lightweight method of target detection algorithm for embedded low-computing power platform as described in any one of claims 1 to 7, including a data acquisition module, a teacher model pre-training and fine-tuning module, a student model distillation training module, and a student model pruning and lightweight module; The data acquisition module is used to collect images of the task objects to obtain a data set with diverse samples. The data set is divided and output into a training set, a validation set, and a test set using a random partitioning method. The teacher model pre-training and fine-tuning module is used to select the Yolov8-x model as the teacher model, perform MIM self-supervised pre-training on the backbone network of the Yolov8-x model, and fine-tune the Yolov8-x model. After the fine-tuning is completed, the teacher model weights are output for the subsequent distillation learning of the student model; The student model distillation training module is used to select MobileOne to replace the ConvModule module in the Yolov8-x model backbone network as the student model, train the student model through the detection head cross distillation method, and output the student model weights after training for subsequent lightweight deployment; The student model pruning and lightweight module is used to use a structured pruning method to prune the backbone network of the student model that has completed distillation learning training, and to lightweight the pruned student model. The lightweight student model is used for deployment on a specified embedded platform.

9. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements a lightweight target detection algorithm for an embedded low-computing-power platform as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for lightweighting a target detection algorithm for an embedded low-computing-power platform according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Target detection method for YOLOv8 partial convolutional network

    CN117173395A

Cited By

  • Multi-mode-based target detection model training method and device, vehicle and medium

    CN120298860A

  • Deployment method, prediction method and system of electrical equipment state prediction network

    CN121561523A