High-speed miniaturized target recognition algorithm based on Yolo v10
By optimizing the YOLO v10 backbone network, constructing a bidirectional feature pyramid network, and deploying an end-to-end dual label allocation strategy, the problem of high computational cost and slow speed of small target detection algorithms on edge devices has been solved, achieving high-precision and high-speed small target recognition, which is suitable for scenarios such as security monitoring, drone inspection, and intelligent transportation.
Patent Information
- Application Number
- CN202511025030.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-11-14
AI Technical Summary
Existing target detection and recognition algorithms suffer from high computational complexity, slow speed, and a tendency to miss detections in small target recognition, and are not suitable for deployment on edge devices, making it difficult to achieve high-precision real-time detection.
A high-speed, miniaturized target recognition algorithm based on YOLO v10 is adopted. This algorithm optimizes the backbone network, constructs a bidirectional feature pyramid network, and deploys an end-to-end dual label assignment strategy, including depthwise separable convolution, bidirectional feature pyramid network, and dual label assignment prediction, to replace the traditional nonmaximum suppression post-processing.
It enables high-speed small target detection on embedded devices, improving recognition accuracy and positioning robustness. It is suitable for target detection tasks with high accuracy but limited computing power, especially in scenarios such as security monitoring, drone inspection and intelligent transportation.
Smart Images

Figure CN120953969A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision target detection technology, and in particular to a high-speed, miniaturized target recognition algorithm based on YOLO v10. Background Technology
[0002] In practical applications such as intelligent security, industrial inspection, and drone monitoring, the rapid detection and recognition of small targets has always been one of the core challenges in the field of computer vision. Small targets typically have problems such as small size, weak features, and susceptibility to background interference, making it difficult for traditional target detection algorithms to achieve both accuracy and speed. Especially in edge device deployment scenarios, higher requirements are placed on the high accuracy, real-time performance, and lightweight nature of the models. With the development of the YOLO series of algorithms, YOLOv10 has achieved significant breakthroughs in structural design and performance optimization, becoming a key engine for promoting the practical application of small target detection.
[0003] YOLOv10 inherits the end-to-end design advantages of previous versions and achieves a new balance between accuracy, speed, and model size. It enhances the localization and classification capabilities of small targets by introducing an improved feature pyramid structure and a more efficient attention mechanism. Simultaneously, in conjunction with MobileNetV4, network pruning, and quantization techniques, it enables efficient model operation on edge computing platforms, significantly reducing deployment costs and computing power requirements. Through these innovations, YOLOv10 can provide higher small target detection accuracy than previous models while maintaining extremely low latency, making it suitable for various scenarios such as smart cameras, automotive terminals, and drone platforms, and possessing broad engineering application prospects. Therefore, research on high-precision, high-speed, and miniaturized small target recognition algorithms based on YOLOv10 has significant theoretical value and practical significance. Summary of the Invention
[0004] The technical problem to be solved by this invention is: addressing the issues of large model computation and CPU time consumption in existing target detection and recognition algorithms, especially the problems of easy missed detection, slow speed, large model size, and unsuitability for edge device deployment in small target recognition. This invention provides a high-speed miniaturized target recognition algorithm based on YOLO v10, which has the functions of high-speed small target detection and recognition and fast localization, and is suitable for edge deployment of target detection tasks with high accuracy but limited computing power.
[0005] The technical solution adopted by this invention to solve its technical problem is: a high-speed, miniaturized target recognition algorithm based on YOLO v10, comprising the following steps:
[0006] Step 1: Optimize the backbone network: Optimize the YOLOv10 backbone network using the Universal Inverted Bottleneck (UIB) structure of MobileNet V4, and reduce the amount of computation through depthwise separable convolutions and dynamic channel adjustment;
[0007] Step 2: Construct a Bidirectional Feature Pyramid Network (BiFPN): Construct a bidirectional feature pyramid network to achieve multi-scale feature fusion, and introduce learnable weights for cross-resolution feature weighted fusion;
[0008] Step 3, Dual Label Assignment Prediction: Deploy an end-to-end dual label assignment strategy, replacing the traditional Non-Maximum Suppression (NMS) post-processing by combining dual-branch prediction with one-to-many supervision and one-to-one matching mechanisms. In previous object detection frameworks, NMS was indispensable, its purpose being to remove a large number of redundant boxes from the network output. However, this operation requires traversing all predicted boxes, resulting in a time complexity of O(N...). 2 Here, N represents the number of predicted boxes to be processed. When encountering dense target scenes, the time consumption increases significantly. Furthermore, the framework relies on a one-to-many label allocation mechanism during training, while NMS forces a one-to-one matching during inference. This mechanism leads to inconsistencies between training and inference targets, making end-to-end deployment impossible.
[0009] Step 1 specifically includes:
[0010] The standard convolution is replaced with depthwise separable convolution, including channel-wise spatial convolution and pointwise cross-channel convolution; the standard convolution kernel of traditional convolution extracts spatial features from all channels of the feature map simultaneously, and its computational cost is as follows: FLOPS = C in ×C out ×K 2 ×W×H
[0011] The input channel is C. in The output channel is C. out The kernel size is K×K, and W×H is the feature map size. This invention selects the depthwise separable convolution in the UIB module and optimizes it as follows: 1. Depthwise convolution - each input channel uses an independent K×K kernel; 2. 1x1 pointwise convolution for cross-channel feature fusion, significantly reducing computational cost: FLOPS = C in ×K 2 ×W×H+C in ×C out ×W×H.
[0012] A lightweight SE attention mechanism is introduced into the residual structure. Spatial information is compressed through global average pooling, and channel weights are dynamically adjusted using a gating mechanism. The residual structure (Inverted Residual) first increases the dimensionality through 1x1 convolutions, then performs depthwise convolutions, and finally compresses the channels, retaining more non-linear features. The number of channels is dynamically adjusted based on the input features to avoid redundant computation caused by a fixed expansion rate. Simultaneously, a lightweight attention mechanism, such as a simplified SE module, is introduced into the residual connections to enhance information interaction between channels. The UIB residual connection expression is as follows:
[0013] UIB(x) = x + F(x), where F(x) represents the nonlinear mapping after passing through the UIB backbone branch, x is the input directly passed by the shortcut connection, and addition represents the residual connection with element-wise addition. The channel expansion rate is dynamically adjusted according to the input characteristics to avoid redundant calculations caused by a fixed expansion rate.
[0014] The SE lightweight attention mechanism consists of the following two parts:
[0015] a) Squeeze compression: Compresses the spatial information of each channel into a single number, extracts the information of the global receptive field, and uses global average pooling to achieve this, as shown in the following formula:
[0016]
[0017] Where W, H, and c are the width, height, and number of channels of the feature layer, respectively, and X is the height of the feature layer. c This refers to the input feature value, Z. c This refers to feature values that have undergone Squeeze compression.
[0018] b) Excitation activation: A gating mechanism is learned to dynamically adjust the channel weights, implemented by two fully connected layers plus an activation function. The specific implementation formula is as follows:
[0019] s=σ(W2*δ(W1*z c ))
[0020] Where W1 and W2 are the weight parameters for dimensionality reduction and dimensionality increase, respectively, r is the compression ratio, δ is the ReLU activation function, and σ is the Sigmoid activation function.
[0021] The channel expansion rate is dynamically adjusted based on input characteristics to avoid redundant calculations caused by a fixed expansion rate.
[0022] The implementation method of the bidirectional feature pyramid network in step 2 is as follows:
[0023] A bidirectional feature propagation path is established, consisting of top-down and bottom-up directions. A BiFPN bidirectional feature pyramid network is introduced to establish a bidirectional link between the top-down and bottom-up paths based on the traditional FPN (Feature Pyramid Network), allowing for more effective flow and fusion of information between features of different scales.
[0024] Assign learnable weights w to features at different resolutions;
[0025] The weighted fusion formula F=∑ i w i ·f i Enhance feature representation ability, where f i For the input features, w i The corresponding weights are used. In traditional FPN networks, all input features are usually treated equally without distinction, which means that features at different resolutions are simply added together. However, BiFPN introduces an additional weight w based on the different resolutions of different input features to learn the importance of each feature, thereby enabling the network to converge faster during the learning process.
[0026] Step 3's dual-label assignment strategy includes a dual-branch prediction structure and dynamic weight fusion with gradient decoupling;
[0027] Two-branch prediction structure:
[0028] One-to-many branch: Following the traditional dense prediction model, each ground truth box (GT box) is assigned multiple prediction boxes to maximize recall;
[0029] One-to-one branch: A new independent branch is added, which uses the Hungarian Algorithm to match a unique predicted box for each ground truth bounding box, forcing redundancy removal; the specific implementation of the Hungarian Algorithm is as follows: Among them l cls For the category loss function, Real frame b i With prediction box IOU, l obj (P j ) represents the confidence loss, y i , For the true category and the predicted category, λ cls , λ iou , λ obj These are weighted coefficients for the class loss, IOU loss, and whether it is a positive sample loss, used to balance the different loss terms;
[0030] Dynamic weight fusion and gradient decoupling:
[0031] A dynamic weight fusion mechanism is introduced. During training, the losses of the one-to-many and one-to-one branches are weighted using learnable weights. In the early stages of training, the one-to-many learning weights are increased to improve the recall rate of the detection model; in the later stages, the one-to-one learning weights are strengthened to improve the accuracy of target localization and recognition. The specific implementation formula is as follows: total =α t *l sparse +(1-α t )*l dense , where l dense It is the loss from dense branch many-to-one, l sparse It is the loss from the one-to-one relationship of sparse branches, α t As the weights become dynamically sparse, α increases during the initial training phase due to adjustments in training progress or gradient information. t Smaller alpha, primarily using intensive supervision to avoid model divergence; α in the later stages of training t The trend tends towards 1, with sparse supervision as the main approach;
[0032] Gradient blocking is used to isolate the gradient influence of dense branches on shared parameters. Sparse and dense branches, due to different supervision objectives and matching strategies, may cause gradient interference, affecting the training of the main branch. A gradient decoupling strategy is used to ensure that the gradients of the two branches do not affect each other during backpropagation, avoiding side effects from one task on the other. Specifically, during backpropagation of the loss of the dense branch, the `.detach()` operation is used to block its gradient propagation to shared parameters, while the sparse branch still uses normal backpropagation for inference optimization. The mathematical expression for gradient decoupling is as follows: ▽ θ L→λ d ·dir(▽ θ L)+λ m ·mag(▽ θ L), where dir(▽) θ L) represents the gradient in the unit direction. mag(▽ θ L) represents the magnitude of the gradient ||▽ θ L||,λ d , λ m These represent the adjustment coefficients for direction and magnitude, respectively, which control the "directional influence" and "magnitude influence" of gradient propagation.
[0033] The matching criterion of the Hungarian algorithm is to minimize the joint loss between the predicted bounding box and the ground truth bounding box, including classification error, confidence error and IoU distance.
[0034] The algorithm of this invention is suitable for deployment on embedded devices and is implemented on the COCO small target dataset.
[0035] The beneficial effects of this invention are that the high-speed miniaturized target recognition algorithm based on YOLOv10 has better recognition accuracy and better localization robustness for detecting and recognizing fast-moving small targets; it optimizes the redundant network structure of YOLOv10, adopts an end-to-end dual label allocation strategy, and eliminates the problem of time consumption on the CPU in traditional nonmaximum suppression (NMS), enabling high-speed small target detection and recognition on embedded devices, thereby achieving fast and accurate localization; it can be widely applied to scenarios with extremely high requirements for small target detection accuracy and speed, such as security monitoring, drone inspection, intelligent transportation, and industrial quality inspection. Attached Figure Description
[0036] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0037] Figure 1 This is an overall flowchart of the high-speed miniaturized target recognition algorithm based on YOLO v10 of the present invention.
[0038] Figure 2 This is a flowchart illustrating the implementation of the optimized backbone network in this invention.
[0039] Figure 3 This is a flowchart illustrating the implementation of the Bidirectional Feature Pyramid Network (BiFPN) of this invention.
[0040] Figure 4 This is a structural diagram of the Bidirectional Feature Pyramid Network (BiFPN) of the present invention.
[0041] Figure 5 This is a flowchart illustrating the implementation of the dual-label allocation prediction method of the present invention.
[0042] Figure 6 This is a flowchart illustrating a specific implementation scheme of the present invention. Detailed Implementation
[0043] The invention will now be described in further detail with reference to the accompanying drawings. It should be emphasized that the following description is merely exemplary and not intended to limit the scope or application of the invention.
[0044] like Figure 1 As shown, a high-speed, miniaturized target recognition algorithm based on YOLO v10 according to the present invention includes the following steps:
[0045] Step 1: Optimize the backbone network: Optimize the YOLOv10 backbone network using the Universal Inverted Bottleneck (UIB) structure of MobileNet V4, and reduce the amount of computation through depthwise separable convolutions and dynamic channel adjustment;
[0046] Step 2: Construct a bidirectional feature pyramid network: Construct a bidirectional feature pyramid network to achieve multi-scale feature fusion, and introduce learnable weights for cross-resolution feature weighted fusion;
[0047] Step 3, Dual Label Assignment Prediction: Deploy an end-to-end dual label assignment strategy, replacing the traditional Non-Maximum Suppression (NMS) post-processing by combining dual-branch prediction with one-to-many supervision and one-to-one matching mechanisms. In previous object detection frameworks, NMS was indispensable, its purpose being to remove a large number of redundant boxes from the network output. However, this operation requires traversing all predicted boxes, resulting in a time complexity of O(N...). 2 Here, N represents the number of predicted boxes to be processed. When encountering dense target scenes, the time consumption increases significantly. Furthermore, the framework relies on a one-to-many label allocation mechanism during training, while NMS forces a one-to-one matching during inference. This mechanism leads to inconsistencies between training and inference targets, making end-to-end deployment impossible.
[0048] like Figure 2 As shown, step 1 specifically includes:
[0049] The standard convolution is replaced with depthwise separable convolution, including channel-wise spatial convolution and pointwise cross-channel convolution; the standard convolution kernel of traditional convolution extracts spatial features from all channels of the feature map simultaneously, and its computational cost is as follows: FLOPS = C in ×C out ×K 2 ×W×H
[0050] The input channel is C. in The output channel is C. out The kernel size is K×K, and W×H is the feature map size. This invention selects the depthwise separable convolution in the UIB module and optimizes it as follows: 1. Depthwise convolution - each input channel uses an independent K×K kernel; 2. 1x1 pointwise convolution for cross-channel feature fusion, significantly reducing computational cost: FLOPS = C in ×K 2 ×W×H+C in ×C out ×W×H.
[0051] A lightweight SE attention mechanism is introduced into the residual structure. Global average pooling is used to compress spatial information, and a gating mechanism is employed to dynamically adjust channel weights. The inverted residual structure first increases the dimensionality through 1x1 convolutions, then performs depthwise convolutions, and finally compresses the channels, preserving more non-linear features. The number of channels is dynamically adjusted based on the input features to avoid redundant computation caused by a fixed expansion rate. Simultaneously, a lightweight attention mechanism, such as a simplified SE module, is introduced into the residual connections to enhance information interaction between channels.
[0052] The SE lightweight attention mechanism consists of the following two parts:
[0053] a) Squeeze compression: Compresses the spatial information of each channel into a single number, extracts the information of the global receptive field, and uses global average pooling to achieve this, as shown in the following formula:
[0054]
[0055] Where W, H, and c are the width, height, and number of channels of the feature layer, respectively, and X is the height of the feature layer. c This refers to the input feature value, Z. c This refers to feature values that have undergone Squeeze compression.
[0056] b) Excitation activation: A gating mechanism is learned to dynamically adjust the channel weights, implemented by two fully connected layers plus an activation function. The specific implementation formula is as follows:
[0057] s=σ(W2*δ(W1*z c ))
[0058] Where W1 and W2 are the weight parameters for dimensionality reduction and dimensionality increase, respectively, r is the compression ratio, δ is the ReLU activation function, and σ is the Sigmoid activation function.
[0059] The channel expansion rate is dynamically adjusted based on input characteristics to avoid redundant calculations caused by a fixed expansion rate.
[0060] like Figure 3 As shown, the implementation method of the bidirectional feature pyramid network in step 2 is as follows:
[0061] A bidirectional feature propagation path is established, consisting of top-down and bottom-up directions. A BiFPN bidirectional feature pyramid network is introduced to establish a bidirectional link between the top-down and bottom-up paths based on the traditional FPN (Feature Pyramid Network), allowing for more effective flow and fusion of information between features of different scales.
[0062] Assign learnable weights w to features at different resolutions;
[0063] The weighted fusion formula F=∑ i w i ·f i Enhance feature representation ability, where f i For the input features, w iThe corresponding weights are used. In traditional FPN networks, all input features are usually treated equally without distinction, which means that features at different resolutions are simply added together. However, BiFPN introduces an additional weight w based on the different resolutions of different input features to learn the importance of each feature, thereby enabling the network to converge faster during the learning process.
[0064] Figure 4 This is a diagram of the Bidirectional Feature Pyramid Network (BiFPN) structure of the present invention, where P3-P7 represent feature maps of different sizes extracted from the backbone network. The smaller the number after P, the larger the feature map size. For example, P3 represents a shallower feature map with a larger size, while P7 represents a deeper feature map with a smaller size.
[0065] like Figure 5 As shown, the dual label assignment strategy in step 3 includes a dual-branch prediction structure and dynamic weight fusion and gradient decoupling.
[0066] Two-branch prediction structure:
[0067] One-to-many branch: Following the traditional dense prediction model, each ground truth box (GT box) is assigned multiple prediction boxes to maximize recall;
[0068] One-to-one branch: A new independent branch is added, which uses the Hungarian Algorithm to match a unique predicted box for each ground truth bounding box, forcing redundancy removal; the specific implementation of the Hungarian Algorithm is as follows: Among them l cls For the category loss function, Real frame b i With prediction box IOU, l obj (P j ) represents the confidence loss, y i , For the true category and the predicted category, λ cls , λ iou , λ obj These are weighted coefficients for the class loss, IOU loss, and whether it is a positive sample loss, used to balance the different loss terms;
[0069] Dynamic weight fusion and gradient decoupling:
[0070] A dynamic weight fusion mechanism is introduced. During training, the losses of the one-to-many and one-to-one branches are weighted using learnable weights. In the early stages of training, the one-to-many learning weights are increased to improve the recall rate of the detection model; in the later stages, the one-to-one learning weights are strengthened to improve the accuracy of target localization and recognition. The specific implementation formula is as follows: total =αt *l sparse +(1-α t )*l dense , where l dense It is the loss from dense branch many-to-one, l sparse It is the loss from the one-to-one relationship of sparse branches, α t As the weights become dynamically sparse, α increases during the initial training phase due to adjustments in training progress or gradient information. t Smaller alpha, primarily using intensive supervision to avoid model divergence; α in the later stages of training t The trend tends towards 1, with sparse supervision as the main approach;
[0071] Gradient blocking is employed to isolate the gradient impact of dense branches on shared parameters. Sparse and dense branches, due to differing supervision objectives and matching strategies, may cause gradient interference, affecting the training of the main branch. A gradient decoupling strategy is used to ensure that the gradients of the two branches do not affect each other during backpropagation, preventing one task from having a side effect on the other. Specifically, during backpropagation of the loss in the dense branch, the `.detach()` operation is used to block its gradient propagation to shared parameters, while the sparse branch continues to use normal backpropagation for inference optimization.
[0072] The matching criterion of the Hungarian algorithm is to minimize the joint loss between the predicted bounding box and the ground truth bounding box, including classification error, confidence error and IoU distance.
[0073] The algorithm of this invention is suitable for deployment on embedded devices. The following is a comparative experiment using the PyTorch framework on the COCO small target dataset, conducted in an RTX3090*4 test environment, comparing YOLOv10s with YOLOv8s PP-YOLOE+s:
[0074] Model mAP Inference delay (ms) Number of parameters (M) FLOPs(G) YOLOv10s 47.6 4.3 6.0 13.2 YOLOv8s 46.2 6.1 11.2 28.6 PP-YOLOE+s 45.9 5.4 7.9 20.5
[0075] In a comprehensive comparison, the YOLOv10s of this invention is optimal in terms of accuracy, speed, and size, making it suitable for applications that require high accuracy and real-time performance for small models.
[0076] The advantages of optimization before and after are compared in several aspects, including network computing, computational overhead, feature fusion, parameter sensitivity, dense target processing, inference consistency, and hardware deployment. The table below shows the advantages of optimization before and after:
[0077]
[0078] The following is a flowchart illustrating the implementation process of small target detection in industrial quality inspection based on YOLOv10s. Figure 6 As shown:
[0079] The model selection and parameter optimization are as follows: The detection model adopts YOLOv10s+UIB+BiFPN+SE attention; the input size is 640×640, that is, the resolution of the input image is 640 pixels wide × 640 pixels high; the optimization strategy is MobileNetV4 backbone + model pruning + Int8 quantization; the target categories include four categories: capacitor, resistor, diode and small solder joint.
[0080] The edge device deployment parameters are as follows: the device platform is HiSilicon Hi3519AV100 (NPU), with an NPU computing power of 3.0 TOPS; the operating system is Linux; and the inference framework is configured with the Hisi NPU acceleration plugin.
[0081] The following table shows a comparison of the measured inference speed and resource consumption of edge devices under the same input size:
[0082]
[0083] The robustness test involves comparing the input device image under the same interference conditions such as blurring, tilting, occlusion, and noise.
[0084] Therefore, it can be seen that the YOLOv10s solution presented here outperforms existing lightweight models in terms of recognition accuracy and robustness when dealing with small, rapidly moving targets in a pipeline. Furthermore, it can stably perform target detection with a high frame rate industrial camera at 40FPS, fully meeting the requirements for online detection at 25+ frames per second.
[0085] Based on the above-described preferred embodiments of the present invention, and through the foregoing description, those skilled in the art can make various changes and modifications without departing from the inventive concept. The technical scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.
Claims
1. A high-speed, miniaturized target recognition algorithm based on YOLO v10, characterized in that, Includes the following steps: Step 1: Optimize the backbone network: Optimize the YOLO v10 backbone network using the general inverted bottleneck structure of MobileNet V4, and reduce the amount of computation through depthwise separable convolutions and dynamic channel adjustment; Step 2: Construct a bidirectional feature pyramid network: Construct a bidirectional feature pyramid network to achieve multi-scale feature fusion, and introduce learnable weights for cross-resolution feature weighted fusion; Step 3, Dual Label Assignment Prediction: Deploy an end-to-end dual label assignment strategy, which replaces the traditional nonmaximum suppression post-processing by combining dual-branch prediction with one-to-many supervision and one-to-one matching mechanism.
2. The high-speed, miniaturized target recognition algorithm based on YOLO v10 as described in claim 1, characterized in that, Step 1 specifically includes: Replace standard convolutions with depthwise separable convolutions, including channel-wise spatial convolutions and pointwise cross-channel convolutions; A lightweight SE attention mechanism is introduced into the residual structure. Spatial information is compressed through global average pooling, and channel weights are dynamically adjusted using a gating mechanism. The UIB residual connection expression is as follows: UIB(x) = x + F(x), where F(x) represents the nonlinear mapping after passing through the UIB backbone branch, x is the input directly passed by the shortcut connection, and addition represents the residual connection with element-wise addition.
3. The high-speed, miniaturized target recognition algorithm based on YOLO v10 as described in claim 1, characterized in that, The implementation method of the bidirectional feature pyramid network in step 2 is as follows: Establish a two-way feature propagation path, both top-down and bottom-up; Assign learnable weights w to features at different resolutions; The weighted fusion formula F=∑ i w i ·f i Enhance feature representation ability, where f i For the input features, w i For the corresponding weights.
4. The high-speed, miniaturized target recognition algorithm based on YOLO v10 as described in claim 1, characterized in that, Step 3's dual-label assignment strategy includes a dual-branch prediction structure and dynamic weight fusion with gradient decoupling; Two-branch prediction structure: One-to-many branching: Employs a dense prediction mode, assigning multiple prediction boxes to each ground truth box; One-to-one branch: A new independent branch is added, and a unique predicted box is matched for each ground truth bounding box using the Hungarian algorithm; the specific implementation of the Hungarian algorithm is as follows: in For the category loss function, Real frame b i With prediction box IOU, For confidence loss, y i , For the true category and the predicted category, λ cls , λ iou , λ obj These are weighted coefficients for the class loss, IOU loss, and whether it is a positive sample loss, used to balance the different loss terms; Dynamic weight fusion and gradient decoupling: A dynamic weight fusion mechanism is introduced. During training, the losses of the one-to-many branch and the one-to-one branch are weighted using learnable weights. The one-to-many learning weights are increased in the early stage of training, and the one-to-one learning weights are strengthened in the later stage. The specific implementation formula is as follows: in It is the loss from dense branch many-to-one relationships. It is the loss from the one-to-one relationship of sparse branches, α t As the weights become dynamically sparse, α increases during the initial training phase due to adjustments in training progress or gradient information. t Smaller alpha, primarily using intensive supervision; alpha in the later stages of training t The trend tends towards 1, with sparse supervision as the main approach; Gradient blocking operations are used to isolate the gradient effects of dense branches on shared parameters.
5. The high-speed, miniaturized target recognition algorithm based on YOLO v10 as described in claim 4, characterized in that, The matching criterion of the Hungarian algorithm is to minimize the joint loss between the predicted bounding box and the ground truth bounding box, including classification error, confidence error and IoU distance.
Citation Information
Patent Citations
Construction site safety helmet detection method, computer equipment and storage medium
CN118644761A
Target area small target detection method based on unmanned aerial vehicle image
CN118968035A
Intelligent community task report generation method and device and related components
CN119294782A
Highway tunnel lining crack detection system based on improved YOLO network
CN119941674A
Multi-scale target detection method and system based on deep learning
CN120147828A