Environment multi-article identification method based on YOLO
By improving the CBAM attention mechanism and adaptive convolution kernel of the YOLO model, combined with ATTS dynamic tag allocation, the detection accuracy and robustness of the YOLO model in complex environments is solved, and efficient and accurate multi-item recognition is achieved.
Patent Information
- Application Number
- CN202510428674.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-08
AI Technical Summary
The existing YOLO model has insufficient recognition ability for small targets, occlusion targets and low-contrast targets in complex environments, unreasonable distribution of training data labels, limited feature extraction effect, and inability to adaptively adjust, resulting in insufficient detection accuracy and robustness.
Combined with the improved YOLO framework, CBAM attention mechanism, adaptive convolution kernel and ATTS dynamic label allocation strategy, the detection accuracy and robustness of the model in complex scenarios are improved through data preprocessing, feature extraction and optimization, multi-scale feature fusion, detection post-processing and other steps.
It significantly improves the detection accuracy in small targets and complex scenarios, reduces missed and missed detection, adapts to the detection needs of multiple categories of items, reduces calculation complexity, and is suitable for embedded devices.
Smart Images

Figure CN120279256A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for multi-object recognition in an environment based on a YOLO model, belonging to the technical fields of computer vision and intelligent recognition. Background Art
[0002] Today, with the continuous development of the technology for multi-object recognition in an environment, the application of single-stage object detection algorithms such as YOLO has been gradually popularized. Such algorithms have received wide attention for their high detection speed, but still face many technical bottlenecks in practical applications. Specifically, in complex scenarios, small target objects have significantly reduced detection accuracy due to their small size, low contrast, and easy occlusion; at the same time, the mutual occlusion and similarity between multi-class targets make it more difficult to distinguish. These challenges not only affect the overall performance of the model, but also to a certain extent limit its actual application effects in fields such as intelligent monitoring and autonomous driving.
[0003] In order to simplify the calculation parameters of the deep learning model and reduce the consumption of computing resources and storage resources, the present invention proposes a lightweight relational gated neural model LR-GRU for the problem of user electricity load prediction. On the one hand, this model can ensure the accuracy of electricity load prediction while solving the problem that traditional deep learning models are not lightweight enough; on the other hand, this model can accurately and efficiently predict the electricity load, and the prediction results can provide a reference for the reasonable allocation of electricity in the power market and the reasonable pricing strategy of electricity in the power market.
[0004] Patent US20200123456A1 proposes a YOLO-based object detection method, and its main technical solution is to improve the detection accuracy of small targets by improving the feature extraction module of the YOLO model. However, this method is mainly aimed at industrial scenarios and is difficult to adapt to the multi-object detection task in a complex natural environment. For example, the detection accuracy is still not ideal in scenarios with large light changes or complex backgrounds.
[0005] In 2018, Joseph published "Real-Time Object Detection with YOLOv3" on arXiv, introducing the excellent detection performance of the YOLOv3 model on the COCO dataset and proposing a multi-scale feature fusion method for small target detection. However, when this technology is deployed on resource-constrained embedded devices, the computational overhead is large, resulting in the inability to meet the real-time requirements of some high-frequency detection needs. In addition, there is still room for improvement in the detection effect of YOLOv3 in multi-object complex occlusion scenarios.
[0006] In 2020, Li Yanghao published "Dynamic Convolution: Attention Over Convolution Kernels" in the CVPR journal, introducing a method to generate adaptive convolution kernels by designing multiple convolution kernel branches and using a weight generation module to assign dynamic weights to each branch. However, this method only optimizes the convolution kernel part and has the drawback of not combining the attention mechanism to further enhance the network's attention ability.
[0007] The current problems are that existing YOLO models are mostly optimized for standard datasets and are difficult to handle actual scenarios such as complex backgrounds and multi-object occlusions. They are limited in terms of real-time performance, with a high computational complexity of the model and low inference efficiency in embedded devices or edge computing environments. Moreover, the convolution kernels in traditional YOLO series algorithms are fixed parameters and cannot be dynamically adjusted according to the input feature map, affecting the quality of feature extraction in complex scenarios. Summary of the Invention
[0008] The present invention discloses a method for multi-object recognition in an environment based on YOLO, aiming to solve the following main problems in the prior art:
[0009] (1) Insufficient detection accuracy: Existing object detection methods have limited recognition capabilities for small objects, occluded objects, and low-contrast objects in complex environments, resulting in low accuracy of detection results.
[0010] (2) Unreasonable distribution of training data labels: Existing methods do not optimize the label distribution in training data and are difficult to meet the detection requirements of multi-category objects, especially in the case of class imbalance.
[0011] (3) Limited feature extraction effect: Some object detection models have insufficient extraction of key features in complex backgrounds, affecting the stability and robustness of detection.
[0012] (4) Unable to adaptively adjust according to the input image features, which easily leads to inaccurate feature extraction, especially in dealing with diverse environments (such as complex backgrounds or object occlusions) with poor effects.
[0013] To solve the above problems, the present invention is implemented through the following technical solutions:
[0014] 1. Method for multi-object recognition in an environment based on YOLO
[0015] The core of the present invention lies in combining an improved YOLO framework, CBAM attention mechanism, adaptive convolution kernel, and ATTS dynamic label assignment strategy to improve the accuracy and robustness of multi-object detection in complex scenarios. The specific technical solutions include the following steps:
[0016] 1.1 Feature Extraction and CBAM Attention Enhancement
[0017] (1) Input Image Preprocessing: Enhance the input image (such as brightness adjustment, random cropping, Mosaic data augmentation) to improve the generalization ability of the model.
[0018] (2) Feature Extraction and Attention Optimization: Use a lightweight YOLO backbone network (such as YOLOv8n) for feature extraction, and embed the CBAM attention module at specific levels of the backbone network to optimize the feature map in the following ways:
[0019] 1) Channel Attention Sub-module: Perform global average pooling and max pooling on the feature map to generate a channel weight matrix, and weight the global importance of each channel of the feature map.
[0020] 2) Spatial Attention Sub-module: Pool the feature map along the channel dimension to generate a spatial weight matrix, and weight the significant regions at each spatial position of the feature map.
[0021] 1.2 Adaptive Convolution Kernel Dynamic Parameter Adjustment, Dynamic Convolution Kernel Generation:
[0022] (1) Based on the scale information of the feature map output by CBAM, adjust the size and dilation rate of the convolution kernel through a dynamic weight matrix to make it adapt to target features of different scales.
[0023] (2) Combine the morphological distribution of the feature map (such as edge density, texture complexity), and use a lightweight neural network to predict the offset of the convolution kernel parameters to achieve morphology-adaptive convolution operations.
[0024] 1.3 Cross-Stage Local Feature Fusion, Multi-Scale Feature Generation:
[0025] (1) Downsample the shallow feature map (high resolution, low semantics), and perform channel alignment with the deep feature map (low resolution, high semantics);
[0026] (2) Use grouped convolution to reduce the computational amount, and fuse the cross-level local features through pointwise convolution to output a multi-scale fused feature map.
[0027] 1.4 ATTS Dynamic Label Assignment and Detection
[0028] (1) Label Assignment Strategy:
[0029] 1) Introduce ATTS in the detection head, and construct a dynamic weight function based on the confidence, spatial position, and size of the target;
[0030] 2) Based on the intersection over union (IoU), classification confidence, and size matching degree between the target and the anchor box, screen positive and negative samples, and assign classification and regression loss weights through a threshold adaptive mechanism.
[0031] (2) Detection head optimization:
[0032] 1) Adopt a decoupled prediction structure to output the confidence of target classification and the coordinates of the bounding box respectively, improving the prediction accuracy;
[0033] 2) In the training stage, use the Focal Loss function to optimize the classification loss of small targets and occluded targets, alleviating the problem of class imbalance.
[0034] 1.5 Optimization of post-detection processing
[0035] (1) Adaptive non-maximum suppression: Dynamically adjust the bounding box overlap threshold according to the target density to reduce missed detections and false detections;
[0036] (2) Dynamic adjustment of the confidence threshold: Flexibly set the confidence threshold according to the target category and scene requirements to filter out low-quality prediction boxes;
[0037] (3) Multi-frame target tracking: Perform temporal correlation on the detection results of consecutive frames, and optimize the class probability assignment by combining attention weights to reduce target jumps or losses.
[0038] 2. Implementation details of key modules
[0039] 2.1 CBAM attention mechanism (corresponding to claim 2)
[0040] (1) Channel attention sub-module: After the input feature map undergoes global average pooling and max pooling, a channel weight matrix is generated through a shared fully connected layer and multiplied with the original feature map channel by channel.
[0041] (2) Spatial attention sub-module: Perform average pooling and max pooling on the feature map along the channel dimension, and after concatenation, generate a spatial weight matrix through a standard convolution and multiply it with the original feature map position by position.
[0042] 2.2 Implementation of adaptive convolution kernel (corresponding to claim 3)
[0043] (1) Generation of dynamic weight matrix: Based on the scale information of the feature map (such as the resolution level), generate the dilation rate and size parameters of the convolution kernel through a linear transformation.
[0044] (2) Prediction of parameter offset: Use a lightweight MLP network to predict the dynamic offset of the convolution kernel parameters according to the morphological distribution of the feature map (such as gradient magnitude, local variance).
[0045] 2.3 Cross-stage feature fusion (corresponding to claim 4)
[0046] (1) Downsampling and alignment: Use a convolution with a stride of 2 to downsample the shallow feature map, and adjust the number of channels through a 1×1 convolution to align with the deep feature map.
[0047] (2) Group fusion strategy: Group the feature maps and perform convolution operations separately to reduce the computational amount, and then integrate the cross-layer feature information through pointwise convolution.
[0048] 2.4 ATTS Label Assignment Strategy (corresponding to Claim 5)
[0049] (1) Dynamic weight function: ω = α·IoU + β·Confidence + γ·Size_Match, where α, β, and γ are learnable parameters that control the weights of the intersection over union, classification confidence, and size match, respectively;
[0050] (2) Adaptive threshold screening: Dynamically adjust the IoU threshold of positive samples according to the target density to ensure the label assignment priority of small targets and occluded targets.
[0051] 3. Model Optimization and Deployment
[0052] (1) Lightweight design: Adopt the YOLOv8n backbone network, combined with model pruning and quantization techniques (such as 8-bit integer quantization), to reduce the model complexity and adapt to the deployment of embedded devices.
[0053] (2) Training strategy: Use Focal Loss to optimize the classification loss, combined with the AdamW optimizer and cosine annealing learning rate scheduling to accelerate the model convergence.
[0054] Advantages of the present invention:
[0055] (1) The ATTS label assignment strategy improves the model's detection ability for small targets and complex scenes;
[0056] (2) After embedding the CBAM attention mechanism, the model's ability to extract key features is enhanced, and it can more accurately focus on the target area, thereby reducing missed detections and false detections;
[0057] (3) The model can effectively extract target features under conditions such as light changes, complex backgrounds, and target occlusion;
[0058] (4) By adopting a lightweight backbone network (such as YOLOv8n) and an efficient dynamic convolution kernel weight generation module, the accuracy of target detection is significantly improved on the premise of low computational overhead;
[0059] (5) This method has good scalability and can be easily migrated to other target detection frameworks or scenarios, such as being embedded in models like RetinaNet and Faster R-CNN. Brief Description of the Drawings
[0060] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the relevant drawings required in the description of the embodiments of the present application:
[0061] Figure 1 It is a system framework diagram, showing the overall framework of the improved YOLO model, including five major modules: data preprocessing, feature extraction and optimization, multi-scale feature fusion, dynamic label assignment, and post-detection processing.
[0062] Figure 2 It is a flowchart of data preprocessing, illustrating the overall process of the data preprocessing module, including steps of loading input data, enhancement operations, and generating final training data.
[0063] Figure 3 It is a structural diagram of the improved YOLOv8 network, showing the detailed network structure after adding ATTS and CBAM improvements. Detailed Implementation Modes
[0064] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0065] 1. Overall System Architecture and Process
[0066] The implementation of the present invention is based on an improved YOLO framework, and the overall process includes five major modules: data preprocessing, feature extraction and optimization, multi-scale feature fusion, dynamic label assignment, and post-detection processing (see the attached Figure 1 ). The specific implementation steps are as follows:
[0067] 2. Data Preprocessing and Label Assignment
[0068] 2.1 Data Augmentation
[0069] (1) The input image adopts the Mosaic augmentation strategy: randomly splice 4 images, adjust brightness, contrast, and saturation to simulate the illumination changes in complex scenarios;
[0070] (2) Add random erasure operations to occluded targets to improve the robustness of the model to occlusion.
[0071] 2.2 ATTS Dynamic Label Assignment
[0072] (1) Label Weight Adjustment:
[0073] 1) Small target priority: For targets with an area less than 5% of the total image pixels, the classification loss weight is increased to 1.5 times.
[0075] 2) Class balance: Dynamically adjust the classification loss weight according to the frequency of each class in the training data. The lower the frequency, the higher the weight.
[0076] 3) Occlusion optimization: For occluded targets with an intersection over union (IoU) below 0.3, the regression loss weight is increased to 2.0 times.
[0077] (2) Label assignment execution: During the training phase, update the weight parameters in the loss function in real time according to the above rules.
[0078] 3. Feature extraction and attention enhancement
[0079] 3.1 Backbone network construction
[0080] Use YOLOv8n as the lightweight backbone network. Regarding the embedding of the CBAM module:
[0081] (1) Insert the CBAM module after the 3rd, 6th, and 9th layers of the backbone network (see the appendix Figure 3 ).
[0082] (2) Channel attention sub-module:
[0083] 1) Perform global average pooling (GAP) and global max pooling (GMP) on the input feature map respectively to obtain two vectors of 1×1×C.
[0084] 2) Input the two vectors into a shared two-layer fully connected network (with a middle layer dimension of C / 16) to generate a channel weight matrix.
[0085] 3) Multiply the weight matrix with the original feature map channel by channel to obtain the channel-weighted feature map.
[0086] (3) Spatial attention sub-module:
[0087] 1) Perform average pooling and max pooling on the feature map along the channel dimension to obtain two feature maps of H×W×1.
[0088] 2) Concatenate the two and generate a spatial weight matrix through a 7×7 convolutional layer.
[0089] 3) Multiply the weight matrix with the channel-weighted feature map position by position and output the final optimized feature map.
[0090] 3.2 Adaptive convolutional kernel implementation
[0091] (1) Dynamic parameter adjustment:
[0092] 1) For the feature maps output by CBAM, generate a dynamic weight matrix according to their resolution levels (such as 1 / 8, 1 / 16, 1 / 32 scales).
[0093] 2) For example, the feature map at the 1 / 8 scale corresponds to a 3×3 convolutional kernel with a dilation rate of 2, and the 1 / 16 scale corresponds to a 5×5 convolutional kernel with a dilation rate of 1.
[0094] (2) Morphological adaptive mechanism:
[0095] 1) Use a lightweight MLP network (2-layer fully connected, with a hidden layer dimension of 32) to predict the offset of the convolutional kernel parameters.
[0096] 2) The input is the local statistical features of the feature map (such as the mean and variance of the gradient magnitude), and the output is the convolutional kernel weight ΔW and bias Δb.
[0097] 4. Cross-stage local feature fusion
[0098] 4.1 Feature map alignment and fusion
[0099] (1) Shallow feature downsampling: Perform a 3×3 convolution with a stride of 2 on the feature map (1 / 8 scale) from the third layer of the backbone network to reduce its resolution to 1 / 16 scale.
[0100] (2) Channel alignment: Adjust the number of channels of the shallow feature map to be the same as that of the deep feature map (1 / 16 scale) through 1×1 convolution.
[0101] (3) Group convolution fusion:
[0102] 1) Group the shallow and deep feature maps along the channel dimension (32 channels per group).
[0103] 2) Perform 3×3 depthwise separable convolution on each group of features to extract local features.
[0104] 3) Integrate all group features through pointwise convolution (1×1 convolution) to output a multi-scale fusion feature map.
[0105] 5. Detection head and dynamic label assignment
[0106] 5.1 Decoupled detection head design
[0107] (1) Classification branch: Output the class confidence of each anchor box, use the Sigmoid activation function, and support multi-label classification.
[0108] (2) Regression branch: Output the bounding box coordinates (center point offset, width and height scaling factors), use the Swish activation function.
[0109] 5.2 Implementation of the ATTS strategy
[0110] (1) Dynamic weight function: ω = 0.6·IoU + 0.3·Confidence + 0.1·Size_Match, where Size_Match is the matching degree between the anchor box and the target size (calculated as the reciprocal of the aspect ratio similarity);
[0111] (2) Adaptive threshold screening: If the target density is higher than the threshold (e.g., the number of detected targets per frame > 50), the IoU threshold for positive samples is reduced from 0.5 to 0.4 to ensure that small targets are effectively assigned.
[0112] 6. Detection post - processing optimization
[0113] 6.1 Adaptive non - maximum suppression
[0114] (1) Dynamically adjust the IoU threshold of NMS according to the target density:
[0115] 1) Low - density scenario (number of targets < 10): The IoU threshold is 0.6.
[0116] 2) High - density scenario (number of targets ≥ 10): The IoU threshold is 0.4.
[0117] 6.2 Confidence optimization and multi - frame tracking
[0118] (1) Dynamic threshold adjustment: For main categories such as pedestrians and vehicles, the confidence threshold is set to 0.5; for small - target categories (such as traffic signs), the threshold is reduced to 0.3.
[0119] (2) Multi - frame association: Use Kalman filter to predict the target motion trajectory, and combine the IoU matching between the detection results of the current frame and the historical trajectory to reduce target jumps.
[0120] Example 1 An environmental multi - object recognition method combining ATTS label assignment strategy and CBAM attention mechanism
[0121] Overall framework design: As Figure 1 shown, this method is based on the YOLOv8 object detection model. By improving the label assignment strategy and introducing the attention mechanism, the multi - object detection performance in complex environments is optimized. The overall system includes a data pre - processing module, a label assignment module, a backbone network optimization module (embedded with CBAM module), and a detection post - processing module.
[0122] As Figure 2 shown, after inputting the original training image data, data augmentation operations such as random cropping, rotation, scaling, and color jittering are performed. The Mosaic method is used to splice multiple images into a large image to improve the diversity of training data. Dynamically assign the label weights of the target boxes: According to the size of the target boxes, the weights of small targets are preferentially increased to solve the problem of small targets being ignored;
[0123] For occluded targets, an adaptive weight adjustment method is adopted. By calculating the occlusion ratio of the target, the significance of the labels in the overlapping area is increased; for the problem of class imbalance, a class balance optimization strategy is used to weight the small-sample classes to ensure the fairness of model training.
[0124] As Figure 3 shown, the YOLOv8 model is used as the backbone network, and an adaptive convolutional kernel is embedded in the feature extraction stage to improve the network's adaptability to dynamic scenes for extracting multi-scale features; the multi-scale features are pyramidally integrated through a feature fusion module to handle target objects of different sizes.
[0125] Channel attention module: Calculate the global average and maximum values of each channel in the feature map, generate a weight matrix and apply it to each channel of the feature map. Spatial attention module: Calculate the saliency weight in the spatial dimension of the feature map to enhance the feature expression of the target area. The CBAM module is embedded in the shallow feature extraction layer and the deep feature integration layer of the backbone network to simultaneously improve the local and global feature expression capabilities.
[0126] Use the Adaptive Non-Maximum Suppression (Adaptive NMS) method: By dynamically adjusting the overlapping threshold, reduce the false detections and missed detections in multi-object detection. Output the optimal results according to the confidence and class sorting of the detection boxes.
[0127] Example 2 Optimization and Application on Embedded Devices
[0128] Model pruning: Remove redundant network structures that have little impact on detection accuracy, reducing the number of model parameters and computational complexity. Model quantization: Use the Post-Training Quantization (PTQ) method to quantize the weights and activation functions of the model to 8-bit to reduce storage and computational resource occupancy. The optimized model is deployed on an embedded device (such as NVIDIA Jetson Nano) and can achieve real-time detection in a limited computing power environment.
[0129] Application scenarios:
[0130] Smart home: Deployed in smart home devices for multi-object detection and recognition in the home environment, such as target classification and position determination of automatic sorting devices. Industrial inspection: Used for rapid classification and detection of items on industrial production lines to improve production efficiency and accuracy. Drone monitoring: Applied to the real-time target detection system of drones for multi-object detection tasks in complex scenarios such as disaster monitoring and environmental survey.
[0131] The above has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification only illustrates the principles of the present invention.
Claims
1. An environmental multi-object recognition method based on YOLO, characterized in that It includes the following steps: (1) Extract features from the enhanced input image through a convolutional neural network, and embed the CBAM attention mechanism during the feature extraction process to optimize the channel weights and spatial weights of the key regions in the feature map; (2) Based on the feature map output by the CBAM module, introduce an adaptive convolutional kernel to dynamically adjust the size and parameters of the convolutional kernel according to the scale and morphology of the current input features; (3) Input the features optimized by CBAM into the cross-stage local feature fusion module, and generate multi-scale fusion features through cross-level feature splicing and convolutional operations; (4) Adopt the ATTS label assignment strategy in the detection head, and dynamically assign classification and localization labels based on the object confidence, spatial position, and size to complete object detection.
2. The method according to claim 1, wherein The CBAM attention mechanism in step (1) includes: 1.1 Channel attention sub-module, which generates channel weights through global average pooling and max pooling, and weights each channel of the feature map; 1.2 Spatial attention sub-module, which generates spatial weights through pooling operations along the channel dimension, and weights each spatial position of the feature map.
3. The method according to claim 1, wherein The adaptive convolutional kernel in step (2) is implemented in the following way: 2.1 Generate a dynamic weight matrix according to the scale information of the feature map to adjust the size and dilation rate of the convolutional kernel; 2.2 Based on the morphological distribution of the feature map, predict the convolutional kernel parameter offset through a lightweight neural network.
4. The method according to claim 1, characterized in that The cross-stage local feature fusion module in step (3) includes: 3.1 Downsample the shallow feature map and align the channels with the deep feature map; 3.2 Adopt grouped convolution and pointwise convolution to achieve cross-level feature fusion and output a multi-scale feature map.
5. The method according to claim 1, characterized in that, The ATTS label assignment strategy in step (4) includes: 4.1 Construct a dynamic weight function based on the intersection over union (IoU), classification confidence, and size matching degree between the object and the anchor box; 4.2 Screen positive and negative samples through a threshold adaptive mechanism, and assign classification and regression loss weights according to the weight function.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: In the model training stage, adopt the Focal Loss function to optimize the classification loss of small objects and occluded objects; Introduce a decoupled prediction structure in the detection head to separately output the object classification confidence and the bounding box coordinates.
Citation Information
Patent Citations
System for conversion of crude oil to petrochemicals and fuel products integrating vacuum residue hydroprocessing
US20200123456A1
Cited By
Cross-scale image bridging decoupling target detection method based on dynamic adjacency
CN121236369A
A dynamic adjacency-based cross-scale graph bridging decoupling target detection method
CN121236369B