Low-quality image-oriented method for detecting graspable target based on EMSDH-YOLO network

By replacing specific modules and optimizing loss functions in the YOLO11 model, the problem of insufficient accuracy of object detection under low-quality image conditions is solved, and the effect of improving detection accuracy and robustness while maintaining real-time performance is achieved.

CN120259847APending Publication Date: 2025-07-04JILIN UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510660336.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing object detection algorithms are difficult to balance between maintaining real-time and detection accuracy in poor lighting or poor visual sensors, especially in the absence of detection accuracy of captureable targets under low-quality image conditions.

Method used

Based on the YOLO11 model, some modules in the replacement backbone network and neck structure are EMS-ConvBlock, LSK-SPPF, DyT-C2PSA and EU-SCMixer modules, and Wise-IoU v3 loss function optimization is used to enhance feature interaction and robustness.

Benefits of technology

It improves detection accuracy and robustness under low-quality image conditions, can better detect occlusions and small objects, and improves detection accuracy of gripping objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259847A_ABST
    Figure CN120259847A_ABST
Patent Text Reader

Abstract

The invention discloses an EMSDH-YOLO network-based detectable target detection method for a low-quality image, and relates to the technical field of target detection in industrial automation. According to the method, partial convolution in a fourth C3k2 module of a backbone network of a YOLO11 model and a C3k2 module at the tail end of a Neck network is replaced by efficient multi-scale convolution, and an EMS-ConvBlock module is constructed; adding large-kernel separable convolution into a spatial pyramid structure of a backbone network of the YOLO11 model, fusing a two-dimensional convolution solution idea, and constructing an LSK-SPPF module; the method comprises the following steps: adding the idea of Dynamic Tanh into a C2PSA module of a backbone network of a YOLO11 model, realizing adaptive calibration of attention weight by using hyperbolic tangent gating, and constructing a DyT-C2PSA module; the method comprises the following steps: adding an idea of channel shuffling and space rolling mixing into an up-sampling architecture in a YOLO11 model Neck network, and constructing an EU-SCMixer module; and a loss function of the YOLO11 model is optimized by adopting a Wise-IoU v3 method. According to the method, the detection precision of the mechanical arm on various grabbable objects can be improved when the mechanical arm carries out grabbing operation under the low-quality imaging condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection technology in industrial automation, and involves technologies such as image processing, deep learning and target detection, and is specifically manifested as a graspable target detection method for low-quality images based on an EMSDH-YOLO network. Background Art

[0002] With the rapid development of artificial intelligence technology, artificial intelligence technology has gradually become the focus of scholars and many companies from all walks of life. Artificial intelligence technology is the product of the deep integration of high-end hardware and intelligent algorithms. It is inseparable from the ingenious coordination of hardware structure and the excellent design of deep neural networks. More than 70% of human environmental information relies on vision, and for artificial intelligence, visual perception is also the main source of most external information. A powerful robot cannot do without the strong support of the visual system, so visual perception as the core capability of artificial intelligence has received widespread attention. The robot's realization of positioning and grasping the grasping target through vision is the key to the practical application of artificial intelligence, and the premise of accurate grasping is the accurate positioning of the target. Therefore, how to improve the accuracy and robustness of graspable target detection has become the only way to promote the development of robotics technology.

[0003] When traditional computer vision algorithms detect graspable objects, they often have good real-time performance but low detection accuracy due to simple network structures, or have high detection accuracy but poor real-time performance due to overly complex network structures. Most target detection algorithms find it difficult to balance detection accuracy and real-time performance. In recent years, deep neural networks have developed rapidly, and target detection algorithms based on deep learning have shown unprecedented advantages, especially the YOLO (You Only Look Once) series of models, which have achieved significant breakthroughs in real-time performance and accuracy. However, in the case of poor lighting or poor visual sensors, the main problem faced in the detection of graspable targets is the poor quality of images generated by visual sensors due to factors such as uneven lighting, blurred shapes of graspable objects, and interference from reflected light. In order to ensure the success rate of robot grasping, it is necessary to enhance the recognition accuracy of deep neural networks when the image quality obtained by visual sensors is relatively poor.

[0004] Therefore, while avoiding the large number of parameters brought by the model and causing a decrease in the model detection speed, this study added some network enhancement and feature fusion techniques on the basis of the YOLO11 model to further improve the robustness of detection. It develops a graspable object detection algorithm that can achieve a balance between detection accuracy and real-time performance and is suitable for poor imaging of visual sensors. This has important practical significance. Summary of the invention

[0005] The purpose of the present invention is to provide a graspable target detection method for low-quality images based on the EMSDH-YOLO network, so as to improve the detection accuracy of common graspable objects in low-quality images.

[0006] To achieve the above object, the present invention provides the following solutions: The present invention first provides an EMSDH-YOLO network for graspable target detection in low-quality images. The EMSDH-YOLO network uses the YOLO11 network as the baseline network, replaces the fourth C3k2 module of the YOLO11 model backbone network and the C3k2 module at the end of the Neck network with the EMS-ConvBlock module, replaces the SPPF module of the YOLO11 model backbone network with the LSK-SPPF module, replaces the C2PSA module with the DyT-C2PSA module, replaces the upsampling module in the YOLO11 model Neck network with the EU-SCMixer module, and optimizes the loss function of the YOLO11 model using the Wise-IoU v3 method; The EMS-ConvBlock module uses hierarchical multi-scale convolution to replace the original convolution, which not only achieves mild lightweighting but also enables better cross-scale interaction of features. The structure of this module is determined by the c3k parameter. If c3k is True, the input feature map first passes through a 1×1 convolutional layer for initial feature extraction, and then is split by Split into two parallel branches. The feature map of one branch passes through n C3k_EMSCB modules for in-depth processing in sequence. The processed feature maps are fused through the Concat splicing operation, and finally pass through a 1×1 convolutional layer for feature integration and dimension adjustment to output the final result. If c3k is False, the C3k_EMSCB module is replaced by the Bottle_EMSCB module. In the C3k-EMSCB module, the input feature map first passes through a 1×1 convolutional layer, then passes through 2 Bottleneck-EMSCB modules, and then performs channel dimension Concat splicing processing with the input feature map that has passed through the 1×1 convolutional layer. Finally, a 1×1 convolutional layer is used to fuse the information of different channels for output. In the Bottleneck-EMSCB module, the input feature map first passes through a 3×3 convolutional layer, then passes through the EMSCB module, and whether to perform a residual connection between the output and the input is controlled by a parameter. In the EMSCB module, the input feature map is first divided into four equal parts along the channel dimension, and then these four groups of feature maps pass through 1×1, 3×3, 5×5, and 7×7 convolutional layers respectively to generate four groups of different feature maps. Finally, a 1×1 convolutional layer is used to fuse the information of different channels for output; The LSK-SPPF module incorporates the idea of integrating large-kernel separable convolution and attention-enhanced cross-level feature focusing into the fast spatial pyramid pooling structure. While increasing the receptive field of the neural network, it avoids the explosion of the number of parameters and enhances the feature extraction ability of the structure. The implementation method of this module is as follows: The input feature map is first processed by a 1×1 convolutional layer to obtain the first feature map, and then successively undergoes three max-pooling operations to obtain the second, third, and fourth feature maps in sequence. Then, the fourth feature map after pooling is concatenated with the first, second, and third feature maps in the channel dimension through Concat. Next, it is processed by the LSKA module, and finally, the information of different channels is fused through a 1×1 convolutional layer for output; in the LSKA module, the input feature map first successively undergoes depthwise separable convolutional layer processing of 1×(2 d -1) and (2 d -1)×1, and then successively undergoes dilated convolutional layer processing of 1× and ×1. Then, it undergoes 1×1 convolution for channel adjustment and information fusion. Finally, the integrated information is weighted and multiplied with the input feature map through residual connection for output; where d represents the dilation rate, k represents the maximum receptive field of the convolutional kernel, represents the floor operation; The DyT-C2PSA module adds Dynamic Tanh for non-linear activation of features, and realizes the balance between feature expression ability and computational efficiency through the collaborative design of dynamic activation, attention mechanism, and multi-scale convolution. The implementation method of this module is as follows: The input feature map is first processed by a 1×1 convolutional layer, and then the feature map undergoes Split segmentation and is evenly divided into two parts along the channel dimension. One part of the feature map successively undergoes processing by n DyT-PSABlock modules. The processed feature map is concatenated with the other unprocessed part of the feature map in the channel dimension. Finally, the concatenated result undergoes final processing through another 1×1 convolutional layer to obtain the output of the module; in the DyT-PSABlock module, the input feature map first undergoes non-linear transformation through a DynamicTanh activation layer, then the importance of the features is weighted through the attention mechanism layer, and further non-linear processing is performed through the second DynamicTanh activation layer. Finally, it successively undergoes two consecutive 1×1 convolutional layers for feature mapping and dimension transformation, and finally outputs the processing result; The EU-SCMixer module adds channel shuffle and spatial rolling mixing operations in the upsampling part, which displace features along the channel (C), height (H), and width (W) dimensions. Without increasing parameters, it promotes information interaction in the spatial and channel dimensions and greatly enhances the diversity of features. The implementation method of this module is as follows: the input feature map is first enlarged in size through an upsampling layer to improve the spatial resolution of the feature map, and then a 3×3 depthwise separable convolutional layer is used to independently extract local features for each channel. Next, the feature interaction of the module is enhanced through channel shuffle and spatial rolling mixing operations. Finally, 1×1 convolution is used for channel fusion and dimension adjustment to output the final feature map; The Wise-IoU v3 method can significantly improve the robustness and detection accuracy of the model when dealing with complex scenarios by using the Wise-IoU v3 loss function to replace the original CIoU localization loss function in the YOLO11 model and utilizing its mechanism of dynamically suppressing low-quality samples.

[0007] The present invention also provides a graspable target detection method for low-quality images based on the EMSDH-YOLO network described in claim 1. The steps of this method are as follows: Step 1: Establish a self-built dataset containing graspable objects, obtain a publicly available dataset containing graspable objects, and process this publicly available dataset. All dataset labels have class labels, and save the best model weights for training; Step 2: Configure the system environment required for model training; Step 3: Construct an EMSDH-YOLO network model; Step 4: Load the constructed EMSDH-YOLO network model into the configured training environment, modify the parameters of the model, and train it on the self-built dataset and the publicly available dataset respectively; Step 5: Detect the graspable target to be detected using the saved best model weights.

[0008] Preferably, the specific steps of establishing the self-built dataset containing graspable objects in step 1 are as follows: (1) Obtain the Cornell grasp dataset; (2) On the basis of retaining the data in the original dataset, extract the RGB images in the Cornell grasp dataset, and define the categories of the graspable objects in this dataset. The number of category types is 71; (3) Obtain supplementary images of the same category through category screening and web crawling, and use annotation tools to annotate the category labels of the above RGB images to obtain the annotation files corresponding to the RGB images; (4)Integrate the above annotation files with the data of the original dataset to obtain a comprehensive dataset that contains both grasping labels and category labels, and name this dataset Cornell-Detection; (5)Divide the RGB images and corresponding annotation files in the Cornell-Detection dataset into a training set, a validation set, and a test set according to the ratio of 7:1:2 for model training.

[0009] Preferably, the specific steps of obtaining a publicly available dataset containing graspable objects and processing this publicly available dataset in Step 1 are as follows: (1)Obtain the COCO dataset; (2)Extract graspable objects from the training set and validation set of the COCO dataset respectively, and extract 36 types of graspable objects from the original 80 types of objects in the COCO dataset; (3)The extracted training set and validation set form a new dataset for model training, and name it COCO-Graspable.

[0010] Preferably, in Step 2, configure the system environment required for model training. In this system environment, Pytorch is used as the deep learning framework, and Anaconda is used as the virtual environment management tool.

[0011] Preferably, in Step 4, load the constructed EMSDH-YOLO network model into the configured training environment, modify the parameters of the model, and train it on the self-built dataset and the publicly available dataset respectively, and save the best model weights of the training. The specific method is as follows: In the model loading process, deploy the EMSDH-YOLO network model on Windows for training, select the Pytorch framework, and use a single GPU for training; For modifying the parameters of the model, the number of training rounds is 500, the batch size is 32, the default size of the input image is 640×640, the optimizer is SGD, and the number of threads is 4; For training on the self-built dataset and the publicly available dataset, perform multiple groups of training on the EMSDH-YOLO model on the Cornell-Detecion self-built dataset and the COCO-Graspable publicly available dataset, and save the best model weights of the training when the model converges within the set batch.

[0012] Preferably, in Step 5, use the saved best model weights to detect the graspable targets to be detected. The specific method is as follows: Detect the graspable object to be detected, import the path of the best training weight into the detection script, and detect the image, video, and graspable object through the camera.

[0013] The present invention has the following characteristics and beneficial effects: The present invention provides a graspable object detection method for low-quality images based on the EMSDH-YOLO network. An EMSDH-YOLO model with high detection accuracy is obtained by improving the YOLO11 model. In some C3k2 modules in the backbone and neck structures of the YOLO11 model, hierarchical multi-scale convolution is used to replace the original convolution to construct an EMS-ConvBlock module, which can not only achieve mild lightweight but also better enable cross-scale interaction of features. In the fast spatial pyramid pooling structure in the backbone of the YOLO11 model, the idea of fusing large-kernel separable convolution and attention-enhanced cross-level feature focusing is added to construct an LSK-SPPF module. While increasing the receptive field of the neural network, it not only avoids the explosion of the number of parameters but also increases the feature extraction ability of this structure. In the C2PSA structure in the backbone of the YOLO11 model, DynamicTanh is added to nonlinearly activate the features to construct a DyT-C2PSA module. Through the collaborative design of dynamic activation, attention mechanism, and multi-scale convolution, the balance between feature expression ability and computational efficiency is achieved. In the upsampling part of the neck structure of the YOLO11 model, channel shuffle and spatial rolling mixing operations are added to construct an EU-SCMixer module, which displaces the features along the channel (C), height (H), and width (W) dimensions, promoting information interaction in the spatial and channel dimensions without increasing parameters and greatly enhancing the diversity of features. The Wise-IoU v3 loss function is selected to optimize the localization loss function of the YOLO11 model, solving the problem of poor bounding box fitting ability of the CIoU loss function in the YOLO11 model. The boundary box regression quality is optimized through the dynamic focusing mechanism, which can better detect occlusions and small objects, improve the utilization rate of low-quality image data, and improve the detection accuracy of the model when the imaging is poor. The EMSDH-YOLO model with high detection accuracy is used for graspable object detection, improving the detection accuracy of common graspable objects in life. Description of the Drawings

[0014] Figure 1 It is a schematic flowchart of a graspable object detection method based on the EMSDH-YOLO network provided by an embodiment of the present invention; Figure 2 It is the overall structure diagram of the EMSDH-YOLO model in an embodiment of the present invention; Figure 3 It is the structure diagram of the EMS-ConvBlock; Figure 4 It is the C3k_EMSCB structure diagram in the EMS-ConvBlock structure diagram; Figure 5 It is the Bottleneck_EMSCB structure diagram in the EMS-ConvBlock structure diagram; Figure 6 It is the schematic diagram of the principle of EMS-ConvBlock; Figure 7 It is the LSK-SPPF structure diagram; Figure 8 It is the LSKA structure diagram in the LSK-SPPF structure diagram; Figure 9 It is the DyT-C2PSA structure diagram; Figure 10 It is the DyT-PSABlock structure diagram in the DyT-C2PSA structure diagram; Figure 11 It is the EU-SCMixer structure diagram; Figure 12 It is the schematic diagram of the principle of channel shuffle; Figure 13 It is the schematic diagram of the mathematical principle of channel shuffle; Figure 14 It is the schematic diagram of the principle of spatial rolling mixing; Figure 15 It is the comparison chart of the detection effects of the YOLO11 model and the EMSDH-YOLO model on the partial test set of the Cornell-Detection dataset (the left figure is the detection effect diagram of the YOLO11 model, and the right figure is the detection effect diagram of the EMSDH-YOLO model); Figure 16 It is the comparison chart of the detection effects of the YOLO11 model and the EMSDH-YOLO model on the partial validation set of the COCO-Graspable dataset (the left figure is the detection effect diagram of the YOLO11 model, and the right figure is the detection effect diagram of the EMSDH-YOLO model). Specific implementation manner

[0015] Next, the technical solution of the present invention will be clearly and completely described in the form of specific embodiments in combination with the accompanying drawings.

[0016] A graspable target detection method for low-quality images and based on the EMSDH-YOLO network provided by the present invention, the technical route of the detection method is as Figure 1 shown, and the network structure of the EMSDH-YOLO model is as Figure 2 shown. The complete detection steps of the method of the present invention are as follows: Step 1: Create a self-built data set containing graspable objects, obtain a public data set containing graspable objects and process the public data set, wherein the data set identifiers all have category labels.

[0017] In step 1, the establishment of a self-built data set specifically includes: First, obtain the download address of the Cornell crawling dataset from the official website, and parse the data file after downloading.

[0018] Subsequently, on the basis of retaining the original data set data, the RGB image files required for the subsequent steps are extracted, and the categories of the graspable objects in the data set are defined. There are 71 categories, and the specific categories are shown in Table 1.

[0019] Table 1 Then, through category screening and web crawling, we obtain complementary images of the same category, summarize the collected RGB images and the RGB images of the original dataset, and use the X-AnyLabeling annotation tool to annotate the above RGB images with category labels. We use the X-AnyLabeling intelligent annotation tool to quickly annotate the relevant categories and export the .json file format required for network training. We integrate the obtained annotation files with the original dataset data to obtain a comprehensive dataset that contains both crawled labels and category labels, and name this dataset Cornell-Detection.

[0020] Finally, use the dataset segmentation script to divide the RGB images of the Cornell-Detection dataset and their corresponding annotation files into training sets, validation sets, and test sets in a 7:1:2 ratio for model training.

[0021] In step 1, the establishment of a public data set specifically includes: First, obtain the COCO dataset from the official website, and then use the script file to convert the training set and validation set files to obtain the COCO dataset in YOLO format to facilitate the training of the YOLO network.

[0022] Then, the types of graspable objects are screened and extracted from the training set and validation set of the COCO dataset respectively. 36 types of graspable objects are screened out from the original 80 types of objects in the COCO dataset. The specific categories are shown in Table 2.

[0023] Table 2 Finally, the extracted training set and validation set form a new dataset for model training and are named COCO - Graspable.

[0024] Step 2: Configure the system environment required for model training.

[0025] In this step 2, the specific system environment includes: Software aspect: The selected computer system is Windows 11, the system virtual environment management tool is Anaconda, the software program compiler is Visual Studio Code, the deep learning framework and version is Pytorch 2.3.1, and the cuda version is 11.8.

[0026] Hardware aspect: The CPU is Intel(R) Core(TM) i9 - 10980XE, the memory is 128GB, and the GPU is NVIDIA GeForce RTX 3090.

[0027] Step 3: Build the EMSDH - YOLO network model; The EMSDH - YOLO network model uses the YOLO11 network as the baseline network. It replaces some convolutions in the fourth C3k2 module of the YOLO11 model backbone network and the C3k2 module at the end of the Neck network with efficient multi - scale convolutions to build the EMS - ConvBlock module; adds large - kernel separable convolutions to the spatial pyramid structure of the YOLO11 model backbone network and incorporates the idea of two - dimensional convolution decomposition to build the LSK - SPPF module; adds the idea of Dynamic Tanh to the C2PSA module of the YOLO11 model backbone network and uses hyperbolic tangent gating to achieve adaptive calibration of attention weights to build the DyT - C2PSA module; adds the idea of channel shuffle and spatial rolling mixing to the up - sampling architecture in the YOLO11 model Neck network to build the EU - SCMixer module; and optimizes the loss function of the YOLO11 model using the Wise - IoU v3 method.

[0028] Among them, the method of building the EMS - ConvBlock module is: For the fourth C3k2 module ( Figure 2 the 8th layer) of the YOLO11 model backbone network and the C3k2 module ( Figure 2 the 22nd layer) at the end of the Neck network, some convolutions ( Figure 5The second convolution in [ ] is replaced by an Efficient Multi-Scale Convolutional Block (EMSCB). The structure of the EMS-ConvBlock module is determined by the c3k parameter. If c3k is True, the core part is C3k_EMSCB; if c3k is False, the core part is Bottle_EMSCB. Figure 2 The c3k parameters of the two EMS-ConvBlock modules in [ ] are both True. The forward propagation process of the EMS-ConvBlock module is as Figure 3 shown. The input feature map first passes through a 1×1 convolutional layer for initial feature extraction, and then is split into two parallel branches by Split. The feature map of one branch passes through n C3k_EMSCB modules for in-depth processing in sequence. The processed feature maps are fused through the Concat operation, and finally pass through a 1×1 convolutional layer for feature integration and dimensionality adjustment to output the final result. Among them, the forward propagation process of the C3k-EMSCB structure in the EMS-ConvBlock module is as Figure 4 shown. The input feature map first passes through a 1×1 convolutional layer, then passes through 2 Bottleneck-EMSCB modules, and then performs a Concat operation on the channel dimension with the input feature map processed by the 1×1 convolutional layer. Finally, the information of different channels is fused through a 1×1 convolutional layer for output. The forward propagation process of the above Bottleneck-EMSCB module is as Figure 5 shown. The input feature map first passes through a 3×3 convolutional layer, and then passes through the EMSCB module. Whether to perform a residual connection between the output and the input is controlled by a parameter.

[0029] Through the collaborative design of multi-scale parallel convolution and lightweight feature fusion, the EMS-ConvBlock module significantly reduces the computational complexity while enhancing the feature expression ability. Its core innovative idea is concentrated in the EMSCB part, as Figure 6 shown. Its core idea is to divide the input feature map into four equal parts along the channel dimension and process them through 1×1, 3×3, 5×5, and 7×7 convolutional kernels respectively to capture multi-scale features from local details to global context in a parallel manner. The small kernel convolutions (1×1, 3×3) can focus on local texture and edge details, while the large kernel convolutions (5×5, 7×7) can integrate long-range spatial dependencies and enhance global semantic information. Subsequently, the features of different scales are fused through a 1×1 convolution, which not only compresses the channel dimension to reduce the computational amount, but also improves the model's ability to model complex spatial patterns through multi-scale complementarity.

[0030] Among them, is the model's computational load, is the number of input channels, is the number of output channels, is the convolutional kernel size, is the height of the feature map, is the width of the feature map.

[0031] It can be calculated from the above formula that this design reduces the computational burden of a single branch through channel splitting. The total computational load of multi-scale convolution is much less than that of single convolution. At the same time, the complementarity of multi-scale convolution is used to enhance the robustness of features. Finally, while maintaining a computational cost similar to that of single-scale convolution, it achieves efficient and comprehensive feature expression, providing a solution that balances computational efficiency and feature quality for lightweight deep learning models.

[0032] In this step 3, depthwise separable convolutions with large kernels are added to the spatial pyramid structure of the backbone network of the YOLO11 model, and the idea of two-dimensional convolution decomposition is incorporated to construct the LSK-SPPF module. Specifically, it includes: adding depthwise separable convolutions with large kernels and two-dimensional convolution decomposition to the spatial pyramid structure ( Figure 2 the 9th layer in Figure 7 ) of the backbone network of the YOLO11 model. The forward propagation process of the LSK-SPPF module is as Figure 7 shown. The input feature map is first processed by a 1×1 convolutional layer to obtain the first feature map, and then successively undergoes three max-pooling operations to obtain the second, third, and fourth feature maps in sequence. Then, the fourth feature map after pooling is concatenated with the first, second, and third feature maps in the channel dimension. Next, it is processed by the LSKA module, and finally, the information of different channels is fused through a 1×1 convolutional layer for output. Among them, the forward propagation process of the LSKA structure in the LSK-SPPF module is as Figure 8 shown. The input feature map is first successively processed by a 1×(2 d -1) depthwise separable convolutional layer in the horizontal direction and a (2 d -1)×1 depthwise separable convolutional layer in the vertical direction. Subsequently, it is successively processed by a 1× dilated convolutional layer in the horizontal direction and a ×1 dilated convolutional layer in the vertical direction. Then, it is adjusted in channels and the information is fused through a 1×1 convolution. Finally, the integrated information is weighted and multiplied with the input feature map through a residual connection for output, where d represents the dilation rate, k represents the maximum receptive field of the convolutional kernel, represents the floor operation; The core innovative idea of the LSK-SPPF module is the LSKA structure. In the LSK-SPPF module, LSKA combines two-dimensional convolution decomposition with large-kernel separable convolution, effectively balancing model performance and computational efficiency. Through cascaded depthwise separable convolutions in the horizontal and vertical directions, it achieves large receptive field coverage with low computational overhead, captures multi-scale spatial information, and preferentially focuses on the overall shape of the target rather than local textures, enhancing the robustness of feature representation. Decomposing traditional two-dimensional large-kernel convolutions into cascaded one-dimensional horizontal and vertical convolutions significantly reduces the number of parameters and memory footprint, alleviating the quadratic growth problem caused by large-kernel convolutions. On the multi-scale feature concatenation result after spatial pyramid pooling (SPPF), dilated convolutions are used to expand the receptive field, and residual connections are combined to achieve dynamic feature weighting, further enhancing the model's ability to model long-range dependencies and context information. This design enables LSKA to effectively improve the computational efficiency and feature expression ability of the LSK-SPPF module while maintaining high performance.

[0033] In step 3, the idea of Dynamic Tanh is added to the C2PSA module in the backbone network of the YOLO11 model, and adaptive calibration of attention weights is achieved using hyperbolic tangent gating to construct the DyT-C2PSA module, which specifically includes: combining the C2PSA module ( Figure 2 layer 10 in) of the backbone network of the YOLO11 model with the idea of Dynamic Tanh to perform non-linear activation on the input features, making the training more stable. The forward propagation process of the DyT-C2PSA module is as Figure 9 shown. The input feature map first passes through a 1×1 convolutional layer, and then the feature map undergoes Split segmentation and is evenly divided into two parts along the channel dimension. One part of the feature map continuously passes through n DyT-PSABlocks, and the processed feature map is concatenated with the other unprocessed part of the feature map in the channel dimension. Finally, the concatenated result passes through another 1×1 convolutional layer for final processing to obtain the output of the module. Among them, the forward propagation process of the DyT-PSABlock structure in the DyT-C2PSA module is as Figure 10 shown. The input feature map first undergoes non-linear transformation through a DynamicTanh activation layer, then the importance of the features is weighted through the attention mechanism layer, further non-linearly processed through a second DynamicTanh activation layer, and finally passes through two consecutive 1×1 convolutional layers for feature mapping and dimensional transformation, and finally outputs the processed result.

[0034] The core innovative idea of the DyT-C2PSA module is to introduce Dynamic Tanh. By introducing a learnable non-linear activation mechanism, Dynamic Tanh achieves dynamic adaptive calibration of the feature map. Specifically, the DynamicTanh layer performs non-linear transformation on the input features through a parameterized hyperbolic tangent function (such as , where , , are learnable parameters), dynamically adjusting the amplitude distribution and response range of the features. This mechanism not only simulates the compression and standardization effects of traditional normalization layers (such as LayerNorm) on the input, but also enhances the model's adaptive ability to different input scales through learnable parameters, alleviating the impact of extreme values on the gradient, thereby improving training stability. In the DyT-PSABlock, Dynamic Tanh works in tandem with the attention mechanism: the former enhances the expressive ability of the features through non-linear transformation, while the latter highlights the features of key regions based on the weighting mechanism. Finally, feature mapping and dimension optimization are achieved through consecutive 1×1 convolutions. This design effectively enhances the fusion efficiency of multi-scale features, while reducing the computational overhead of traditional normalization layers, ultimately improving the detection accuracy and robustness of the model in complex scenarios.

[0035] In the third step, the idea of channel shuffle and spatial rolling mixing is added to the upsampling architecture in the Neck network of the YOLO11 model to construct the EU-SCMixer module, which specifically includes: performing channel shuffle and spatial rolling mixing on the features input to the upsampling architecture ( Figure 2 , layers 11 and 14) in the Neck network of the YOLO11 model. The forward propagation process of the EU-SCMixer module is as Figure 11 shown. The input feature map is first enlarged in size through the upsampling layer to improve the spatial resolution of the feature map, and then local features are independently extracted for each channel through a 3×3 depthwise separable convolution layer. Subsequently, the feature interaction of the module is enhanced through channel shuffle and spatial rolling mixing operations. Finally, channel fusion and dimension adjustment are performed using 1×1 convolutions to output the final feature map.

[0036] The core innovative idea of the DyT-C2PSA module is to introduce the strategy of channel shuffle and spatial rolling mixing in the upsampling part. As Figure 12 , Figure 13 shown, channel shuffle breaks the inter-group isolation caused by channel grouping operations by rearranging the order of the feature map channels, enabling random combination of feature channels from different groups in subsequent layers, thereby promoting information interaction and fusion of cross-group features and enhancing the model's learning ability for multi-scale and diverse features. As Figure 14As shown, spatial rolling mixing rearranges the feature responses at different positions by performing periodic translation operations (such as rolling up and down or left and right) on the feature map along the spatial dimension, forcing the model to focus on the perturbed local spatial correlations, enhancing the context interaction and robustness of features in the spatial domain, and at the same time enhancing the model's ability to model the global structure of the feature map. In this process, in the EU-SCMixer module, through upsampling, depthwise separable convolution, and rearrangement of the channel dimension and spatial position on the input feature map, more efficient feature extraction and fusion are achieved, effectively enhancing the model's ability to capture and express multi-scale information, thereby optimizing the quality of the feature map.

[0037] In the third step, the Wise-IoU v3 method is used to optimize the loss function of the YOLO11 model, specifically including: replacing the CIoU loss function in the localization loss function of the YOLO11 model with the Wise-IoU v3 loss function. The Wise-IoU v3 loss function can significantly improve the robustness and detection accuracy of the model through the mechanism of dynamically suppressing low-quality samples.

[0038] From IoU to CIoU and then to Wise-IoU v3, the bounding box regression loss function in object detection has experienced an evolutionary process from single overlap measurement to comprehensive geometric constraint optimization. IoU uses the overlap ratio (intersection over union) of the predicted box and the ground truth box as the loss, but there are problems such as gradient disappearance when there is no overlap and ignoring shape differences. Subsequently, GIoU introduced a closure rectangle penalty term to solve the no-overlap problem, DIoU accelerated convergence through the center point distance, and CIoU further introduced an aspect ratio penalty term on the basis of DIoU, combining the overlap area, center point distance, and shape matching, significantly improving the localization accuracy. Its formula is as follows: Among them, is the intersection over union, measuring the overlap degree between the predicted box and the target box, represent the center points of the predicted box and the ground truth box respectively, is the Euclidean distance between the center points of the predicted box and the ground truth box, is the diagonal length of the smallest enclosing box (the smallest rectangle covering both boxes), is the aspect ratio difference between the target box and the predicted box (measuring shape similarity) and is defined as , is the weight coefficient and is defined as .

[0039] Wise-IoU v3 is a loss function designed for the bounding box regression problem in object detection tasks. It introduces a dynamic non-monotonic focusing mechanism on the basis of the traditional IoU to optimize the impact of low-quality samples on model training. Its formula is as follows: Among them, and are the horizontal and vertical coordinates of the ground truth box, and are the dimensions of the minimum bounding box of the predicted box and the ground truth box, represents separating and from the computational graph (not participating in the subsequent backpropagation process), is an intermediate parameter, is a non-monotonic focusing coefficient, defining the outlier degree to describe the quality of the anchor box, represents the monotonic focus coefficient, represents the moving average, is a manually set parameter, and the parameter can be adjusted to achieve the best effect, is the outlier degree coefficient.

[0040] Compared with CIoU, Wise-IoU v3 further considers the quality differences of anchor boxes. By defining the concept of "outlier degree" to measure the quality of anchor boxes, it dynamically adjusts the gradient gain of each anchor box accordingly. Wise-IoU v3 assigns a smaller gradient gain to high-quality anchor boxes to avoid over-optimization, while assigns an even smaller gradient gain to low-quality anchor boxes to reduce their negative impacts, so that the model can pay more attention to anchor boxes of ordinary quality and improve the overall localization accuracy. This approach helps to process datasets containing a large number of low-quality annotations, improving the model performance while also accelerating the convergence speed.

[0041] Step 4: Load the constructed EMSDH-YOLO network model into the configured training environment, modify the parameters of the model, and train it on the self-built dataset and the public dataset respectively, and save the best model weights of the training.

[0042] In the fourth step, the EMSDH-YOLO network model constructed in the third step is trained in the software and hardware environment mentioned in the second step. The main parameter settings during model training include: the input image size imgsz is 640×640, the number of training iterations epochs is 500, the batch size batch for each iteration is 32, the number of parallel threads workers for data loading is 4, the optimizer selects the SGD optimizer, the early stopping count patience is 50, and the pre-trained weight file is not used. After setting the above parameters, on the premise of ensuring the same parameters, the EMSDH-YOLO network is trained on the Cornell-Detection dataset and the COCO-Graspable dataset respectively, and the performance of the network model is evaluated according to the obtained evaluation metrics, and the model weight with the best evaluation metric during training is saved as the best model weight.

[0043] To evaluate the performance of the EMSDH-YOLO model, accuracy (P), recall (R), mean average precision (mAP@0.5, mAP@0.5:0.95), model computational cost (GFLOPs), number of model parameters (Params), detection time (time), and refresh rate (FPS) when batch is 1 are used as metrics to evaluate the model. Accuracy is used to describe the ratio of predicted positive examples to all positive examples, and the calculation formula is as follows: Among them, in the formula P represents accuracy, TP represents the number of correctly predicted positive samples, FP is the number of samples that predict negative examples as positive examples in the samples.

[0044] Recall is used to describe the ratio of predicted positive examples to actual positive examples, and the calculation formula is as follows: Among them, in the formula R represents recall, FN represents the number of samples that predict positive examples as negative examples in the samples.

[0045] After that, the present invention introduces a parameter to integrate the above two parameters, that is mAP , which is defined as the average of the APs of all objects and is applicable to multi-label image classification and detection, AP and mAP The calculation formula is as follows: Among them, N represents the number of classification categories.

[0046] In addition, the parameters selected for the present invention are as follows: the model computation amount (GFLOPs) and the number of model parameters (Params) of the overall model are the sum of the computation amounts and the number of parameters of each layer, and the detection time (time) is the sum of the preprocessing time, the inference time, and the postprocessing time. The above can be obtained by running the verification script. Regarding the frame rate (FPS) when the calculation batch is 1, the method of taking the average of ten calculations is used to eliminate the error caused by a single calculation, and its value is the reciprocal of the single time.

[0047] The EMSDH-YOLO network model is trained on the Cornell-Detection dataset and the COCO-Graspable dataset to obtain the respective best model weight files (best.pt). The test set of the Cornell-Detection dataset and the validation set of the COCO-Graspable dataset are used to test the model performance, and the model evaluation indicators shown in Table 3 and Table 4 are obtained; Table 3 Table 4 The EMSDH-YOLO network model has improved the mAP50 on the Cornell-Detection dataset by 1.6% compared with the YOLO11 network model, and the EMSDH-YOLO network model has improved the mAP50 on the COCO-Graspable dataset by 1.2% compared with the YOLO11 network model. Therefore, it can be proved that the EMSDH-YOLO network model has improved the detection accuracy on the premise of not increasing too many parameters and computation amounts. The EMSDH-YOLO model has enhanced the feature extraction ability of the network in low-quality images, separated useful target information from a large amount of feature data, has strong anti-interference ability, and can accurately identify graspable objects.

[0048] Step Five: Use the saved best model weights to detect the graspable target to be detected.

[0049] In this step five, the best.pt file obtained by training is used as the best weight to detect the graspable target to be detected.

[0050] The detection effect is as Figure 15 、 Figure 16 shown, Figure 15 The left picture shows the detection effect of using the best weights generated by training the YOLO11 network model on a part of the test set of the Cornell-Detection dataset, Figure 15The right picture shows the detection effect of using the optimal weights generated by training the EMSDH-YOLO network model on a partial test set of the Cornell-Detection dataset. Figure 16 The left picture shows the detection effect of using the optimal weights generated by training the YOLO11 network model on a partial validation set of the COCO-Graspable dataset. Figure 16 The right picture shows the detection effect of using the optimal weights generated by training the EMSDH-YOLO network model on a partial validation set of the COCO-Graspable dataset. Since both of these datasets have problems such as excessive categories and poor image quality, the overall accuracy of the network models is not high. However, by comparing the detection maps of the improved model with those of the original model, it can be found that there are missed detections and misdetections in the detection map of the YOLO11 network, while the EMSDH-YOLO network rarely has the above problems and can better perform the detection task of low-quality images. It can be seen from the comparison pictures that the improved model has better detection performance than the original model in cases of target occlusion and blurred target features when the image quality is low.

[0051] The embodiments given above are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts fall within the scope of protection of the present invention.

Claims

1. An EMSDH-YOLO network for graspable object detection in low-quality images, characterized in that, The EMSDH-YOLO network uses the YOLO11 network as the baseline network. The fourth C3k2 module in the backbone network of the YOLO11 model and the C3k2 module at the end of the Neck network are replaced with the EMS-ConvBlock module. The SPPF module in the backbone network of the YOLO11 model is replaced with the LSK-SPPF module, and the C2PSA module is replaced with the DyT-C2PSA module. The upsampling module in the Neck network of the YOLO11 model is replaced with the EU-SCMixer module. The loss function of the YOLO11 model is optimized using the Wise-IoU v3 method. The EMS-ConvBlock module has its structure determined by the c3k parameter. If c3k is True, the input feature map first undergoes initial feature extraction through a 1×1 convolutional layer and is then split by Split into two parallel branches. The feature map of one of the branches passes through n C3k_EMSCB modules for in-depth processing in sequence. The processed feature maps are fused through the Concat splicing operation and finally undergo feature integration and dimension adjustment through a 1×1 convolutional layer to output the final result. If c3k is False, the C3k_EMSCB module is replaced by the Bottle_EMSCB module. In the C3k-EMSCB module, the input feature map first undergoes processing through a 1×1 convolutional layer, then through 2 Bottleneck-EMSCB modules, and then undergoes Concat splicing processing in the channel dimension with the input feature map that has undergone 1×1 convolutional layer processing. Finally, the information of different channels is fused through a 1×1 convolutional layer for output. In the Bottleneck-EMSCB module, the input feature map first passes through a 3×3 convolutional layer, then through the EMSCB module, and whether to perform a residual connection between the output and the input is controlled by a parameter. In the EMSCB module, the input feature map is first split into four equal parts along the channel dimension, and then these four groups of feature maps are respectively processed through 1×1, 3×3, 5×5, and 7×7 convolutional layers to generate four groups of different feature maps. Finally, the information of different channels is fused through a 1×1 convolutional layer for output. In the LSK-SPPF module, the input feature map is first processed by a 1×1 convolutional layer to obtain a first feature map. Subsequently, it undergoes max pooling three times in a row to obtain a second feature map, a third feature map, and a fourth feature map in sequence. Then, the pooled fourth feature map is concatenated with the first, second, and third feature maps in the channel dimension through Concat. Next, it is processed by the LSKA module. Finally, a 1×1 convolutional layer is used to fuse the information of different channels for output. In the LSKA module, the input feature map first undergoes depthwise separable convolutional layer processing of 1×(2 d -1) and (2 d -1)×1 in sequence. Subsequently, it undergoes dilated convolutional layer processing of 1× and ×1 in sequence. Then, a 1×1 convolution is used to adjust the channels and fuse the information. Finally, the integrated information is weighted and multiplied with the input feature map through a residual connection for output. Among them, d represents the dilation rate, k represents the maximum receptive field of the convolutional kernel, represents the floor operation; In the DyT-C2PSA module, the input feature map is first processed by a 1×1 convolutional layer, and then the feature map is split along the channel dimension into two parts through Split processing. One part of the feature map is continuously processed by n DyT-PSABlock modules. After processing, the feature map is concatenated with the other unprocessed part of the feature map in the channel dimension. Finally, the concatenated result is further processed by a 1×1 convolutional layer to obtain the output of the module. In the DyT-PSABlock module, the input feature map is first subjected to a non-linear transformation through a DynamicTanh activation layer, then the importance of the features is weighted through an attention mechanism layer, and further non-linear processing is performed through a second DynamicTanh activation layer. Finally, it passes through two consecutive 1×1 convolutional layers for feature mapping and dimensional transformation, and finally outputs the processing result. In the EU-SCMixer module, the input feature map is first upsampled by an upsampling layer to improve the spatial resolution of the feature map. Subsequently, local features are independently extracted for each channel through a 3×3 depthwise separable convolutional layer. Then, the feature interaction of the module is enhanced through channel shuffle and spatial rolling mixing operations. Finally, 1×1 convolution is used for channel fusion and dimensional adjustment to output the final feature map. The method of optimizing the loss function of the YOLO11 model using the Wise-IoU v3 method is to replace the original CIoU localization loss function in the YOLO11 model with the Wise-IoU v3 loss function.

2. A graspable object detection method based on the EMSDH-YOLO network for low-quality images as described in claim 1, characterized in that, The steps of this method are as follows: Step 1: Establish a self-built dataset containing graspable objects, obtain a publicly available dataset containing graspable objects, and process the publicly available dataset. All dataset labels have class labels. Step 2: Configure the system environment required for model training. Step 3: Construct the EMSDH-YOLO network model. Step 4: Load the constructed EMSDH-YOLO network model into the configured training environment, modify the parameters of the model, and train on the self-built dataset and the publicly available dataset respectively, and save the best model weights of the training. Step 5: Detect the graspable target to be detected using the saved best model weights.

3. According to the detection method described in claim 2, characterized in that The specific steps of establishing the self-built dataset containing graspable objects in step 1 are as follows: (1) Obtain the Cornell grasp dataset. (2) On the basis of retaining the data of the original dataset, extract the RGB images in the Cornell grasp dataset, and define the categories of the graspable objects in the dataset. The number of category types is 71. (3) Obtain supplementary images of the same category through category screening and web crawling, and use annotation tools to annotate the category labels of the above RGB images to obtain the annotation files corresponding to the RGB images. Integrate the above annotation files with the original dataset data to obtain a comprehensive dataset that contains both grasping labels and category labels, and name this dataset Cornell-Detection; Divide the RGB images and corresponding annotation files in the Cornell-Detection dataset into a training set, a validation set, and a test set according to the ratio of 7:1:2 for model training.

4. The detection method according to claim 1, wherein The specific steps for obtaining a publicly available dataset containing graspable objects and processing this publicly available dataset in Step 1 are as follows: (1) Obtain the COCO dataset; (2) Extract graspable objects from the training set and validation set of the COCO dataset respectively, and extract 36 types of graspable objects from the original 80 types of objects in the COCO dataset; (3) The extracted training set and validation set form a new dataset for model training, and name it COCO-Graspable.

Citation Information

Cited By

  • Optical fiber anomaly detection method, device and equipment and storage medium

    CN121502685A

  • Visual classification method and system based on dynamic multi-scale space blocks

    CN121600330A