An Object Detection Fusion Method for Improving Yolov10

By introducing the context fusion CFM module into the neck network of YOLOv10, the limitations of YOLOv10 in complex scenarios and multi-object detection are solved, and higher detection accuracy and robustness are achieved.

CN119723272BActive Publication Date: 2025-06-17SHANDONG ARTIFICIAL INTELLIGENCE INSTITUTE +1

Patent Information

Application Number
CN202510198951.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-17
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

YOLOv10 has limitations in dealing with the detection of complex scenarios, different scales and multiple targets, resulting in poor detection performance of small or low-contrast targets and lacks effective context information integration strategies and dynamic adjustment mechanisms.

Method used

Four context fusion CFM modules are introduced into the neck network. The CFM module includes a channel adjustment convolution module, a feature splicing module, a CBAM attention module, a weighted feature recombination module and an output feature generation module. Through the combination of these modules, the fusion weight of each scale feature is dynamically generated, and the feature fusion is adaptively adjusted.

Benefits of technology

It significantly improves the accuracy of object detection, enhances the robustness of the model, suppresses the influence of redundant features, and can better handle object detection tasks of different scales and complex backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723272B_ABST
    Figure CN119723272B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical fields of computer vision and object detection, and specifically relates to an object detection fusion method for improving Yolov10, which is as follows: Select data to construct a data set, and then preprocess the data in the data set to obtain a preprocessed data set; Enhance the preprocessed data set, construct an object recognition network model based on the lightweight Yolov10n network, and improve the Yolov10n network. The improved Yolov10n network includes an input end, a backbone network, a neck network, and an output end. Among them, four context fusion CFM modules are introduced in the neck network. The CFM module includes a channel adjustment convolution module, a feature splicing module, an attention module, a weighted feature recombination module, and an output feature generation module; Input the image data in the preprocessed data set into the improved constructed object recognition network model to obtain the final detection result. The present invention can better process targets of different scales and diversities, thereby improving the robustness of the model to complex backgrounds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and target detection, and in particular to a target detection fusion method for improving Yolov10. Background Art

[0002] Feature fusion is an important part of the target detection algorithm. It improves the performance of target detection by integrating feature information from different levels and scales. Most target detection algorithms use simple weighted summation, splicing or layer-by-layer superposition to perform feature fusion. These methods cannot fully capture the complex relationship between multiple layers of features, resulting in limited feature fusion effects.

[0003] Existing feature fusion methods, such as feature pyramid network (FPN) and path aggregation network (PANet), can effectively fuse multi-scale features, but there are still some shortcomings: simple feature fusion may lead to information redundancy or conflict between features of different scales, reducing the discriminative ability of features. Traditional fusion methods are usually susceptible to complex background interference, especially when the contrast between the target and the background is low, the target edge is blurred, or the background texture is complex. In the target detection task, the location and category of the target often depend on contextual information. The effective fusion of contextual information can help the model better understand the background, but many feature fusion methods fail to fully integrate global contextual information and only focus on local features. However, the design of YOLOv10 is biased towards local feature extraction, lacks an effective contextual information integration strategy, cannot be adaptively adjusted according to the characteristics of the input image, fails to effectively integrate global and local information, and limits the generalization ability of the model. Moreover, the feature fusion strategy of YOLOv10 is usually static and fails to dynamically adjust the contribution of different features. YOLOv10 lacks this dynamic adaptation mechanism and cannot flexibly respond to changes in different scenes.

[0004] The existing Yolov10 algorithm still has certain limitations when dealing with complex scenes, different scales and multiple targets, resulting in unsatisfactory detection performance for small targets or low-contrast targets. In complex scenes, especially when the target scale varies greatly and is severely occluded, its detection performance still has room for improvement.

[0005] Therefore, the present invention provides a target detection fusion method for improving Yolov10 to solve the above problems. Summary of the invention

[0006] In view of the shortcomings of the prior art, the present invention develops a target detection fusion method for improving Yolov10. The main purpose of the present invention is to better handle targets of different scales and diversity, thereby improving the robustness of the model to complex backgrounds and further improving the detection effect of Yolov10.

[0007] The technical solution for the present invention to solve the technical problem is a target detection fusion method for improving Yolov10, including the following steps:

[0008] S1. Select data from the DeepPCB dataset in the publicly available intelligent manufacturing workshop to construct a dataset, and then preprocess the data in the dataset to obtain a preprocessed dataset;

[0009] S2. Augment the preprocessed dataset, construct a target recognition network model based on the lightweight Yolov10n network, and improve the Yolov10n network. The improved Yolov10n network includes an input end, a backbone network, a neck network, and an output end. Among them, four context fusion CFM modules are introduced into the neck network. The CFM module includes a channel adjustment convolution module, a feature splicing Concat module, an attention module CBAM, a weighted feature recombination module, and an output feature generation module;

[0010] S3. Input the image data in the preprocessed dataset into the improved constructed target recognition network model to obtain the final detection result.

[0011] S1 is specifically as follows:

[0012] S1.1. Construct a dataset:

[0013] The DeepPCB dataset is an existing PCB circuit board defect dataset. The DeepPCB dataset contains 1500 groups of DeepPCB images. Each group of images contains a defect-free template image and a corresponding test image. Construct a dataset according to the existing DeepPCB dataset;

[0014] S1.2. Preprocess the data in the dataset, clear the redundant information in the dataset, verify the accuracy of the data of the original images in the dataset, and then perform image enhancement and bounding box enhancement on the image data and perform flipping to obtain a preprocessed dataset , , the dataset contains images, represents the th preprocessed DeepPCB image in the dataset, ;

[0015] Specifically, it is cropped through the getRectSubPix function in the OpenCV library in Python, the image is smoothed through the GaussianBlur function in the OpenCV library in Python, and the coordinate axis is vertically flipped by 0 through the flip function in the OpenCV library in Python.

[0016] S2 is specifically as follows:

[0017] Construct a target recognition network model. The target recognition network model is essentially the YOLOv10 network. When training the target recognition network model, set the number of iterations to 300, the learning rate to 0.0001, the learning speed to decay to half of the original every 20 rounds of iteration, and the initial learning rate to , and the batch size is set to 16;

[0018] Improve the lightweight network Yolov10n. The improved lightweight network Yolov10n is divided into four parts: the input end, the backbone network, the neck network, and the output end;

[0019] The input end includes Mosaic data augmentation, adaptive anchor boxes, and adaptive image scaling;

[0020] The backbone network includes four convolutional modules Conv, four feature extraction layer modules C2f, two downsampling modules SCDown, and one adaptive information fusion integration module AIFI;

[0021] The neck network is located between the backbone network and the output end, and includes two upsampling modules Upsample, four context fusion modules CFM, three feature extraction layer modules C2f, one convolutional module Conv, one context awareness module C2fCIB, and one downsampling module SCDown;

[0022] The output end includes a detection head module v10Detect;

[0023] Among them, the convolutional module Conv includes a two-dimensional convolution Conv2d with a 3 ×3 convolution kernel, a batch normalization layer BatchNormal, and an activation function SiLU. The padding of the convolution kernel is 1 and the stride is 1;

[0024] The context fusion CFM module, which includes a channel adjustment convolution module, a feature splicing Concat module, a CBAM attention module, a weighted feature recombination module, and an output feature generation module.

[0025] S3 is specifically as follows:

[0026] Process the data in the pre - processed dataset according to the improved Yolov10n object detection network model in the dataset The image data in the dataset is input into the improved Yolov10n object detection network model, and passes through the input end, backbone network, neck network, and output end in sequence;

[0027] S3.1. The operation process in the input end:

[0028] The data in the dataset is input into the input end, and the data in the dataset is adjusted through the Mosaic data augmentation module, adaptive anchor box module, and adaptive image scaling module in sequence to obtain the adjusted dataset , denotes the th adjusted DeepPCB image in the dataset;

[0029] The specific steps are as follows: The data in the dataset first enters the Mosaic data augmentation module. The Mosaic data augmentation module randomly selects 3 images for each image in the dataset, performs splicing processing, crops and scales the image and the selected 3 images so that their sizes are suitable for the splicing operation. The 4 processed images are spliced into a complete image according to a 2×2 layout, and the information of the complete image is updated according to the splicing position and scaling ratio to obtain the dataset of the complete image. The updated information includes the target box and class label; then the images in the complete image dataset are input into the adaptive anchor box module. The widths and heights of all the annotation boxes in the complete images are statistically analyzed to generate a histogram of the aspect ratio, and then clustering is performed based on the aspect ratio of the target box to generate anchor boxes. The size of each anchor box is the clustering center, representing the main size distribution of the targets in the complete image dataset. The anchor box parameters are updated according to the clustering results, and the updated anchor box sizes are aligned with the target boxes; finally, the complete image with updated information enters the adaptive image scaling module to obtain the original width of the image and height of the image , and the complete image is adjusted to generate the adjusted DeepPCB image , and further obtain the adjusted dataset , denotes the th Adjusted DeepPCB image.

[0030] S3.2, Operational Process in the Backbone Network:

[0031] The backbone network includes four convolutional modules Conv, four feature extraction layer modules C2f, two downsampling modules SCDown, and one module AIFI;

[0032] The four convolutional modules Conv are the first convolutional module , the second convolutional module , the third convolutional module , and the fourth convolutional module . The four feature extraction layers C2f are the first feature extraction layer , the second feature extraction layer , the third feature extraction layer , and the fourth feature extraction layer . The two downsampling SCDown modules include the first downsampling layer and the second downsampling layer ;

[0033] Input the image data in the dataset into the backbone network. First, read the image through the imread function in the OpenCV library in Python, and then input the read image into the first convolutional module . The number of output channels increases from 3 to 64 to obtain the feature map . The feature map passes through the second convolutional module . The number of output channels increases from 64 to 128 to obtain the feature map . Then the feature map passes through the first feature extraction layer to obtain the feature map . The feature map passes through the third convolutional module . The number of output channels increases from 128 to 256 to obtain the feature map . The feature map passes through the second feature extraction layer to obtain the feature map . The feature map passes through the first downsampling layer . The number of output channels increases from 256 to 512 to obtain the feature map . The feature map passes through the third feature extraction layer to obtain the feature map with 512 output channels , feature map After the second downsampling layer , the number of output channels increases from 512 to 1024, obtaining a feature map , the feature map passes through the fourth feature extraction layer , obtaining a feature map with 1024 output channels , the feature map then passes through the fourth convolution module for feature extraction, the output channels are reduced from 1024 to 512, and the extracted features are input into the AIFI module. Through the AIFI module, scale-internal interaction of high-level semantic features is performed, obtaining a feature map with 512 output channels .

[0034] S3.3. Operation process in the network:

[0035] The neck network is located between the backbone network and the output end, including two upsampling modules Upsample, four context fusion modules CFM, three feature extraction layer modules C2f, one convolution module Conv, one context awareness module C2fCIB, and one downsampling module SCDown;

[0036] Among them, the two upsampling modules Upsample are the first upsampling module and the second upsampling module , the four context fusion modules CFM are the first context fusion module , the second context fusion module , the third context fusion module and the fourth context fusion module , the three feature extraction layer modules C2f respectively include the fifth feature extraction layer , the sixth feature extraction layer and the seventh feature extraction layer , the convolution module Conv includes the fifth convolution module , the context awareness module C2fCIB includes the first context awareness module , the downsampling module SCDown includes the third downsampling layer ;

[0037] Input the final output feature map of the Backbone backbone network into the Neck network. The feature map first passes through the first upsampling module for upsampling, the feature map size is doubled, obtaining a feature map with 512 output channels , the feature map is combined with the feature map Through the first context fusion module perform fusion splicing to output a feature map with 1024 channels , the feature map passes through the fifth feature extraction layer , and outputs a feature map with 512 channels , the feature map passes through the second upsampling module , and outputs a feature map with 512 channels , the feature map and the feature map through the second context fusion module perform fusion splicing to output a feature map with 1024 channels , and the feature map passes through the sixth feature extraction layer , and outputs a feature map with 256 channels , and the feature map is input into the fifth convolution module , and outputs a feature map with 256 channels , and the feature map and the feature map through the third context fusion module perform fusion splicing to output a feature map with 768 channels , the feature map passes through the seventh feature extraction layer , and outputs a feature map with 768 channels , the feature map passes through the third downsampling layer , and outputs a feature map with 512 channels , and the feature map and the feature map through the fourth context fusion module perform fusion splicing to output a feature map with 1024 channels , the feature map passes through the first context-aware module , and outputs a feature map with 1024 channels .

[0038] S3.4. Operation process at the output end;

[0039] The feature map performs object detection through the v10Detect detection module at three different spatial resolutions of scales. The feature map It is processed independently at the spatial resolution of each scale and then input into a dedicated convolutional layer for further feature extraction. Using a convolution with a kernel size of 1×1, the feature maps of each scale are mapped to the number of channels required for detection, which is , , where and represent the coordinates of the target box, and represent the width and height of the feature maps of each scale, represents the confidence, represents the number of classes. Before making predictions, pairwise feature fusion is performed for each scale. After the feature maps of each scale are processed, two branches are formed, namely the classification branch and the regression branch. The classification branch represents the class prediction of the target and outputs the confidence of the class. The regression branch represents the regression of the coordinates of the target box, predicts the position and size of the target box, and outputs the feature maps of each scale. The output resolutions of the three scales are , and , and represent the height and width of the feature map respectively. Finally, the feature maps of the three scales are merged to generate the complete detection results. The output of each feature map includes the position of the target box, the confidence, and the class probability.

[0040] The working process of the CFM module is as follows:

[0041] (1) The CFM module receives two feature maps, which are the low-resolution, high-semantic feature map and the high-resolution, low-semantic feature map , , , represents the batch size, , and represent the number of channels, height, and width of the feature map respectively, , and represent the number of channels, height, and width of the feature map respectively. Then, the number of channels of the feature map is adjusted through the channel adjustment convolutional module to make its number of channels consistent with that of the feature map . Then, the feature map with aligned channels and the feature map are concatenated along the channel dimension through the feature concatenation module Concat, , and the concatenated feature , determine whether there is and before performing the splicing operation. If or , adjust the resolution of the feature map through interpolation or downsampling operations;

[0042] (2) Then input the spliced feature into the CBAM module, and perform global pooling operation on the spatial dimension of the spliced feature through the channel attention module of the CBAM module. The global pooling operation on the spatial dimension includes global average pooling operation GAP and global max pooling operation GMP, and the calculation formulas are as follows:

[0043] ,

[0044] ,

[0045] where and respectively represent the height and width of the spliced feature , represents the index of the height , represents the index of the width , represents taking the maximum height and width, represents the result of performing the global average pooling operation GAP on the spliced feature , represents the result of performing the global max pooling operation GMP on the spliced feature ;

[0046] Then input the results of GAP and GMP into the shared two-layer fully connected layer MLP to generate the channel attention weight , and the calculation formula is as follows:

[0047] ,

[0048] where represents the Sigmoid activation function used for normalization, represents the fully connected layer;

[0049] Then perform max pooling operation on the channel dimension of the spliced feature and average pooling operation through the spatial attention module of the CBAM module, and the calculation formulas are as follows:

[0050] ,

[0051] ,

[0052] Among them, represents the number of channels of the splicing feature , represents the index of the number of channels , represents the result of performing max pooling operation on the splicing feature ; represents the result of performing average pooling operation on the splicing feature ;

[0053] Then, the result of and are spliced, and after two-dimensional convolution operation and normalization operation, the spatial attention weight is obtained. The calculation formula is as follows:

[0054] ,

[0055] According to the spatial attention weight and the channel attention weight calculate the enhanced attention feature map finally output by the CBAM module. The calculation formula is as follows:

[0056] ;

[0057] (3) Input the enhanced attention feature map into the weight feature recombination module for weight-based feature recombination. The weight feature recombination module splits the enhanced attention feature map and the feature map into the feature map and the feature map according to the weights of the feature map , and then add the feature map and the feature map item by item to obtain the feature map . Add the feature map and the feature map item by item to obtain the feature map ;

[0058] (4) Input the feature map and the feature map into the output feature generation module, splice the two features, and obtain the feature map finally output by the CFM module. The calculation formula is as follows:

[0059] ,

[0060] Among them, represents a splicing operation.

[0061] The effects provided in the invention content are only the effects of the embodiments, rather than all the effects of the invention. The above technical solutions have the following advantages or beneficial effects:

[0062] The present invention uses the multi-scale feature maps extracted by Yolov10n as input, and through operations such as convolution, activation functions, and attention mechanisms, dynamically generates the fusion weights of each scale feature, so as to adaptively adjust according to the content of the input image under multi-scale targets and complex backgrounds. Finally, the weighted and fused feature maps are input into the detection head of Yolov10n for object detection, which can significantly improve the accuracy of object detection, enhance the robustness of the model, and suppress the influence of redundant features;

[0063] The present invention solves the problem of the feature fusion mechanism of Yolov10n that cannot effectively distinguish targets from backgrounds and cannot dynamically adjust context information in some scenarios, and proposes an improvement and enhancement of the context fusion module CFM. The core goal of the CFM module is to fuse feature maps from multiple resolutions and achieve efficient feature enhancement through the guidance of context information. This fusion mechanism can integrate high-semantic low-resolution features and high-resolution low-semantic features. The low-resolution feature maps contain more extensive context and semantic information, which is suitable for extracting the category and global relationship of targets. The high-resolution feature maps contain more details and local feature information, which is suitable for capturing local features such as the boundaries, details, and shapes of targets. The high-level low-resolution feature maps emphasize the global position, category, and large-range background information of the targets in the image, which helps to provide context understanding of the scene. The low-level high-resolution feature maps contain fine local information of the image, such as edges, textures, etc., which helps to more accurately locate small targets and targets with complex shapes, enabling the model to have both spatial detail information and global semantic information in the object detection task, and thus improving the detection accuracy. The CFM module makes the high-semantic features supplement the semantic information of the low-semantic features through guided fusion, and allows the low-resolution features to absorb the detail information of the high-resolution features;

[0064] The CFM module also introduces the CBAM attention mechanism to achieve efficient fusion. Through the CBAM attention mechanism, the module can capture and utilize important context information during the feature fusion process, thereby enhancing the effectiveness of feature representation and effectively guiding the model to learn the information of the detection targets, thus improving the detection accuracy of the model. Through the weighted feature recombination operation, the module can enhance important features while suppressing unimportant features and improve the discriminative ability of the feature maps.

[0065] In summary, the present invention effectively improves the object detection effect, the model can more accurately locate information, and enhances the information flow. Brief Description of the Drawings

[0066] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention.

[0067] Figure 1 It is a schematic diagram of the method flow of the present invention.

[0068] Figure 2 It is a diagram of the detection result of the present invention. Detailed Embodiments

[0069] In order to clearly illustrate the technical features of the present solution, the present invention will be described in detail below through specific embodiments and in conjunction with its drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, the components and settings of specific examples are described below.

[0070] Embodiment 1

[0071] A target detection fusion method for improving Yolov10 includes the following steps:

[0072] S1. Select data in the publicly available DeepPCB dataset in the intelligent manufacturing workshop to construct a dataset, and then preprocess the data in the dataset to obtain a preprocessed dataset;

[0073] S2. Enhance the preprocessed dataset, construct a target recognition network model based on the lightweight Yolov10n network, and improve the Yolov10n network. The improved Yolov10n network includes an input end, a backbone network, a neck network, and an output end. Among them, four context fusion CFM modules are introduced into the neck network. The CFM module includes a channel adjustment convolution module, a feature splicing Concat module, an attention module CBAM, a weighted feature recombination module, and an output feature generation module;

[0074] The Yolov10n network is a variant of the Yolov10 network;

[0075] S3. Input the image data in the preprocessed dataset into the improved constructed target recognition network model to obtain the final detection result.

[0076] S1 is specifically as follows:

[0077] S1.1. Construct a dataset:

[0078] The DeepPCB dataset is an existing PCB circuit board defect dataset. The DeepPCB dataset contains 1500 groups of DeepPCB images. Each group of images includes a defect-free template image and a corresponding test image. A dataset is constructed based on the existing DeepPCB dataset;

[0079] S1.2. Preprocess the data in the dataset, clear the redundant information in the dataset, verify the accuracy of the data of the original images in the dataset, and then perform image enhancement and bounding box enhancement on the image data and perform flipping to obtain the preprocessed dataset , ,the dataset contains images, denotes the th preprocessed DeepPCB image in the dataset, ;

[0080] Specifically, it is cropped through the getRectSubPix function in the OpenCV library in Python, smoothed through the GaussianBlur function in the OpenCV library in Python, and vertically flipped with the axis 0 through the flip function in the OpenCV library in Python.

[0081] S2 is specifically as follows:

[0082] Construct a target recognition network model. The target recognition network model is essentially the YOLOv10 network. When training the target recognition network model, set the number of iterations to 300, the learning rate to 0.0001, the learning speed to decay to half of the original every 20 rounds of iteration, and the initial learning rate to , and the batch size is set to 16;

[0083] Improve the lightweight network Yolov10n. The improved lightweight network Yolov10n is divided into four parts: the input end, the backbone network, the neck network, and the output end;

[0084] The input end includes Mosaic data augmentation, adaptive anchor boxes, and adaptive image scaling;

[0085] The backbone network includes four convolutional modules Conv, four feature extraction layer modules C2f, two downsampling modules SCDown, and one adaptive information fusion integration module AIFI;

[0086] The neck network is located between the backbone network and the output end, and includes two upsampling modules Upsample, four context fusion modules CFM, three feature extraction layer modules C2f, one convolution module Conv, one context-aware module C2fCIB, and one downsampling module SCDown;

[0087] The output end includes a detection head module v10Detect;

[0088] Among them, the convolution module Conv includes a two-dimensional convolution Conv2d with a 3 ×3 convolution kernel, a batch normalization layer BatchNormal, and an activation function SiLU. The padding of the convolution kernel is 1, and the stride is 1;

[0089] The context fusion CFM module, which includes a channel adjustment convolution module, a feature concatenation Concat module, a CBAM attention module, a weighted feature recombination module, and an output feature generation module.

[0090] S3 is specifically as follows:

[0091] Process the data in the preprocessed dataset according to the improved Yolov10n object detection network model in the dataset input the image data in the dataset into the improved Yolov10n object detection network model, and sequentially pass through the input end, the backbone network, the neck network, and the output end;

[0092] S3.1. The operation process in the input end:

[0093] Input the data in the dataset into the input end, and sequentially adjust the data in the dataset through the Mosaic data augmentation module, the adaptive anchor box module, and the adaptive image scaling module to obtain the adjusted dataset where , represents the th adjusted DeepPCB image in the dataset;

[0094] The specific steps are that the data in the dataset first enters the Mosaic data augmentation module. The Mosaic data augmentation module randomly selects 3 images for each image in the dataset, performs splicing processing, and crops and scales the image and the selected 3 images so that their sizes are suitable for and the selected 3 images to make their sizes suitable for The splicing operation is to splice the 4 processed images into a complete image according to a 2×2 layout, update the information of the complete image according to the splicing position and scaling ratio to obtain a dataset of the complete image, and the updated information includes the target box and class label; then input the images in the complete image dataset into the adaptive anchor box module, count the width and height distributions of the annotation boxes of all complete images, generate a histogram of the aspect ratio, and then perform clustering based on the aspect ratio of the target box to generate a number of anchor boxes, where the size of each anchor box is the clustering center, representing the main size distribution of the targets in the complete image dataset, update the anchor box parameters according to the clustering result, and align the updated anchor box size with the target box; finally, input the complete image with updated information into the adaptive image scaling module to obtain the original width and height of the complete image, adjust the complete image to generate an adjusted DeepPCB image and further obtain an adjusted dataset where represents the nth adjusted DeepPCB image in the dataset.

[0095] S3.2. Operation process in the backbone network:

[0096] The backbone network includes four convolutional modules Conv, four feature extraction layer modules C2f, two downsampling modules SCDown, and one module AIFI;

[0097] The four convolutional modules Conv are the first convolutional module , the second convolutional module , the third convolutional module , and the fourth convolutional module . The four feature extraction layers C2f are the first feature extraction layer , the second feature extraction layer , the third feature extraction layer , and the fourth feature extraction layer . The two downsampling SCDown modules include the first downsampling layer and the second downsampling layer ;

[0098] Input the image data in the dataset into the backbone network. First, read the image through the imread function in the OpenCV library in Python, and then input the read image into the first convolutional module Among them, the number of output channels is increased from 3 to 64 to obtain a feature map , the feature map passes through the second convolution module , the number of output channels is increased from 64 to 128 to obtain a feature map , and then the feature map passes through the first feature extraction layer , to obtain a feature map , the feature map passes through the third convolution module , the number of output channels is increased from 128 to 256 to obtain a feature map , the feature map passes through the second feature extraction layer , to obtain a feature map , the feature map passes through the first downsampling layer , the number of output channels is increased from 256 to 512 to obtain a feature map , the feature map passes through the third feature extraction layer , to obtain a feature map with 512 output channels , the feature map passes through the second downsampling layer , the number of output channels is increased from 512 to 1024 to obtain a feature map , the feature map passes through the fourth feature extraction layer , to obtain a feature map with 1024 output channels , the feature map Then, it passes through the fourth convolution module for feature extraction. The output channels are reduced from 1024 to 512, and the extracted features are input into the AIFI module. The AIFI module performs scale-internal interaction on the high-level semantic features to obtain a feature map with 512 output channels .

[0099] S3.3. Operation process in the network:

[0100] The neck network is located between the backbone network and the output end, and includes two upsampling modules Upsample, four context fusion modules CFM, three feature extraction layer modules C2f, one convolution module Conv, one context awareness module C2fCIB, and one downsampling module SCDown;

[0101] Among them, the two upsampling modules Upsample are the first upsampling module and the second upsampling module , and the four context fusion modules CFM are the first context fusion module , the second context fusion module , the third context fusion module and the fourth context fusion module , and the three feature extraction layer modules C2f respectively include the fifth feature extraction layer , the sixth feature extraction layer and the seventh feature extraction layer , the convolution module Conv includes the fifth convolution module , the context awareness module C2fCIB includes the first context awareness module , the downsampling module SCDown includes the third downsampling layer ;

[0102] Input the final output feature map of the Backbone main network into the Neck network. The feature map is first upsampled through the first upsampling module , the size of the feature map is doubled, and a feature map with an output channel number of 512 is output . The feature map is fused and concatenated with the feature map through the first context fusion module , and a feature map with an output channel number of 1024 is output . The feature map passes through the fifth feature extraction layer , and a feature map with an output channel number of 512 is output . The feature map passes through the second upsampling module , and a feature map with an output channel number of 512 is output . The feature map is fused and concatenated with the feature map through the second context fusion module , and a feature map with an output channel number of 1024 is output . The feature map passes through the sixth feature extraction layer , and a feature map with an output channel number of 256 is output . The feature map is input into the fifth convolution module , and a feature map with an output channel of 256 is output . The feature map and the feature map are fused and concatenated through the third context fusion module , and a feature map with an output channel number of 768 is output . The feature map passes through the seventh feature extraction layer The feature map with 768 output channels The feature map Passes through the third downsampling layer The feature map with 512 output channels The feature map And the feature map Are fused and stitched through the fourth context fusion module To output a feature map with 1024 output channels The feature map Passes through the first context-aware module To output a feature map with 1024 output channels .

[0103] S3.4, The operation process in the output end;

[0104] The feature map Undergoes object detection by the v10Detect detection module at three different spatial resolutions of scales. The feature map Is independently processed at the spatial resolution of each scale and then input into a dedicated convolutional layer for further feature extraction. A convolution with a kernel size of 1×1 is used, and the feature map of each scale is mapped to the number of channels required for detection, which is , , where And Represent the coordinates of the target box, And Represent the width and height of the feature map of each scale, Represents the confidence, Represents the number of classes. Before prediction, pairwise feature fusion is performed for each scale. After the feature map of each scale is processed, two branches are formed, namely the classification branch and the regression branch. The classification branch represents class prediction of the target and outputs the confidence of the class. The regression branch represents regression of the coordinates of the target box and predicts the position and size of the target box, outputting the feature map of each scale. The output resolutions of the three scales are , And , And Respectively represent the height and width of the feature map . Finally, the feature maps of the three scales are merged to generate the complete detection result. The output of each feature map includes the position of the target box, the confidence, and the class probability.

[0105] The working process of the CFM module is as follows:

[0106] (1) The CFM module receives two feature maps, namely a low-resolution and high-semantic feature map and a high-resolution and low-semantic feature map , , , represents the batch size, , and respectively represent the number of channels, height, and width of the feature map . Then, the number of channels of the feature map , and respectively represent the number of channels, height, and width of the feature map . Then, the number of channels of the feature map is adjusted by the channel adjustment convolution module to make its number of channels consistent with that of the feature map . Then, the feature map with aligned channels and the feature map are concatenated along the channel dimension through the feature concatenation module Concat , obtaining the concatenated feature . Before performing the concatenation operation, it is determined whether there are and . If or , the resolution of the feature map is adjusted through interpolation or downsampling operations;

[0107] By fusing and , the semantic information of high-level features and the detailed information of low-level features can be utilized simultaneously, thereby improving the performance of the model in object detection. provides semantic information to help understand the category and background environment of the object, and at the same time provides semantic support for the background context. provides detailed information to help accurately locate the boundaries and contours of the object. At the same time, the ability to detect small objects is enhanced. Through the fusion of upper and lower layer features, the model can more comprehensively understand the environment where the object is located and improve the robustness to complex scenes;

[0108] In complex backgrounds, semantic information helps to separate the object from the background. When detecting small objects with blurred edges or low resolution, detailed information helps to improve the accuracy. Its significance lies in comprehensively utilizing the complementary advantages of the two types of feature maps, thereby improving the accuracy and robustness of object detection;

[0109] High-semantic and low-resolution feature map ( Function of [[ ]]: After deep convolutional processing, it contains global context information and higher-level semantic features, enabling rich semantic information and being more suitable for identifying the complex background of the scene. At the same time, the low resolution is selected because the receptive field of each pixel is large, which can capture a larger range of global information and is very important for detecting large targets or providing an overall semantic understanding of the scene background;

[0110] Feature map with low semantics and high resolution ( ) Function: It is rich in detail information, contains more spatial details and edge information, can capture low-level features such as the shape and contour of the target, and is more conducive to the detection of small targets and fine-grained targets. The high resolution is selected because the receptive field of each pixel is small, but it can represent more refined local features and is more suitable for detecting small targets and complex detailed structures;

[0111] (2) Then input the concatenated features into the CBAM module. Through the channel attention module of the CBAM module, perform global pooling operation on the concatenated features in the spatial dimension. The global pooling operation in the spatial dimension includes global average pooling operation GAP and global maximum pooling operation GMP. The calculation formulas are as follows:

[0112]

[0113]

[0114] Among them, and respectively represent the height and width of the concatenated features , represents the index of the height , represents the index of the width , represents taking the maximum height and width, represents the result of performing global average pooling operation GAP on the concatenated features , represents the result of performing global maximum pooling operation GMP on the concatenated features ;

[0115] Then input the results of GAP and GMP into a shared two-layer fully connected layer MLP to generate channel attention weights , and the calculation formula is as follows:

[0116]

[0117] Among them, represents the Sigmoid activation function used for normalization, represents the fully connected layer;

[0118] Then, through the spatial attention module of the CBAM module, the concatenated features are subjected to a max-pooling operation in the channel dimension and an average-pooling operation , and the calculation formulas are as follows:

[0119] ,

[0120] ,

[0121] where, represents the number of channels of the concatenated features , represents the index of the number of channels , represents the result of the max-pooling operation on the concatenated features , and represents the result of the average-pooling operation on the concatenated features ; Then, the results of

[0122] and are concatenated, and after a two-dimensional convolution operation and a normalization operation, the spatial attention weight is obtained, and the calculation formula is as follows: ;

[0123] ,

[0124] According to the spatial attention weight and the channel attention weight , the enhanced attention feature map finally output by the CBAM module is calculated, and the calculation formula is as follows:

[0125] ;

[0126] (3) Input the enhanced attention feature map into the weight feature recombination module for weighted feature recombination. The weight feature recombination module splits the enhanced attention feature map and the feature map into the feature map and the feature map according to the weights of the feature map , and then adds the feature map and the feature map item by item to obtain the feature map , and adds the feature map and the feature map item by item to obtain the feature map ;

[0127] (4) Input the feature map and the feature map into the output feature generation module, splice the two features, and obtain the feature map finally output by the CFM module. The calculation formula is as follows:

[0128] ,

[0129] where represents the splicing operation.

[0130] The principle of feature map fusion:

[0131] In terms of the complementarity of semantic information and detail information, high-semantic features and low-resolution feature maps provide the large-scale semantic information and background information of the image, while low-semantic features and high-resolution feature maps provide the detail information and boundary information in the image. By fusing these two types of information, the model can not only understand the global background and target classification of the image, but also locate the specific boundaries and details of the target, enabling the model to accurately identify and locate the target under the guidance of the global context;

[0132] In terms of the fusion of feature scales, the global information captured by the low-resolution feature map through a large receptive field needs to be combined with the local information of the high-resolution feature map . The low-resolution feature map usually has stronger global semantic perception due to coarser sampling, while the high-resolution feature map can capture the local details of the image. In cross-scale feature fusion, the fusion of low-resolution and high-resolution features can combine global semantics and local fine structures, enabling the model to consider both the global background and local details of the target when performing object detection;

[0133] In terms of the balance of information flow, by introducing the fusion of upper and lower layer features in the network and effectively fusing feature maps of different scales, the flow of global information and local information can be made more balanced. The advantage of this balance is that it provides a strong global understanding ability for large targets and the background, and a more accurate local detection ability for small targets and details.

[0134] Selecting low-resolution, high-semantic feature maps and high-resolution, low-semantic feature maps for fusion is actually fusing two different levels of feature information, which respectively represent global semantic information and local detail information. One is from the low-resolution feature map Generally, it can capture a wide range of context information of the image. These feature maps go through deeper network layers and have strong semantic abstraction capabilities, enabling the recognition and understanding of large objects, background information, and the relationships between objects. Low-resolution feature maps cover a wider area in the receptive field of each pixel, making it more sensitive to large objects, environmental information, and context in the image. The other type is from high-resolution feature maps which contain more details and local information, such as the edges, textures, shapes, and fine features of the object. These feature maps are relatively sparse, capturing local details but having weak semantic information. Since high-resolution feature maps only cover a small area in the receptive field of each pixel, it is more sensitive to the detailed structure of the image and is suitable for capturing small objects or subtle differences in the image.

[0135] The significance of choosing the combination of these two types of feature maps is that low-resolution feature maps provide strong global understanding, capable of recognizing large objects, background environments, and the relationships between objects in the image, while high-resolution feature maps can capture small objects and local details, helping the model accurately locate the object boundaries and fine features. By combining these two types of feature maps, the model can simultaneously understand the global context and local details of the image, providing a more comprehensive and accurate object detection ability. It has better compatibility for large and small objects: low-resolution feature maps can help identify and process large-scale object information, so they are suitable for detecting large objects. High-resolution feature maps can better capture small objects, especially in cases where the background is complex or small objects are partially occluded, providing more detailed information.

[0136] The advantages of doing so are as follows:

[0137] Firstly, it can enhance the accuracy of object detection. By combining low-resolution semantic feature maps and high-resolution detail feature maps, the model can comprehensively integrate information at multiple levels, avoiding information loss that may occur when relying solely on features at a certain level; for example, low-resolution feature maps help locate large objects, while high-resolution feature maps ensure the precise location of small objects. Secondly, it can improve the detection ability of small objects and fine-grained objects. High-resolution feature maps have an advantage in detecting small objects because they can capture more detailed information and avoid detail loss when dealing with complex scenes. By combining this small-object detection ability with global semantic information, the CFM module can improve the localization accuracy of small objects and enhance the robustness of the model. Low-resolution feature maps can help the model better understand background information, distinguish between objects and the background, thereby improving the ability to handle complex backgrounds; high-resolution feature maps It can effectively retain details in the image, help the model avoid occlusion problems or the influence of blurred targets, and increase the robustness of the model to complex scenarios; finally, it can improve the training and inference efficiency. The computational cost of low-resolution feature maps is relatively small, and it can provide rich semantic information without increasing the computational burden; by combining with high-resolution feature maps , the model can balance computational efficiency and accuracy. Feature map fusion can maintain computational efficiency through weighted and concatenation operations, and effectively utilize the advantages of low-resolution and high-resolution feature maps.

[0138] Implementation of the fusion process: First, for the feature map input, the input feature maps and are respectively low-resolution, high-semantic feature maps and high-resolution, low-semantic feature maps. Then, feature map fusion is performed. Usually, the low-resolution feature map and the high-resolution feature map will be fused through a weighted fusion operation. First, the feature maps are upsampled or downsampled as needed to ensure that their spatial resolutions are the same. After aligning the spatial dimensions, the two types of features are combined using a weighted fusion operation to obtain a richer feature representation. Secondly, a convolution operation is performed. A 1×1 convolution operation is performed on the fused feature map to adjust the number of channels of the feature map and further extract more refined features. This step usually helps to reduce redundant information. The 1×1 convolution can reduce the number of channels of the feature map, avoid information redundancy, and can also enhance the feature expression ability. More representative features are extracted through convolution, enabling the model to better perform object detection. Then, a branching operation is performed. Specifically, the category and category confidence of the target are output through the classification head, and the coordinates x, y, w, h of the target box are output through the regression head. Finally, the fusion output is performed. After the processing of all scales is completed, the final output results are integrated through multi-scale fusion to obtain the final object detection result.

[0139] Example 2

[0140] To prove the beneficial effects of the present invention, the method of the present invention is compared with existing methods. First, select the image to be detected, and select the image containing annotation information from the existing DeepPCB dataset as the image to be detected. Then, parse these annotation files and convert them into YOLO-format annotations. The YOLO format usually includes the number of the target category and the relative coordinates of the target box; to improve the generalization ability of the model, some common enhancement operations can be performed on the image, such as flipping, rotation, brightness adjustment, scaling, etc., which can increase the diversity of training data and reduce overfitting;

[0141] Select this method for configuration, set the hyperparameters of this method's network, input the images and corresponding labels into the network for training. Since the DeepPCB dataset contains different electronic components and each target has a corresponding class label, the training process will learn the features of these targets and output their positions and classes. The training process optimizes the network weights by minimizing these loss functions and uses the SGD optimizer to adjust the parameters. After each training epoch, the performance of the model is evaluated through the validation set, and the trained model is evaluated using the test set. By analyzing the detection results, check whether there are problems of missed detection or false detection in the model. After completing the training, use this model to perform object detection on new PCB images. The trained model can be embedded into the PCB automatic detection system to achieve real-time detection and component positioning; similarly, select other existing methods to process the images in the DeepPCB dataset.

[0142] As shown in Table 1, the method of the present invention and existing methods (Image Processing, Single Shot Detector (SSD), Faster Region-based Convolutional Neural Network (Faster R-CNN)) are respectively compared in terms of the mean Average Precision (mAP) of open circuit, short circuit, burr, pseudo copper, mouse bite mark, and hole. The present invention has a higher average accuracy in each problem, which proves that the method of the present invention is superior to the other three methods.

[0143] Table 1 Detection comparison results between the method of the present invention and existing methods

[0144]

[0145] Example 3

[0146] As Figure 2 shown, the detection result of the present invention Figure 2 is the detection result of the DeepPCB image to be detected composed of 16 local images. Among them, "open" represents an open circuit, "short" represents a short circuit, "mousebite" represents a mouse bite mark, "spur" represents a burr, "pin-hole" represents a hole, and "copper" represents pseudo copper. The above are the problems existing in the detected DeepPCB image, and the number after the problem represents the detected accuracy rate.

[0147] Although the specific implementation manners of the invention are described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A target detection fusion method for improving Yolov10, characterized in that: The following steps are involved: S1. Select the data in the DeepPCB dataset in the public intelligent manufacturing workshop to construct a dataset, and then preprocess the data in the dataset to obtain a preprocessed dataset; S2. Enhance the preprocessed data set, build a target recognition network model based on the lightweight Yolov10n network, and improve the Yolov10n network. The improved Yolov10n network includes an input end, a backbone network, a neck network, and an output end. The neck network introduces four context fusion CFM modules. The CFM module includes a channel adjustment convolution module, a feature concatenation Concat module, an attention module CBAM, a weighted feature reorganization module and an output feature generation module; S3. Input the image data in the preprocessed data set into the improved target recognition network model to obtain the final detection result.

2. The target detection fusion method for improving Yolov10 according to claim 1, characterized in that: S1 is as follows: S1.

1. Constructing the dataset: The DeepPCB dataset is an existing PCB circuit board defect dataset. The DeepPCB dataset contains 1,500 sets of DeepPCB images. Each set of images contains a defect-free template image and a corresponding test image. The dataset is constructed based on the existing DeepPCB dataset. S1.

2. Preprocess the data in the dataset, remove redundant information in the dataset, verify the accuracy of the original image data in the dataset, and then perform image enhancement and bounding box enhancement on the image data, and flip it to obtain the preprocessed dataset. , , dataset Included images, Representation dataset Middle A preprocessed DeepPCB image. ; Specifically, the getRectSubPix function in the OpenCV library in Python is used for cropping, the GaussianBlur function in the OpenCV library in Python is used for smoothing the image, and the flip function in the OpenCV library in Python is used for vertical flipping with the coordinate axis set to 0.

3. The target detection fusion method for improving Yolov10 according to claim 2, characterized in that: S2 is as follows: Build a target recognition network model. The target recognition network model is essentially a YOLOv10 network. When training the target recognition network model, set the number of iterations to 300, the learning rate to 0.0001, the learning speed to decay to half of the original speed every 20 iterations, and the initial learning rate to , the batch size is set to 16; Improve the lightweight network Yolov10n. The improved lightweight network Yolov10n is divided into four parts: input end, backbone network, neck network and output end; The input includes Mosaic data enhancement, adaptive anchor boxes, and adaptive image scaling; The backbone network includes four convolution modules Conv, four feature extraction layer modules C2f, two downsampling modules SCDown and one adaptive information fusion integration module AIFI; The neck network is located between the backbone network and the output end, and includes two upsampling modules Upsample, four context fusion modules CFM, three feature extraction layer modules C2f, a convolution module Conv, a context perception module C2fCIB and a downsampling module SCDown; The output end includes a detection head module v10Detect; Among them, the convolution module Conv includes a 3 3 convolution kernels of two-dimensional convolution Conv2d, a batch normalization layer BatchNormal and activation function SiLU, the padding of the convolution kernel is 1, and the stride is 1; The context fusion CFM module, which includes a channel adjustment convolution module, a feature concatenation Concat module, a CBAM attention module, a weighted feature reorganization module and an output feature generation module.

4. The target detection fusion method for improving Yolov10 according to claim 3, characterized in that: S3 is as follows: According to the improved Yolov10n target detection network model, the preprocessed dataset Process the data in the dataset Image data in Input into the target detection network model after improving Yolov10n, passing through the input end, backbone network, neck network and output end in sequence; S3.

1. Operation process at the input end: Dataset The data is input to the input end, and the data set is successively enhanced by the Mosaic data enhancement module, the adaptive anchor frame module, and the adaptive image scaling module. Adjust the data in the , Representation dataset Middle Adjusted DeepPCB image; The specific steps are as follows: The data first enters the Mosaic data enhancement module, which extracts For each image Randomly select 3 images and stitch them together to make the images The selected 3 images are cropped and scaled to fit Stitching operation: stitch the 4 processed images into a complete image in a 2×2 layout, and update the information of the complete image according to the stitching position and scaling ratio to obtain a complete image dataset. The updated information includes the target box and category label. Then, the images in the complete image dataset are input into the adaptive anchor box module, the width and height distribution of the annotation boxes of all complete images are counted, and a histogram of the aspect ratio is generated. Clustering, generation Anchor frames, each of which has a size of cluster center, represent the main size distribution of the target in the complete image dataset. The anchor frame parameters are updated according to the clustering results, and the updated anchor frame size is aligned with the target frame. Finally, the complete image with updated information enters the adaptive image scaling module to obtain the image The original width and height , adjust the complete image and generate the adjusted DeepPCB image , and then get the adjusted data set , Representation dataset Middle The adjusted DeepPCB image.

5. The target detection fusion method for improving Yolov10 according to claim 4, characterized in that: S3 is as follows: S3.

2. Operation process in the backbone network: The backbone network consists of four convolution modules Conv, four feature extraction layer modules C2f, two downsampling modules SCDown and one module AIFI; The four convolution modules Conv are the first convolution module , the second convolution module , the third convolution module And the fourth convolutional module , the four feature extraction layers C2f are the first feature extraction layer , the second feature extraction layer , the third feature extraction layer and the fourth feature extraction layer , the two downsampling SCDown modules include the first downsampling layer and the second downsampling layer ; The dataset The image data is input into the backbone network. First, the image is read through the imread function in the OpenCV library in Python, and then the read image is Input to the first convolution module In the example, the number of output channels increases from 3 to 64, and the feature map is obtained. , feature map After the second convolution module , the number of output channels increases from 64 to 128, and the feature map is obtained , then the feature map After the first feature extraction layer , get the feature map , feature map After the third convolution module , the number of output channels increases from 128 to 256, and the feature map is obtained , feature map After the second feature extraction layer , get the feature map , feature map After the first downsampling layer , the number of output channels increases from 256 to 512, and the feature map is obtained , feature map After the third feature extraction layer , and get a feature map with an output channel number of 512 ,Feature map After the second downsampling layer , the number of output channels increases from 512 to 1024, and the feature map is obtained , feature map After the fourth feature extraction layer , and get a feature map with an output channel number of 1024 , feature map Then through the fourth convolution module Feature extraction is performed, and the output channels are reduced from 1024 to 512. The extracted features are input into the AIFI module, and the high-level semantic features are interacted within the scale through the AIFI module to obtain a feature map with an output channel number of 512 .

6. The target detection fusion method for improving Yolov10 according to claim 5, characterized in that: S3 is as follows: S3.

3. Operation process in the network: The neck network is located between the backbone network and the output end, and includes two upsampling modules Upsample, four context fusion modules CFM, three feature extraction layer modules C2f, one convolution module Conv, one context perception module C2fCIB, and one downsampling module SCDown; Among them, the two upsampling modules Upsample are the first upsampling module and the second upsampling module The four context fusion modules CFM are the first context fusion module , Second context fusion module , the third context fusion module and the fourth context fusion module , the three feature extraction layer modules C2f include the fifth feature extraction layer , the sixth feature extraction layer and the seventh feature extraction layer , the convolution module Conv includes the fifth convolution module , the context-aware module C2fCIB includes the first context-aware module The downsampling module SCDown includes the third downsampling layer ; The final output feature map of the Backbone network Input to the Neck network, feature map First, through the first upsampling module Upsampling is performed, the feature map size is doubled, and the output channel number is 512. , the feature map With feature map Through the first context fusion module Fusion and splicing are performed to output a feature map with 1024 channels. , feature map After the fifth feature extraction layer , the output channel number is 512 feature map , feature map After the second upsampling module , the output channel number is 512 feature map , feature map With feature map Through the second context fusion module Fusion and splicing are performed to output a feature map with 1024 channels. , the feature map After the sixth feature extraction layer , the output channel number is 256 feature maps , the feature map Input to the fifth convolution module , the output channel is 256 feature maps , the feature map and feature map Through the third context fusion module Fusion splicing is performed to output a feature map with 768 channels , feature map After the seventh feature extraction layer , the output channel number is 768 feature map , feature map After the third downsampling layer , the output channel number is 512 feature map , the feature map and feature map Through the fourth context fusion module Fusion and splicing are performed to output a feature map with 1024 channels. , feature map Through the first context-aware module , the output channel number is 1024 feature map .

7. The target detection fusion method for improving Yolov10 according to claim 6, characterized in that S3 The details are as follows: S3.4, the operation process in the output end; Feature Map After the v10Detect detection module performs target detection at three different spatial resolutions, the feature map The spatial resolution of each scale is processed independently and then input into a dedicated convolutional layer for further feature extraction. Using a convolution with a kernel size of 1×1, the feature map of each scale is mapped to the number of channels required for detection. , ,in and Represents the coordinates of the target box, and Represents the width and height of each scale feature map, Indicates confidence, Represents the number of categories. Before prediction, the features of each scale are fused in pairs. After each scale feature map is processed, two branches are formed, namely the classification branch and the regression branch. The classification branch represents the category prediction of the target and outputs the confidence of the category. The regression branch represents the regression of the coordinates of the target box, predicts the position and size of the target box, and outputs each scale feature map. The output resolution of the three scales is , and , and Represents the feature maps The height and width of the grid are calculated, and finally the feature maps of the three scales are merged to generate a complete detection result. The output of each feature map includes the target box position, confidence, and category probability.

8. The target detection fusion method for improving Yolov10 according to claim 6, characterized in that CF The M module workflow is as follows: (1) The CFM module receives two feature maps, which are low-resolution and high-semantic feature maps. and high-resolution, low-semantic feature maps , , , represents the batch size, , and Represents the feature maps The number of channels, height and width, , and Represents the feature maps The number of channels, height and width are then adjusted by the channel convolution module Adjusting feature maps The number of channels is the same as the feature map The consistency of the channel alignment feature map and feature map The concatenation is performed along the channel dimension through the feature concatenation module Concat. , get the splicing features , before performing the splicing operation, determine whether there is and ,like or , then the feature map is interpolated or downsampled Adjust the resolution; (2) Then the splicing features Input to the CBAM module, and the concatenated features are processed by the channel attention module of the CBAM module. Perform a global pooling operation in the spatial dimension. The global pooling operation in the spatial dimension includes the global average pooling operation GAP and the global maximum pooling operation GMP. The calculation formula is as follows: , , in, and Respectively represent the splicing features The height and width of Indicates height The index of Indicates width The index of Indicates the maximum height and width. Represents splicing features The result of the global average pooling operation GAP is: Represents splicing features The result of performing the global maximum pooling operation GMP; Then input the results of GAP and GMP into the shared two-layer fully connected layer MLP to generate channel attention weights , the calculation formula is as follows: , in, represents the Sigmoid activation function used for normalization, represents a fully connected layer; Then the concatenated features are processed by the spatial attention module of the CBAM module. Perform the maximum pooling operation on the channel dimension and average pooling operation , the calculation formula is as follows: , , in, Represents splicing features The number of channels, Indicates the number of channels The index of Represents splicing features Perform the maximum pooling operation As a result, Represents splicing features Perform average pooling operation Result; Then The results and The results are concatenated, and after two-dimensional convolution and normalization operations, the spatial attention weight is obtained. , the calculation formula is as follows: , According to the spatial attention weight and channel attention weights Calculate the enhanced attention feature map of the final output of the CBAM module , the calculation formula is as follows: ; (3) Enhanced attention feature map Input to the weighted feature reorganization module for weighted feature reorganization. The weighted feature reorganization module is based on the feature map. and feature map The weights will enhance the attention feature map Split into feature maps and feature map , and then the feature map With feature map Add item by item to get the feature map , the feature map With feature map Add item by item to get the feature map ; (4) Feature map and feature map Input to the output feature generation module, concatenate the two features, and obtain the feature map of the final output of the CFM module , the calculation formula is as follows: , in, Represents a concatenation operation.

Citation Information

Patent Citations

  • Strawberry fruit identification method based on improved YOLOv5s

    CN118397427A

  • Lightweight parking detection method based on multi-scale attention mechanism

    CN119314141A

Cited By

  • Cover plate appearance defect automatic detection system and detection method

    CN120976132A

  • A warehouse damage detection method embedding MDA and WBFF mechanisms

    CN122675806A