Target detection method and system based on multi-scale frequency feature fusion

By introducing multi-scale frequency feature fusion and self-integrated attention mechanisms in the YOLO model, the boundary information loss and inconsistency within categories of the YOLO model in complex scenarios are solved, and the accuracy of target detection and the detection effect of occluding targets are improved.

CN120495812APending Publication Date: 2025-08-15FOSHAN UNIVERSITY +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510482346.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In complex scenarios, such as multi-object stacking, target occlusion, and blurred boundary problems, there are problems such as boundary information loss and inconsistent characteristics within categories, resulting in a decrease in target detection accuracy.

Method used

The object detection method based on multi-scale frequency feature fusion is adopted. By introducing the frequency feature fusion module and the self-integrated attention mechanism module, a multi-scale frequency feature pyramid structure is constructed, combining adaptive low-pass high-pass filtering and upsampling operations to enhance attention to the occlusion target, and smooth feature fluctuations within the category through the self-integrated attention mechanism.

Benefits of technology

In complex scenarios, rich detailed information is retained, the accuracy of target detection and detection ability of occluded targets are improved, and the consistency of features within the category is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495812A_ABST
    Figure CN120495812A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection method and system based on multi-scale frequency feature fusion, and the method comprises the steps: obtaining a to-be-processed target image, carrying out the image data preprocessing, and obtaining a sample image data set; based on a YOLOv11 network model, introducing a frequency feature fusion module and a self-integrated attention mechanism module, and constructing a target detection model based on multi-scale frequency feature fusion; and based on a multi-scale frequency feature fusion target detection model, carrying out target detection identification processing on the sample image data set to obtain a sample image target identification result. According to the method, richer detail information can be reserved in the frequency feature fusion process of different scales, the attention of the model to the shielded target is enhanced, and the target detection precision is improved. The target detection method and system based on multi-scale frequency feature fusion can be widely applied to the technical field of target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection technology, and in particular to a target detection method and system based on multi-scale frequency feature fusion. Background Art

[0002] Object detection is a key task in computer vision, widely used in fields such as autonomous driving, security monitoring, and industrial inspection. The YOLO (You Only Look Once) family of object detection algorithms has garnered widespread attention due to its speed and high accuracy. However, existing YOLO models still have limitations in complex scenarios, such as those involving multiple objects, occlusion, and blurred boundaries.

[0003] Traditional convolutional neural networks (CNNs) typically use multiple layers of downsampling to gradually reduce the size of feature maps. While this helps extract high-level, abstract features, it also leads to a loss of boundary information, which is crucial for accurate object localization in complex object detection tasks. To address this issue, existing methods typically employ feature fusion layers, combining deep, high-level abstract features with lower-level, high-resolution features to form a multi-scale feature pyramid structure. Well-known methods include the Bidirectional Feature Pyramid (BiFPN), Adaptive Spatial Feature Fusion (ASFF), and Asymptotic Feature Pyramid (AFPN). However, in mainstream standard feature fusion processes, high-level abstract features are often upsampled using nearest neighbor or bilinear interpolation before being added or concatenated with high-resolution features. Furthermore, during feature extraction, feature values within an object can vary or change rapidly. This high-frequency interference can lead to low feature similarity within a category, resulting in inconsistent features within the category. Using bilinear upsampling can upsample a single inconsistent feature to multiple pixels, exacerbating this problem. Furthermore, simple interpolation often oversmoothes features, resulting in boundary displacement. Summary of the Invention

[0004] In order to solve the above technical problems, the purpose of the present invention is to provide a target detection method and system based on multi-scale frequency feature fusion, which can retain richer detail information and enhance the model's attention to occluded targets in the process of fusion of frequency features at different scales, thereby improving the accuracy of target detection.

[0005] The first technical solution adopted by the present invention is: a target detection method based on multi-scale frequency feature fusion, comprising the following steps:

[0006] Acquire the target image to be processed and perform image data preprocessing to obtain a sample image data set;

[0007] Based on the YOLOv11 network model, the frequency feature fusion module and the self-integrated attention mechanism module are introduced to build an object detection model based on multi-scale frequency feature fusion;

[0008] Based on the target detection model of multi-scale frequency feature fusion, target detection and recognition processing is performed on the sample image dataset to obtain the target recognition results of the sample image.

[0009] Furthermore, the target detection model based on multi-scale frequency feature fusion specifically includes a backbone network module, a multi-scale frequency feature pyramid module, a self-integrated attention mechanism module and a head network module, wherein the backbone network module, the multi-scale frequency feature pyramid module, the self-integrated attention mechanism module and the head network module are connected in sequence, wherein:

[0010] The backbone network module includes a first convolutional layer, a second convolutional layer, a first C3K2 feature extraction module, a third convolutional layer, a second C3K2 feature extraction module, a fourth convolutional layer, a third C3K2 feature extraction module, a fifth convolutional layer, a fourth C3K2 feature extraction module, an SPPF pooling pyramid module and a C2PSA attention module;

[0011] The multi-scale frequency feature pyramid module includes a sixth convolutional layer, a seventh convolutional layer, an eighth convolutional layer, a ninth convolutional layer, a first frequency feature fusion module, a fifth C3K2 feature extraction module, a second frequency feature fusion module, a sixth C3K2 feature extraction module, a third frequency feature fusion module, a seventh C3K2 feature extraction module, a tenth convolutional layer, a first feature fusion and splicing module, an eighth C3K2 feature extraction module, a second feature fusion and splicing module, a ninth C3K2 feature extraction module, an eleventh convolutional layer, a third feature fusion and splicing module and a tenth C3K2 feature extraction module;

[0012] The self-integrated attention mechanism module includes a first self-integrated attention mechanism layer, a second self-integrated attention mechanism layer and a third self-integrated attention mechanism layer;

[0013] The head network module includes a first output layer, a second output layer and a third output layer.

[0014] Furthermore, the first frequency feature fusion module, the second frequency feature fusion module and the third frequency feature fusion module all have the same network structure, wherein the first frequency feature fusion module, the second frequency feature fusion module and the third frequency feature fusion module all include a first convolution kernel, a second convolution kernel, a first adaptive low-pass filter, a first adaptive high-pass filter, a first upsampling operation module, a second adaptive low-pass filter, a second adaptive high-pass filter and a second upsampling operation module, the output end of the first convolution kernel and the output end of the first adaptive low-pass filter are connected to the input end of the first upsampling operation module, the output end of the second convolution kernel is respectively connected to the input end of the first adaptive low-pass filter and the input end of the first adaptive high-pass filter, the output end of the first adaptive high-pass filter and the output end of the first upsampling operation module are both connected to the input end of the second adaptive low-pass filter and the input end of the second adaptive high-pass filter, and the output end of the second adaptive low-pass filter is connected to the input end of the second upsampling operation module.

[0015] Furthermore, the first self-integrated attention mechanism layer, the second self-integrated attention mechanism layer and the third self-integrated attention mechanism layer all have the same network structure, wherein the first self-integrated attention mechanism layer, the second self-integrated attention mechanism layer and the third self-integrated attention mechanism layer all include an input layer, a first channel spatial mixing module, a second channel spatial mixing module, a third channel spatial mixing module, a fully connected layer and an output layer, the output end of the input layer is respectively connected to the input end of the first channel spatial mixing module, the input end of the second channel spatial mixing module and the input end of the third channel spatial mixing module, the output end of the first channel spatial mixing module, the output end of the second channel spatial mixing module and the output end of the third channel spatial mixing module are all connected to the input end of the fully connected layer, and the output end of the fully connected layer is connected to the input end of the output layer.

[0016] Furthermore, the loss function of the target detection model based on multi-scale frequency feature fusion includes a target classification loss function, a bounding box regression loss function, and a distribution loss function. The expression of the target classification loss function is specifically as follows:

[0017]

[0018] In the above formula, Classification Loss represents the target classification loss function, S represents the grid size, Indicates whether the i-th network unit contains the target, p i (c) represents the probability that the target in the i-th grid cell belongs to category c, is the true label, indicating whether the target in the i-th grid cell belongs to category c;

[0019] The expression of the bounding box regression loss function is as follows:

[0020]

[0021] In the above formula, L CIOU represents the bounding box regression loss function, IOU represents the intersection-over-union ratio, α represents the weight function, b, b gt Represent the center points of the predicted box and the real box respectively, ρ represents the Euclidean distance between the two center points, c represents the diagonal distance of the minimum closure area containing both the predicted box and the real box, and v represents the similarity used to measure the aspect ratio;

[0022] The expression of the distribution loss function is specifically as follows:

[0023]

[0024] In the above formula, DFL represents the distribution loss function, N represents the number of samples, C represents the number of categories, and y ic represents the true label of the i-th sample, p ic represents the predicted probability that the i-th sample belongs to category c, α represents the balancing factor, and τ represents the focusing parameter.

[0025] Furthermore, the target detection model based on multi-scale frequency feature fusion performs target detection and recognition processing on the sample image dataset to obtain the sample image target recognition result, which specifically includes:

[0026] Input the sample image dataset into the target detection model with multi-scale frequency feature fusion;

[0027] The backbone network module of the target detection model based on multi-scale frequency feature fusion performs feature extraction processing on the sample image dataset to obtain a sample image feature dataset;

[0028] The multi-scale frequency feature pyramid module of the target detection model based on multi-scale frequency feature fusion performs multi-scale feature fusion processing on the sample image feature dataset to obtain the fused sample image feature dataset;

[0029] A self-integrated attention mechanism module of the target detection model based on multi-scale frequency feature fusion performs dimension adjustment on the fused sample image feature dataset to obtain an adjusted sample image feature dataset;

[0030] The head network module of the target detection model based on multi-scale frequency feature fusion performs target detection and recognition processing on the adjusted sample image feature dataset to obtain the sample image target recognition result.

[0031] Furthermore, the multi-scale frequency feature pyramid module of the target detection model based on multi-scale frequency feature fusion performs multi-scale feature fusion processing on the sample image feature dataset to obtain a fused sample image feature dataset, which specifically includes:

[0032] Inputting the sample image feature dataset into the multi-scale frequency feature pyramid module of the target detection model with multi-scale frequency feature fusion;

[0033] Based on the convolution layer of the multi-scale frequency feature pyramid module, the sample image feature dataset is convolved to obtain the convolved sample image feature dataset;

[0034] The frequency feature fusion module based on the multi-scale frequency feature pyramid module performs frequency feature fusion on the convolved sample image feature dataset to obtain a sample image frequency feature fusion dataset;

[0035] The C3K2 feature extraction module based on the multi-scale frequency feature pyramid module performs deep feature extraction processing on the sample image frequency feature fusion dataset to obtain the sample image deep feature dataset;

[0036] Based on the feature fusion and splicing module of the multi-scale frequency feature pyramid module, the deep feature dataset of the sample image is fused and spliced to obtain the fused sample image feature dataset.

[0037] Furthermore, the frequency feature fusion module based on the multi-scale frequency feature pyramid module performs frequency feature fusion on the convolved sample image feature dataset to obtain the sample image frequency feature fusion dataset, which specifically includes:

[0038] The convolved sample image feature dataset is input into the frequency feature fusion module of the multi-scale frequency feature pyramid module, and the convolved sample image feature dataset is split into fusion features and backbone features;

[0039] Based on the first convolution kernel and the second convolution kernel of the frequency feature fusion module, the fusion features and the backbone features are adjusted respectively to obtain the initial fusion features and the initial backbone features;

[0040] Based on the first adaptive low-pass filter of the frequency feature fusion module, the initial backbone features are adaptively low-pass filtered to obtain the first low-pass filtered features;

[0041] Based on the first adaptive high-pass filter of the frequency feature fusion module, the initial backbone features are adaptively high-pass filtered to obtain the first high-pass filtered features;

[0042] Performing a convolution operation on the first low-pass filter feature and the initial fusion feature and inputting the convolution operation module into a first upsampling operation module for upsampling to obtain a first upsampling feature;

[0043] Perform convolution operation on the first high-pass filter feature and the initial backbone feature and perform feature addition calculation to obtain a first added feature;

[0044] Add the first added feature and the first up-sampled feature to obtain an initial fusion feature;

[0045] Based on the second adaptive low-pass filter and the second adaptive high-pass filter of the frequency feature fusion module, the initial fusion features are adaptively low-pass filtered and adaptive high-pass filtered respectively to obtain second low-pass filtered features and second high-pass filtered features;

[0046] Performing a convolution operation on the second low-pass filter feature and the initial fusion feature, and then inputting the result into a second upsampling operation module for upsampling to obtain a second upsampling feature;

[0047] Perform convolution operation on the second high-pass filter feature and the initial backbone feature and perform feature addition calculation to obtain the second added feature;

[0048] The second up-sampled feature and the second added feature are added to obtain a sample image frequency feature fusion dataset.

[0049] Furthermore, the self-integrated attention mechanism module of the target detection model based on multi-scale frequency feature fusion performs dimension adjustment processing on the fused sample image feature dataset to obtain an adjusted sample image feature dataset, which specifically includes:

[0050] The fused sample image feature dataset is input into the self-integrated attention mechanism module of the target detection model that fused multi-scale frequency features;

[0051] Based on the input layer of the self-integrated attention mechanism module, the fused sample image feature dataset is obtained;

[0052] Based on the first channel spatial mixing module, the second channel spatial mixing module and the third channel spatial mixing module of the self-integrated attention mechanism module, the channel dimension adjustment processing is performed on the fused sample image feature dataset respectively to obtain the first adjusted sample image feature dataset, the second adjusted sample image feature dataset and the third adjusted sample image feature dataset;

[0053] Adding feature maps of the first adjusted sample image feature dataset, the second adjusted sample image feature dataset, the third adjusted sample image feature dataset, and the fused sample image feature dataset to obtain an initial adjusted sample image feature dataset;

[0054] Based on the fully connected layer of the self-integrated attention mechanism module, feature mapping is performed on the initially adjusted sample image feature dataset to obtain a mapped sample image feature dataset;

[0055] Performing point-by-point exponential operation on the mapped sample image feature dataset and then multiplying the feature map with the fused sample image feature dataset to obtain an adjusted sample image feature dataset;

[0056] Based on the output layer of the self-integrated attention mechanism module, the adjusted sample image feature dataset is output.

[0057] The second technical solution adopted by the present invention is: a target detection system based on multi-scale frequency feature fusion, comprising:

[0058] The first module is used to obtain the target image to be processed and perform image data preprocessing to obtain a sample image data set;

[0059] The second module is used to build a target detection model based on multi-scale frequency feature fusion by introducing the frequency feature fusion module and the self-integrated attention mechanism module based on the YOLOv11 network model;

[0060] The third module is used to perform target detection and recognition processing on the sample image dataset based on the target detection model of multi-scale frequency feature fusion to obtain the sample image target recognition result.

[0061] The beneficial effects of the method and system of the present invention are as follows: the present invention obtains a sample image data set by acquiring a target image to be processed and performing image data preprocessing, and further introduces a frequency feature fusion module and a self-integrated attention mechanism module based on the YOLOv11 network model to construct a target detection model based on multi-scale frequency feature fusion. The frequency feature fusion module can retain richer detail information in the process of frequency feature fusion at different scales, and the self-integrated attention mechanism module can smooth the fluctuation of eigenvalues within a category while enhancing the contour edge details of the category and improving the feature consistency within the category. Finally, based on the target detection model fused with multi-scale frequency features, target detection and recognition processing is performed on the sample image data set to obtain a sample image target recognition result, thereby enhancing the model's attention to occluded targets and improving the detection effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 This is a flowchart of the steps of a target detection method based on multi-scale frequency feature fusion according to the present invention;

[0063] Figure 2 This is a structural block diagram of a target detection system based on multi-scale frequency feature fusion according to the present invention;

[0064] Figure 3 1 is a schematic diagram of the structure of a target detection model fused with multi-scale frequency features provided by a specific embodiment of the present invention;

[0065] Figure 4 It is a structural diagram of a frequency feature fusion module provided in a specific embodiment of the present invention;

[0066] Figure 5 2 is a schematic diagram of the structure of the self-integrated attention mechanism module provided by a specific embodiment of the present invention;

[0067] Figure 6 It is a flowchart of improved target detection provided by a specific embodiment of the present invention. DETAILED DESCRIPTION

[0068] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The step numbers in the following embodiments are provided for ease of description only and do not limit the order of the steps. The order of execution of the steps in the embodiments can be adaptively adjusted based on the understanding of those skilled in the art.

[0069] Reference Figure 1 The present invention provides a target detection method based on multi-scale frequency feature fusion, which includes the following steps:

[0070] S100, obtaining a target image to be processed and performing image data preprocessing to obtain a sample image data set;

[0071] In this embodiment, the PASCAL VOC2012 public dataset is used as the model experiment dataset, and the dataset is preprocessed. Image enhancement operations are performed through random rotation, vertical transformation, and horizontal transformation. Pixel value normalization is performed to obtain 20 categories of image sample data sets and label data sets, which are divided into training sets and test sets.

[0072] S200, based on the YOLOv11 network model, introduces the frequency feature fusion module and the self-integrated attention mechanism module to build a target detection model based on multi-scale frequency feature fusion;

[0073] Specifically, if Figure 3As shown, the target detection model based on multi-scale frequency feature fusion constructed in an embodiment of the present invention specifically includes a backbone network module, a multi-scale frequency feature pyramid module, a self-integrated attention mechanism module and a head network module, and the backbone network module, the multi-scale frequency feature pyramid module, the self-integrated attention mechanism module and the head network module are connected in sequence, wherein the backbone network module includes a first convolutional layer, a second convolutional layer, a first C3K2 feature extraction module, a third convolutional layer, a second C3K2 feature extraction module, a fourth convolutional layer, a third C3K2 feature extraction module, a fifth convolutional layer, a fourth C3K2 feature extraction module, an SPPF pooling pyramid module and a C2PSA attention module; the multi-scale frequency feature pyramid module includes a first convolutional layer, a second convolutional layer, a first C3K2 feature extraction module, a third convolutional layer, a second C3K2 feature extraction module, a fourth convolutional layer, a third C3K2 feature extraction module, a fifth convolutional layer, a fourth C3K2 feature extraction module, a SPPF pooling pyramid module and a C2PSA attention module; the multi-scale frequency feature pyramid module includes a first The sixth convolutional layer, the seventh convolutional layer, the eighth convolutional layer, the ninth convolutional layer, the first frequency feature fusion module, the fifth C3K2 feature extraction module, the second frequency feature fusion module, the sixth C3K2 feature extraction module, the third frequency feature fusion module, the seventh C3K2 feature extraction module, the tenth convolutional layer, the first feature fusion splicing module, the eighth C3K2 feature extraction module, the second feature fusion splicing module, the ninth C3K2 feature extraction module, the eleventh convolutional layer, the third feature fusion splicing module and the tenth C3K2 feature extraction module; the self-integrated attention mechanism module includes the first self-integrated attention mechanism layer, the second self-integrated attention mechanism layer and the third self-integrated attention mechanism layer; the head network module includes the first output layer, the second output layer and the third output layer.

[0074] Furthermore, it should be noted that if Figure 4 As shown, the first frequency feature fusion module, the second frequency feature fusion module and the third frequency feature fusion module all have the same network structure, wherein the first frequency feature fusion module, the second frequency feature fusion module and the third frequency feature fusion module all include a first convolution kernel, a second convolution kernel, a first adaptive low-pass filter, a first adaptive high-pass filter, a first upsampling operation module, a second adaptive low-pass filter, a second adaptive high-pass filter and a second upsampling operation module, the output end of the first convolution kernel and the output end of the first adaptive low-pass filter are connected to the input end of the first upsampling operation module, the output end of the second convolution kernel is respectively connected to the input end of the first adaptive low-pass filter and the input end of the first adaptive high-pass filter, the output end of the first adaptive high-pass filter and the output end of the first upsampling operation module are both connected to the input end of the second adaptive low-pass filter and the input end of the second adaptive high-pass filter, and the output end of the second adaptive low-pass filter is connected to the input end of the second upsampling operation module.

[0075] Furthermore, Figure 5As shown, the first self-integrated attention mechanism layer, the second self-integrated attention mechanism layer and the third self-integrated attention mechanism layer all have the same network structure, wherein the first self-integrated attention mechanism layer, the second self-integrated attention mechanism layer and the third self-integrated attention mechanism layer all include an input layer, a first channel spatial mixing module, a second channel spatial mixing module, a third channel spatial mixing module, a fully connected layer and an output layer, the output end of the input layer is respectively connected to the input end of the first channel spatial mixing module, the input end of the second channel spatial mixing module and the input end of the third channel spatial mixing module, the output end of the first channel spatial mixing module, the output end of the second channel spatial mixing module and the output end of the third channel spatial mixing module are all connected to the input end of the fully connected layer, and the output end of the fully connected layer is connected to the input end of the output layer.

[0076] In this embodiment, the original architecture of YOLOv11 includes a backbone network (Backbone) for feature extraction, a neck (Neck) for feature fusion, and a head (Head) for final prediction; the backbone network generates multi-scale feature maps by stacking convolutional layers and modules, in which the C3K2 module is introduced to replace the C2f module of the previous version, and two smaller convolution kernels are used to improve computational efficiency; in addition, YOLOv11 retains the SPPF module and adds a C2PSA module, which enhances the focus on important areas of the feature map through the spatial attention mechanism; the neck structure aggregates features of different resolutions and passes them to the head for prediction, and is responsible for outputting the object positioning and classification results through the Detect output layer.

[0077] In the improved network YOLO-FFEAM, which is a target detection model based on multi-scale frequency feature fusion constructed in an embodiment of the present invention, the original backbone network Backbone of YOLOv11 is retained, and improvements are made to the neck network and the head network. Furthermore, the neck network of feature fusion is improved as follows: all upsampling modules Upsample and concatenation fusion modules Concat from bottom to top of the neck network are replaced by frequency feature fusion modules (Frequency feature Fusion), which is named as frequency feature fusion module FreqFusion in the embodiment of the present invention; on the basis of the original neck network, the second layer of features in the backbone network is additionally introduced as a small target fusion layer to enhance the improved network's detection ability for objects of different sizes, and at the same time, the top-down feature fusion process is adjusted to satisfy the further fusion of bottom-up features. In the embodiment of the present invention, the improved neck layer is named as multi-scale frequency feature pyramid FreqFPN (FrequencyFeaturePyramidNetworks); the final predicted head network is improved as follows: before the features are input into the final predicted head network, the self-integrated attention mechanism SEAM is introduced to strengthen the detailed processing of spatial dimensions and channels, and improve the network's attention to and ability to capture the features of occluded objects.

[0078] Furthermore, the training parameters are set, the data set is used to train and verify the target detection model, and the training is optimized through the loss function, specifically using the target classification loss function, the bounding box regression loss function and the distribution loss function to supervise the training of the improved YOLO-FFEAM target detection model.

[0079] The classification regression loss is calculated using cross entropy loss, which is used to measure the difference between the model's predicted probability distribution and the true label. The expression of the target classification loss function is as follows:

[0080]

[0081] In the above formula, Classification Loss represents the target classification loss function, S represents the grid size, Indicates whether the i-th network unit contains the target, p i (c) represents the probability that the target in the i-th grid cell belongs to category c, is the true label, indicating whether the target in the i-th grid cell belongs to category c.

[0082] The bounding box evaluation index IOU (Intersection over Union) uses CIOU, and the evaluation principle is:

[0083]

[0084] In the above formula, α represents the weight function, b, b gt Denote the center points of the predicted box and the true box respectively, ρ denotes the Euclidean distance between the two center points, c denotes the diagonal distance of the minimum closure area containing both the predicted box and the true box, and υ denotes the similarity used to measure the aspect ratio, which is defined as:

[0085]

[0086] In summary, the expression of the bounding box regression loss function is as follows:

[0087]

[0088] In the above formula, L CIOU represents the bounding box regression loss function, IOU represents the intersection-over-union ratio, α represents the weight function, b, b gt Represent the center points of the predicted box and the real box respectively, ρ represents the Euclidean distance between the two center points, c represents the diagonal distance of the minimum closure area that contains both the predicted box and the real box, and v represents the similarity used to measure the aspect ratio.

[0089] In order to solve the problem of class imbalance in target detection and improve the performance of the model when dealing with small targets and difficult samples, a distribution loss function is introduced. The expression of the distribution loss function is as follows:

[0090]

[0091] In the above formula, DFL represents the distribution loss function, N represents the number of samples, C represents the number of categories, and y ic represents the true label of the i-th sample, p ic represents the predicted probability that the i-th sample belongs to category c, α represents the balance factor, which is used to adjust the weight between positive and negative samples, and τ represents the focus parameter, which is used to control the degree of attention to difficult samples.

[0092] S300, performing target detection and recognition processing on the sample image data set based on the target detection model of multi-scale frequency feature fusion to obtain the target recognition result of the sample image.

[0093] S310, inputting the sample image dataset into a target detection model fused with multi-scale frequency features;

[0094] S320, a backbone network module of a target detection model based on multi-scale frequency feature fusion performs feature extraction processing on the sample image dataset to obtain a sample image feature dataset;

[0095] S330, a multi-scale frequency feature pyramid module of a target detection model based on multi-scale frequency feature fusion performs multi-scale feature fusion processing on the sample image feature dataset to obtain a fused sample image feature dataset;

[0096] Specifically, the sample image feature dataset is input into the multi-scale frequency feature pyramid module of the target detection model of multi-scale frequency feature fusion; based on the convolution layer of the multi-scale frequency feature pyramid module, the sample image feature dataset is convolved to obtain the convolved sample image feature dataset; based on the frequency feature fusion module of the multi-scale frequency feature pyramid module, the convolved sample image feature dataset is subjected to frequency feature fusion to obtain the sample image frequency feature fusion dataset; based on the C3K2 feature extraction module of the multi-scale frequency feature pyramid module, the sample image frequency feature fusion dataset is subjected to deep feature extraction to obtain the sample image deep feature dataset; based on the feature fusion and splicing module of the multi-scale frequency feature pyramid module, the sample image deep feature dataset is subjected to fusion and splicing to obtain the fused sample image feature dataset.

[0097] In this embodiment, first, the features of the second, fourth, sixth and tenth layers of the backbone network are extracted and input into the ninth, eighth, seventh and sixth convolutions respectively to unify the number of channels, and the backbone network feature X is obtained. 2 、X 4 、X 6 、X 10 , first change X 10 and X 6 Input to the FreqFusion first feature fusion module to obtain the first frequency fusion feature Y 6 , then Y 6 Input to the fifth C3K2 feature extraction module to obtain feature Y 5 , the backbone feature X 4 and feature Y 5 Input to the second feature fusion module of FreqFusion to obtain the second frequency fusion feature Y 4 , the fusion feature Y 4 Input to the sixth C3K2 feature extraction module to obtain feature Y 3 , the feature Y 3 and backbone features X 2 Input to the third feature fusion module of FreqFusion to obtain the third frequency fusion feature Y 2 , completing the bottom-up frequency feature fusion process. Then from the fusion feature Y 2Starting from top to bottom, the C3K2 feature extraction module is combined with the bottom-up fusion feature input of each level Concat splicing fusion module to complete the further top-down fusion of features, and finally realize the multi-scale frequency feature pyramid FreqFPN feature fusion process.

[0098] In the top-down feature fusion process, the features of the eighth C3K2 feature extraction module, the ninth C3K2 feature extraction module, and the tenth C3K2 feature extraction module are extracted and combined with the SEAM self-integrated attention mechanism to enhance the model's ability to capture objects of different sizes in stacked and densely occluded scenes.

[0099] Furthermore, it should be noted that, for the frequency feature fusion module of the multi-scale frequency feature pyramid module, the convolved sample image feature data set is input into the frequency feature fusion module of the multi-scale frequency feature pyramid module, and the convolved sample image feature data set is split into fusion features and backbone features; based on the first convolution kernel and the second convolution kernel of the frequency feature fusion module, the fusion features and the backbone features are adjusted and processed respectively to obtain initial fusion features and initial backbone features; based on the first adaptive low-pass filter of the frequency feature fusion module, the initial backbone features are adaptively low-pass filtered to obtain first low-pass filtered features; based on the first adaptive high-pass filter of the frequency feature fusion module, the initial backbone features are adaptively high-pass filtered to obtain first high-pass filtered features; the first low-pass filtered features are convolved with the initial fusion features and input into the first upsampling operation The module performs upsampling to obtain a first upsampling feature; the first high-pass filter feature is convolved with the initial backbone feature and the feature is added to obtain a first added feature; the first added feature is added to the first upsampling feature to obtain an initial fusion feature; based on the second adaptive low-pass filter and the second adaptive high-pass filter of the frequency feature fusion module, the initial fusion feature is adaptively low-pass filtered and adaptively high-pass filtered respectively to obtain a second low-pass filter feature and a second high-pass filter feature; the second low-pass filter feature is convolved with the initial fusion feature and then input into the second upsampling operation module for upsampling to obtain a second upsampling feature; the second high-pass filter feature is convolved with the initial backbone feature and the feature is added to obtain a second added feature; the second upsampling feature is feature added to the second added feature to obtain a sample image frequency feature fusion data set.

[0100] In this embodiment, the neck network is a multi-scale frequency feature fusion network. By introducing the frequency feature fusion module FreqFusion, the neck network is redesigned to propose a multi-scale frequency pyramid FreqFPN fusion network structure. Furthermore, the image is first input into the backbone network for preliminary feature extraction. Then, the features of the second, fourth, sixth, and tenth layers of the backbone network are extracted and input into a 1x1 convolutional layer with an adjusted channel number. Starting from the bottom feature map from the tenth layer, the frequency features of the feature maps are fused upwards. The basic principle of frequency feature fusion is as follows:

[0101]

[0102] In the above formula, X l ∈R C×2H×2W Represents the lth feature generated by the skeleton, and the feature map size is C×2H×2W, Y l+1 ∈R C×H×W Represents the fusion feature of the lth layer, the feature map size is C×H×W, F LP represents the low-pass filter predicted by the adaptive low-pass filter generator, (u, v) represents the offset value of the feature coordinate at (i, j) predicted by the offset generator, and F HP They represent the high-pass filter predicted by the adaptive high-pass filter generator, F UP is the upsampling operation.

[0103] In order to effectively generate the low-pass filter F LP , offset value (u,v) and high-pass filter F HP , before frequency feature fusion, we need to compress X l and Y l+1 And fuse them into the adaptive high-pass and low-pass generators. This process is called initial fusion. The process of initial fusion can be expressed as:

[0104] Z l =F UP (Conv 1×1 (Y l+1 ))+Conv 1×1 (X l )

[0105] where Z l ∈R C / r×2H×2W represents the fused compressed features, and r is the channel reduction rate that reduces the computational cost of the generator.

[0106] Therefore, the adaptive low-pass filter generator will initially fuse the Z l As input, it predicts the spatially varying low-pass filter. The adaptive low-pass filter generator consists of a 3×3 convolutional layer and a softmax layer, expressed as:

[0107]

[0108] in represents the spatially variable filter weights, where Represents the kernel size of the low-pass filter. After downsampling by 3x3 convolution, Contains the pixel location of each feature map filter; Ω represents the size After normalizing the filter by kernel-wise softmax, the result is The middle one is a smoothing low-pass filter.

[0109] Next, the width and height of the generated low-pass filter are halved, and the generated low-pass filter is used as the weight to divide the l-th layer fusion feature channel into four groups for low-pass filtering, smoothing the eigenvalue fluctuations within the category, and then rearranged to form the upsampled feature. Prepare for upward integration. The specific process is as follows:

[0110]

[0111] Similarly, the adaptive high-pass filter generator converts the initial fusion Z l As input, it predicts the spatially variable high-pass filter. The adaptive high-pass filter generator consists of a 3×3 convolution layer and a softmax layer, expressed as:

[0112]

[0113] in Contains the initial kernel at each position (i, j); Indicates the kernel size of the high-pass filter; in order to ensure that the kernel is finally generated is high-pass, first obtain low-pass kernels with per-kernel softmax, then invert the kernels by subtracting them from the identity kernel E, when When , the weights of the identity kernel E are [[0,0,0],[0,1,0],[0,0,0]].

[0114] After applying the high-pass filter and adding the residual, the high-frequency enhancement result of the lth feature generated by the backbone network is obtained, which is expressed as:

[0115]

[0116] In the above formula, Represents the high-frequency enhancement result of the lth feature generated by the backbone network.

[0117] S340, a self-integrated attention mechanism module of the target detection model based on multi-scale frequency feature fusion, performing dimension adjustment processing on the fused sample image feature dataset to obtain an adjusted sample image feature dataset;

[0118] Specifically, the fused sample image feature dataset is input into the self-integrated attention mechanism module of the target detection model for multi-scale frequency feature fusion; based on the input layer of the self-integrated attention mechanism module, the fused sample image feature dataset is obtained; based on the first channel space mixing module, the second channel space mixing module and the third channel space mixing module of the self-integrated attention mechanism module, the channel dimension adjustment processing is performed on the fused sample image feature dataset respectively to obtain the first adjusted sample image feature dataset, the second adjusted sample image feature dataset and the third adjusted sample image feature dataset; the feature maps of the first adjusted sample image feature dataset, the second adjusted sample image feature dataset and the third adjusted sample image feature dataset and the fused sample image feature dataset are added to obtain the initial adjusted sample image feature dataset; based on the fully connected layer of the self-integrated attention mechanism module, the initial adjusted sample image feature dataset is feature mapped to obtain the mapped sample image feature dataset; the mapped sample image feature dataset is subjected to point-by-point exponential operation and then multiplied with the fused sample image feature dataset by feature map calculation to obtain the adjusted sample image feature dataset; based on the output layer of the self-integrated attention mechanism module, the adjusted sample image feature dataset is output.

[0119] In this embodiment, in the attention mechanism module SEAM, the input image first passes through the channel and spatial mixing module CSMM to adjust the channel and spatial dimensions of the image. The Patch Embedding layer can change according to the set size and number of channels. After processing by the BatchNorm layer, the importance of different channels is learned through depthwise separable convolution, important channels are retained, and irrelevant parameters are removed; the input before depthwise separable convolution and the input after channel separation are combined through 1x1 convolution. After the input image is processed by three CSMM modules of different scales, the main channels of the occluded images in different dimensions are highlighted while compensating for the loss of details between channels. Subsequently, Average Pooling is performed to remove redundant information, and then all information is fused through two fully connected layers to enhance the connection between channel information. Finally, the output of the SEAM module is multiplied with the original feature as attention.

[0120] S350, a head network module of a target detection model based on multi-scale frequency feature fusion performs target detection and recognition processing on the adjusted sample image feature data set to obtain a sample image target recognition result.

[0121] Finally, the embodiment of the present invention is experimentally verified on the PASCAL VOC 2012 dataset, setting the training round epoch = 300, the batch size batch size = 16, the network initial learning rate to 0.001, using the stochastic gradient descent (SGD) optimizer, and the input image size to be uniformly scaled to 640×640. Table 1 below shows the detection results of the improved model YOLO-FFEAM proposed in this embodiment and the original YOLOv11-S and YOLOv11-M on the VOC 2012 public dataset. As can be seen from Table 1, the improved algorithm has a significant improvement in detection performance compared to YOLOv11-S, and is slightly worse than YOLOv11-M, but the computational complexity and parameter count are only half of YOLOv11-M.

[0122] Table 1 Algorithm comparison experimental data table

[0123] Model mAP0.5 mAP0.5-0.95 Params(M) FLOPs(G) YOLOv11-S 70.5 54.7 9.43 21.6 YOLO-FFEAM 72.0 56.4 9.97 29.0 YOLOv11-M 73.1 57.3 20.07 68.3

[0124] In summary, if Figure 6 As shown, an embodiment of the present invention proposes an improved YOLOv11 complex target detection method based on multi-scale frequency feature fusion, which enhances information of different frequencies through multi-scale feature extraction, helps to restore lost high-frequency details, and at the same time smoothes feature fluctuations within image categories, improves the problem of feature inconsistency within image categories, improves the robustness of the model in detection under multi-object occlusion environments, and enhances the detailed processing of spatial dimensions and channels. Among them, the frequency feature pyramid structure FreqFPN retains richer detail information in the process of fusion of frequency features at different scales, introduces adaptive high-pass filtering and adaptive low-pass filtering, and smoothes the fluctuation of eigenvalues within the category while enhancing the contour edge details of the category and improving the feature consistency within the category; the introduction of the SEAM attention mechanism enhances the model's attention to occluded targets and improves the detection effect.

[0125] Reference Figure 2 , an object detection system based on multi-scale frequency feature fusion, comprising:

[0126] The first module 201 is used to obtain a target image to be processed and perform image data preprocessing to obtain a sample image data set;

[0127] The second module 202 is used to introduce a frequency feature fusion module and a self-integrated attention mechanism module based on the YOLOv11 network model to build an object detection model based on multi-scale frequency feature fusion;

[0128] The third module 203 is used to perform target detection and recognition processing on the sample image dataset based on the target detection model of multi-scale frequency feature fusion to obtain the sample image target recognition result.

[0129] The contents of the above method embodiments are all applicable to the present system embodiments. The functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0130] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A target detection method based on multi-scale frequency feature fusion, characterized in that: The following steps are involved: Acquire the target image to be processed and perform image data preprocessing to obtain a sample image data set; Based on the YOLOv11 network model, the frequency feature fusion module and the self-integrated attention mechanism module are introduced to build an object detection model based on multi-scale frequency feature fusion; Based on the target detection model of multi-scale frequency feature fusion, target detection and recognition processing is performed on the sample image dataset to obtain the target recognition results of the sample image.

2. The target detection method based on multi-scale frequency feature fusion according to claim 1, characterized in that: The target detection model based on multi-scale frequency feature fusion specifically includes a backbone network module, a multi-scale frequency feature pyramid module, a self-integrated attention mechanism module and a head network module. The backbone network module, the multi-scale frequency feature pyramid module, the self-integrated attention mechanism module and the head network module are connected in sequence, wherein: The backbone network module includes a first convolutional layer, a second convolutional layer, a first C3K2 feature extraction module, a third convolutional layer, a second C3K2 feature extraction module, a fourth convolutional layer, a third C3K2 feature extraction module, a fifth convolutional layer, a fourth C3K2 feature extraction module, an SPPF pooling pyramid module and a C2PSA attention module; The multi-scale frequency feature pyramid module includes a sixth convolutional layer, a seventh convolutional layer, an eighth convolutional layer, a ninth convolutional layer, a first frequency feature fusion module, a fifth C3K2 feature extraction module, a second frequency feature fusion module, a sixth C3K2 feature extraction module, a third frequency feature fusion module, a seventh C3K2 feature extraction module, a tenth convolutional layer, a first feature fusion and splicing module, an eighth C3K2 feature extraction module, a second feature fusion and splicing module, a ninth C3K2 feature extraction module, an eleventh convolutional layer, a third feature fusion and splicing module and a tenth C3K2 feature extraction module; The self-integrated attention mechanism module includes a first self-integrated attention mechanism layer, a second self-integrated attention mechanism layer and a third self-integrated attention mechanism layer; The head network module includes a first output layer, a second output layer and a third output layer.

3. The target detection method based on multi-scale frequency feature fusion according to claim 2, characterized in that: The first frequency feature fusion module, the second frequency feature fusion module and the third frequency feature fusion module all have the same network structure, wherein the first frequency feature fusion module, the second frequency feature fusion module and the third frequency feature fusion module all include a first convolution kernel, a second convolution kernel, a first adaptive low-pass filter, a first adaptive high-pass filter, a first upsampling operation module, a second adaptive low-pass filter, a second adaptive high-pass filter and a second upsampling operation module, the output end of the first convolution kernel and the output end of the first adaptive low-pass filter are connected to the input end of the first upsampling operation module, the output end of the second convolution kernel is respectively connected to the input end of the first adaptive low-pass filter and the input end of the first adaptive high-pass filter, the output end of the first adaptive high-pass filter and the output end of the first upsampling operation module are both connected to the input end of the second adaptive low-pass filter and the input end of the second adaptive high-pass filter, and the output end of the second adaptive low-pass filter is connected to the input end of the second upsampling operation module.

4. The target detection method based on multi-scale frequency feature fusion according to claim 3, characterized in that: The first self-integrated attention mechanism layer, the second self-integrated attention mechanism layer and the third self-integrated attention mechanism layer all have the same network structure, wherein the first self-integrated attention mechanism layer, the second self-integrated attention mechanism layer and the third self-integrated attention mechanism layer all include an input layer, a first channel spatial mixing module, a second channel spatial mixing module, a third channel spatial mixing module, a fully connected layer and an output layer, the output end of the input layer is respectively connected to the input end of the first channel spatial mixing module, the input end of the second channel spatial mixing module and the input end of the third channel spatial mixing module, the output end of the first channel spatial mixing module, the output end of the second channel spatial mixing module and the output end of the third channel spatial mixing module are all connected to the input end of the fully connected layer, and the output end of the fully connected layer is connected to the input end of the output layer.

5. The target detection method based on multi-scale frequency feature fusion according to claim 4, characterized in that: The loss function of the target detection model based on multi-scale frequency feature fusion includes a target classification loss function, a bounding box regression loss function, and a distribution loss function. The expression of the target classification loss function is specifically as follows: In the above formula, Classification Loss represents the target classification loss function, S represents the grid size, Indicates whether the i-th network unit contains the target, p i (c) represents the probability that the target in the i-th grid cell belongs to category c, is the true label, indicating whether the target in the i-th grid cell belongs to category c; The expression of the bounding box regression loss function is as follows: In the above formula, L CIOU represents the bounding box regression loss function, IOU represents the intersection-over-union ratio, α represents the weight function, b, b gt Represent the center points of the predicted box and the real box respectively, ρ represents the Euclidean distance between the two center points, c represents the diagonal distance of the minimum closure area containing both the predicted box and the real box, and v represents the similarity used to measure the aspect ratio; The expression of the distribution loss function is specifically as follows: In the above formula, DFL represents the distribution loss function, N represents the number of samples, C represents the number of categories, and y ic represents the true label of the i-th sample, p ic represents the predicted probability that the i-th sample belongs to category c, α represents the balancing factor, and τ represents the focusing parameter.

6. The target detection method based on multi-scale frequency feature fusion according to claim 5, characterized in that: The target detection model based on multi-scale frequency feature fusion performs target detection and recognition processing on the sample image dataset to obtain the target recognition result of the sample image, which specifically includes: Input the sample image dataset into the target detection model with multi-scale frequency feature fusion; The backbone network module of the target detection model based on multi-scale frequency feature fusion performs feature extraction processing on the sample image dataset to obtain a sample image feature dataset; The multi-scale frequency feature pyramid module of the target detection model based on multi-scale frequency feature fusion performs multi-scale feature fusion processing on the sample image feature dataset to obtain the fused sample image feature dataset; A self-integrated attention mechanism module of the target detection model based on multi-scale frequency feature fusion performs dimension adjustment on the fused sample image feature dataset to obtain an adjusted sample image feature dataset; The head network module of the target detection model based on multi-scale frequency feature fusion performs target detection and recognition processing on the adjusted sample image feature dataset to obtain the sample image target recognition result.

7. The target detection method based on multi-scale frequency feature fusion according to claim 6, characterized in that: The multi-scale frequency feature pyramid module of the target detection model based on multi-scale frequency feature fusion performs multi-scale feature fusion processing on the sample image feature dataset to obtain a fused sample image feature dataset, which specifically includes: Inputting the sample image feature dataset into the multi-scale frequency feature pyramid module of the target detection model with multi-scale frequency feature fusion; Based on the convolution layer of the multi-scale frequency feature pyramid module, the sample image feature dataset is convolved to obtain the convolved sample image feature dataset; The frequency feature fusion module based on the multi-scale frequency feature pyramid module performs frequency feature fusion on the convolved sample image feature dataset to obtain a sample image frequency feature fusion dataset; The C3K2 feature extraction module based on the multi-scale frequency feature pyramid module performs deep feature extraction processing on the sample image frequency feature fusion dataset to obtain the sample image deep feature dataset; Based on the feature fusion and splicing module of the multi-scale frequency feature pyramid module, the deep feature dataset of the sample image is fused and spliced to obtain the fused sample image feature dataset.

8. The target detection method based on multi-scale frequency feature fusion according to claim 7, characterized in that: The frequency feature fusion module based on the multi-scale frequency feature pyramid module performs frequency feature fusion on the convolved sample image feature dataset to obtain the sample image frequency feature fusion dataset. This step specifically includes: The convolved sample image feature dataset is input into the frequency feature fusion module of the multi-scale frequency feature pyramid module, and the convolved sample image feature dataset is split into fusion features and backbone features; Based on the first convolution kernel and the second convolution kernel of the frequency feature fusion module, the fusion features and the backbone features are adjusted respectively to obtain the initial fusion features and the initial backbone features; Based on the first adaptive low-pass filter of the frequency feature fusion module, the initial backbone features are adaptively low-pass filtered to obtain the first low-pass filtered features; Based on the first adaptive high-pass filter of the frequency feature fusion module, the initial backbone features are adaptively high-pass filtered to obtain the first high-pass filtered features; Performing a convolution operation on the first low-pass filter feature and the initial fusion feature and inputting the convolution operation module into a first upsampling operation module for upsampling to obtain a first upsampling feature; Perform convolution operation on the first high-pass filter feature and the initial backbone feature and perform feature addition calculation to obtain a first added feature; Add the first added feature and the first up-sampled feature to obtain an initial fusion feature; Based on the second adaptive low-pass filter and the second adaptive high-pass filter of the frequency feature fusion module, the initial fusion features are adaptively low-pass filtered and adaptive high-pass filtered respectively to obtain second low-pass filtered features and second high-pass filtered features; Performing a convolution operation on the second low-pass filter feature and the initial fusion feature, and then inputting the result into a second upsampling operation module for upsampling to obtain a second upsampling feature; Perform convolution operation on the second high-pass filter feature and the initial backbone feature and perform feature addition calculation to obtain the second added feature; The second up-sampled feature and the second added feature are added to obtain a sample image frequency feature fusion dataset.

9. The target detection method based on multi-scale frequency feature fusion according to claim 8, characterized in that: The self-integrated attention mechanism module of the target detection model based on multi-scale frequency feature fusion performs dimension adjustment processing on the fused sample image feature dataset to obtain the adjusted sample image feature dataset, which specifically includes: The fused sample image feature dataset is input into the self-integrated attention mechanism module of the target detection model that fused multi-scale frequency features; Based on the input layer of the self-integrated attention mechanism module, the fused sample image feature dataset is obtained; Based on the first channel spatial mixing module, the second channel spatial mixing module and the third channel spatial mixing module of the self-integrated attention mechanism module, the channel dimension adjustment processing is performed on the fused sample image feature dataset respectively to obtain the first adjusted sample image feature dataset, the second adjusted sample image feature dataset and the third adjusted sample image feature dataset; Adding feature maps of the first adjusted sample image feature dataset, the second adjusted sample image feature dataset, the third adjusted sample image feature dataset, and the fused sample image feature dataset to obtain an initial adjusted sample image feature dataset; Based on the fully connected layer of the self-integrated attention mechanism module, feature mapping is performed on the initially adjusted sample image feature dataset to obtain a mapped sample image feature dataset; Performing point-by-point exponential operation on the mapped sample image feature dataset and then multiplying the feature map with the fused sample image feature dataset to obtain an adjusted sample image feature dataset; Based on the output layer of the self-integrated attention mechanism module, the adjusted sample image feature dataset is output.

10. A target detection system based on multi-scale frequency feature fusion, characterized in that: Includes the following modules: The first module is used to obtain the target image to be processed and perform image data preprocessing to obtain a sample image data set; The second module is used to build a target detection model based on multi-scale frequency feature fusion by introducing the frequency feature fusion module and the self-integrated attention mechanism module based on the YOLOv11 network model; The third module is used to perform target detection and recognition processing on the sample image dataset based on the target detection model of multi-scale frequency feature fusion to obtain the sample image target recognition result.

Citation Information

Cited By

  • Unmanned aerial vehicle pilot signal identification method and device and computer equipment

    CN122594993A