An object detection method based on partial convolution embedding and aggregation distribution mechanism
By introducing partial convolutional embedding and aggregation distribution mechanisms into YOLOv8's object detection model, the problem of backbone network feature extraction delay and insufficient fusion of context information is solved, and more efficient object detection and stronger small object detection capabilities are achieved.
Patent Information
- Application Number
- CN202311537095.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-11-17
AI Technical Summary
The single-stage object detection technology represented by YOLOv8 has a long delay and inference time for feature extraction in its backbone backbone network, and it fails to efficiently fuse context information in the neck layer, resulting in low detection efficiency.
The object detection method based on partial convolution embedding and aggregation distribution mechanism is adopted to improve feature extraction capabilities by embedding partial convolution in the spatial feature extraction module, and the aggregation distribution mechanism is used in the multi-scale information fusion module to enhance multi-scale information fusion capabilities.
The detection efficiency and small object detection capabilities of the object detection model are improved, the overall accuracy is improved, and the calculation complexity and reasoning time are reduced.
Smart Images

Figure CN117671414B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and relates to a target detection method based on partial convolution embedding and aggregation distribution mechanism. Background Art
[0002] In recent years, with the continuous deepening of deep learning related theories and the large-scale improvement of computer computing power, target detection technology based on deep learning has gradually matured. Target detection aims to find the category and location of specified targets in images, and has been widely used in various fields, such as autonomous driving, remote sensing images, video surveillance, and medical testing. YOLO (You Only Look Once) is a classic one-stage target detection algorithm. Its advantages lie in high real-time performance, simplicity and efficiency, multi-scale detection, global context information utilization, and multi-task learning. These features make it perform well in fast target detection and real-time application scenarios.
[0003] After continuous version updates, from YOLOv1 to YOLOv8, it has become a typical representative of single-stage target detection methods. However, at present, the single-stage target detection technology represented by YOLOv8 has a long delay and reasoning time for feature extraction in the backbone network, and fails to efficiently integrate context information in the neck layer, resulting in low detection efficiency, which needs to be improved. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a target detection method based on partial convolution embedding and clustered distribution mechanism to solve the technical problems of single-stage target detection technology represented by YOLOv8, whose backbone network has a long delay and reasoning time for feature extraction, and fails to efficiently fuse context information at the neck layer, resulting in low detection efficiency.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] A target detection method based on partial convolution embedding and aggregate distribution mechanism, the method comprising the following steps:
[0007] S1: Obtain a public target detection dataset, convert the annotation files in the public target detection dataset into the YOLO format, and divide the public target dataset converted from the annotation files into the YOLO format into a training set, a validation set, and a test set in proportion;
[0008] S2: Normalize the images in the training set and synthesize the normalized images and their corresponding target labels into batches;
[0009] S3: Use YOLOv8 as the basic framework of the target detection network, and build a target detection network model based on partial convolution embedding and aggregation distribution mechanism. The target detection model includes a spatial feature extraction module and a multi-scale information fusion module. The spatial feature extraction module embeds partial convolution into the backbone network to increase the extraction capability of spatial feature information. The multi-scale information fusion module uses the aggregation distribution mechanism to increase the multi-scale information fusion capability of the model.
[0010] S4: input the batch synthesized in S2 into the target detection network model based on partial convolution embedding and cluster distribution mechanism built in S3 for training, and obtain the trained model weights;
[0011] S5: Use the training model weights obtained by S4 to test it in the test set to obtain the detection results.
[0012] Furthermore, in S1, the public target detection dataset is converted into a YOLO format using a Python code for converting the data format and splitting the dataset, and the dataset annotation file is divided into a training set, a validation set, and a test set in a ratio of 7:1:2.
[0013] Furthermore, in S3, the spatial feature extraction module embeds part of the convolution into the backbone network to increase the extraction capability of spatial feature information, specifically including:
[0014] Build a spatial feature extraction module, replace the C2f module corresponding to the backbone of YOLOv8 with a partially convolutional Fasternet module, where Fasternet includes a basic network submodule, a fast feature fusion submodule, and an efficient upsampling submodule, and each submodule is connected by embedding partial convolutions;
[0015] The basic network submodule includes a convolution layer, a batch normalization layer and an activation function layer, which are used to perform feature extraction and nonlinear activation on the image;
[0016] The fast feature fusion submodule is responsible for fusing features from different levels;
[0017] The efficient upsampling module is used to implement upsampling of feature maps;
[0018] Partial convolution uses the redundancy of feature maps and systematically applies conventional convolution to some input channels. Partial convolution has a lower total floating-point operation number FLOPs and a higher floating-point operation number FLOPS than general network structures. It extracts spatial features by reducing redundant calculations and memory accesses at the same time. The FLOPs calculation formula of partial convolution is:
[0019]
[0020] Where h is the height of the feature map, w is the width of the feature map, k is the size of the convolution kernel, and c is p is the number of channels for conventional convolution; The FLOPs of partial convolution is only 1 / 2 of that of regular convolution. The memory access of partial convolution is expressed as:
[0021]
[0022] Where h is the height of the feature map, w is the width of the feature map, k is the size of the convolution kernel, and c is p is the number of channels for regular convolution, and the number of memory accesses for partial convolution is the same as that for regular convolution. The rest (cc p ) channels do not participate in the calculation, and some convolutions do not need to access memory;
[0023] Each Fasternet module has a partial convolution layer, followed by two 1*1 dimensional convolutions. The three form an inverted residual architecture, which increases the number of channels in the middle layer, places a shortcut connection to reuse the input features, and places the normalization layer and activation layer after the middle layer.
[0024] Furthermore, in S3, the multi-scale information fusion module uses an aggregation distribution mechanism to increase the multi-scale information fusion capability of the model, specifically including:
[0025] Build a multi-scale information fusion module and introduce the goldyolo detection head at the detection head of the YOLOv8 model. The goldyolo detection head includes a feature alignment module FAM, an information fusion module IFM and an information injection module Inject. The feature alignment module FAM, the information fusion module IFM and the information injection module Inject constitute an aggregation distribution mechanism.
[0026] The feature alignment module FAM collects feature maps of different scales of the backbone part and aligns them by upsampling or downsampling;
[0027] The information fusion module IFM fuses the aligned features to generate global features, which are divided into two parts through a split operation, and then targeted distribution operations are performed on other scales;
[0028] The information injection module Inject uses the attention operation that enhances branch detection capabilities to split the global features and distribute them to each level;
[0029] Assuming the input image shape is N×3×H×W, there are four multi-scale features obtained from the backbone, namely B2, B3, B4, and B5, that is, Where M represents batch-size, Indicates the number of channels of feature maps of different scales, Represents the height and width of feature maps of different scales;
[0030] The feature alignment module FAM takes B4 as the benchmark, downsamples the large feature maps B2 and B3 by average pooling, and upsamples the small feature map B5 by bilinear interpolation, which is expressed as:
[0031] F align =FAM([B2, B3, B4, B5]) (3)
[0032] The combined feature representation obtained by concat is:
[0033]
[0034] The information fusion module IFM design includes Conv, RepBlock modules, and Split operations:
[0035] F fuse =RepBlock(F align ) (5)
[0036] F inj_P3 , F inj_P4 =Split(F fuse ) (6)
[0037] The aligned and concat features F align Input into the RepBlock module to get F fuse Fusion features, while using Conv to adjust the channel to adapt to the size of different models, F fuse Split into F on the channel via Split inj_P3 and F inj_P4 , and then perform the next step of feature fusion with different levels;
[0038] The information injection module Inject adopts the form of self-attention, and its input is the current scale x_local (Flocal) to be distributed, and the global feature x_global (F inj ), and finally the fusion information P is further obtained through ReBlock processing i , the calculation formula is as follows:
[0039]
[0040] P i =ReBlock(F global_pi ) (8).
[0041] Furthermore, the detection head of the YOLOv8 is a symmetrical structure, and a cluster distribution GD mechanism is added to the detection head. The GD mechanism enters the detection head through two different network paths respectively. A decoupled head structure is adopted in the detection head. Two parallel branches extract category features and position features respectively, and each branch uses a 1x1 convolution to complete its respective task.
[0042] The beneficial effects of the present invention are:
[0043] First, the target detection model of the present invention has a stronger ability to detect small targets while improving detection efficiency and has higher overall accuracy.
[0044] Second, the Fasternet in the present invention includes a basic network submodule, a fast feature fusion submodule and an efficient upsampling submodule; the basic network submodule includes a convolution layer, a batch normalization layer and an activation function layer, which are used to extract features and perform nonlinear activation on the image; the convolution layer is responsible for learning local features in the image, the batch normalization layer is used to accelerate the training process and enhance the robustness of the network, and the activation function layer introduces nonlinear factors to increase the expression ability of the network; the fast feature fusion submodule is responsible for fusing features from different levels, which can improve the feature expression ability while ensuring speed; the efficient upsampling module is used to realize upsampling of feature maps to achieve accurate positioning of the target position. Upsampling realizes accurate positioning of the target position by restoring high-resolution feature maps, which can improve positioning accuracy while ensuring speed.
[0045] Third, the partial convolution in the present invention utilizes the redundancy of feature maps and systematically applies conventional convolution on some input channels without affecting the remaining input channels. Therefore, the partial convolution has lower FLOPs (total floating-point operations) and higher FLOPS (floating-point operations per second) than the general network structure. By reducing redundant calculations and memory accesses at the same time, spatial features can be extracted more efficiently.
[0046] Fourth, each Fasternet module in the present invention has a partial convolution layer, followed by two 1*1 dimensional convolutions, and the three constitute an inverted residual architecture, so that the number of channels in the middle layer is larger, and a shortcut connection is placed to reuse the input features. At the same time, in order to maintain the diversity of features and reduce latency, the normalization layer and the activation layer are placed after the middle layer.
[0047] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:
[0049] Figure 1 It is a flow chart of the present invention;
[0050] Figure 2 This is the network structure diagram of the existing YOLOv8 model;
[0051] Figure 3 It is a structural diagram of the network backbone of the present invention;
[0052] Figure 4 Part of the convolution structure diagram and Fasternet module structure diagram; Figure 4 (a) is a partial convolution structure diagram; Figure 4 (b) is the Fasternet module structure diagram;
[0053] Figure 5 This is a schematic diagram of the aggregation and distribution mechanism of the present invention;
[0054] Figure 6 It is a connection schematic diagram of the detection head of the present invention;
[0055] Figure 7 A comparison chart of the actual detection differences between the improved YOLOv8 model and the original model; Figure 7 (a) Detection diagram of the original YOLOv8 model; Figure 7 (b) Model detection diagram of the improved YOLOv8 model;
[0056] Figure 8 A comparison chart of the actual detection differences between the improved YOLOv8 model and the original model; Figure 8 (a) Detection diagram of the original YOLOv8 model; Figure 8 (b) Model detection diagram of the improved YOLOv8 model;
[0057] Fig. 9 A comparison chart of the actual detection differences between the improved YOLOv8 model and the original model; Fig. 9 (a) Detection diagram of the original YOLOv8 model; Fig. 9 (b) Model detection diagram of the improved YOLOv8 model. DETAILED DESCRIPTION
[0058] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0059] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0060] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0061] See also Figures 1 to 9 , which is an object detection method based on partial convolution embedding and aggregation distribution mechanism.
[0062] Its flow chart is as follows Figure 1 As shown, it includes the following steps:
[0063] S1: Obtain a public target detection dataset, convert the annotation files in the public target detection dataset into the YOLO format, and divide the public target dataset converted from the annotation files into the YOLO format into a training set, a validation set, and a test set in proportion;
[0064] In this embodiment, the target detection image dataset selects the VOC2012 Train / Validation public dataset, and uses the self-made conversion format and dataset segmentation python code to convert the dataset annotation file into the YOLO format and divide it into training set, validation set and test set in a ratio of 7:1:2.
[0065] S2: normalize the images in the training set, and synthesize the normalized images and their corresponding target labels into batches;
[0066] S3: Use YOLOv8 as the basic framework of the target detection network, and build a target detection network model based on partial convolution embedding and aggregation distribution mechanism. The target detection model includes a spatial feature extraction module and a multi-scale information fusion module. The spatial feature extraction module embeds partial convolution into the backbone network to increase the extraction capability of spatial feature information. The multi-scale information fusion module uses the aggregation distribution mechanism to increase the multi-scale information fusion capability of the model.
[0067] The detailed steps to build the spatial feature extraction module are as follows:
[0068] like Figure 3 As shown, in the backbone network structure diagram of the present invention, the C2f module corresponding to the backbone of YOLOv8 is replaced by a partially convolutional Fasternet module, wherein the Fasternet includes a basic network submodule, a fast feature fusion submodule, and an efficient upsampling submodule, and each submodule is connected by embedding a partial convolution.
[0069] In the partially convolutional Fasternet module: the basic network submodule includes convolutional layer, batch normalization layer and activation function layer, which are used to extract features and perform nonlinear activation on the image; the convolutional layer is responsible for learning local features in the image, the batch normalization layer is used to accelerate the training process and enhance the robustness of the network, and the activation function layer introduces nonlinear factors to increase the network's expressive power; the fast feature fusion submodule is responsible for fusing features from different levels, which can improve feature expression capabilities while ensuring speed; the efficient upsampling module is used to implement upsampling of feature maps to achieve accurate positioning of the target position. Upsampling achieves accurate positioning of the target position by restoring high-resolution feature maps, which can improve positioning accuracy while ensuring speed.
[0070] The structure of partial convolution (PConv) is as follows Figure 4 As shown in (a), partial convolution uses the redundancy of feature maps to systematically apply regular convolution on some input channels without affecting the remaining input channels. Therefore, partial convolution has lower FLOPs (total floating-point operations) and higher FLOPS (floating-point operations per second) than general network structures. By reducing redundant calculations and memory accesses at the same time, spatial features can be extracted more efficiently. The FLOPs calculation formula for partial convolution is:
[0071]
[0072] Where h is the height of the feature map, w is the width of the feature map, k is the size of the convolution kernel, and c is p is the number of channels for conventional convolution, which is usually Therefore, the FLOPs of partial convolution is only 1 / 2 of that of regular convolution. The memory access of partial convolution can be expressed as:
[0073]
[0074] Where h is the height of the feature map, w is the width of the feature map, k is the size of the convolution kernel, and c is p is the number of channels for regular convolution. The number of memory accesses for partial convolution is only equal to that of regular convolution. The rest (cc p ) channels do not participate in the calculation, so some convolutions do not require memory access.
[0075] Fasternet module structure Figure 4 As shown in (b), each Fasternet module has a partial convolution layer followed by two 1*1 dimensional convolutions. The three form an inverted residual architecture, which increases the number of channels in the middle layer and places a shortcut connection to reuse input features. At the same time, in order to maintain feature diversity and reduce latency, the normalization layer and activation layer are placed after the middle layer.
[0076] The detailed steps for building a multi-scale information fusion module are as follows:
[0077] like Figure 5 As shown in the figure, the goldyolo detection head is introduced at the detection head of the model. The goldyolo detection head includes a feature alignment module (FAM), an information fusion module (IFM) and an information injection module (Inject) to form an aggregation and distribution mechanism. The feature alignment module (FAM) collects feature maps of different scales of the backbone part and aligns them by upsampling or downsampling; the information fusion module (IFM) fuses the aligned features to generate global features, which are then divided into two parts through a split slicing operation, and then targeted distribution operations are performed on other scales; finally, the information injection module (Inject) uses the attention operation that enhances the branch detection capability to split the global features and distribute them to each level.
[0078] Goldyolo detection head uses a clustering distribution mechanism to improve the multi-scale feature fusion capability of the model. The specific implementation process is as follows:
[0079] Assuming the input image shape is N×3×H×W, there are four multi-scale features obtained from the backbone, namely B2, B3, B4, and B5, that is, Where N represents the batch-size, Indicates the number of channels of feature maps of different scales, Represents the height and width of feature maps of different scales;
[0080] In the first step, the module (FAM) uses B4 as the benchmark, downsamples the large feature maps B2 and B3 by average pooling, and upsamples the small feature map B5 by bilinear interpolation:
[0081] F align =FAM([B2, B3, B4, B5]) (3)
[0082] Unify the size of the feature map, and then concat to get the merged features
[0083] The second step is to design the module (IFM) including Conv, RepBlock modules, and Split operations:
[0084] F fuse =RepBlock(F align ) (4)
[0085] F inj_P3 , F inj_P4 =Split(F fuse ) (5)
[0086] The aligned and concat features F align Input into the RepBlock module to get F fuse Fusion features, while using Conv to adjust the channel to adapt to the size of different models, F fuse Split into F on the channel via Split inj_P3 and F inj_P4 , and then perform the next step of feature fusion with different levels;
[0087] The third step is the module (Inject), which uses the self-attention form and takes as input the features x_local (Flocal) at the current scale to be distributed, such as B3 and B4, and the global features x_global (F inj ) Such as F inj_P3 , F inj_P4 Finally, the fusion information P is further obtained through ReBlock processing i , the formula is as follows:
[0088]
[0089] P i =ReBlock(Fglobal_pi ) (7)
[0090] The detection head of YOLOv8 is a symmetrical structure, and the detection head of the present invention adds a GD (aggregation distribution) mechanism on this basis, such as Figure 6 As shown in the figure, the GD mechanism enters the detection head through two different network paths. The detection head adopts the structure of a decoupled head. Two parallel branches extract category features and position features respectively, and then use a layer of 1*1 convolution to complete the classification and positioning tasks.
[0091] S4: Input the batch synthesized in S2 into the target detection network model based on partial convolution embedding and cluster distribution mechanism built in S3 for training to obtain corresponding training model weights.
[0092] S5 uses the training model weights obtained by training in S4 to test it in the test set to obtain the detection result.
[0093] Based on the training set training, the improved model is trained to obtain the optimal target detection model weight. The training parameter configuration table of this embodiment is shown in Table 1:
[0094] Table 1:
[0095] Optimizer SGD Batch-size 8 Epochs 300 Image output size (Inputsize) 640*640 Learning rate 0.01
[0096] In order to verify the effectiveness of the improved solution, this embodiment sets up an ablation experiment to explore the impact of the proposed improved method on the performance of the YOLOv8n model.
[0097] In terms of performance indicators, in the YOLO series of models, the indicators for evaluating its network performance are mainly the following: Precision (P), Recall (R), Mean Average Precision (mAP). In this implementation, mAP50 and mAP50:95 are used as performance reference indicators. mAP@0.5 and mAP@0.5:0.95 represent the mAP value when the IOU threshold is 0.5 and the average mAP value when the IOU starts from 50% and increases to 95% with a step size of 0.05. The larger the mean average precision mAP, the higher the overall accuracy of the model. The calculation formulas for each indicator are as follows:
[0098]
[0099]
[0100]
[0101]
[0102] Among them, TP is the number of correctly predicted positive samples, FN is the number of incorrectly predicted negative samples, and FP is the number of incorrectly predicted positive samples.
[0103] In addition, in general performance evaluation indicators, the number of model parameters and the amount of calculation must also be considered. Therefore, it is also necessary to introduce two model architecture detail parameters: FLOPs (Floating Point Operations) and parameter quantity (Params).
[0104] First, we used YOLOv8n as a baseline algorithm to conduct experiments on the VOC2012 Train / Validation public dataset. After the experiment, we found that YOLOv8n has good detection effects on medium and high-impact targets, but there is still room for improvement in small target detection. Therefore, in the subsequent experiments that introduced some convolution modules and clustering distribution mechanisms, we put it on the small target detection head for performance testing. The ablation experiment results for the comparison of target detection algorithms are shown in Table 2:
[0105] Table 2
[0106] Experimental content mAP@0.5% mAP@0.5: 0.95% Params / M FLOPs / G Epochs YOLOv8n 62.8 45.9 3.0 8.1 300 This embodiment 69.3 51.3 4.1 10.7 300
[0107] As can be seen from Table 2, the performance of this embodiment is improved to a certain extent compared with the baseline.
[0108] The detection effect of YOLOv8n on small target detection has certain shortcomings. After being improved by the method of the present invention, the model's small target detection effect has been effectively improved, and its small target performance comparison is shown in Table 3. The overall mAP@0.5% in small target detection has increased by about 6%, and each small target detection project has improved.
[0109] Table 3
[0110] Experimental content Dog Bird Sheep Cow Cat YOLOv8n 77.9 52.7 70.5 54.7 87.2 This embodiment 82.8 67.9 76.4 72.2 92.7
[0111] Figure 7 to Figure 9 The comparison chart of actual detection differences between the improved YOLOv8 model and the original model is shown. Figure 7 (a) and Figure 7 (b) The results of detecting the image numbered 001073 using the original model and the improved model respectively. Figure 7 (a) The original model mistakenly thinks the dense cow on the far left is a sheep, while Figure 7 In (b), the cow is correctly detected and the person in the middle of the picture can also be detected; Figure 8 In Figure 8 (a) and Figure 8(b) The image detected is the image numbered 002662 in the original dataset. It is not difficult to see that the improved model is Figure 8 In (b), the confidence score for the person class is increased to 0.86, and the chair class can be detected in complex scenes. Fig. 9 middle Fig. 9 (a) and 9(b) detect the image numbered 009005, and can also more clearly identify the bicycle class in a complex scene. It can be seen that compared with YOLOv8n, the improved model of this embodiment has better small target detection ability and higher overall accuracy.
[0112] Experimental results show that compared with YOLOv8n, the improved target detection model based on partial convolution embedding and clustered distribution mechanism has better accuracy performance on small targets, and also has overall performance improvement, and has excellent performance in improving algorithm performance.
[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.
Claims
1. A target detection method based on partial convolution embedding and cluster distribution mechanism, characterized by: The method comprises the following steps: S1: Obtain a public target detection dataset, convert the annotation files in the public target detection dataset into the YOLO format, and divide the public target dataset converted from the annotation files into the YOLO format into a training set, a validation set, and a test set in proportion; S2: Normalize the images in the training set and synthesize the normalized images and their corresponding target labels into batches; S3: Use YOLOv8 as the basic framework of the target detection network, and build a target detection network model based on partial convolution embedding and aggregation distribution mechanism. The target detection network model includes a spatial feature extraction module and a multi-scale information fusion module. The spatial feature extraction module embeds partial convolution into the backbone network to increase the extraction capability of spatial feature information. The multi-scale information fusion module uses the aggregation distribution mechanism to increase the multi-scale information fusion capability of the model. S4: input the batch synthesized in S2 into the target detection network model based on partial convolution embedding and cluster distribution mechanism built in S3 for training, and obtain the trained model weight; S5: Use the training model weights obtained by S4 to test it in the test set to obtain the test results; In S3, the spatial feature extraction module embeds part of the convolution into the backbone network to increase the extraction capability of spatial feature information, specifically including: Build a spatial feature extraction module, replace the C2f module corresponding to the backbone of YOLOv8 with a partially convolutional Fasternet module, where Fasternet includes a basic network submodule, a fast feature fusion submodule, and an efficient upsampling submodule, and each submodule is connected by embedding partial convolutions; The basic network submodule includes a convolution layer, a batch normalization layer and an activation function layer, which are used to extract features and perform nonlinear activation on the image; The fast feature fusion submodule is responsible for fusing features from different levels; The efficient upsampling submodule is used to implement upsampling of feature maps; Partial convolution uses the redundancy of feature maps to apply regular convolution on some input channels. The total floating-point operation number FLOPs of partial convolution is calculated as: Where h is the height of the feature map, w is the width of the feature map, k is the size of the convolution kernel, and c is p is the number of channels for conventional convolution; The FLOPs of partial convolution is only 1 / 2 of that of regular convolution. The memory access of partial convolution is expressed as: Where h is the height of the feature map, w is the width of the feature map, k is the size of the convolution kernel, and c is p is the number of channels for regular convolution, and the number of memory accesses for partial convolution is the same as that for regular convolution. The rest (cc p ) channels do not participate in the calculation, and some convolutions do not need to access memory; Each Fasternet module has a partial convolution layer followed by two 1*1 dimensional convolutions. The three form an inverted residual architecture, which increases the number of channels in the middle layer and places a shortcut connection to reuse input features. At the same time, the normalization layer and activation layer are placed after the middle layer. In S3, the multi-scale information fusion module uses an aggregation distribution mechanism to increase the multi-scale information fusion capability of the model, specifically including: Build a multi-scale information fusion module and introduce the goldyolo detection head at the detection head of the YOLOv8 model. The goldyolo detection head includes a feature alignment module FAM, an information fusion module IFM and an information injection module Inject. The feature alignment module FAM, the information fusion module IFM and the information injection module Inject constitute an aggregation distribution mechanism. The feature alignment module FAM collects feature maps of different scales of the backbone part and aligns them by upsampling or downsampling; The information fusion module IFM fuses the aligned features to generate global features, which are divided into two parts through a split operation, and then targeted distribution operations are performed on other scales; The information injection module Inject uses the attention operation that enhances branch detection capabilities to split the global features and distribute them to each level; Assuming the input image shape is N×3×H×W, there are four multi-scale features obtained from the backbone, namely B2, B3, B4, and B5, that is, Where N represents the batch-size, Indicates the number of channels of feature maps of different scales, Represents the height and width of feature maps of different scales; The feature alignment module FAM takes B4 as the benchmark, downsamples the large feature maps B2 and B3 by average pooling, and upsamples the small feature map B5 by bilinear interpolation, which is expressed as: F align =FAM([B2, B3, B4, B5]) (3) concat to get the combined feature representation: The information fusion module IFM design includes Conv, RepBlock modules, and Split operations: F fuse =RepBlock(F align ) (5) F inj_P3 ,F inj_P4 =Split(F fuse ) (6) The aligned and concat features F align Input into the RepBlock module to get F fuse Fusion features, while using Conv to adjust the channel to adapt to the size of different models, F fuse Split into F on the channel via Split inj_P3 and F inj_P4 , and then perform the next step of feature fusion with different levels; The information injection module Inject adopts the form of self-attention, and its input is the current scale x_local (Flocal) to be distributed, and the global feature x_global (F inj ), and finally the fusion information P is further obtained through ReBlock processing i .
2. According to claim 1, a target detection method based on partial convolution embedding and aggregation distribution mechanism is characterized in that: In S1, the public target detection dataset is converted into a YOLO format using the Python code for converting the data format and splitting the dataset, and the dataset annotation file is divided into a training set, a validation set, and a test set in a ratio of 7:1:
2.
3. The target detection method based on partial convolution embedding and aggregate distribution mechanism according to claim 1, characterized in that: The detection head of the YOLOv8 is a symmetrical structure. A GD mechanism is added to the detection head. The GD mechanism enters the detection head through two different network paths. A decoupled head structure is adopted in the detection head. Two parallel branches extract category features and position features respectively. Each branch uses a 1x1 convolution to complete its respective task.
Citation Information
Patent Citations
Road target detection method and system based on improved YOLOv8
CN117037119A
Cited By
A SAR image target detection method and device based on data augmentation
CN120563796B