Gray level image target detection method and system based on improved YOLOv8 model
By introducing the SAConv transform cavity convolution layer, EMA multi-scale attention layer and RFAConv region feature adaptive convolution layer in the YOLOv8 model, the problem of low detection accuracy of grayscale image objects is solved, and more efficient feature extraction and detection accuracy is achieved.
Patent Information
- Application Number
- CN202510114135.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-16
AI Technical Summary
The existing YOLOv8 model is difficult to detect and has low accuracy when detecting grayscale image objects.
The SAConv transform hollow convolution layer was introduced into the backbone backbone network of the YOLOv8 model, the EMA multi-scale attention layer was introduced into the Neck neck network, and the RFAConv region feature adaptive convolution layer was used in the head detection head network to enhance the feature extraction capability of grayscale images.
Through the improved YOLOv8 model, multi-scale features can be captured more effectively, attention to the details of grayscale image textures, and improvement of accuracy of grayscale image object detection.
Smart Images

Figure CN120014238A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection, and more specifically, relates to a grayscale image target detection method and system based on an improved YOLOv8 model. Background Art
[0002] With the rapid development of deep learning technology, great progress has been made in the fields of image processing and computer vision. At present, deep learning-based object detection algorithms are divided into two-stage algorithms and single-stage algorithms. Among them, two-stage object detection algorithms such as regional convolutional neural networks, fast regional convolutional neural networks, and masked regional convolutional neural networks have excellent accuracy, but due to their complex algorithm structure, the training speed is slow and the resource consumption is large. In contrast, single-stage object detection algorithms such as single-step multi-box detectors, YOLO series, and real-time Transformer detectors directly use convolutional neural networks to extract features from images and predict the location and category of the target at the same time. It is more efficient in detection speed and the network structure is relatively simple.
[0003] At the same time, most current image target detection technologies use color images as the basis for detection. Color image detection has good results, but is easily affected by environmental factors such as weather and light intensity. In contrast, grayscale images, with their single-channel grayscale information characteristics, simplify target features and reduce environmental noise interference, so that the model can greatly reduce the impact of harsh environments such as smoke, haze, rain and snow in target detection tasks. However, due to the lack of color information description, grayscale images have fewer texture details, low contrast and signal-to-noise ratio, and blurred imaging, which affects the grayscale image detection effect. Due to the characteristics of grayscale images with less target feature information and unclear texture details, the existing YOLOv8 model is difficult to detect when used for grayscale image target detection, and the accuracy of grayscale image target detection is low. Summary of the invention
[0004] In order to solve the problem that the existing YOLOv8 model is difficult to detect and has low accuracy when used for grayscale image target detection, the present invention proposes a grayscale image target detection method based on an improved YOLOv8 model to improve the accuracy of grayscale image target detection.
[0005] In order to achieve the above technical effects, the technical solution of the present invention is as follows:
[0006] In a first aspect, the present invention proposes a grayscale image target detection method based on an improved YOLOv8 model, wherein the YOLOv8 model includes a backbone network, a neck network, and a head detection network connected in sequence, and comprises the following steps:
[0007] S1: Get the grayscale image to be detected;
[0008] S2: Build an improved YOLOv8 model; in the backbone network, use SAConv to transform the hole convolution layer instead of the standard convolution layer to adapt to the grayscale image features of different scales; in the neck network, introduce the EMA multi-scale attention layer to improve the attention to the texture details of the grayscale image; in the head detection network, use the RFAConv regional feature adaptive convolution layer instead of the standard convolution layer to enhance the attention to the small target area of the grayscale image;
[0009] S3: Train the improved YOLOv8 model to obtain a trained improved YOLOv8 model;
[0010] S4: Input the grayscale image to be detected into the trained improved YOLOv8 model to obtain the target detection result.
[0011] Furthermore, the Backbone network includes: a first Conv convolution module, a first combination module, a second combination module, a third combination module, a fourth combination module and a first SPPF module connected in sequence;
[0012] Any one of the first combination module, the second combination module, the third combination module, and the fourth combination module includes: a first SAConv transformation hole convolution layer and a C2f layer;
[0013] The Neck network comprises: a first C2f module, a second C2f module, a first ELEN-M module, a second ELEN-M module and a third ELEN-M module connected in sequence; the output end of the second combination module is connected to the input end of the second C2f module; the output end of the third combination module is connected to the input end of the first C2f module; the output end of the first C2f module is connected to the input end of the second ELEN-M module; the output end of the second C2f module is connected to the input end of the first ELEN-M module;
[0014] Any of the first ELEN-M module, the second ELEN-M module and the third ELEN-M module includes: a first EMA multi-scale attention layer and a first Conv layer; the output end of the first SPPF module is connected to the input end of the first C2f module and the input end of the third ELEN-M module respectively;
[0015] The Head network includes: a first super-resolution reconstruction detection head module, a second super-resolution reconstruction detection head module and a third super-resolution reconstruction detection head module; the output end of the first ELEN-M module is connected to the first super-resolution reconstruction detection head module, the output end of the second ELEN-M module is connected to the second super-resolution reconstruction detection head module, and the output end of the third ELEN-M module is connected to the third super-resolution reconstruction detection head module.
[0016] Furthermore, the first SAConv transformation hole convolution layer includes: a first GlobalAvgPool structure, a first Conv structure, a first AvgPool structure, a second Conv structure, a second Global AvgPool structure and a third Conv structure connected in sequence; and also includes: a fourth Conv structure and a fifth Conv structure; the output end of the first Conv structure is respectively connected to the input end of the fourth Conv structure and the input end of the fifth Conv structure; the output end of the fourth Conv structure and the output end of the fifth Conv structure are multiplied by the output end of the second Conv structure; after multiplication, they are spliced and connected to the input end of the second Global AvgPool structure; the first Global AvgPool structure and the first Conv structure extract the features of the input grayscale image, and then input them to the input end of the fourth Conv structure and the fifth Conv structure for feature weight adjustment, and finally input them to the second GlobalAvgPool structure and the third Conv structure to screen important features for output.
[0017] According to the above technical features, adding the SAConv transformation hole convolution layer to the Backbone network can adaptively adjust the receptive field of the convolution, thereby more effectively capturing the multi-scale features in the image and improving the accuracy of grayscale image target detection.
[0018] Furthermore, the first EMA multi-scale attention layer includes: a first Groups structure, a first XAvgPool structure, a first Y AvgPool structure, a first Concat_Conv structure, a first Sigmoid structure, and a second Sigmoid structure; a first Re-weight structure, a first GroupNorm structure, a second AvgPool structure, a first Softmax structure, a first Matmul structure, a third Sigmoid structure, and a second Re-weight structure connected in sequence; a sixth Conv structure, a third AvgPool structure, a second Softmax structure, and a second Matmul structure connected in sequence; the output end of the first Groups structure is respectively connected to the input ends of the first X AvgPool structure, the first Y AvgPool structure, the sixth Conv structure, the first Re-weight structure, and the second Re-weight structure; the first X AvgPool structure and the first Y The output ends of the AvgPool structure are all connected to the input ends of the first Concat_Conv structure; the output ends of the first Concat_Conv structure are respectively connected to the input ends of the first Sigmoid structure and the second Sigmoid structure; the output ends of the first Sigmoid structure and the second Sigmoid structure are both connected to the input end of the first Re-weight structure; the output end of the first Re-weight structure is connected to the input end of the second Matmul structure; the output end of the sixth Conv structure is connected to the input end of the first Matmul structure; the output end of the second Matmul structure and the output end of the third Sigmoid structure are spliced and connected to the input end of the third Sigmoid structure.
[0019] According to the above technical means, an EMA multi-scale attention layer is added to the Neck network. The EMA multi-scale attention layer focuses on the local details and global structure of the grayscale image, integrates feature information from different spatial positions, further enhances the feature representation capability, and thus improves the accuracy of grayscale image target detection.
[0020] Furthermore, in the EMA multi-scale attention layer: the input feature map is input into the first Groups structure, and is divided into sub-feature maps according to the channel dimension of the feature map. The calculation expression of the division process is:
[0021] X=[X0,X i ,...,X G-1 ],X i ∈R C / / G×H×W ,G<<C
[0022] In the formula, X represents the feature map, G represents the number of sub-features, and X i represents the i-th sub-feature;
[0023] In the horizontal and vertical directions, the first X AvgPool structure and the first Y AvgPool structure are used to perform one-dimensional global average pooling on the channel. The calculation expression of global average pooling is:
[0024]
[0025]
[0026] In the formula, represents the output of the Cth channel with height H, Represents the output of the Cth channel with width w;
[0027] The outputs of the first X AvgPool structure and the first Y AvgPool structure are concatenated and decomposed by the first Concat_Conv structure, and then multiplied and aggregated after linear fitting by the first Sigmoid structure and the second Sigmoid structure; at the same time, the third AvgPool structure is used to perform a two-dimensional global average pooling operation, and the calculation expression is:
[0028]
[0029] In the formula, Z C Represents the output of the Cth channel.
[0030] Furthermore, any super-resolution reconstruction detection head module of the first super-resolution reconstruction detection head module, the second super-resolution reconstruction detection head module and the third super-resolution reconstruction detection head module includes: a first RFAConv regional feature adaptive convolution layer, a first BOX layer and a first CLS layer; the features output by the first RFAConv regional feature adaptive convolution layer are input into the first BOX layer for bounding box prediction of target detection, and are input into the first CLS layer for classification.
[0031] Furthermore, the first RFAConv regional feature adaptive convolution layer includes: a first Group Conv structure, a fourth AvgPool structure, a seventh Conv structure, a third Softmax structure, a fourth Softmax structure, a fifth Softmax structure, a third Re-weight structure and an eighth Conv structure; the output end of the fourth AvgPool structure is connected to the input end of the seventh Conv structure; the output end of the seventh Conv structure is respectively connected to the input ends of the third Softmax structure, the fourth Softmax structure and the fifth Softmax structure; the output ends of the third Softmax structure, the fourth Softmax structure, the fifth Softmax structure and the first Group Conv structure are connected to the input end of the third Re-weight structure; the output end of the third Re-weight structure is connected to the input end of the seventh Conv structure; the feature information output by the fourth AvgPool structure is input into the seventh Conv structure for information exchange, and then input into the third Softmax structure, the fourth Softmax structure and the fifth Softmax structure for normalization operation, and the eighth Conv structure performs information interaction to output the final feature information.
[0032] According to the above technical means, the RFAConv regional feature adaptive convolution layer is added to the super-resolution reconstruction detection head module, focusing on the spatial features within the receptive field, learning the local features in the grayscale image more carefully, and improving the accuracy of feature extraction.
[0033] Furthermore, the SAConv transform hole convolution layer dynamically adjusts the weights of the fourth Conv structure and the fifth Conv structure, and the expression of the dynamic adjustment is:
[0034] Output=S(x)×Conv(x,y,1)+(1-S(x))×Conv(x,y,Δy,r)
[0035] In the formula, Output represents the output information, Conv represents the convolution structure, r represents the hyperparameter of the SAConv transformation hole convolution layer, Δy represents the trainable value, x represents the input information, y represents the calculation weight of the convolution kernel, and S(x) represents the switching function.
[0036] Furthermore, in the RFAConv regional feature adaptive convolution layer: let X∈R C×H×W Represents the input grayscale image feature map. The input grayscale image feature map is processed by combining the receptive field adaptive convolution and the spatial attention mechanism, and the receptive field spatial attention feature information is output. The process satisfies the calculation expression:
[0037] F = Softmax(g 1×1(AvgPool(X)))×ReLU(Norm(g k×k (X)))
[0038] =Arf×Frf
[0039] In the formula, F represents the feature output, X represents the feature input, Arf represents the weight information of the attention map, Frf represents the feature information of the receptive field space, and g 1×1 Represents 1×1 grouped convolution, AvgPool represents average pooling, Norm represents normalization operation, ReLU represents activation function, and Softmax represents normalized exponential function.
[0040] The present invention also provides a grayscale image target detection system based on an improved YOLOv8 model, comprising:
[0041] An image acquisition module, used for acquiring a grayscale image to be detected;
[0042] The model building module is used to build an improved YOLOv8 model. In the backbone network, the SAConv dilated convolution layer is used to replace the standard convolution layer to adapt to the grayscale image features of different scales. In the neck network, the EMA multi-scale attention layer is introduced to improve the attention to the texture details of the grayscale image. In the head detection network, the RFAConv regional feature adaptive convolution layer is used to replace the standard convolution layer to enhance the attention to the small target area of the grayscale image.
[0043] A training module is used to train the improved YOLOv8 model to obtain a trained improved YOLOv8 model;
[0044] The detection module is used to input the grayscale image to be detected into the trained improved YOLOv8 model to obtain the target detection result.
[0045] Compared with the prior art, the beneficial effects of this method are:
[0046] The present invention provides a grayscale image target detection method and system based on an improved YOLOv8 model. First, a grayscale image to be detected is obtained. Then, an improved YOLOv8 model is constructed; in the backbone network, a SAConv transformation hole convolution layer is added to capture feature information of different scales; in the neck network, an EMA multi-scale attention layer is added to enhance the feature representation capability and improve the attention to the texture details of the grayscale image; in the head detection network, a super-resolution reconstruction detection head is added to improve the detail recovery capability and spatial perception capability. The present invention uses an improved YOLOv8 model as a whole to improve the accuracy of grayscale image target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 A flowchart showing a grayscale image target detection method based on an improved YOLOv8 model proposed in an embodiment of the present invention;
[0048] Figure 2 The overall structure diagram of the improved YOLOv8 model proposed in the embodiment of the present invention is shown;
[0049] Figure 3 A schematic diagram showing the structure of the SAConv transformation hole convolution layer proposed in an embodiment of the present invention;
[0050] Figure 4 A schematic diagram showing the structure of the EMA multi-scale attention layer proposed in an embodiment of the present invention;
[0051] Figure 5 A schematic diagram showing the structure of the RFAConv regional feature adaptive convolution layer proposed in an embodiment of the present invention;
[0052] Figure 6 A comparison diagram of grayscale image target detection results based on the original YOLOv8 model and the improved YOLOv8 model proposed in an embodiment of the present invention is shown;
[0053] Figure 7 The figure shows a structural diagram of a grayscale image target detection system based on an improved YOLOv8 model proposed in an embodiment of the present invention. DETAILED DESCRIPTION
[0054] The drawings are for illustrative purposes only and should not be construed as limiting the present patent;
[0055] In order to better illustrate the present embodiment, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the actual size;
[0056] It is understandable to those skilled in the art that descriptions of certain well-known contents in the drawings may be omitted.
[0057] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0058] The positional relationships described in the drawings are only for illustrative purposes and should not be construed as limiting the present patent;
[0059] Example 1
[0060] This embodiment proposes a grayscale image target detection method based on an improved YOLOv8 model, wherein the YOLOv8 model includes a backbone network, a neck network, and a head detection network connected in sequence. Figure 1The grayscale image target detection method proposed in this embodiment generally includes the following steps:
[0061] S1: Get the grayscale image to be detected;
[0062] S2: Build an improved YOLOv8 model; in the backbone network, use SAConv to transform the hole convolution layer instead of the standard convolution layer to adapt to the grayscale image features of different scales; in the neck network, introduce the EMA multi-scale attention layer to improve the attention to the texture details of the grayscale image; in the head detection network, use the RFAConv regional feature adaptive convolution layer instead of the standard convolution layer to enhance the attention to the small target area of the grayscale image;
[0063] S3: Train the improved YOLOv8 model to obtain a trained improved YOLOv8 model;
[0064] S4: Input the grayscale image to be detected into the trained improved YOLOv8 model to obtain the target detection result.
[0065] This embodiment specifically describes the composition of the improved YOLOv8 model. In this embodiment, Figure 2 As shown in the overall structure diagram of the improved YOLOv8 model, the Backbone network includes: a first Conv convolution module, a first combination module, a second combination module, a third combination module, a fourth combination module and a first SPPF module connected in sequence.
[0066] Any one of the first combination module, the second combination module, the third combination module, and the fourth combination module includes: a first SAConv transformation hole convolution layer and a C2f layer.
[0067] like Figure 3The structural schematic diagram of the SAConv transformation hole convolution layer shown in the figure, the first SAConv transformation hole convolution layer includes: a first Global AvgPool structure, a first Conv structure, a first AvgPool structure, a second Conv structure, a second Global AvgPool structure and a third Conv structure connected in sequence; it also includes: a fourth Conv structure and a fifth Conv structure; the output end of the first Conv structure is respectively connected to the input end of the fourth Conv structure and the input end of the fifth Conv structure; the output end of the fourth Conv structure and the output end of the fifth Conv structure are multiplied by the output end of the second Conv structure; after multiplication, they are spliced and connected to the input end of the second Global AvgPool structure; the first Global AvgPool structure and the first Conv structure extract the features of the input grayscale image, and then input them to the input end of the fourth Conv structure and the fifth Conv structure for feature weight adjustment, and finally input them to the second Global AvgPool structure and the third Conv structure to screen important features for output.
[0068] In this embodiment, the image is input to the Backbone network, and first passes through the first Conv convolution module, which consists of Conv convolution, BN (batchnorm2d) and silu activation function, downsamples the feature map, normalizes the grayscale image data, and the silu activation function increases the nonlinearity of the data. Then, the feature map processed by the first Conv convolution module enters four combination modules. The combination module consists of the first SAConv transformation hole convolution layer and the C2f layer.
[0069] In this embodiment, in the first SAConv transformation hole convolution layer, the input feature map is first subjected to the first Global AvgPool structure and the first Conv structure, the features of the feature map are extracted, the global information is obtained, and the global information is input into the hole convolution, the input information is observed twice, two hole rates are assigned to the input information, and two different hole convolutions are performed. At the same time, in order to avoid the weight loss problem caused by a large hole rate, a switch function is set after the hole convolution, the initialization weight is stored in the switch function, the parallel convolution result is introduced into the switch function, and the output weight and feature after the convolution fusion is determined by the switch function, and finally the output information is input into the second Global AvgPool structure and the third Conv structure, the important features in the hole convolution layer are summarized and screened, and comprehensive grayscale image information is formed to reduce the difficulty of model detection.
[0070] The SAConv transform hole convolution layer dynamically adjusts the weights of the fourth Conv structure and the fifth Conv structure. The expression of dynamic adjustment is:
[0071] Output=S(x)×Conv(x,y,1)+(1-S(x))×Conv(x,y,Δy,r)
[0072] In the formula, Output represents the output information, Conv represents the convolution structure, r represents the hyperparameter of the SAConv transformation hole convolution layer, Δy represents the trainable value, x represents the input information, y represents the calculation weight of the convolution kernel, and S(x) represents the switching function.
[0073] In this embodiment, a SAConv transformation dilated convolution layer is added to the Backbone network, different dilated rates are applied to the input features to calculate the dilated convolution, the convolution results are weighted and shared, and the Global AvgPool structure and the Conv structure are combined to coordinate global information, thereby improving the flexibility of the network and its adaptability to fuzzy information.
[0074] In this embodiment, the C2f layer adopts the efficient aggregation network structure of the ELEN module in the existing YOLOv7 model, enriches the gradient flow of the model through multi-branch cross-layer links, improves parameter utilization, and increases network depth.
[0075] Finally, after passing through four groups of combination modules, it enters the SPPF module to process feature maps of different pixel sizes to complete feature fusion.
[0076] The neck network includes: a first C2f module, a second C2f module, a first ELEN-M module, a second ELEN-M module and a third ELEN-M module connected in sequence; the output end of the second combination module is connected to the input end of the second C2f module; the output end of the third combination module is connected to the input end of the first C2f module; the output end of the first C2f module is connected to the input end of the second ELEN-M module; the output end of the second C2f module is connected to the input end of the first ELEN-M module.
[0077] Any of the first ELEN-M module, the second ELEN-M module and the third ELEN-M module includes: a first EMA multi-scale attention layer and a first Conv layer; the output end of the first SPPF module is respectively connected to the input end of the first C2f module and the input end of the third ELEN-M module.
[0078] like Figure 4Schematic diagram of the structure of the EMA multi-scale attention layer shown in the figure, the first EMA multi-scale attention layer includes: a first Groups structure, a first X AvgPool structure, a first Y AvgPool structure, a first Concat_Conv structure, a first Sigmoid structure and a second Sigmoid structure; a first Re-weight structure, a first GroupNorm structure, a second AvgPool structure, a first Softmax structure, a first Matmul structure, a third Sigmoid structure and a second Re-weight structure connected in sequence; a sixth Conv structure, a third AvgPool structure, a second Softmax structure and a second Matmul structure connected in sequence; the output end of the first Groups structure is respectively connected to the input end of the first X AvgPool structure, the first Y AvgPool structure, the sixth Conv structure, the first Re-weight structure and the second Re-weight structure; the first X AvgPool structure and the first Y The output ends of the AvgPool structure are all connected to the input ends of the first Concat_Conv structure; the output ends of the first Concat_Conv structure are respectively connected to the input ends of the first Sigmoid structure and the second Sigmoid structure; the output ends of the first Sigmoid structure and the second Sigmoid structure are both connected to the input end of the first Re-weight structure; the output end of the first Re-weight structure is connected to the input end of the second Matmul structure; the output end of the sixth Conv structure is connected to the input end of the first Matmul structure; the output end of the second Matmul structure and the output end of the third Sigmoid structure are spliced and connected to the input end of the third Sigmoid structure.
[0079] After the image passes through the Backbone network to obtain deep feature information, it enters the Neck network. The Neck network adopts the PAFPN structure, which consists of two parts: Feature Pyramid Networks (FPN) and Path Aggregation Network (PAN), and introduces the EMA multi-scale attention layer.
[0080] In this embodiment, the EMA multi-scale attention layer first divides the input feature map of G×H×W into sub-feature maps according to the channel dimension direction according to the processing method of the channel attention mechanism. The calculation expression of the division process is:
[0081] X=[X0,X1,...,X G-1 ],X i ∈R C / / G×H×W ,G<<C
[0082] In the formula, X represents the feature map, G represents the number of sub-features, and X i represents the i-th sub-feature.
[0083] Next, each channel information is decomposed into two 1×1 feature encoding branches and one 3×3 feature encoding branch. On the two 1×1 feature encoding branches, a one-dimensional global average pooling operation is performed on the channel along the horizontal and vertical directions respectively;
[0084] The calculation expression of the one-dimensional global average pooling in the horizontal direction is:
[0085]
[0086] In the formula, Represents the output of the Cth channel with height H.
[0087] The calculation expression of one-dimensional global average pooling in the vertical direction is:
[0088]
[0089] In the formula, Represents the output of the Cth channel with width w.
[0090] Then, the feature codes of the two branches are concatenated along the height direction, decomposed into two vectors through a 1×1 convolution operation, and linearly fitted using a nonlinear Sigmoid function. The feature information of the 1×1 branch is then aggregated together using multiplication to form a new branch to achieve cross-channel feature information interaction. At the same time, the 3×3 feature encoding branch obtains semantic information in the local area through a 3×3 convolution. The two branches simultaneously enter the cross-spatial information part, and the 1×1 branch first performs a two-dimensional global average pooling operation and a Softmax function linear transformation.
[0091] The two-dimensional global average pooling structure formula is as follows:
[0092]
[0093] In the formula, Z C Represents the output of the Cth channel.
[0094] Finally, the output features of the 1×1 branch and the 3×3 branch are multiplied by matrix dot product operation, and the output dimension is converted to The 3×3 branch first performs two-dimensional global average pooling on the cross-spatial information part, transforms the pooling result linearly with Softmax, performs matrix dot product operation with the 1×1 branch output feature, and converts the output dimension into The output features of the two branches are combined to form a feature map with dual spatial attention weights, which helps the model capture pixel-level relationships in the image and enhances the ability of feature representation.
[0095] The Head network includes: a first super-resolution reconstruction detection head module, a second super-resolution reconstruction detection head module and a third super-resolution reconstruction detection head module; the output end of the first ELEN-M module is connected to the first super-resolution reconstruction detection head module, the output end of the second ELEN-M module is connected to the second super-resolution reconstruction detection head module, and the output end of the third ELEN-M module is connected to the third super-resolution reconstruction detection head module.
[0096] Any super-resolution reconstruction detection head module of the first super-resolution reconstruction detection head module, the second super-resolution reconstruction detection head module and the third super-resolution reconstruction detection head module includes: a first RFAConv regional feature adaptive convolution layer, a first BOX layer and a first CLS layer; the features output by the first RFAConv regional feature adaptive convolution layer are input into the first BOX layer for bounding box prediction of target detection, and are input into the first CLS layer for classification.
[0097] like Figure 5 The structural schematic diagram of the RFAConv regional feature adaptive convolution layer shown in the figure, the first RFAConv regional feature adaptive convolution layer includes: a first Group Conv structure, a fourth AvgPool structure, a third Softmax structure, a fourth Softmax structure, a fifth Softmax structure, a second Re-weight structure and a seventh Conv structure; the output end of the fourth AvgPool structure is respectively connected to the input end of the third Softmax structure, the fourth Softmax structure and the fifth Softmax structure; the output ends of the third Softmax structure, the fourth Softmax structure, the fifth Softmax structure and the first Group Conv structure are connected to the input end of the second Re-weight structure; the output end of the second Re-weight structure is connected to the input end of the seventh Conv structure.
[0098] The feature information output by the fourth AvgPool structure is input into the seventh Conv structure for information exchange, and then input into the third Softmax structure, the fourth Softmax structure and the fifth Softmax structure for normalization operation. The eighth Conv structure performs information exchange and outputs the final feature information.
[0099] In this embodiment, the output results of the three ELEN-M modules of the Neck network are input to the Head network. The input passes through the RepVGG structure, which adopts a residual structure and obtains feature information of different receptive fields by applying multiple different branches and convolution kernels to the model.
[0100] The output features of the Neck part are used as the input of the super-resolution reconstruction detection head. The super-resolution reconstruction detection head is based on the decoupling head. It combines convolution and spatial attention mechanism to replace the standard convolution with RFAConv regional feature adaptive convolution layer to enhance the detail recovery ability and spatial perception ability of the detection head, thereby improving the robustness of the model to multi-scale information recognition.
[0101] The RFAConv regional feature adaptive convolution layer is divided into two branches. The input feature map X of size C×H×W enters the right branch through the receptive field adaptive convolution. The receptive field adaptive convolution takes a 3×3 large-scale convolution kernel as the core, introduces a sliding window of the same size to dynamically focus on important information in the image, and then multiplies the convolution kernel information with the window information to obtain feature information. The left branch aggregates the global information of the receptive field features through average pooling, and then uses 1×1 convolution for information interaction, and the normalized structure focuses on the important information weights. A parameter sharing strategy is adopted to combine the weights in the attention map with the receptive field convolution to extract features, adjust the feature size, and output the receptive field spatial attention feature information.
[0102] In this embodiment, in the RFAConv regional feature adaptive convolution layer: let X∈R C×H×W Represents the input grayscale image feature map. The input grayscale image feature map is processed by combining the receptive field adaptive convolution and the spatial attention mechanism, and the receptive field spatial attention feature information is output. The process satisfies the calculation expression:
[0103] F = Softmax(g 1×1 (AvgPool(X)))×ReLU(Norm(g k×k (X)))=Arf×Frf
[0104] In the formula, F represents the feature output, X represents the feature input, Arf represents the weight information of the attention map, Frf represents the feature information of the receptive field space, and g 1×1 Represents 1×1 grouped convolution, AvgPool represents average pooling, Norm represents normalized structure, ReLU represents activation function, and Softmax represents normalized exponential function.
[0105] Example 2
[0106] In this embodiment, grayscale images in the open source grayscale dataset NEU-DET are used as experimental images. The NEU-DET dataset is a steel surface defect detection dataset consisting of a total of 1,800 images with 6 types of labeled objects, including: crazing, patches, inclusion, pitted_surface, rolled-in_scale, and scratches.
[0107] In this embodiment, the grayscale images are randomly divided into a training set, a test set, and a validation set in a ratio of 7:2:1.
[0108] In this embodiment, when training the improved YOLOv8 model, the basic platform and parameters are as follows:
[0109] The improved YOLOv8 model network was trained in the PyTroch environment, with the training settings of learning rate 0.01, weight decay coefficient 0.0005, image size 640×640, batch size 8, and training cycle 300. The grayscale images in the dataset were input into the improved YOLOv8 model for training. Figure 6 As shown in the figure, the comparison of grayscale image target detection results based on the original YOLOv8 model and the improved YOLOv8 model. In the image, the steel surface defect target is marked with a rectangular frame. The results of the improved YOLOv8 model are compared with the results of the YOLOv8 model. The results show that the improved YOLOv8 model can accurately grasp the global and local information, and the detection effect is better than the original YOLOv8 model.
[0110] Example 3
[0111] This embodiment provides a grayscale image target detection system based on an improved YOLOv8 model. The structure diagram of the system is as follows: Figure 7 As shown, including:
[0112] An image acquisition module, used for acquiring a grayscale image to be detected;
[0113] The model building module is used to build an improved YOLOv8 model. In the backbone network, the SAConv dilated convolution layer is used to replace the standard convolution layer to adapt to the grayscale image features of different scales. In the neck network, the EMA multi-scale attention layer is introduced to improve the attention to the texture details of the grayscale image. In the head detection network, the RFAConv regional feature adaptive convolution layer is used to replace the standard convolution layer to enhance the attention to the small target area of the grayscale image.
[0114] A training module is used to train the improved YOLOv8 model to obtain a trained improved YOLOv8 model;
[0115] The detection module is used to input the grayscale image to be detected into the trained improved YOLOv8 model to obtain the target detection result.
[0116] The embodiments are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.
Claims
1. A grayscale image target detection method based on an improved YOLOv8 model, wherein the YOLOv8 model comprises a backbone network, a neck network and a head detection network connected in sequence, characterized in that: The following steps are involved: S1: Get the grayscale image to be detected; S2: Build an improved YOLOv8 model; in the backbone network, use SAConv to transform the hole convolution layer instead of the standard convolution layer to adapt to the grayscale image features of different scales; in the neck network, introduce the EMA multi-scale attention layer to improve the attention to the texture details of the grayscale image; in the head detection network, use the RFAConv regional feature adaptive convolution layer instead of the standard convolution layer to enhance the attention to the small target area of the grayscale image; S3: Train the improved YOLOv8 model to obtain a trained improved YOLOv8 model; S4: Input the grayscale image to be detected into the trained improved YOLOv8 model to obtain the target detection result.
2. A grayscale image target detection method based on an improved YOLOv8 model according to claim 1, characterized in that: The Backbone network includes: a first Conv convolution module, a first combination module, a second combination module, a third combination module, a fourth combination module and a first SPPF module connected in sequence; Any one of the first combination module, the second combination module, the third combination module, and the fourth combination module includes: a first SAConv transformation hole convolution layer and a C2f layer; The Neck network comprises: a first C2f module, a second C2f module, a first ELEN-M module, a second ELEN-M module and a third ELEN-M module connected in sequence; the output end of the second combination module is connected to the input end of the second C2f module; the output end of the third combination module is connected to the input end of the first C2f module; the output end of the first C2f module is connected to the input end of the second ELEN-M module; the output end of the second C2f module is connected to the input end of the first ELEN-M module; Any of the first ELEN-M module, the second ELEN-M module and the third ELEN-M module includes: a first EMA multi-scale attention layer and a first Conv layer; the output end of the first SPPF module is connected to the input end of the first C2f module and the input end of the third ELEN-M module respectively; The Head network includes: a first super-resolution reconstruction detection head module, a second super-resolution reconstruction detection head module and a third super-resolution reconstruction detection head module; the output end of the first ELEN-M module is connected to the first super-resolution reconstruction detection head module, the output end of the second ELEN-M module is connected to the second super-resolution reconstruction detection head module, and the output end of the third ELEN-M module is connected to the third super-resolution reconstruction detection head module.
3. A grayscale image target detection method based on an improved YOLOv8 model according to claim 2, characterized in that: The first SAConv transformation hole convolution layer includes: a first Global AvgPool structure, a first Conv structure, a first AvgPool structure, a second Conv structure, a second Global AvgPool structure and a third Conv structure connected in sequence; and also includes: a fourth Conv structure and a fifth Conv structure; the output end of the first Conv structure is respectively connected to the input end of the fourth Conv structure and the input end of the fifth Conv structure; the output end of the fourth Conv structure and the output end of the fifth Conv structure are multiplied by the output end of the second Conv structure; after multiplication, they are spliced and connected to the input end of the second Global AvgPool structure; the first GlobalAvgPool structure and the first Conv structure extract the features of the input grayscale image, and then input them to the input end of the fourth Conv structure and the fifth Conv structure for feature weight adjustment, and finally input them to the second Global AvgPool structure and the third Conv structure to screen important features for output.
4. A grayscale image target detection method based on an improved YOLOv8 model according to claim 2, characterized in that: The first EMA multi-scale attention layer includes: a first Groups structure, a first X AvgPool structure, a first YAvgPool structure, a first Concat_Conv structure, a first Sigmoid structure, and a second Sigmoid structure; a first Re-weight structure, a first GroupNorm structure, a second AvgPool structure, a first Softmax structure, a first Matmul structure, a third Sigmoid structure, and a second Re-weight structure connected in sequence; a sixth Conv structure, a third AvgPool structure, a second Softmax structure, and a second Matmul structure connected in sequence; the output end of the first Groups structure is respectively connected to the input ends of the first XAvgPool structure, the first Y AvgPool structure, the sixth Conv structure, the first Re-weight structure, and the second Re-weight structure; the first X AvgPool structure and the first Y The output ends of the AvgPool structure are all connected to the input ends of the first Concat_Conv structure; the output ends of the first Concat_Conv structure are respectively connected to the input ends of the first Sigmoid structure and the second Sigmoid structure; the output ends of the first Sigmoid structure and the second Sigmoid structure are both connected to the input end of the first Re-weight structure; the output end of the first Re-weight structure is connected to the input end of the second Matmul structure; the output end of the sixth Conv structure is connected to the input end of the first Matmul structure; the output end of the second Matmul structure and the output end of the third Sigmoid structure are spliced and connected to the input end of the third Sigmoid structure.
5. A grayscale image target detection method based on an improved YOLOv8 model according to claim 4, characterized in that: In the EMA multi-scale attention layer: the input feature map is input into the first Groups structure and divided into sub-feature maps according to the channel dimension of the feature map. The calculation expression of the division process is: X=[X0,X i ,...,X G-1 ],X i ∈R C / / G×H×W ,G< <C In the formula, X represents the feature map, G represents the number of sub-features, and X i represents the i-th sub-feature; In the horizontal and vertical directions, the first X AvgPool structure and the first Y AvgPool structure are used to perform one-dimensional global average pooling on the channel. The calculation expression of global average pooling is: In the formula, represents the output of the Cth channel with height H, Represents the output of the Cth channel with width w; The outputs of the first X AvgPool structure and the first Y AvgPool structure are concatenated and decomposed by the first Concat_Conv structure, and then multiplied and aggregated after linear fitting by the first Sigmoid structure and the second Sigmoid structure; at the same time, the third AvgPool structure is used to perform a two-dimensional global average pooling operation, and the calculation expression is: In the formula, Z C Represents the output of the Cth channel.
6. A grayscale image target detection method based on an improved YOLOv8 model according to claim 2, characterized in that: Any super-resolution reconstruction detection head module of the first super-resolution reconstruction detection head module, the second super-resolution reconstruction detection head module and the third super-resolution reconstruction detection head module includes: a first RFAConv regional feature adaptive convolution layer, a first BOX layer and a first CLS layer; the features output by the first RFAConv regional feature adaptive convolution layer are input into the first BOX layer for bounding box prediction of target detection, and are input into the first CLS layer for classification.
7. A grayscale image target detection method based on an improved YOLOv8 model according to claim 6, characterized in that: The first RFAConv regional feature adaptive convolution layer includes: a first GroupConv structure, a fourth AvgPool structure, a seventh Conv structure, a third Softmax structure, a fourth Softmax structure, a fifth Softmax structure, a third Re-weight structure and an eighth Conv structure; the output end of the fourth AvgPool structure is connected to the input end of the seventh Conv structure; the output end of the seventh Conv structure is respectively connected to the input ends of the third Softmax structure, the fourth Softmax structure and the fifth Softmax structure; the output ends of the third Softmax structure, the fourth Softmax structure, the fifth Softmax structure and the first Group Conv structure are connected to the input end of the third Re-weight structure; the output end of the third Re-weight structure is connected to the input end of the seventh Conv structure; the feature information output by the fourth AvgPool structure is input into the seventh Conv structure for information exchange, and then input into the third Softmax structure, the fourth Softmax structure and the fifth Softmax structure for normalization operation, and the eighth Conv structure performs information interaction to output the final feature information.
8. The grayscale image target detection method based on the improved YOLOv8 model according to claim 3, characterized in that: The SAConv transformation hole convolution layer dynamically adjusts the weights of the fourth Conv structure and the fifth Conv structure, and the expression of dynamic adjustment is: Output=S(x)×Conv(x,y,1)+(1-S(x))×Conv(x,y,Δy,r) In the formula, Output represents the output information, Conv represents the convolution structure, r represents the hyperparameter of the SAConv transformation hole convolution layer, Δy represents the trainable value, x represents the input information, y represents the calculation weight of the convolution kernel, and S(x) represents the switching function.
9. A grayscale image target detection method based on an improved YOLOv8 model according to claim 7, characterized in that: In the RFAConv regional feature adaptive convolution layer: Let X∈R C×H×W Represents the input grayscale image feature map. The input grayscale image feature map is processed by combining the receptive field adaptive convolution and the spatial attention mechanism, and the receptive field spatial attention feature information is output. The process satisfies the calculation expression: F=Softmax(g 1×1 (AvgPool(X)))×ReLU(Norm(g k×k (X))) =Arf×Frf In the formula, F represents the feature output, X represents the feature input, Arf represents the weight information of the attention map, Frf represents the feature information of the receptive field space, and g 1×1 Represents 1×1 grouped convolution, AvgPool represents average pooling, Norm represents normalization operation, ReLU represents activation function, and Softmax represents normalized exponential function.
10. A grayscale image target detection system based on an improved YOLOv8 model, characterized in that: include: An image acquisition module, used for acquiring a grayscale image to be detected; The model building module is used to build an improved YOLOv8 model. In the backbone network, the SAConv dilated convolution layer is used to replace the standard convolution layer to adapt to the grayscale image features of different scales. In the neck network, the EMA multi-scale attention layer is introduced to improve the attention to the texture details of the grayscale image. In the head detection network, the RFAConv regional feature adaptive convolution layer is used to replace the standard convolution layer to enhance the attention to the small target area of the grayscale image. A training module is used to train the improved YOLOv8 model to obtain a trained improved YOLOv8 model; The detection module is used to input the grayscale image to be detected into the trained improved YOLOv8 model to obtain the target detection result.
Citation Information
Cited By
Tomato image real-time detection method, tomato image real-time detection system, tomato picking method and tomato picking system
CN120388367A
Real-time tomato image detection method and system, tomato picking method and system
CN120388367B
Medical image super-resolution method based on dynamic attention and implicit neural representation
CN120450963A