Image detection method based on improved YOLO algorithm

By introducing an improved multi-channel spatial pyramid pooling module and embedded attention mechanism in the YOLO algorithm, as well as an improved bounding box regression loss function, the problems of unreliable and poor lesion detection in MRI medical images in the prior art are solved, and higher detection accuracy and robustness are achieved.

CN120088208AActive Publication Date: 2025-06-03GUIZHOU UNIV

Patent Information

Application Number
CN202510127200.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-01
Publication Date
2025-06-03
Estimated Expiration
2045-02-01

AI Technical Summary

Technical Problem

The existing object detection model is unreliable in MRI medical imaging, the model is poorly robust and has a low accuracy rate.

Method used

Based on the image detection method of improved YOLO algorithm, the image feature extraction and feature fusion capabilities are enhanced by introducing an improved multi-channel spatial pyramid pooling module SPPMC and an embedded attention mechanism UECA module, and the improved bounding box regression loss function L (iS-IoU) is used to improve positioning accuracy.

Benefits of technology

It significantly improves the robustness and detection performance of the model, enhances the characterization and processing capabilities of complex data characteristics, and improves the accuracy and reliability of lesion detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088208A_ABST
    Figure CN120088208A_ABST
Patent Text Reader

Abstract

The invention discloses an image detection method based on an improved YOLO algorithm. The method comprises the following steps: collecting a nuclear magnetic resonance image with a focus; a target detection network model is constructed: basic framework improvement is carried out based on YOLOv8, an improved backbone network comprises a C2f module, a convolution module and an improved multichannel spatial pyramid pooling module SPPMC, and the image feature extraction capability is enhanced; the improved neck network comprises an up-sampling layer, a splicing layer, a C2f module, a convolution layer and an embedded attention mechanism UECA module, multi-scale feature map fusion is carried out, and the feature fusion efficiency is improved; the improved header network adopts an improved bounding box regression loss function L (iS-IoU), the convergence speed of the model is increased, and the positioning precision is improved; inputting a training image set into the target detection network model for training; and performing image focus detection by using the trained target detection network model. The method has the characteristic of effectively improving the robustness and the detection performance of the target detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to an image detection method based on an improved YOLO algorithm. Background Art

[0002] Magnetic resonance imaging (MRI) technology, with advantages such as non-invasive, non-ionizing radiation, high resolution for soft tissues, accurate positioning, no need to use contrast agents, multi-plane and multi-parameter imaging, and non-invasive in vivo chemical analysis, is widely used in the medical field for medical imaging of various human tissues and organs, playing an irreplaceable role in the diagnosis and treatment of various diseases. Object detection is an important research direction in computer vision, aiming to study how to automatically and accurately detect, locate, and identify target objects from images or videos. In recent years, with the remarkable achievements of deep learning technology in computer vision tasks, object detection technology has also achieved leapfrog development. Applying object detection technology to detect and identify lesions in cardiac MRI medical images provides a solution to the above problems. The object detection model can deeply learn and extract various lesion features on medical images, and accurately and quickly detect and identify the diseased tissues or organs in medical images. This can not only help doctors improve the accuracy and efficiency of disease diagnosis, but also help medical researchers conduct in-depth research on diseases and promote the development of medical technology.

[0003] In the prior art, Sibo Qiao et al. proposed a residual learning diagnosis system (RLDS) based on a convolutional neural network (CNN) for fetal congenital heart disease (coronary heart disease), and conducted experiments on a self-made fetal echocardiogram dataset. The detection accuracy and recall rate of the RLDS model for fetal coronary heart disease both reached 93%. Elhanashi et al. proposed a comprehensive model combining multi-object detection models to detect multiple abnormalities in chest X-ray images. Compared with the state-of-the-art methods, the proposed method achieved good results in multi-classification and localization of abnormalities including COVID-19. Ji Zhanlin proposed a YOLOv5-CASP model based on YOLOv5s for the detection of pulmonary nodules, introduced an improved convolutional attention module, replaced the SPPF module with an improved ASPP module, and introduced a context transformer module, achieving certain detection effects. Yuan Zizhong proposed a YOLOv5-GDF model based on the detection research of laterally spreading tumors of the large intestine with improved YOLO and achieved good detection results. Su Yipeng proposed an automatic detection method for rib fractures in chest CT images based on a center network with a heatmap pyramid structure, realizing accurate detection of rib fractures. However, the existing object detection models have problems such as poor robustness, low accuracy, and unreliable detection results. Summary of the Invention

[0004] The object of the present invention is to overcome the above-mentioned drawbacks and propose an image detection method based on an improved YOLO algorithm, which has a significant improvement effect on the model performance and effectively improves the robustness and detection performance of the model.

[0005] An image detection method based on an improved YOLO algorithm according to the present invention, wherein: the method comprises the following steps:

[0006] Step 1: Collect nuclear magnetic resonance images with lesions, and divide the images into a training image set, a test image set and a validation image set;

[0007] Step 2: Construct a target detection network model: The target detection network model is based on an improved YOLO algorithm network architecture. The improved YOLO algorithm network architecture is improved based on the basic framework of YOLOv8. The improved backbone network includes a C2f module, a convolutional module, and an improved multi-channel spatial pyramid pooling module SPPMC to enhance the image feature extraction ability; the improved neck network includes an upsampling layer, a splicing layer, a C2f module, a convolutional layer, and an embedded attention mechanism UECA module for multi-scale feature map fusion to improve the feature fusion efficiency; the improved head network uses an improved bounding box regression loss function L (iS-IoU) , to accelerate the convergence speed of the model and improve the positioning accuracy;

[0008] The improved multi-channel spatial pyramid pooling module SPPMC: First, use 3 convolutional layers to capture the spatial features of the input image and generate feature maps. Subsequently, 4 pooling layers with pooling kernel sizes of 5, 9, 13, and 17 are connected in parallel to the network with a residual network structure. Then, through convolutional layers with convolutional kernel sizes of 1 and 3 in the network, after convolutional operations, feature information is output; the pooling layers and the parallel residual network structure in the network generate receptive fields of sizes 1×1, 5×5, 7×7, 11×11, 15×15, and 19×19, enabling the network to extract more scale image feature information, obtain richer context information, and enhance the network's perception ability. The calculation process of the receptive field is as follows:

[0009] R 0 = 1; R 1 = k 1

[0010]

[0011] wherein, Rn represents the receptive field size of the nth layer of the neural network, k n represents the convolutional kernel or pooling kernel size of the nth layer of the neural network, and s i represents the convolutional stride or pooling stride of the ith layer of the neural network;

[0012] The embedded attention mechanism UECA module is located at the last layer of the improved neck network. It is improved by combining the attention mechanism SE and the attention mechanism CA. It is formed by the weighted parallel connection of the squeeze-and-excitation sub-module of the attention mechanism SE and the coordinate information embedding sub-module of the attention mechanism CA, and can effectively process the feature information in the channel and spatial dimensions. This UECA module is an independent computing unit, and the computing process can be represented by the process of enhancing and converting the input tensor X into the output tensor Y:

[0013] Y = X * (L 1 + L 2 )

[0014]

[0015] Among them, represents the input three-dimensional tensor with the shape of C′×H′×W′; represents the output three-dimensional tensor with the shape of C×H×W; C′, H′, W′ and C, H, W respectively represent the number of channels of the input tensor X and the output tensor Y, the number of pixels of the image in the vertical direction, and the number of pixels of the image in the horizontal direction; L 1 and L 2 respectively represent the outputs of the squeeze-and-excitation sub-module and the coordinate information embedding sub-module;

[0016] For the squeeze-and-excitation sub-module, first, perform the squeeze operation on the input tensor X in the horizontal and vertical spatial dimensions H×W, which is expressed as:

[0017]

[0018] Among them, z represents the result of the channel output after the squeeze operation in the vertical and horizontal spatial dimensions H×W; z C represents the output result of the C-th channel after the squeeze operation;

[0019] Then, capture the channel dependence relationship for the output of the squeeze operation. This process uses a gated mechanism function with a Sigmoid activation. The output result L 1 of this sub-module is expressed as:

[0020] L 1 = σ(T 1 δ(T 2 z))

[0021] Among them, σ is the Sigmoid function; is the linear transformation parameter, which is used to learn and capture the importance of each channel; r represents the reduction ratio that controls the block size; δ represents the ReLU function;

[0022] The coordinate information embedding sub-module first encodes each channel of an arbitrary input tensor X by using pooling kernels along the horizontal coordinate (H, 1) and the vertical coordinate (1, W) respectively to generate feature information z C h (h) and z C w (w), which is expressed as:

[0023] z C h (h) = (1 / W) * ∑ (0≤i≤W) x C (h, i)

[0024] z C w (w) = (1 / H) * ∑ (0≤j≤H) x C (j, w)

[0025] where z C h (h) represents the output of the C-th channel at height h; z C w (w) represents the output of the C-th channel at width w;

[0026] Secondly, the generated feature information is concatenated and subjected to convolution transformation to obtain a feature map f, which is expressed as:

[0027]

[0028] where [z h , z w represents the concatenation operation in the spatial dimension, F 1 represents the 1×1 convolution transformation function, is the non-linear activation function;

[0029] Then, the obtained feature map f is decomposed into f h and f w in the spatial dimensions H and W, and convolution transformations are respectively performed to obtain feature vectors g h and g w , which is expressed as:

[0030] g h = σ(F h (f h ))

[0031] g w = σ(F w (f w ))

[0032] where σ represents the sigmoid function; F hand F w represents a 1×1 convolutional transform in the spatial dimension;

[0033] Finally, for the feature vectors g h and g w perform weighted integration to obtain the output L of the coordinate information embedding sub-module 2 , expressed as:

[0034] L 2= g h *g w

[0035] The improved bounding box regression loss function L (iS-IoU) , uses an auxiliary bounding box to participate in calculating the intersection over union (IoU) of the ground truth bounding box GT and the anchor box. At the same time, the distance loss Δ and the shape loss Ω are introduced into the bounding box regression loss function, which is defined as follows:

[0036] L (iS-IoU) = 1 - IoU in + Δ + Ω

[0037] where, IoU in is the intersection over union of the auxiliary bounding box;

[0038] Step 3: Train the object detection network model: Input the training image set into the object detection network model for training; first, adjust the size of each image in the training image set to be the same, and then perform grid partitioning on each training image. When the center point of the target to be detected exists in the grid of the partition, predict the type and location information of the target to be detected, that is, the pathology; use the test image set and the validation image set to evaluate and test the trained object detection network model;

[0039] Step 4: Image detection: Use the trained object detection network model to perform image lesion detection.

[0040] An image detection system based on an improved YOLO algorithm, which includes:

[0041] A nuclear magnetic resonance image acquisition module, which acquires nuclear magnetic resonance images that need to determine the pathology;

[0042] Target detection network model construction module: The target detection network model is built based on the improved YOLO algorithm network architecture. The improved YOLO algorithm network architecture is improved based on the basic framework of YOLOv8, including: Input part: The input image is processed with adaptive size and adjusted to an RGB format image with a size of 640×640 pixels, and then input into the improved backbone network for further processing; The improved backbone network includes C2f module, convolutional module, and improved multi-channel spatial pyramid pooling module SPPMC to enhance the image feature extraction ability; The improved neck network includes path aggregation network - feature pyramid network PAN-FPN and embedded attention mechanism UECA module for multi-scale feature map fusion to improve the feature fusion efficiency; The improved head network is a prediction output module, which uses the improved bounding box regression loss function L (iS-IoU) , to accelerate the convergence speed and improve the positioning accuracy;

[0043] Lesion detection module: Load the trained target detection network model, input the target image to be detected, and perform lesion detection.

[0044] The above image detection method based on the improved YOLO algorithm, wherein: In step 1, the nuclear magnetic resonance image is labeled using the labeling tool LabelImg, and the labeled nuclear magnetic resonance images are divided into a training image set, a test image set, and a validation image set according to the ratio of 7:2:1.

[0045] The above image detection method based on the improved YOLO algorithm, wherein: In step 2, the IoU in is the intersection over union of the auxiliary bounding box, and its calculation formula is:

[0046] IoU in =(in_inter) / (in_union)

[0047] in_inter=(min(b r gt ,b r )-max(b l gt ,b l ))*(min(b b gt ,b b )-max(b t gt ,b t ))

[0048] in_union=(w gt *h gt )*r 2 +(w*h)*r 2 -in_inter

[0049] b l gt = x c gt -(w gt * r) / 2; b r gt = x c gt +(w gt * r) / 2

[0050] b t gt = y c gt -(h gt * r) / 2; b b gt = y c gt +(h gt * r) / 2

[0051] b l = x c -(w * r) / 2; b r = x c +(w * r) / 2

[0052] b t = y c -(h * r) / 2; b b = y c +(h * r) / 2

[0053] Wherein, b gt is the center point of the GT box and the auxiliary GT box, with coordinates (x c gt , y c gt ); the coordinates of the upper left corner and the lower right corner of the auxiliary GT box are (b l gt , b t gt ) and (b r gt , b b gt ); b are respectively the center points of the anchor box and the auxiliary anchor box, with coordinates (x c , y c ); the coordinates of the upper left corner and the lower right corner of the auxiliary anchor box are respectively (b l , b t ) and (b r , b b ); r is the scale factor; in_inter is the intersection area of the auxiliary GT box and the auxiliary anchor box; in_union is the union area of the auxiliary GT box and the auxiliary anchor box.

[0054] The above image detection method based on the improved YOLO algorithm, where: the r is a scale factor, which is related to the size of the lesion area in the collected nuclear magnetic resonance image.

[0055] The above image detection method based on the improved YOLO algorithm, where: in step 2, the distance loss Δ describes the difference in the central positions of the GT box and the anchor box in the horizontal and vertical directions, and the calculation formula is:

[0056] Δ = hh * (x c - x c gt ) 2 / c 2 + ww * (y c - y c gt ) 2 / c 2

[0057] ww = 2 * (w gt ) s / [(w gt ) s + (h gt ) s

[0058] hh = 2 * (h gt ) s / [(w gt ) s + (h gt ) s

[0059] where ww and hh are the weights in the horizontal and vertical directions respectively, which are related to the size of the GT box; s is the scale factor, and c is the diagonal distance of the smallest closed bounding box between b and b gt .

[0060] The above image detection method based on the improved YOLO algorithm, where: the s is a scale factor, which is related to the size of the lesion area in the collected nuclear magnetic resonance image.

[0061] The above image detection method based on the improved YOLO algorithm, where: in step 2, the shape loss Ω describes the difference in the shape sizes of the GT box and the anchor box, and the calculation formula is:

[0062] Ω = (1 / 2) * [(1 - e wa ) Θ + (1 - e wb ) Θ

[0063] wa = hh * (|w - w​​​gt |) / (max(w, w gt ))

[0064] wb = ww * (|h - h gt |) / (max(h, h gt ))

[0065] Wherein, w, h, and w gt , h gt are respectively the width and height of the GT box and the anchor box, and Θ represents the attention to the shape loss.

[0066] In the above image detection method based on the improved YOLO algorithm, in step 3, the size of each image in the training image set is adjusted to 640×640.

[0067] In the above image detection method based on the improved YOLO algorithm, in step 4, the trained object detection network model is used for image lesion detection: first, the trained object detection network model is loaded, the target image to be detected is input, after obtaining all output candidate detection boxes for lesion detection, non-maximum suppression operation is performed on all output candidate boxes to suppress redundant detection boxes and perform the final output.

[0068] Compared with the prior art, the present invention has obvious beneficial effects. As can be seen from the above solution, the present invention is an object detection algorithm model improved based on YOLOv8. On the basis of the baseline model, a more advanced bounding box regression loss function L (iS-IoU) is introduced, and two innovative modules, namely the spatial pyramid pooling network SPPMC that supports multi-level feature interaction and the unified attention and channel attention network UECA proposed by the present invention, are combined. The SPPMC network inherits the advantages of networks such as SPP and SPPCSPC, and further expands the receptive field of the network. The model can more easily capture and process feature information at more scales, improves the network's ability to represent and process complex data features, and alleviates problems such as mutual interference or degradation of feature information. The UECA network inherits the advantages of the SE and CA networks, takes into account the attention in both the channel dimension and the spatial dimension, expands the perception range of the network, enhances the representation of important features, and promotes the fusion of feature information. The main advantages of the present invention are as follows:

[0069] (1) The multi-channel spatial pyramid pooling structure (SPPMC) that supports multi-level feature interaction in the present invention has a simple structure and can be well transplanted into the object detection model. This structure parallels multiple max-pooling layers and convolutional layers with different sampling rates, enriching the scales of receptive fields in the network. This design enables the network to extract more comprehensive feature information from the input feature map, covering more fine and coarse features. In addition, this structure also adopts a strategy of separately processing and then fusing the extracted global or local feature information, which helps to reduce the interference and information redundancy between global and local information, enhances the model's effective representation ability of the input data features, and improves the model's processing efficiency for complex data.

[0070] (2) The unified embedded attention mechanism (UECA) in the present invention is formed by the weighted parallel connection of a squeeze-and-excitation sub-module and a coordinate information embedding sub-module, comprehensively considering the attention in the channel dimension and the spatial dimension. This design effectively improves the network's ability to obtain target context information and perception range, and helps to better handle long-range dependencies. In addition, this mechanism learns to weightedly adjust the channel and spatial weights, enhances the representation of important features, suppresses the interference of irrelevant features, and makes the network more focused on key feature information.

[0071] (3) Combine SPPMC, UECA with the single-stage object detection network architecture YOLOv8, and optimize the bounding box regression loss function to solve problems such as the unreliable detection of lesions in MRI medical images by the object detection model and the poor robustness of the model.

[0072] In summary, the present invention has the characteristics of effectively improving the robustness and detection performance of the object detection model.

[0073] The beneficial effects of the present invention are further illustrated below through specific embodiments. Brief Description of the Drawings

[0074] Figure 1 is the flow structure diagram of the present invention;

[0075] Figure 2 is the schematic diagram of the object detection network model structure of the present invention;

[0076] Figure 3 is the schematic diagram of the improved multi-channel spatial pyramid pooling module structure of the present invention;

[0077] Figure 4 is the schematic diagram of the embedded attention mechanism UECA module structure of the present invention;

[0078] Figure 5 is the schematic diagram of the calculation principle of the improved bounding box regression loss function L (iS-IoU) of the present invention;

[0079] Figure 6 It is the process structure diagram of the embodiment. Specific implementation manner

[0080] The following combines the accompanying drawings and preferred embodiments to detail the specific implementation manner, features and effects of an image detection method based on an improved YOLO algorithm proposed according to the present invention.

[0081] See Figure 1 , for the image detection method based on the improved YOLO algorithm of the present invention, wherein: the method includes the following steps:

[0082] Step 1: Collect nuclear magnetic resonance images with lesions, and divide the images into a training image set, a test image set and a validation image set; the nuclear magnetic resonance images are labeled using the labeling tool LabelImg. The labeled nuclear magnetic resonance images are divided into a training image set, a test image set and a validation image set according to a ratio of 7:2:1.

[0083] Step 2: Construct a target detection network model: The target detection network model is based on an improved YOLO algorithm network architecture, and the improved YOLO algorithm network architecture is improved based on the basic framework of YOLOv8 (such as Figure 2 ), the improved backbone network includes a C2f module, a convolutional module, and an improved multi-channel spatial pyramid pooling module SPPMC to enhance the image feature extraction ability; the improved neck network includes an upsampling layer, a concatenation layer (Concat), a C2f module, a convolutional layer, and an embedded attention mechanism UECA module for multi-scale feature map fusion to improve the feature fusion efficiency; the improved head network uses an improved bounding box regression loss function L (iS-IoU) to accelerate the convergence speed of the model and improve the positioning accuracy;

[0084] The improved multi-channel spatial pyramid pooling module SPPMC (such as Figure 3 ): First, 3 convolutional layers are used to capture the spatial features of the input image and generate feature maps. Subsequently, 4 pooling layers with pooling kernel sizes of 5, 9, 13, and 17 are connected in parallel to the network together with a residual network structure. Then, convolutional layers with convolutional kernel sizes of 1 and 3 in the network are used for convolutional operations and then output feature information; the pooling layers and the parallel residual network structure in the network generate receptive fields of sizes 1×1, 5×5, 7×7, 11×11, 15×15, and 19×19, enabling the network to extract more scale image feature information, obtain richer context information, and enhance the network's perception ability. The calculation process of the receptive field is as follows:

[0085] R 0 = 1; R1 = k 1

[0086]

[0087] Among them, Rn represents the receptive field size of the nth layer of the neural network, kn represents the size of the convolutional kernel or pooling kernel of the nth layer of the neural network, and s i represents the convolutional stride or pooling stride of the ith layer of the neural network.

[0088] The embedded attention mechanism UECA module (such as Figure 4 ), is located at the last layer and the fifth layer from the bottom of the improved neck network, and is improved by integrating the attention mechanism SE and the attention mechanism CA. It is composed of the squeeze-and-excitation sub-module of the attention mechanism SE and the coordinate information embedding sub-module of the attention mechanism CA in parallel with weighting, and can effectively process the feature information in the channel and spatial dimensions; this UECA module is an independent computing unit, and the computing process can be represented by the process of enhancing and converting the input tensor X into the output tensor Y.

[0089] Y = X * (L 1 + L 2 )

[0090]

[0091] Among them, represents the input three-dimensional tensor with the shape of C′×H′×W′; represents the output three-dimensional tensor with the shape of C×H×W; C′, H′, W′ and C, H, W respectively represent the number of channels of the input tensor X and the output tensor Y, the number of pixels of the image in the vertical direction, and the number of pixels of the image in the horizontal direction. L 1 and L 2 respectively represent the outputs of the squeeze-and-excitation sub-module and the coordinate information embedding sub-module.

[0092] The squeeze-and-excitation sub-module first performs a squeezing operation on the input tensor X in the horizontal and vertical spatial dimensions H×W, which is expressed as:

[0093]

[0094] C where z represents the result of the channel output after the squeezing operation in the vertical and horizontal spatial dimensions H×W; z

[0095] Then, capture the channel dependence relationship for the output of the squeezing operation. This process uses a gated mechanism function with Sigmoid activation, and the output result L 1 of this sub-module is expressed as:

[0096] L 1 = σ(T 1 δ(T 2 z))

[0097] where σ is the Sigmoid function; are linear transformation parameters used to learn and capture the importance of each channel; r represents the reduction ratio that controls the block size; δ represents the ReLU function;

[0098] In the coordinate information embedding sub-module, first, for any input tensor X, the pooling kernel encodes each channel along the horizontal coordinate (H, 1) and the vertical coordinate (1, W) respectively to generate the feature information z C h (h) and z C w (w), expressed as:

[0099] z C h (h) = (1 / W) * ∑ (0≤i≤W) x C (h, i)

[0100] z C w (w) = (1 / H) * ∑ (0≤j≤H) x C (j, w)

[0101] where z C h (h) represents the output of the C-th channel at height h; z C w (w) represents the output of the C-th channel at width w;

[0102] Secondly, the generated feature information is concatenated and then subjected to a convolution transformation to obtain the feature map f, expressed as:

[0103]

[0104] where [z h , z w represents the concatenation operation in the spatial dimension, and F1 represents the 1×1 convolution transformation function; is a non-linear activation function;

[0105] Then, the obtained feature map f is decomposed into f h and f w in the spatial dimensions H and W, and they are respectively subjected to a convolution transformation to obtain the feature vectors g h and g w , expressed as:

[0106] gh = σ(F h (f h ))

[0107] g w = σ(F w (f w ))

[0108] where σ represents the sigmoid function; F h and F w represent the 1×1 convolution transformation in the spatial dimension.

[0109] Finally, the feature vectors g h and g w are weighted and integrated to obtain L 2 , which is expressed as:

[0110] L 2= g h * g w

[0111] The improved bounding box regression loss function L (iS-IoU) , as Figure 5 shown, uses the auxiliary bounding box to participate in calculating the intersection over union (IoU) of the ground truth bounding box (GT box) and the anchor box. At the same time, the distance loss Δ and the shape loss Ω are introduced into the bounding box regression loss function, which is defined as follows:

[0112] L (iS-IoU) = 1 - IoU in + Δ + Ω

[0113] where IoU in is the intersection over union of the auxiliary bounding box, and its calculation formula is:

[0114] IoU in = (in_inter) / (in_union)

[0115] in_inter = (min(b r gt , b r ) - max(b l gt , b l )) * (min(b b gt , b b ) - max(b t gt , b t ))

[0116] in_union = (w gt * h gt ) * r2 +(w * h) * r 2 -in_inter

[0117] b l gt = x c gt -(w gt * r) / 2; b r gt = x c gt +(w gt * r) / 2

[0118] b t gt = y c gt -(h gt * r) / 2; b b gt = y c gt +(h gt * r) / 2

[0119] b l = x c -(w * r) / 2; b r = x c +(w * r) / 2

[0120] b t = y c -(h * r) / 2; b b = y c +(h * r) / 2

[0121] where b gt is the center point of the GT box and the auxiliary GT box, with coordinates (x c gt , y c gt ); the coordinates of the upper left corner and the lower right corner of the auxiliary GT box are (b l gt , b t gt ) and (b r gt , b b gt ); b are the center points of the anchor box and the auxiliary anchor box, with coordinates (x c , y c ); the coordinates of the upper left corner and the lower right corner of the auxiliary anchor box are respectively (b l , b t ) and (b r , b b); r is the scale factor, which is related to the size of the lesion area in the collected nuclear magnetic resonance image; in_inter is the area where the auxiliary GT box and the auxiliary anchor box intersect; in_union is the union area of the auxiliary GT box and the auxiliary anchor box.

[0122] The distance loss Δ describes the difference in the central positions of the GT box and the anchor box in the horizontal and vertical directions, and the calculation formula is:

[0123] Δ = hh * (x c - x c gt ) 2 / c 2 + ww * (y c - y c gt ) 2 / c 2

[0124] ww = 2 * (w gt ) s / [(w gt ) s + (h gt ) s )

[0125] hh = 2 * (h gt ) s / [(w gt ) s + (h gt ) s )

[0126] Among them, ww and hh are the weights in the horizontal and vertical directions respectively, which are related to the size of the GT box; s is the scale factor, which is related to the size of the lesion area in the collected nuclear magnetic resonance image; c is the diagonal distance of the smallest closed bounding box between b and b gt ;

[0127] The shape loss Ω describes the difference in the shape and size of the GT box and the anchor box, and the calculation formula is:

[0128] Ω = (1 / 2) * [(1 - e wa ) Θ + (1 - e wb ) Θ )

[0129] wa = hh * (|w - w gt |) / (max(w, w gt ))

[0130] wb = ww * (|h - h gt |) / (max(h, h gt ))

[0131] Among them, w, h, and w gt , h gt are the widths and heights of the GT box and the anchor box respectively, and Θ represents the attention to the shape loss.

[0132] Step 3: Train the object detection network model: Input the training image set into the object detection network model for training; The training of the object detection network model: First, adjust the size of each image in the training image set to be consistent, and then perform grid partitioning on each training image. When the center point of the target to be detected exists in the grid of the partition, predict the type and location information of the target to be detected, that is, the lesion; Use the test image set and the validation image set to evaluate and test the trained object detection network model; Adjust the size of each image in the training image set to 640×640.

[0133] Step 4: Image detection: Use the trained object detection network model to detect image lesions.

[0134] The use of the trained object detection network model to detect image lesions: First, load the trained object detection network model, input the image of the target to be detected. After obtaining all the output candidate detection boxes for lesion detection, perform non-maximum suppression operation on all the output candidate boxes to suppress redundant detection boxes and perform the final output.

[0135] An image detection system based on an improved YOLO algorithm, where: it includes:

[0136] A nuclear magnetic resonance image acquisition module, which acquires the nuclear magnetic resonance image of the lesion to be determined;

[0137] An object detection network model construction module: The object detection network model is built based on the improved YOLO algorithm network architecture. The improved YOLO algorithm network architecture is improved based on the basic framework of YOLOv8, including: Input part: Perform adaptive size processing on the input image, adjust it to an RGB format image with a size of 640×640 pixels, and input it into the improved backbone network for further processing; The improved backbone network includes a C2f module, 5 convolutional layers, and an improved multi-channel spatial pyramid pooling module SPPMC to enhance the image feature extraction ability and extract feature maps of 80×80, 40×40, and 20×20; The improved neck network includes a path aggregation network - feature pyramid network PAN-FPN and an embedded attention mechanism UECA module to perform multi-scale feature map fusion and improve the feature fusion efficiency; The improved head network is a prediction output module, which uses an improved bounding box regression loss function L (iS-IoU) , to accelerate the convergence speed and improve the positioning accuracy;

[0138] Lesion Detection Module: Load the trained object detection network model, input the target image to be detected, and perform lesion detection.

[0139] Specifically, as Figure 6 shown below, taking the magnetic resonance imaging (MRI) of the heart for heart disease as an example, the working process of the image detection method based on the improved YOLO algorithm is described as follows:

[0140] Step 1: Obtain the patient's heart MRI image. In this invention, the patient's heart MRI image is used as the original dataset, and the size of each image is 640*640. Convert the MRI image in nii.gz format to the JPG format required by the object detection model for the training of the object detection network model.

[0141] Step 2: Image annotation. Use the LabelImg image annotation tool to annotate the heart disease lesions in the image. The datasets used in this invention are all annotated under the guidance of professional doctors. The obtained data is divided into a training set, a validation set, and a test set according to the ratio of 7:2:1 for the training and evaluation of the object detection network model.

[0142] Step 3: Construction of the object detection network model: Based on the basic framework of the improved YOLOv8, build a neural network model for heart disease lesion detection. As Figure 2 shown, the improvements mainly include using the multi-channel spatial pyramid pooling module SPPMC in the last layer of the backbone network, which enhances the network's ability to extract multi-scale features. The fused attention mechanism UECA module is used in the neck network, which improves the network's attention ability to key feature information and enhances the feature fusion efficiency. In addition, the object detection network model uses the improved bounding box regression loss function L (iS-IoU) to optimize the training process of the model and improve the positioning accuracy.

[0143] (1) The improved backbone network is built using the C2f module, the convolutional module, and the improved multi-channel spatial pyramid pooling module SPPMC for feature extraction of the heart MRI image to obtain a shared feature map. The backbone network is specifically composed of 5 convolutional layers, 4 C2f modules, and 1 SPPMC module, and finally outputs 3 types of feature maps with scales of 80×80, 40×40, and 20×20 for subsequent feature enhancement and fusion in the neck network. The SPPMC module is a spatial pyramid pooling network module with stronger perception and adaptation capabilities. The key idea is that this module enhances the network's ability to capture multi-scale feature information and its perception ability by expanding the receptive field size of the network model.

[0144] As Figure 3 shown, the overall architecture of the SPPMC module mainly includes the following steps:

[0145] First, convolutional layers with convolution kernels of 1, 3, and 1 respectively are used to obtain the spatial features (such as edges, textures, etc.) of the input image and generate feature maps. Subsequently, the feature maps output by the convolutional layers will be processed by 4 max-pooling layers with different sizes in parallel to extract the feature information of the image at different scales. Then, the network module splices and fuses the multi-scale feature information extracted by the pooling layer with the feature information extracted by the residual network processing, and inputs it into convolutional layers with convolution kernels of 1 and 3 respectively for further feature extraction, fusion, and feature map dimensionality reduction, etc., and converts the obtained feature information into a fixed-size feature vector to facilitate the subsequent classification and regression tasks.

[0146] The pooling layer and the parallel residual network structure in the network generate receptive fields with sizes of 1×1, 5×5, 7×7, 11×11, 15×15, and 19×19, enabling the network to extract more scale of image feature information, obtain richer context information, and enhance the network's perception ability. The calculation process of the receptive field is as follows:

[0147] R 0 = 1; R 1 = k 1

[0148]

[0149] where Rn represents the receptive field size of the nth layer of the neural network, k n represents the size of the convolution kernel or pooling kernel of the nth layer of the neural network, and s i represents the convolution stride or pooling stride of the ith layer of the neural network.

[0150] (2) The improved neck network consists of 2 upsampling layers (Upsample), 4 concatenation layers (Concat), 4 C2f modules, 2 convolutional layers, and 2 fusion attention mechanism UECA modules. Among them, the sampling layer, concatenation layer, C2f module, and convolutional layer are the original structures of YOLOv8, and the UECA module is a newly added and improved module, which is placed in the last layer and the fifth layer from the bottom of the neck network to increase the network's processing ability for the feature information of the effective processing channels and spatial dimensions, better focus on the key feature information, and promote the effective fusion of features.

[0151] As Figure 4 shown, the overall architecture of the UECA module mainly includes the following steps:

[0152] After the UECA module obtains the input information, it will first be separately processed by the squeeze-and-excitation sub-module and the coordinate information embedding sub-module with weights, and then the information processed by the above two sub-modules will be weighted and fused. The calculation process can be represented by converting the input tensor X to the output tensor Y:

[0153] Y = X * (L 1 + L 2 )

[0154]

[0155] where represents the input three - dimensional tensor with shape C'×H'×W'; represents the output three - dimensional tensor with shape C×H×W; C', H', W' and C, H, W represent the number of channels of the input tensor X and the output tensor Y, the number of pixels of the image in the vertical direction, and the number of pixels of the image in the horizontal direction respectively. L 1 and L 2 represent the outputs of the squeeze - excitation sub - module and the coordinate - information embedding sub - module respectively.

[0156] a. Squeeze - excitation sub - module:

[0157] First, a squeeze operation on the input tensor X is performed in the horizontal and vertical spatial dimensions H×W:

[0158]

[0159] where z represents the result of the channel output after the squeeze operation in the vertical and horizontal spatial dimensions H×W; z C represents the output result of the C - th channel after the squeeze operation.

[0160] Then, to capture the channel - dependence relationship of the squeeze - operation output, a gated - mechanism function with Sigmoid activation is used, and the output result L 1 of this sub - module is expressed as:

[0161] L 1 = σ(T 1 δ(T 2 z))

[0162] where σ is the Sigmoid function; is the linear - transformation parameter used to learn and capture the importance of each channel; r represents the reduction ratio controlling the block size; δ represents the ReLU function.

[0163] b. Coordinate - information embedding sub - module:

[0164] First, for any input tensor X, the pooling kernel encodes each channel along the horizontal coordinate (H, 1) and the vertical coordinate (1, W) respectively to generate the feature information z C h (h) and z C w (w), which is expressed as:

[0165] z C h (h) = (1 / W) * ∑ (0≤i≤W) x C (h, i)

[0166] z C w (w) = (1 / H) * ∑ (0≤j≤H) x C (j, w)

[0167] where z C h (h) represents the output of the C-th channel at height h; z C w (w) represents the output of the C-th channel at width w;

[0168] Secondly, the generated feature information is concatenated and subjected to convolutional transformation to obtain the feature map f:

[0169]

[0170] where [z h , z w represents the concatenation operation in the spatial dimension, and F1 represents the 1×1 convolutional transformation function; is the non-linear activation function;

[0171] Then, the obtained feature map f is decomposed into f h and f w in the spatial dimensions H and W, and convolutional transformations are respectively performed to obtain the feature vectors g h and g w :

[0172] g h = σ(F h (f h ))

[0173] g w = σ(F w (f w ))

[0174] where σ represents the sigmoid function; F h and F w represent the 1×1 convolutional transformation in the spatial dimension.

[0175] Finally, the feature vectors g h and g w are weighted and integrated to obtain L 2 :

[0176] L 2= gh *g w

[0177] (3) The improved object detection network model uses an improved bounding box regression loss function L (iS-IoU) to optimize the training process of the model, accelerate convergence, and improve the accuracy of object detection. The bounding box regression loss function L (iS-IoU) is calculated through the following steps:

[0178] The loss function L (iS-IoU) , by using the Figure 5 auxiliary bounding box shown in the figure, participates in calculating the intersection over union (IoU) of the ground truth (GT) box and the anchor box, and considers the influence brought by the distance loss and the shape loss. The calculation of the loss function L (iS-IoU) is as follows:

[0179] L (iS-IoU) = 1 - IoU in + Δ + Ω

[0180] where IoU in is the intersection over union of the auxiliary bounding box, Δ is the distance loss; Ω is the shape loss.

[0181] a. The intersection over union IoU of the auxiliary bounding box in is mainly used to assist in measuring the overlap degree between the predicted bounding box and the true bounding box. The higher its value, the more accurate the prediction:

[0182] IoU in = (in_inter) / (in_union)

[0183] in_inter = (min(b r gt , b r ) - max(b l gt , b l )) * (min(b b gt , b b ) - max(b t gt , b t ))

[0184] in_union = (w gt * h gt ) * r 2 + (w * h) * r 2 - in_inter

[0185] b l gt = x c gt - (wgt *r) / 2; b r gt = x c gt +(w gt *r) / 2

[0186] b t gt = y c gt -(h gt *r) / 2; b b gt = y c gt +(h gt *r) / 2

[0187] b l = x c -(w * r) / 2; b r = x c +(w * r) / 2

[0188] b t = y c -(h * r) / 2; b b = y c +(h * r) / 2

[0189] Among them, b gt is the center point of the GT box and the auxiliary GT box, with coordinates (x c gt , y c gt ); the coordinates of the upper left corner and the lower right corner of the auxiliary GT box are (b l gt , b t gt ) and (b r gt , b b gt ); b are respectively the center points of the anchor box and the auxiliary anchor box, with coordinates (x c , y c ); the coordinates of the upper left corner and the lower right corner of the auxiliary anchor box are respectively (b l , b t ) and (b r , b b ); r is the scale factor, which is related to the size of the lesion area in the collected nuclear magnetic resonance images, in_inter is the intersection area of the auxiliary GT box and the auxiliary anchor box; in_union is the union area of the auxiliary GT box and the auxiliary anchor box.

[0190] b. The distance loss Δ is mainly used to reduce the impact caused by the difference in the central positions of the GT box and the anchor box in the horizontal and vertical directions:

[0191] Δ = hh * (x c - x c gt ) 2 / c 2 + ww * (y c - y c gt ) 2 / c 2

[0192] ww = 2 * (w gt ) s / [(w gt ) s + (h gt ) s )

[0193] hh = 2 * (h gt ) s / [(w gt ) s + (h gt ) s )

[0194] Among them, ww and hh are the weights in the horizontal and vertical directions respectively, related to the size of the GT box; s is the scale factor, related to the size of the lesion area in the collected nuclear magnetic resonance images; c is the diagonal distance of the smallest closed bounding box between b and b gt ;

[0195] c. The shape loss Ω is mainly used to weaken the impact caused by the difference in the shape and size of the GT box and the anchor box:

[0196] Ω = (1 / 2) * [(1 - e wa ) Θ + (1 - e wb ) Θ )

[0197] wa = hh * (|w - w gt |) / (max(w, w gt ))

[0198] wb = ww * (|h - h gt |) / (max(h, h gt ))

[0199] Among them, w, h and w gt , h gt are the widths and heights of the GT box and the anchor box respectively, and Θ represents the attention degree to the shape loss

[0200] Step 4: Model training. Import the training set data divided in Step 2 into the improved YOLO neural network for training. Do not use the pre-trained model during training. Set the number of training iterations to 200 rounds and the batch size to 8. After training, select the training model with the best performance in the 200 rounds of iterations for the next evaluation.

[0201] Step 5: Model verification. Use the validation dataset obtained in Step 2 to compare the anchor boxes identified by the improved YOLO model with the anchor boxes labeled under the guidance of professional doctors. The neural network model generates corresponding reports based on the results of the comparison and verification. Metrics such as precision (P), recall (R), F1-score, average precision (AP), mean average precision (mAP), and floating-point operations (GFLOPs) can be used to evaluate the overall performance of the model. The calculation methods for the relevant metrics are as follows:

[0202] P = TP / (TP + FP)

[0203] R = TP / (TP + FN)

[0204] F1-score = 2 * (P * R) / (P + R)

[0205] AP = the area enclosed by the P-R curve

[0206]

[0207] Among them, TP represents true positives, which refers to the number of samples correctly predicted as positive classes by the model; FP represents false positives, which refers to the number of samples incorrectly predicted as positive classes by the model; FN represents false negatives, which refers to the number of samples incorrectly predicted as negative classes by the model; N is the number of types of heart diseases in the dataset.

[0208] Compared with directly incorporating the data into the YOLOv8 neural network for training, the present invention has better detection performance for various types of heart diseases.

[0209] Step 6: Heart disease detection. Use the trained optimal model to detect the heart MRI images of patients, and identify the heart disease lesions in the MRI images in the form of rectangular boxes. It can assist radiologists in analyzing heart MRI images and improve their diagnostic efficiency. This method improves the accuracy of the object detection network model for heart disease detection, and the values of various evaluation metrics are better than those of the original YOLOv8 model.

[0210] The above description is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. An image detection method based on an improved YOLO algorithm, characterized in that: The method comprises the following steps: Step 1: Collect MRI images with lesions and divide the images into a training image set, a test image set, and a validation image set; Step 2: Construct a target detection network model: The target detection network model is based on the improved YOLO algorithm network architecture. The improved YOLO algorithm network architecture is improved based on the basic framework of YOLOv8. The improved backbone network includes a C2f module, a convolution module, and an improved multi-channel spatial pyramid pooling module SPPMC to enhance the image feature extraction capability; the improved neck network includes an upsampling layer, a splicing layer, a C2f module, a convolution layer, and an embedded attention mechanism UECA module to perform multi-scale feature map fusion and improve feature fusion efficiency; the improved head network adopts an improved bounding box regression loss function L (iS-IoU) , speed up the convergence speed of the model and improve the positioning accuracy; The improved multi-channel spatial pyramid pooling module SPPMC: firstly uses three convolutional layers to capture the spatial features of the input image and generate a feature map, then connects four pooling layers with pooling kernel sizes of 5, 9, 13, and 17 respectively to the network in parallel with a residual network structure, and then outputs feature information after convolution operation through convolutional layers with convolution kernel sizes of 1 and 3 respectively in the network; the pooling layers in the network and the parallel residual network structure generate receptive fields of sizes of 1×1, 5×5, 7×7, 11×11, 15×15, and 19×19, so that the network can extract image feature information of more scales, obtain richer context information, and enhance the perception ability of the network. The calculation process of the receptive field is as follows: R0=1;R1=k1 Among them, Rn represents the receptive field size of the nth layer of the neural network, k n Represents the size of the convolution kernel or pooling kernel of the nth layer of the neural network, s i Represents the convolution step or pooling step of the i-th layer neural network; The embedded attention mechanism UECA module is located in the last layer of the improved neck network, and is improved by combining the attention mechanism SE and the attention mechanism CA. It is composed of the weighted parallel connection of the squeeze excitation submodule of the attention mechanism SE and the coordinate information embedding submodule of the attention mechanism CA, and effectively processes the feature information of the channel and spatial dimensions; the UECA module is an independent computing unit, and the computing process can be represented by the process of enhancing the conversion of the input tensor X to the output tensor Y: Y=X*(L1+L2) in, Represents the input 3D tensor of shape C′×H′×W′; represents the output three-dimensional tensor of shape C×H×W; C′, H′, W′ and C, H, W represent the number of channels of input tensor X and output tensor Y, the number of pixels of the image in the vertical direction, and the number of pixels of the image in the horizontal direction respectively; L1 and L2 represent the outputs of the squeeze excitation submodule and the coordinate information embedding submodule respectively; The squeeze excitation submodule first performs a squeeze operation of the horizontal and vertical spatial dimensions H×W on the input tensor X, which is expressed as: Where z represents the result of the channel output after the vertical and horizontal spatial dimensions H×W are squeezed; z C Indicates the output result of the Cth channel after the squeeze operation; Then, the channel dependency is captured on the output of the squeeze operation. This process uses a gating mechanism function with Sigmoid activation. The output result L1 of this submodule is expressed as: L1=σ(T1δ(T2 z)) Among them, σ is the Sigmoid function; is a linear transformation parameter used to learn and capture the importance of each channel; r represents the reduction ratio of the control block size; δ represents the ReLU function; The coordinate information embedding submodule first encodes each channel of any input tensor X along the horizontal coordinate (H, 1) and the vertical coordinate (1, W) to generate feature information z C h (h) and z C w (w), expressed as: z C h (h)=(1 / W)*∑ (0≤i≤W) x C (h,i) z C w (w)=(1 / H)*∑ (0≤j≤H) x C (j,w) Among them, z C h (h) represents the output of the Cth channel at height h; z C w (w) represents the output of the Cth channel with width w; Secondly, the generated feature information is connected and convolution is performed to obtain the feature map f, which is expressed as: Among them, [z h ,z w ] represents the splicing operation of the spatial dimension, F1 represents the 1×1 convolution transformation function, is a nonlinear activation function; Then, the obtained feature map f is decomposed into f in the spatial dimensions H and W h and f w , and perform convolution transformation respectively to obtain the feature vector g h and g w , expressed as: g h =σ(F h (f h )) g w =σ(F w (f w )) Where σ represents the sigmoid function; F h and F w Represents a 1×1 convolution transformation on the spatial dimension; Finally, for the eigenvector g h and g w The output L2 of the coordinate information embedding submodule is obtained by weighted integration, which is expressed as: L2=g h *g w The improved bounding box regression loss function L (iS-IoU) , use the auxiliary bounding box to participate in the calculation of the intersection-over-union ratio of the real bounding box GT and the anchor box, and introduce the distance loss Δ and shape loss Ω into the bounding box regression loss function, which is defined as follows: L (iS-IoU) =1-IoU in +D+O Among them, IoU in is the intersection-over-union ratio of the auxiliary bounding box; Step 3: training the target detection network model: input the training image set into the target detection network model for training; firstly adjust the size of each image in the training image set to be consistent, then divide each training image into grid blocks, and when there is a center point of the target to be detected in the grid of the block, predict the type and location information of the target to be detected, i.e., the pathology; use the test image set and the verification image set to evaluate and verify the trained target detection network model; Step 4: Image detection: Use the trained target detection network model to perform image lesion detection.

2. The image detection method based on the improved YOLO algorithm as claimed in claim 1, characterized in that: In step 1, the nuclear magnetic resonance image is labeled using a labeling tool LabelImg, and the labeled nuclear magnetic resonance image is divided into a training image set, a test image set, and a verification image set in a ratio of 7:2:

1.

3. The image detection method based on the improved YOLO algorithm as claimed in claim 1, characterized in that: In step 2, the IoU in is the intersection-over-union ratio of the auxiliary bounding box, and its calculation formula is: IoU in =(in_inter) / (in_union) in_inter=(min(b r gt ,b r )-max(b l gt ,b l ))*(min(b b gt ,b b )-max(b t gt ,b t )) in_union=(w gt *h gt )*r 2 +(w*h)*r 2 -in_inter b l gt =x c gt -(w gt *r) / 2;b r gt =x c gt +(w gt *r) / 2 b t gt =y c gt -(h gt *r) / 2;b b gt =y c gt +(h gt *r) / 2 b l =x c -(w*r) / 2;b r =x c +(w*r) / 2 b t =y c -(h*r) / 2;b b =y c +(h*r) / 2 Among them, b gt It is the center point of the GT frame and the auxiliary GT frame, with coordinates (x c gt ,y c gt ); the coordinates of the upper left corner and lower right corner of the auxiliary GT box are (b l gt , b t gt ) and (b r gt , b b gt ); b are the center points of the anchor box and the auxiliary anchor box, with coordinates (x c ,y c ); the coordinates of the upper left corner and lower right corner of the auxiliary anchor box are (b l , b t ) and (b r , b b ), r is the scaling factor, in_inter is the area of ​​the intersection of the auxiliary GT box and the auxiliary anchor box, and in_union is the area of ​​the union of the auxiliary GT box and the auxiliary anchor box.

4. The image detection method based on the improved YOLO algorithm as claimed in claim 3, characterized in that: The r is a scaling factor related to the size of the lesion area in the collected MRI images.

5. The image detection method based on the improved YOLO algorithm as claimed in claim 1, characterized in that: In step 2, the distance loss Δ describes the difference in the center position of the GT box and the anchor box in the horizontal and vertical directions, and the calculation formula is: Δ=hh*(x c -x c gt ) 2 / c 2 +ww*(y c -y c gt ) 2 / c 2 ww=2*(w gt ) s / [(w gt ) s +(h gt ) s ] hh=2*(h gt ) s / [(w gt ) s +(h gt ) s ] Among them, ww and hh are the weights in the horizontal and vertical directions, respectively, which are related to the size of the GT box; s is the scale factor, and c is the ratio of b and b gt The diagonal distance of the minimum enclosing bounding box between .

6. The image detection method based on the improved YOLO algorithm as claimed in claim 5, characterized in that: The s is a scale factor, which is related to the size of the lesion area in the collected nuclear magnetic resonance image.

7. The image detection method based on the improved YOLO algorithm as claimed in claim 1, characterized in that: In step 2, the shape loss Ω describes the difference in shape size between the GT box and the anchor box, and the calculation formula is: Ω=(1 / 2)*[(1-e wa ) Θ +(1-e wb ) Θ ] wa=hh*(|ww gt |) / (max(w,w gt )) wb=ww*(|h-h gt |) / (max(h,h gt )) Among them, w, h and w gt 、h gt are the width and height of the GT box and the anchor box respectively, and Θ represents the attention paid to the shape loss.

8. The image detection method based on the improved YOLO algorithm as claimed in claim 1, characterized in that: In step 3, the size of each image in the training image set is adjusted to 640×640.

9. The image detection method based on the improved YOLO algorithm as claimed in claim 1, characterized in that: The method uses a trained target detection network model to perform image lesion detection: first load the trained target detection network model, input the target image to be detected, and after obtaining all output candidate detection frames for lesion detection, perform non-maximum suppression operation on all output candidate frames to suppress redundant detection frames and perform final output.

10. An image detection system based on an improved YOLO algorithm, characterized in that: Using the method according to any one of claims 1 to 9, comprising: A nuclear magnetic resonance image acquisition module, which acquires the nuclear magnetic resonance image required for pathology determination; Target detection network model building module: The target detection network model is built based on the improved YOLO algorithm network architecture, which is improved based on the basic framework of YOLOv8, including: input part: the input image is adaptively resized, adjusted to an RGB format image of 640×640 pixels, and input into the improved backbone network for further processing; the improved backbone network includes a C2f module, a convolution module, and an improved multi-channel spatial pyramid pooling module SPPMC to enhance the image feature extraction capability; the improved neck network includes a path aggregation network-feature pyramid network PAN-FPN and an embedded attention mechanism UECA module to perform multi-scale feature map fusion and improve feature fusion efficiency; the improved head network is a prediction output module, which uses an improved bounding box regression loss function L (iS-IoU) , speed up the convergence speed and improve the positioning accuracy; Lesion detection module: load the trained target detection network model, input the target image to be detected, and perform lesion detection.

Citation Information

Patent Citations

  • Chip defect detection method based on improved YOLOv3 model

    CN117274775A

  • Remote sensing image target detection method based on attention mechanism weighted feature fusion

    CN117611994A

  • Space target identification method

    CN118230118A

  • Colorectal polyp detection method, device and equipment and storage medium

    CN118333942A

  • Sewage biological phase target automatic identification detection and instance segmentation method

    CN118351535A

Cited By

  • Pedestrian clothing identification method, device and equipment based on improved YOLOv8 model and storage medium

    CN120656043A

  • Endoscope image focus detection method

    CN120931602A

  • Electric power target detection method based on state space model and high-frequency information enhancement

    CN120931897A

  • A Power Target Detection Method Based on State-Space Model and High-Frequency Information Enhancement

    CN120931897B

  • Defect detection method and device based on improved YOLO model

    CN121169795A