Brain MRI image tumor detection method based on improved YOLOv7
By improving the YOLOv7 model, using partial convolution, three-dimensional spatial attention mechanism and dynamic attention loss function, the problems of low feature extraction accuracy and data imbalance in brain tumor detection are solved, and higher detection accuracy and speed are achieved.
Patent Information
- Application Number
- CN202510120534.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-25
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art has problems in brain tumor detection with low feature extraction accuracy, challenges in processing complexity and diversity, and data imbalances lead to decreased detection accuracy.
Improved YOLOv7 model by optimizing convolution operations using partial convolution (PConv), introducing three-dimensional spatial attention mechanism (SimAM) and dynamic attention BBR loss function (WIoU) to improve the accuracy of feature extraction and bounding box regression.
The accuracy and speed of brain tumor detection are improved, especially when dealing with tumors with high similarity and complex morphology, and the detection performance of the model is significantly improved.
Smart Images

Figure CN119963535A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image processing and target detection, and in particular relates to a brain MRI image tumor detection method based on improved YOLOv7. Background Art
[0002] Brain tumors are abnormal tissue growths caused by uncontrolled cell proliferation, and these tissues have no physiological function in the brain. The presence of tumors not only increases the volume and pressure of the brain, but can also cause swelling, which can trigger various neurological symptoms. Magnetic resonance imaging (MRI) imaging detection of brain tumors has important clinical significance, which can improve the diagnostic efficiency of brain tumors, optimize resource allocation, and promote the progress of medical research and clinical practice. As a widely used brain imaging technology, MRI has high soft tissue resolution and multi-plane imaging capabilities, which can provide rich information about brain tissue and help locate tumors more accurately.
[0003] However, the complexity and diversity of brain tumor images make tumor detection and identification challenging. Common types of brain tumors, such as gliomas, meningiomas, and pituitary tumors, have different manifestations in MRI images, which increases the difficulty of detection. Traditional brain tumor detection methods mostly rely on manually designed features and combine them with classifiers (such as KNN, SVM, etc.) for differentiation. These methods improve detection accuracy through steps such as image preprocessing and feature extraction, but still have certain limitations when dealing with complex texture, edge, and internal structure features.
[0004] With the rapid development of deep learning technology, detection methods based on deep learning have gradually become mainstream. Deep learning methods can automatically learn feature representations from image data, avoiding the limitations of traditional manual feature design. As an efficient target detection algorithm, the YOLO series of algorithms are widely used in medical image detection because they can provide high accuracy and speed in a single-stage detection framework. Although the YOLOv7 algorithm performs well in target detection, it still faces the need to further improve accuracy and speed when facing complex detection tasks of brain tumor images. The existing technology has the following shortcomings: 1) Feature extraction problem to deal with similarities between brain tumors: When processing brain tumor images, the existing technology, especially when dealing with tumor types with high similarity such as glioma and normal brain tissue, the traditional convolution operation fails to effectively distinguish the subtle differences between different tissues. Due to the blurred boundary between tumor and normal tissue, the traditional convolution network cannot specifically adjust the convolution kernel action area when extracting features, resulting in low feature extraction accuracy, which in turn affects the detection results.
[0005] 2) The complexity and diversity of brain tumors pose challenges to feature extraction: The diversity of brain tumors in morphology, location, and edges increases the difficulty of detection. Traditional feature extraction methods, especially when dealing with tumors with complex morphology and irregular boundaries, often fail to effectively capture the key features of tumors. For example, meningiomas and pituitary tumors may have very regular shapes and are closely related to surrounding brain tissue, making them difficult to distinguish from other tissues.
[0006] 3) The detection accuracy decreases due to data imbalance: In the training data of brain tumors, the image signals of small tumors (such as microadenomas) are weak and easily missed. Larger tumors are usually easier to detect, which may cause the model to be biased towards detecting significant large tumors and ignore those small tumors that are difficult to identify. Therefore, when faced with unbalanced data sets, existing technologies often have low detection accuracy, especially in the detection of small tumors. These shortcomings indicate that there is still room for improvement in the accuracy, speed and ability to process complex features of brain tumor detection in existing methods, and more advanced and sophisticated algorithms are needed to overcome these challenges. Therefore, it is necessary to propose a brain MRI image tumor detection method based on improved YOLOv7 to solve the above problems. Summary of the invention
[0007] The technical problem to be solved by the present invention is to provide a brain MRI image tumor detection method based on improved YOLOv7, aiming to solve the problem of feature extraction of similarities between brain tumors, the problem of feature extraction of complexity and diversity of brain tumors, and the problem of decreased detection accuracy caused by data imbalance, so as to further improve the detection accuracy and speed of the model.
[0008] In order to achieve the above technical effects, the technical solution adopted by the present invention is: A brain MRI image tumor detection method based on improved YOLOv7, comprising the following steps: S1, collect brain MRI images as raw data, preprocess the collected raw data and input them into the original YOLOv7 model; S2, convolution optimization of the original YOLOv7 model: Use partial convolution PConv to optimize computational efficiency and extract image features from preprocessed data; replace traditional convolutional layers with partial convolutional layers, perform convolution operations only on valid areas, and implement forward and backward propagation calculations of partial convolution to improve computational efficiency; make the optimized convolutional layers compatible with other layers in the model, and enable smooth training and inference; S3, design and implement a three-dimensional spatial attention mechanism to process the extracted features, combining the attention in the spatial and channel dimensions, and integrate the SimAM module in the specific layer of the YOLOv7 model after convolution optimization in step S2 for feature selection and weighting; S4, design and implement the dynamic attention BBR loss function WIoU, combined with the features processed in step S2, for bounding box regression; dynamically adjust the loss weight according to the similarity between the anchor box and the target box to improve the accuracy of bounding box regression; S5, select the final detection result by non-maximum suppression method; S6, evaluates the performance of the improved YOLOv7 model.
[0009] Preferably, the data preprocessing method in step S1 is: The brain MRI images were preprocessed, including image standardization and normalization.
[0010] Furthermore, since deep learning models usually require input images to have a fixed size, the model needs to resize the MRI image to the input size of 640×640 pixels required by the model before input.
[0011] Furthermore, in step S2, the model uses a simple PConv partial convolution technique, in which partial convolution adjusts the action area of the convolution kernel according to the validity of the data, and only performs convolution on the valid part of each convolution window; unlike conventional convolution, convolution is only applied to part of the input features for spatial feature extraction, while other channels remain unchanged. In this way, partial convolution can effectively reduce the amount of calculation, thereby reducing the overall detection time of the model and improving the speed.
[0012] Preferably, the convolution calculation delay time Delay is determined by the calculation amount and the calculation speed, as shown in the following formula: ; In the formula, It is the number of floating point calculations, including multiplication and addition, which is only related to the model itself; It is the number of floating-point operations per second, i.e., the computing speed, which is affected by the number of memory accesses; The model proposes to use partial convolution PConv to reduce the amount of calculation without reducing the calculation speed, thereby reducing the calculation time to improve the detection speed of the entire model.
[0013] Preferably, for partial convolution, only a few input channels The convolution kernel is applied to the remaining channels while keeping them unchanged; partial convolution for: ; The memory access amount is: ; This reduces the total number of calculations without increasing the amount of memory access. H is the vertical dimension of the image,W is the horizontal dimension of the image, K is the size of the convolution kernel, c is the number of output channels; The ordinary convolution is replaced by partial convolution plus 1x1 convolution, residual is added to prevent gradient explosion and promote model convergence, and a BN layer and activation function are added between two 1x1 convolutions to maintain the feature diversity of the feature map while reducing complexity.
[0014] Preferably, in step S3, when processing the extracted features, a three-dimensional spatial attention mechanism is used, and attention is paid to the spatial domain and channel domain of the feature map at the same time, so as to quickly find features that need to be focused on in tumors with diverse and complex shapes, sizes, and locations.
[0015] Furthermore, in order to simultaneously focus on the changing attention of channels and space and reduce the complexity of the structure, a three-dimensional attention weight is assigned to each neuron without adding additional parameters; the neurological spatial inhibition principle shows that the neurons with the richest information are usually those that show unique discharge patterns to surrounding neurons, and active neurons may also inhibit the activity of surrounding neurons; based on this principle, the model uses the three-dimensional spatial attention mechanism SimAM to minimize the energy function of each neuron, find the linear separability of the target neuron and other neurons, and assign high-weighted attention to neurons that show obvious spatial inhibition effects; based on neurological and mathematical theories, the energy function of each neuron is defined as the following formula to estimate the importance of each neuron: ; in is the number of neurons on the channel, Represents input features Target neurons in a single channel, Represents input features Other neurons in a single channel, is the spatial dimension index. and yes and The linear transformation of is shown in the formula: and are the weights and biases of the transformation; ; right and Using binary labels, -1 and 1 respectively; we can minimize: ; Use the above formula to find the same channel and Linear separability between: when ,all When the energy of the target neuron is Get the minimum value; In order to prevent overfitting, Adding the regularization expression yields the following: ; pass and The quick analytical formula simplifies the above formula: ; in and The expression is as follows: ; and are the mean and variance of all neurons except the target neuron on the same channel. Since the fast analytical formula is calculated on a single channel, it is assumed that the neurons on the same channel are independent and identically distributed. Therefore, the mean of all neurons can be calculated. and variance The calculation formula is: ; Simplifying the above formula can get the simplest form of neuron energy As follows: ; Since the linear separability of each neuron from other neurons is negatively correlated with the energy of the neuron, the importance of each neuron is considered to be 1 / ; The single neuron attention expression of the three-dimensional spatial attention mechanism is: ; in is the input feature, is the output feature, It's neuronal energy.
[0016] Furthermore, the above formula can quickly and accurately obtain three-dimensional spatial attention and extract more important features from complex brain tumor images, thereby improving the overall detection accuracy of the model.
[0017] Preferably, in step S4, in the target detection task, the BBR loss function has a great influence on the accuracy of the final result, and the ideal loss function should satisfy that when the anchor box and the actual target box tend to overlap, the amplitude of the gradient converges to zero.
[0018] DIoU is a distance-based extension of IoU that considers the center point distance between detection boxes and better captures the spatial relationship between target detection boxes. DIoU can more accurately measure the overlap between two detection boxes than traditional IoU. CIoU is an improvement on DIoU. In addition to considering the center point distance, it also considers the difference in aspect ratio and length-to-width ratio, which can more accurately describe the overlap between target detection boxes and perform better when dealing with targets of different shapes and sizes. SIoU is a complex bounding box regression method that solves the limitations of traditional loss functions by integrating angle considerations and scale sensitivity. It can more comprehensively consider factors such as the position, angle, and shape of the target, thereby achieving significant improvements in prediction accuracy. However, in the bounding box regression task, due to the sparsity of the target, there is a problem of unbalanced training instances. The geometric factors introduced by the previous methods, such as distance and aspect ratio, will increase the penalty for low-quality data, produce excessively large gradients, and be harmful to the network training process. In order not to affect the overall training results of the network and solve the problem of imbalanced training instances, examples of normal quality should be given greater attention and assigned larger gradients, while outlier examples with high outliers and high-quality examples with small outliers should be assigned small gradient gains.
[0019] Preferably, the gradient gain of the anchor box in Focal-EIoU changes with the curve, and its gradient change trend is as follows: ; in, is a parameter that controls the degree of outlier suppression; It consists of three parts: IoU loss, distance loss and height-width loss. Although this method can adjust the attention to examples of different qualities, it is a static attention and ignores the change in the quality of the anchor box, that is, the quality distribution of the anchor box changes dynamically in comparison with each other.
[0020] Use dynamic attention WIoU to make the model focus on anchor boxes of normal quality, quantify the quality distribution of anchor box changes, and give different gradients to anchor boxes of different qualities; Construct a penalty term to reduce the regression penalty of geometric factors such as distance and aspect ratio on low-quality data, and weaken the penalty of geometric factors when the anchor box and the target box overlap well: ; The distance attention mechanism loss function multiplied by is shown as follows: ; exist middle, and is the width and height of the smallest enclosing boundary, ( , )and( , ) are the center coordinates of the anchor box and the target box, respectively. In order to speed up the overall convergence of the model and reduce the penalty of geometric factors on low-quality examples, no new metrics such as aspect ratio are introduced here; middle, [0,1], used to reduce the high quality anchorbox , reduce the attention paid to it; [1, e), used to significantly enlarge the normal quality anchor box , increase its attention; In order to dynamically represent the outlier degree of the anchor box, construct , and introduce the gradient gain parameter As shown in the formula, and : ; For low-quality anchor boxes with high outliers and high-quality anchor boxes with small outliers, a smaller gradient gain is assigned to them, while a larger gradient gain is assigned to the ordinary quality anchor box at the intermediate level, so as to focus the work of bounding box regression on the ordinary quality anchor box. and Multiplying together gives the following formula: ; Where r represents the gradient gain parameter; When detecting brain tumors with diverse morphologies and complex structures, it can dynamically allocate attention, reduce the impact of low-quality data on the results, and effectively improve the accuracy of the test results.
[0021] Preferably, in step S5, non-maximum suppression (NMS) is a technique commonly used in post-processing of target detection results, which is used to remove overlapping prediction boxes and retain the most representative target box; selecting the final detection result by the non-maximum suppression method specifically includes: NMS sorts all the boxes according to the confidence of each candidate box and selects the box with the highest confidence as the current detection box; Calculate the intersection over union (IoU) between the current box and other boxes. If the IoU exceeds the preset threshold, the two boxes are considered redundant and the boxes with overlap higher than the set threshold are removed. Repeat this process until there are no remaining boxes. Output target boxes whose confidence is higher than the set threshold and whose overlap with other boxes is less than the set threshold; through NMS, redundant boxes can be effectively removed to reduce false positives and false negatives in brain tumor detection results.
[0022] Preferably, step S6 comprises: The efficiency of model detection is evaluated from the two aspects of accuracy and speed. The accuracy is expressed as the average accuracy of each category. The mean To measure, the speed is measured in FPS, the number of images processed per second; The expression is shown as follows: ; represents the total number of categories, Indicates The average precision of each category depends on the precision under different thresholds. and recall : ; Indicates how many of the samples predicted by the model as positive examples are true positive examples, which is the proportion of objects actually contained in the bounding box predicted by the model: ; in, (True Positives) is the number of samples correctly predicted as positive examples, (False Positives) is the number of samples that are incorrectly predicted as positive examples; It indicates the proportion of positive samples that the model can correctly detect, which is the proportion of real objects that the model successfully detects: ; (False Negatives) indicates the number of samples that were not correctly predicted as positive examples.
[0023] Further, The model's detection performance on different categories is comprehensively considered, and the comprehensive performance of the model on the entire dataset is given. It can comprehensively evaluate the overall performance of the model in the target detection task, rather than focusing on the performance of a single category. The inference speed metric measures the processing speed of the model on a given hardware, usually in terms of the number of images processed per second. In practical applications, higher This means that the model can process images faster and achieve real-time object detection.
[0024] The beneficial effects of the present invention are as follows: 1. Aiming at the needs of brain MRI image tumor detection, the present invention makes multiple improvements on the basis of the YOLOv7 model. First, in view of the similarity of brain tumors, it is proposed to use PConv in feature extraction to eliminate redundant calculations and improve the detection speed; secondly, in view of the complexity and diversity of brain tumors, the SimAM three-dimensional spatial attention mechanism is introduced into the feature extraction process to enhance the focus on important features. Finally, in view of the imbalance problem of brain tumor data, WIoU is used in the BBR stage to increase the focus on the anchor box of ordinary quality; through experiments, it is found that compared with the original YOLOv7 model, the mAP value and FPS value of the improved model on two public data sets are improved, which proves that the improved method of the present invention can effectively improve the accuracy and speed of brain tumor detection; through ablation experiments, it is found that when PConv, SimAM and WIoU are added at the same time, although the overall detection accuracy and speed of the model are improved, the speed is lower than that of adding only PConv.
[0025] 2. The present invention proposes to use partial convolution (PConv) to replace the traditional convolution operation for feature extraction in brain tumor images. By adjusting the area of action of the convolution kernel according to the validity of the data, only the more effective part of the input data is convolved, thereby avoiding excessive processing of invalid information. This method can improve the sensitivity to subtle differences, enhance the model's detection accuracy and speed for different tumor types and morphologies, and solve the problem of feature extraction of similarities between brain tumors.
[0026] 3. In order to solve the problem of complexity and diversity of brain tumor features, the present invention introduces a three-dimensional spatial attention mechanism into the YOLOv7 model, which focuses on the attention of both the spatial domain and the channel domain, and can adaptively extract key information from brain tumor images. By paying more attention to neurons with rich information in the feature map, the network can more effectively extract discriminative tumor features, and improve the accuracy of the model in processing complex and morphologically variable tumors.
[0027] 4. In order to solve the problem of data imbalance and the decrease in detection accuracy, the present invention proposes to use a dynamic attention loss function in the bounding box regression (BBR) process. By dynamically adjusting the gradient gain, the bounding boxes of image data of different qualities can be reasonably weighted. This method reduces the negative impact of low-quality bounding boxes on model training, thereby improving the detection capability of brain tumors and improving the performance of the overall model. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a schematic diagram of a flow chart of the present invention; Figure 2 is a schematic diagram of partial convolution in an embodiment of the present invention; Figure 32 is a schematic diagram of optimizing the use of partial convolution in an embodiment of the present invention; Figure 4 is a three-dimensional spatial attention map in an embodiment of the present invention; Figure 5 is a diagram of an anchor frame and an actual target frame in an embodiment of the present invention; Figure 6 is a gradient distribution trend diagram in an embodiment of the present invention; Figure 7 is a schematic diagram of brain tumors in an embodiment of the present invention, from left to right respectively pituitary tumor, glioma and meningioma; Figure 8 It is a schematic diagram of the detection effect of different models in the embodiments of the present invention on various types of brain tumors. DETAILED DESCRIPTION
[0029] Embodiment 1: like Figure 1 As shown, a brain MRI image tumor detection method based on improved YOLOv7 includes the following steps: S1, collect brain MRI images as raw data, preprocess the collected raw data and input them into the original YOLOv7 model; S2, convolution optimization of the original YOLOv7 model: Use partial convolution PConv to optimize computational efficiency and extract image features from preprocessed data; replace traditional convolutional layers with partial convolutional layers, perform convolution operations only on valid areas, and implement forward and backward propagation calculations of partial convolution to improve computational efficiency; make the optimized convolutional layers compatible with other layers in the model, and enable smooth training and inference; S3, design and implement a three-dimensional spatial attention mechanism to process the extracted features, combining the attention in the spatial and channel dimensions, and integrate the SimAM module in the specific layer of the YOLOv7 model after convolution optimization in step S2 for feature selection and weighting; S4, design and implement the dynamic attention BBR loss function WIoU, combined with the features processed in step S2, for bounding box regression; dynamically adjust the loss weight according to the similarity between the anchor box and the target box to improve the accuracy of bounding box regression; S5, select the final detection result by non-maximum suppression method; S6, evaluates the performance of the improved YOLOv7 model.
[0030] Preferably, the data preprocessing method in step S1 is: The brain MRI images were preprocessed, including image standardization and normalization.
[0031] Furthermore, since deep learning models usually require input images to have a fixed size, the model needs to resize the MRI image to the input size of 640×640 pixels required by the model before input.
[0032] like Figure 2 As shown, further, in step S2, the model uses a simple PConv partial convolution technique, in which the partial convolution adjusts the action area of the convolution kernel according to the validity of the data, and only convolves the valid part of each convolution window; unlike conventional convolution, convolution is only applied to part of the input features for spatial feature extraction, while other channels remain unchanged. In this way, partial convolution can effectively reduce the amount of calculation, thereby reducing the overall detection time of the model and improving the speed.
[0033] like Figure 3 As shown, preferably, the convolution calculation delay time Delay is determined by the calculation amount and the calculation speed, as shown in the following formula: ; In the formula, It is the number of floating point calculations, including multiplication and addition, which is only related to the model itself; It is the number of floating-point operations per second, i.e., the computing speed, which is affected by the number of memory accesses; The model proposes to use partial convolution PConv to reduce the amount of calculation without reducing the calculation speed, thereby reducing the calculation time to improve the detection speed of the entire model.
[0034] Preferably, for partial convolution, only a few input channels The convolution kernel is applied to the remaining channels while keeping them unchanged; partial convolution for: ; The memory access amount is: ; This reduces the total number of calculations without increasing the amount of memory access. H is the vertical dimension of the image, W is the horizontal dimension of the image, K is the size of the convolution kernel, c is the number of output channels; The ordinary convolution is replaced by partial convolution plus 1x1 convolution, residual is added to prevent gradient explosion and promote model convergence, and a BN layer and activation function are added between two 1x1 convolutions to maintain the feature diversity of the feature map while reducing complexity.
[0035] like Figure 4As shown, preferably, in step S3, when processing the extracted features, a three-dimensional spatial attention mechanism is adopted, and the spatial domain and channel domain attention of the feature map are paid attention to simultaneously, so as to quickly find the features that need to be focused on in tumors with diverse and complex shapes, sizes, and locations.
[0036] Furthermore, in order to simultaneously focus on the changing attention of channels and space and reduce the complexity of the structure, a three-dimensional attention weight is assigned to each neuron without adding additional parameters; the neurological spatial inhibition principle shows that the neurons with the richest information are usually those that show unique discharge patterns to surrounding neurons, and active neurons may also inhibit the activity of surrounding neurons; based on this principle, the model uses the three-dimensional spatial attention mechanism SimAM to minimize the energy function of each neuron, find the linear separability of the target neuron and other neurons, and assign high-weighted attention to neurons that show obvious spatial inhibition effects; based on neurological and mathematical theories, the energy function of each neuron is defined as the following formula to estimate the importance of each neuron: ; in is the number of neurons on the channel, Represents input features Target neurons in a single channel, Represents input features Other neurons in a single channel, is the spatial dimension index. and yes and The linear transformation of is shown in the formula: and are the weights and biases of the transformation; ; right and Using binary labels, -1 and 1 respectively; we can minimize: ; Use the above formula to find the same channel and Linear separability between: when ,all When the energy of the target neuron is Get the minimum value; In order to prevent overfitting, Adding the regularization expression yields the following: ; pass and The quick analytical formula simplifies the above formula: ; in and The expression is as follows: ; and are the mean and variance of all neurons except the target neuron on the same channel. Since the fast analytical formula is calculated on a single channel, it is assumed that the neurons on the same channel are independent and identically distributed. Therefore, the mean of all neurons can be calculated. and variance The calculation formula is: ; Simplifying the above formula can get the simplest form of neuron energy As follows: ; Since the linear separability of each neuron from other neurons is negatively correlated with the energy of the neuron, the importance of each neuron is considered to be 1 / ; The single neuron attention expression of the three-dimensional spatial attention mechanism is: ; in is the input feature, is the output feature, It's neuronal energy.
[0037] Furthermore, the above formula can quickly and accurately obtain three-dimensional spatial attention and extract more important features from complex brain tumor images, thereby improving the overall detection accuracy of the model.
[0038] like Figure 5 As shown, preferably, in step S4, in the target detection task, the BBR loss function has a great influence on the accuracy of the final result, and the ideal loss function should satisfy that when the anchor box and the actual target box tend to overlap, the amplitude of the gradient converges to zero.
[0039] DIoU is a distance-based extension of IoU that considers the center point distance between detection boxes and better captures the spatial relationship between target detection boxes. DIoU can more accurately measure the overlap between two detection boxes than traditional IoU. CIoU is an improvement on DIoU. In addition to considering the center point distance, it also considers the difference in aspect ratio and length-to-width ratio, which can more accurately describe the overlap between target detection boxes and perform better when dealing with targets of different shapes and sizes. SIoU is a complex bounding box regression method that solves the limitations of traditional loss functions by integrating angle considerations and scale sensitivity. It can more comprehensively consider factors such as the position, angle, and shape of the target, thereby achieving significant improvements in prediction accuracy. However, in the bounding box regression task, due to the sparsity of the target, there is a problem of unbalanced training instances. The geometric factors introduced by the previous methods, such as distance and aspect ratio, will increase the penalty for low-quality data, produce excessively large gradients, and be harmful to the network training process. In order not to affect the overall training results of the network and solve the problem of imbalanced training instances, examples of normal quality should be given greater attention and assigned larger gradients, while outlier examples with high outliers and high-quality examples with small outliers should be assigned small gradient gains.
[0040] like Figure 6 As shown in the figure, the gradient gain of the anchor box in Focal-EIoU changes with the change of the gradient trend as shown below: ; in, is a parameter that controls the degree of outlier suppression; It consists of three parts: IoU loss, distance loss and height-width loss. Although this method can adjust the attention to examples of different qualities, it is a static attention and ignores the change in the quality of the anchor box, that is, the quality distribution of the anchor box changes dynamically in comparison with each other.
[0041] Preferably, use dynamic attention WIoU to make the model focus on anchor boxes of normal quality, quantify the quality distribution of anchor box changes, and give different gradients to anchor boxes of different qualities; Construct a penalty term to reduce the regression penalty of geometric factors such as distance and aspect ratio on low-quality data, and weaken the penalty of geometric factors when the anchor box and the target box overlap well: ; The distance attention mechanism loss function multiplied by is shown as follows: ; exist middle, and is the width and height of the smallest enclosing boundary, ( , )and( , ) are the center coordinates of the anchor box and the target box, respectively. In order to speed up the overall convergence of the model and reduce the penalty of geometric factors on low-quality examples, no new metrics such as aspect ratio are introduced here; middle, [0,1], used to reduce the high quality anchorbox , reduce the attention paid to it; [1, e), used to significantly enlarge the normal quality anchor box , increase its attention; In order to dynamically represent the outlier degree of the anchor box, construct , and introduce the gradient gain parameter As shown in the formula, and : ; For low-quality anchor boxes with high outliers and high-quality anchor boxes with small outliers, a smaller gradient gain is assigned to them, while a larger gradient gain is assigned to the ordinary quality anchor box at the intermediate level, so as to focus the work of bounding box regression on the ordinary quality anchor box. and Multiplying together gives the following formula: ; Where r represents the gradient gain parameter; When detecting brain tumors with diverse morphologies and complex structures, it can dynamically allocate attention, reduce the impact of low-quality data on the results, and effectively improve the accuracy of the test results.
[0042] Preferably, in step S5, non-maximum suppression (NMS) is a technique commonly used in post-processing of target detection results, which is used to remove overlapping prediction boxes and retain the most representative target box; selecting the final detection result by the non-maximum suppression method specifically includes: NMS sorts all the boxes according to the confidence of each candidate box and selects the box with the highest confidence as the current detection box; Calculate the intersection over union (IoU) between the current box and other boxes. If the IoU exceeds the preset threshold, the two boxes are considered redundant and the boxes with overlap higher than the set threshold are removed. Repeat this process until there are no remaining boxes. Output target boxes whose confidence is higher than the set threshold and whose overlap with other boxes is less than the set threshold; through NMS, redundant boxes can be effectively removed to reduce false positives and false negatives in brain tumor detection results.
[0043] Preferably, step S6 comprises: The efficiency of model detection is evaluated from the two aspects of accuracy and speed. The accuracy is expressed as the average accuracy of each category. The mean To measure, the speed is measured in FPS, the number of images processed per second; The expression is shown as follows: ; represents the total number of categories, Indicates The average precision of each category depends on the precision under different thresholds. and recall : ; Indicates how many of the samples predicted by the model as positive examples are true positive examples, which is the proportion of objects actually contained in the bounding box predicted by the model: ; in, (True Positives) is the number of samples correctly predicted as positive examples, (False Positives) is the number of samples that are incorrectly predicted as positive examples; It indicates the proportion of positive samples that the model can correctly detect, which is the proportion of real objects that the model successfully detects: ; (False Negatives) indicates the number of samples that were not correctly predicted as positive examples.
[0044] Further, The model's detection performance on different categories is comprehensively considered, and the comprehensive performance of the model on the entire dataset is given. It can comprehensively evaluate the overall performance of the model in the target detection task, rather than focusing on the performance of a single category. The inference speed metric measures the processing speed of the model on a given hardware, usually in terms of the number of images processed per second. In practical applications, higher This means that the model can process images faster and achieve real-time object detection.
[0045] Embodiment 2: In order to test the robustness of the model while testing the speed and accuracy of the model, two public datasets were selected for the experiment: Brain_Tumor from Insure Chain and Glioma_of_test from Roboflow. Brain_Tumor contains 2161 brain MRI images with a resolution of 640×640 pixels, which can be divided into 4 categories: "glioma", "meningioma", "pituitary tumor" and "no tumor", and the corresponding number of images in each category is 492, 483, 587, and 599 respectively. Gliomaof test contains 2338 brain MRI images with a resolution of 640×640 pixels, and the categories are the same as Brain_Tumor, and the corresponding number of images in each category is 760, 555, 405, and 618 respectively. In order to train and verify the improved model under the Pytorch framework, the label information is annotated in txt format, and Brain_Tumor is divided into a training set containing 1878 images and a validation set of 188 images, and Glioma_of_test is divided into a training set containing 1392 images and a validation set of 478 images.
[0046] After verification on the data set, the comparison between the proposed model and other mainstream models is obtained. Value and The values are shown in Table 1 and Table 2.
[0047] Table 1: Experimental results comparing the model of this embodiment with other models;
[0048] Table 2: Ablation experiment results;
[0049] From Table 1, Table 2 and Figure 7 , Figure 8 It can be seen that the model of this embodiment is on Brain_Tumor Value and The values were 96.9% and 162.7 respectively; Value and The values are 92.8% and 158.1 respectively, which shows a certain improvement in the accuracy and speed of brain tumor detection compared with other models.
[0050] The randomly selected original image and the test result image can more intuitively show the effect of the test. The improved model can detect glioma, meningioma, pituitary tumor and no tumor (the whole brain area) with a higher confidence. At the same time, the model can eliminate the error interference of meningioma and glioma and detect pituitary tumor results. The average detection accuracy is significantly higher than that of the original model. In summary, the improved model has obvious improvements in the accuracy and speed of brain tumor detection and has comprehensive advantages.
[0051] In order to further reflect the effectiveness of each improvement point in improving the overall model efficiency, only a single improvement point PConv, SimAM, WIoU and a combination of any two of the three improvement points were added to the original YOLOv7 model, and experiments were conducted on Brain_Tumor. Each improvement method has different degrees of improvement. Adding PConv can significantly improve the FPS index and greatly improve the speed of the model. In terms of the mAP index, the model with SimAM and WIoU added performs better. When PConv and other improvement points are added to the original model at the same time, the detection speed and accuracy of the model are improved, which reflects the effectiveness of each improvement point.
Claims
1. A brain MRI image tumor detection method based on improved YOLOv7, characterized in that: The following steps are involved: S1, collect brain MRI images as raw data, preprocess the collected raw data and input them into the original YOLOv7 model; S2, convolution optimization of the original YOLOv7 model: Use partial convolution PConv to optimize computational efficiency and extract image features from preprocessed data; replace traditional convolutional layers with partial convolutional layers, perform convolution operations only on valid areas, and implement forward and backward propagation calculations of partial convolutions; make the optimized convolutional layers compatible with other layers in the model, and enable smooth training and reasoning; S3, design and implement a three-dimensional spatial attention mechanism to process the extracted features, combining the attention in the spatial and channel dimensions, and integrate the SimAM module in the specific layer of the YOLOv7 model after convolution optimization in step S2 for feature selection and weighting; S4, design and implement the dynamic attention BBR loss function WIoU, combined with the features processed in step S2, for bounding box regression; dynamically adjust the loss weight according to the similarity between the anchor box and the target box to improve the accuracy of bounding box regression; S5, select the final detection result by non-maximum suppression method; S6, evaluates the performance of the improved YOLOv7 model.
2. According to claim 1, a brain MRI image tumor detection method based on improved YOLOv7 is characterized in that: The data preprocessing method in step S1 is: The brain MRI images were preprocessed, including image standardization and normalization.
3. The method for detecting brain MRI tumors based on improved YOLOv7 according to claim 1, characterized in that: In step S2, the convolution calculation delay time Delay is determined by the calculation amount and calculation speed, as shown in the following formula: ; In the formula, It is the number of floating point calculations, including multiplication and addition, which is only related to the model itself; It is the number of floating-point operations per second, i.e., the computing speed, which is affected by the number of memory accesses; Partial convolution for: ; The memory access amount is: ; This reduces the total number of calculations without increasing the amount of memory access. H is the vertical dimension of the image, W is the horizontal dimension of the image, K is the size of the convolution kernel, c is the number of output channels; The ordinary convolution is replaced by partial convolution plus 1x1 convolution, residual is added to prevent gradient explosion and promote model convergence, and a BN layer and activation function are added between two 1x1 convolutions to maintain the feature diversity of the feature map while reducing complexity.
4. The method for detecting brain MRI tumors based on improved YOLOv7 according to claim 1, characterized in that: In step S3, when processing the extracted features, a three-dimensional spatial attention mechanism is used to simultaneously focus on the spatial domain and channel domain attention of the feature map, so as to quickly find the features that need to be focused on in tumors with diverse and complex shapes, sizes, and locations; Without adding additional parameters, a three-dimensional attention weight is assigned to each neuron to simultaneously focus on the changes in channel and space and reduce the complexity of the structure. The three-dimensional spatial attention mechanism SimAM is used to minimize the energy function of each neuron, find the linear separability of the target neuron and other neurons, and assign high-weighted attention to neurons that show obvious spatial inhibition effects. Based on neurological and mathematical theories, the energy function of each neuron is defined as the following formula to estimate the importance of each neuron: ; in is the number of neurons on the channel, Represents input features Target neurons in a single channel, Represents input features Other neurons in a single channel, is the spatial dimension index; in and yes and The linear transformation is shown in the formula: ; In the formula, and are the weights and biases of the transformation.
5. The method for detecting brain MRI tumors based on improved YOLOv7 according to claim 4, characterized in that: Find the same channel by minimizing and Linear separability between: ; Among them and Use binary labels, -1 and 1 respectively; when ,all When the energy of the target neuron is Get the minimum value.
6. The method for detecting brain MRI tumors based on improved YOLOv7 according to claim 5, characterized in that: In order to prevent overfitting, Adding the regularization expression yields the following: ; pass and The quick analytical formula simplifies the above formula: ; in and The expression is as follows: ; and are the mean and variance of all neurons except the target neuron on the same channel. Since the fast analytical formula is calculated on a single channel, it is assumed that the neurons on the same channel are independent and identically distributed. Therefore, the mean of all neurons can be calculated. and variance The calculation formula is: ; Simplifying the above formula can get the simplest form of neuron energy As follows: ; Since the linear separability of each neuron from other neurons is negatively correlated with the energy of the neuron, the importance of each neuron is considered to be 1 / ; The attention expression of a single neuron in the three-dimensional spatial attention mechanism is: ; in is the input feature, is the output feature, It's neuronal energy.
7. The method for detecting brain MRI tumors based on improved YOLOv7 according to claim 1, characterized in that: In step S4, in the target detection task, the gradient gain of the anchor box in Focal-EIoU changes with the change curve, and its gradient change trend is as follows: ; in, is a parameter that controls the degree of outlier suppression; It consists of three parts: IoU loss, distance loss and height-width loss; Using dynamic attention WIoU makes the model focus on anchor boxes of common quality, quantifies the quality distribution of anchor box changes, and gives different gradients to anchor boxes of different quality.
8. The method for detecting brain MRI tumors based on improved YOLOv7 according to claim 7, characterized in that: It also includes constructing penalty items to reduce the regression penalty of geometric factors such as distance and aspect ratio on low-quality data, and weaken the penalty of geometric factors when the anchor box coincides well with the target box: ; The distance attention mechanism loss function multiplied by is shown as follows: ; exist middle, and is the width and height of the smallest enclosing boundary, ( , )and( , ) are the center coordinates of the anchor box and the target box respectively; middle, [0,1], used to reduce the high quality anchor box , reduce the attention paid to it; [1, e), used to significantly enlarge the normal quality anchor box , increase its attention; In order to dynamically represent the outlier degree of the anchor box, construct , and introduce the gradient gain parameter As shown in the formula, and : ; For low-quality anchor boxes with high outliers and high-quality anchor boxes with small outliers, a smaller gradient gain is assigned to them, while a larger gradient gain is assigned to the middle-level ordinary-quality anchor boxes, so as to focus the work of bounding box regression on the ordinary-quality anchor boxes. and Multiplying together gives the following formula: ; Where r represents the gradient gain parameter.
9. The method for detecting brain MRI tumors based on improved YOLOv7 according to claim 1, characterized in that: In step S5, selecting the final detection result by the non-maximum suppression method specifically includes: NMS sorts all the boxes according to the confidence of each candidate box and selects the box with the highest confidence as the current detection box; Calculate the intersection over union (IoU) between the current box and other boxes. If the IoU exceeds the preset threshold, the two boxes are considered redundant and the boxes with overlap higher than the set threshold are removed. Repeat this process until there are no remaining boxes. Output target boxes whose confidence is higher than the set threshold and whose overlap with other boxes is less than the set threshold.
10. The method for detecting brain MRI tumors based on improved YOLOv7 according to claim 1, characterized in that: Step S6 includes: The efficiency of model detection is evaluated from the two aspects of accuracy and speed. The accuracy is expressed as the average accuracy of each category. The mean To measure, the speed is measured in FPS, the number of images processed per second; The expression is as shown in the formula: ; represents the total number of categories, Indicates The average precision of each category depends on the precision under different thresholds. and recall : ; Indicates how many of the samples predicted by the model as positive examples are true positive examples, which is the proportion of objects actually contained in the bounding box predicted by the model: ; in, is the number of samples correctly predicted as positive examples, is the number of samples that are incorrectly predicted as positive examples; It indicates the proportion of positive samples that the model can correctly detect, which is the proportion of real objects that the model successfully detects: ; Indicates the number of samples that were not correctly predicted as positive examples.
Citation Information
Patent Citations
Lightweight multi-environment tomato detection method based on improved yolov8
CN117557787A
Unmanned aerial vehicle remote sensing small target detection method based on improved YOLOv7
CN118411634A
YOLOv8 unmanned aerial vehicle image target detection method based on optimization and improvement
CN118711089A