A method and system for detecting brain tumors in nuclear magnetic resonance images
By improving the YOLOv8 model and combining DivC2f, PVSS and RP-LSKM networks, the accuracy and efficiency of brain tumor detection in nuclear magnetic images are improved, and the problem of low detection accuracy and efficiency in existing MRI technologies is solved.
Patent Information
- Application Number
- CN202510552928.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-29
AI Technical Summary
Existing MRI technology has problems of accuracy and low detection efficiency in brain tumor detection.
Using the improved YOLOv8 model, by constructing a dual convolutional cross-stage division (DivC2f) network and a partial visual state space (PVSS) network, combined with the reparameterized dynamic, large-core convolution (RP-LSKM) network can be selected to improve the model's performance and feature extraction capabilities, iterative optimization training, and output the brain tumor detection results of nuclear magnetic images.
It improves the accuracy and efficiency of brain tumor detection, can conduct context modeling in channel dimensions and global space, and enhances the model's state space modeling and global representation learning ability.
Smart Images

Figure CN120070455B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and in particular, to a method and system for detecting brain tumors in nuclear magnetic resonance (MRI) images. Background Art
[0002] Accurate detection of nuclear magnetic resonance (MRI) images has become the gold standard for screening brain diseases due to its superior imaging characteristics. Compared with imaging technologies such as CT, MRI has higher contrast in soft tissue imaging, can clearly show the boundary between tumors and normal tissues, enabling doctors to accurately judge the nature, size, and location of tumors. Accurately detecting brain tumors in MRI images is of great significance in improving the early diagnosis rate, optimizing treatment plans, promoting personalized medicine, and clinical research.
[0003] Although existing MRI technologies have many advantages in brain tumor detection, there are still problems of low detection accuracy and efficiency. Summary of the Invention
[0004] The present invention aims to at least improve one of the technical problems existing in the prior art. For this purpose, the present invention proposes a method and system for detecting brain tumors in MRI images.
[0005] The technical solution of the present invention is as follows:
[0006] A method for detecting brain tumors in MRI images, which includes:
[0007] S1, obtaining an initial image dataset of nuclear magnetic resonance brain tumors, performing annotation and dividing it into a training set, a validation set, and a test set;
[0008] S2, constructing a Dual Convolution Cross Stage Division (DivC2f) network and a Partial Visual State Space (PVSS) network, using the Dual Convolution Cross Stage Division (DivC2f) network as the backbone network in the YOLOv8 model, using the Partial Visual State Space (PVSS) network as the feature enhancement network after SPPF in the YOLOv8 model, and adjusting the backbone network of the YOLOv8 model to lightweight the YOLOv8 model while improving performance; the Dual Convolution Cross Stage Division network includes pointwise convolution, batch normalization, SiLU activation function, ReLU activation function, feature segmentation, feature splicing, channel attention layer (SElayer) processing, and a Feature Interaction Dual Division (DivOpera) network that performs pointwise division of feature signals;
[0009] S3. In the YOLOv8 model with improved performance, in the neck structure, the reparameterized dynamic selectable large kernel convolution (RP-LSKM) network constructed is used as a multi-scale feature interaction network, and at the same time, reparameterization is utilized to further improve the performance of the YOLOv8 model to obtain an improved YOLOv8 model;
[0010] S4. Use the initial image dataset to perform iterative optimization training on the improved YOLOv8 model to obtain a nuclear magnetic resonance (NMR) image brain tumor detection model;
[0011] S5. Input the initial image dataset into the NMR image brain tumor detection model for detection, and output the NMR brain tumor detection result.
[0012] In a possible technical solution, further, in S2, the processing process of the double convolutional cross-stage division network includes:
[0013] The input features pass through a separable convolution module composed of a pointwise convolution with a convolution kernel of 1, batch normalization, and a SiLU activation function to evenly divide the input features into feature map A and feature map B in the channel dimension, and the feature map A is processed by a channel attention layer (SE layer) to obtain high-value channel signal features;
[0014] A three-convolution module is constructed, and the feature map B enters the three-convolution modules in their respective branches through two branches in parallel to control the range of the kernel function and perform low-dimensional representation and integration of the feature signals while integrating the input features;
[0015] The feature signals in the first branch of the two branches are respectively processed by a ReLU activation function and a Sigmoid function, and the feature signals in the second branch are mirrored;
[0016] The second branch uses the output of the first branch after being processed by the ReLU activation function as the dividend and the output of the second branch after being processed by the Sigmoid activation function as the divisor, and divides the results of the two branches to obtain a first output result, and mirrors the first branch in the double-stream division network to obtain a second output result;
[0017] The first output result, the second output result, and the high-value channel signal features are concatenated in the channel dimension to obtain an output channel signal;
[0018] The output channel signal is convolved and then output.
[0019] In a possible technical solution, further, in S2, the processing of the feature map A by the channel attention layer to obtain high-value channel signal features is specifically:
[0020] Receive the feature map A and perform global average pooling processing to compress and integrate it to obtain global features;
[0021] Pass the global features through a weighted fusion module to interact the channel signals of the global features and perform weighted fusion. The weighted fusion module consists of a pointwise convolution with a kernel size of 1, a batch normalization regularization function, a SiLU activation function, and a pointwise convolution with a kernel size of 1;
[0022] Apply the HardSigmoid activation function to the channel signals after weighted fusion to obtain the high-value channel signal feature weights;
[0023] Multiply the high-value channel signal feature weights with the input features to extract the high-value channel signal features.
[0024] In a possible technical solution, further, the processing process of the partial vision state space network includes:
[0025] Receive the features processed by the double convolutional cross-stage division network, obtain the output features after passing through a convolutional layer with a kernel size of 1, and pass the output features through the residual network block in the partial vision state space network to obtain the fused multi-scale features, where the residual network block includes a Mamba layer and a multi-layer perceptron layer;
[0026] Concatenate the multi-scale features with the output features, and then output the enhanced output features after passing through a convolutional layer with a kernel size of 1.
[0027] In a possible technical solution, further, in S3, the processing process of the reparameterized dynamic selectable large kernel convolution includes:
[0028] Receive the feature map output in S2 to the image patch embedding module (PatchEmbed) to obtain the signal features after expanding the number of channels. Specifically, in the image patch embedding module, expand the number of channels of the input feature map to the number of hidden layer channels through a pointwise convolutional layer and a depth convolutional layer and integrate to obtain the signal features after expanding the number of channels;
[0029] Pass the signal features after expanding the number of channels through n cascaded random projection local semantic kernel modules (RP-LSKBlock), where RP-LSKBlock contains a network layer with a large selective kernel based on the receptive field pyramid (RP-LSKNet layer) and a depth separable multi-layer perceptron layer (DwMlp layer) to enhance the feature expression ability of the model.
[0030] In a possible technical solution, further, the process of enabling the signal features after the extended channel number to enter the RP-LSKNet layer includes:
[0031] The signal features after the extended channel number are concatenated and passed through a pointwise convolution with a kernel size of 1 and a GELU activation function for non-linear integration;
[0032] The signal features after non-linear integration are passed through a depthwise convolution composed of n parallel convolution kernels with a kernel size of 5 to output feature map C;
[0033] Feature map C is passed through a dilated convolution composed of n parallel convolution kernels with a kernel size of 7 and a dilation rate of 3 to output feature map D, which is used to dynamically expand the receptive field of the convolution;
[0034] Feature map C and feature map D are compressed through a pointwise convolution to obtain the compressed feature map C and the compressed feature map D, where the number of compressed channels is half of the input;
[0035] The compressed feature map C and feature map D are first concatenated to obtain the first concatenated feature map;
[0036] The first concatenated feature map is passed through an average pooling layer and a max pooling layer in parallel to output their respective pooling results, and a second Concat (concatenation) is performed according to their respective pooling results to obtain the second concatenated feature map, where both average pooling and max pooling are performed in the channel dimension;
[0037] The second concatenated feature map is passed through a depthwise convolution with a kernel size of 7 and a Sigmoid activation function to generate a heat map,
[0038] The heat map is evenly divided in the channel dimension, each part is multiplied by the compressed feature map C and feature map D respectively, and the results after multiplication are added together to complete the signal integration of different receptive fields at the same spatial position under different granularity information;
[0039] The heat map after signal integration of different receptive fields is passed through a pointwise convolution with a kernel size of 1 for signal fusion, and the fusion result is multiplied by the signal features after the extended channel number to complete the screening of high-value regions.
[0040] According to the nuclear magnetic resonance (NMR) image brain tumor detection method of the present invention, by improving the C2f structural block in the YOLOv8n model, a double-convolution cross-stage division idea is proposed, aiming to enhance the modeling ability of potential high-dimensional feature information, increase the attention to the semantic information of high-value channels, and further extract the semantic information of high-value channels in the subsequent division block. And to avoid NaN in division, the two divisor branches are first processed by the Sigmoid function to avoid the situation of divisor being 0; the large-kernel convolution in the large-kernel dynamic selectable convolution structural block is improved by the reparameterization technology of the reparameterized dynamic selectable large-kernel convolution network to improve the information extraction ability; the PVSS module enhances the model state space modeling and model global representation learning ability at a lower computational cost, and the linear self-attention mechanism used avoids the excessive overhead brought by the quadratic computational complexity of self-attention. Thus, the improved YOLOv8n model proposed in this application can not only perform context modeling at different scales in the channel dimension but also perform context modeling in the global space representation. The NMR image brain tumor detection model based on the improved YOLOv8n model can improve the accuracy and efficiency of brain tumor detection.
[0041] A nuclear magnetic resonance (NMR) image brain tumor detection system, which is used to implement the detection method as described above, includes:
[0042] An acquisition module, which is used to acquire the initial image dataset of the NMR brain tumor, perform annotation and divide it into a training set, a validation set, and a test set;
[0043] A first construction module, which is used to construct a double-convolution cross-stage division (DivC2f) network and a partial visual state space (PVSS) network. The double-convolution cross-stage division (DivC2f) network is used as the backbone network in the YOLOv8 model, and the partial visual state space (PVSS) network is used as the feature enhancement network after the SPPF in the YOLOv8 model, and the backbone network of the YOLOv8 model is adjusted to lightweight the YOLOv8 model while improving performance; the double-convolution cross-stage division network includes Pointwise convolution, batch normalization, SiLU activation function, ReLU activation function, feature segmentation, feature splicing, channel attention layer (SElayer) processing, and a feature interaction dual-stream division (DivOpera) network that divides the feature signals pointwise;
[0044] A second construction module, which is used to use the constructed reparameterized dynamic selectable large-kernel convolution (RP-LSKM) network as a multi-scale feature interaction network in the neck structure of the YOLOv8 model with improved performance, and at the same time use reparameterization to further improve the performance of the YOLOv8 model to obtain an improved YOLOv8 model;
[0045] A training module for iteratively optimizing and training the improved YOLOv8 model based on the initial image dataset to obtain a nuclear magnetic resonance (NMR) image brain tumor detection model;
[0046] An output module for inputting the initial image dataset into the NMR image brain tumor detection model for detection and outputting the NMR brain tumor detection result.
[0047] A computer device, comprising a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the NMR image brain tumor detection method as described above is implemented.
[0048] A computer storage medium, where instructions are stored in the computer storage medium, and when the instructions are executed on a computer, the computer is caused to execute the NMR image brain tumor detection method as described above.
[0049] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings
[0050] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0051] Figure 1 is a flowchart of the NMR image brain tumor detection method according to an embodiment of the present invention;
[0052] Figure 2 is a schematic diagram of the DivC2f network of the NMR image brain tumor detection method according to an embodiment of the present invention;
[0053] Figure 3 is a schematic diagram of the DivOpera network of the NMR image brain tumor detection method according to an embodiment of the present invention;
[0054] Figure 4 is a schematic diagram of the SELayer network of the NMR image brain tumor detection method according to an embodiment of the present invention;
[0055] Figure 5 is a schematic diagram of the RP-LSKM network of the NMR image brain tumor detection method according to an embodiment of the present invention;
[0056] Figure 6 is a schematic diagram of the PVSS network of the NMR image brain tumor detection method according to an embodiment of the present invention;
[0057] Figure 7 Schematic diagram of the nuclear magnetic resonance image brain tumor detection system proposed in the second embodiment of the present application. Detailed implementation manners
[0058] The embodiments of the present invention will be described in detail below. The embodiments described with reference to the accompanying drawings are exemplary. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0059] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there can also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time.
[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the description of the present invention in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0061] The terms "first", "second", "third", etc. in the description and claims of the present application and the accompanying drawings are used to distinguish different objects and are not used to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a series of steps or units are included, or optionally, steps or units not listed are also included, or optionally, other steps or units inherent to these processes, methods, products or devices are also included.
[0062] Only parts related to the present application are shown in the drawings, not all of the content. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be performed in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. When the operations are completed, the process can be terminated, but there can also be additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0063] As used in this specification, the terms "component", "module", "system", "unit", etc. are used to represent computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a unit may be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or distributed between two or more computers. In addition, these units can be executed from various computer-readable media storing various data structures. A unit can communicate, for example, through local and / or remote processes according to a signal having one or more data packets (e.g., data from a second unit interacting with a local system, a distributed system, and / or another unit between networks. For example, the Internet interacting with other systems through a signal).
[0064] Embodiment 1
[0065] As Figures 1 to 6 shown, this embodiment provides a method for detecting brain tumors in nuclear magnetic resonance (NMR) images, which includes:
[0066] S1. Obtain the initial image dataset of the NMR brain tumor, perform annotation, and divide it into a training set, a validation set, and a test set;
[0067] S2. Construct a DivC2f network and a PVSS network. Use the DivC2f network as the backbone network in the YOLOv8 model, and use the PVSS network as the feature enhancement network after SPPF in the YOLOv8 model, and adjust the backbone network of the YOLOv8 model to lightweight the YOLOv8 model while improving performance; the DivC2f network includes pointwise convolution, batch normalization, SiLU activation function, ReLU activation function, feature segmentation, feature splicing, SElayer processing, and a DivOpera network that divides the feature signals pointwise;
[0068] S3. In the YOLOv8 model with improved performance, use the constructed RP-LSKM network as the multi-scale feature interaction network in the neck structure, and at the same time use the reparameterization technique to further improve the performance of the YOLOv8 model to obtain an improved YOLOv8 model;
[0069] S4. Use the initial image dataset to perform iterative optimization training on the improved YOLOv8 model to obtain a brain tumor detection model for NMR images;
[0070] S5. Input the initial image dataset into the brain tumor detection model for NMR images for detection, and output the brain tumor detection result for NMR.
[0071] It should be noted that in this embodiment, S1 also includes format conversion of the divided dataset.
[0072] Taking an open brain tumor MRI image dataset as an example, images are screened from the image dataset to ensure that diverse brain tumor samples such as benign and malignant ones are covered in the screened images.
[0073] Next, the format of the label files of the screened images is converted. The original image dataset format is converted into a YOLO format annotation file to meet the subsequent training and verification requirements for testing. In this embodiment, according to the scale of the image dataset, a sampling scheme based on a ratio of 7:2:1 and adopting the cross-validation method (GroupKFold) for maintaining grouped data is used to sample in different category data folders, and the image dataset is divided into a training set, a validation set, and a test set to ensure a reasonable data distribution among the three.
[0074] To obtain a more accurate and efficient training set, data cleaning is performed on the training set and the validation set. Images without annotation boxes or data with duplicate annotation boxes in the YOLO format annotation files are removed, and unreasonable annotation boxes are re-annotated manually through human intervention. No additional processing is required for the test set.
[0075] During the process of optimizing the training set, a variety of image enhancement techniques are used to expand the diversity and richness of the data, which specifically include the following:
[0076] I. The mosaic enhancement method is adopted: Four images are randomly selected from the image dataset, and independent data augmentation operations are performed on each image. Subsequently, the four images after the data augmentation operations are spliced into one image, effectively increasing the complexity and diversity of the dataset.
[0077] II. The image mixing technique is adopted: Two images are randomly selected from the image dataset and mixed and classified according to a preset ratio, and the classification results obtained are also distributed according to the corresponding ratio to achieve data enhancement.
[0078] III. The copy-paste strategy is adopted: Two images are randomly selected from the image dataset for augmentation processing, and a random subset of the objects in one of the images is randomly selected and pasted to a random position on the other image.
[0079] IV. The random flipping technique is adopted: The randomly selected images from the image dataset are flipped horizontally or vertically to simulate image changes at different angles.
[0080] V. The random scaling technique is adopted: The sizes of the randomly selected images from the image dataset are randomly adjusted, which can be either reduced or enlarged to simulate objects at different scales.
[0081] VI. Random affine transformation technology is adopted: randomly perform an affine transformation and translation operation on the selected images from the image dataset, including rotation, translation, scaling, shearing, etc. These operations can simulate various transformations of objects in three-dimensional space, further enhancing the complexity and generalization ability of the dataset.
[0082] It should be noted that in S2, the processing process of the double convolutional cross-stage division network includes:
[0083] In the DivOpera network: the input features pass through a separable convolution module composed of a pointwise convolution with a convolution kernel of 1, batch normalization, and a SiLU activation function, evenly divide the input features into feature map A and feature map B in the channel dimension, and perform SElayer processing on feature map A to obtain high-value channel signal features;
[0084] Construct a triple convolution module, and let feature map B enter the triple convolution modules in their respective branches through a double-branch parallel structure. The triple convolution module corresponds to Figure 3 the C3 module in it to control the range of the kernel function and perform low-dimensional representation and integration of the feature signals while integrating the input features;
[0085] Process the feature signals of the first branch in the double branch through a ReLU activation function and a Sigmoid function respectively, and perform a mirror operation on the feature signals of the second branch;
[0086] In the double-stream division network, the second branch uses the output of the first branch after being processed by the ReLU activation function as the dividend, and the output of the second branch after being processed by the Sigmoid activation function as the divisor, and divides the results of the two branches to obtain the first output result. Perform a mirror operation on the first branch in the double-stream division network to obtain the second output result. In this embodiment, the mirroring means that this branch repeats the steps of the other branch, that is, the two branches execute the same steps;
[0087] Concatenate the first output result, the second output result, and the high-value channel signal features in the double-stream division network in the channel dimension to obtain the output channel signal;
[0088] Pass the output channel signal through a pointwise convolution and a depth convolution to perform interactive integration output on the output channel signal to complete the processing process of the input features in the double convolutional cross-stage division network.
[0089] It should be noted that in this embodiment, the triple convolution module includes:
[0090] The feature map B passes through a low-dimensional representation module, which includes a pointwise convolution with a convolution kernel of 1, a batch normalization regularization function, a SiLU activation function, a dilated convolution with a convolution kernel of 3 and a dilation rate of 3, a pointwise convolution with a convolution kernel of 1, and a batch normalization regularization function, so as to control the range size of the kernel function while integrating the input features, perform low-dimensional representation on the feature signals, and integrate them.
[0091] It should be noted that in S2, the process of performing SElayer processing on the feature map A to obtain high-value channel signal features is specifically as follows:
[0092] Receive the feature map A and perform global average pooling processing for compression and integration to obtain global features;
[0093] Pass the global features through a weighted fusion module to interact and perform weighted fusion on the channel signals of the global features. The weighted fusion module consists of a pointwise convolution with a convolution kernel of 1, a batch normalization regularization function, a SiLU activation function, and a pointwise convolution with a convolution kernel of 1;
[0094] The channel signals after weighted fusion enter the HardSigmoid activation function to obtain the weights of high-value channel signal features, where HardSigmoid is similar to the Sigmoid function but has better computational efficiency;
[0095] Multiply the high-value channel signal feature weights with the input features to extract high-value channel signal features.
[0096] It should be noted that the processing process of the partial vision state space network includes:
[0097] Receive the features processed by the double convolution cross-stage division network, obtain the output features after passing through a convolution layer with a convolution kernel of 1, and pass the output features through the residual network block in the partial vision state space network to obtain the fused multi-scale features, where the residual network block includes a Mamba layer and a multi-layer perceptron layer;
[0098] Concatenate the multi-scale features with the output features, and then pass through a convolution layer with a convolution kernel of 1 to output the enhanced output features.
[0099] It should be noted that in S3, the processing process of the reparameterized dynamically selectable large kernel convolution includes:
[0100] The feature map output in S2 is received by the PatchEmbed module to obtain the signal features after expanding the number of channels. Specifically, in the PatchEmbed module, the number of channels of the input feature map is expanded to the number of hidden layer channels through a pointwise convolution layer and a depth convolution, and then integrated to obtain the signal features after expanding the number of channels.
[0101] The signal features after expanding the number of channels are passed through n cascaded Random Projection Local Semantic Kernel (RP-LSK) blocks. The RP-LSK block contains a network layer based on a receptive field pyramid with large selective kernels (RP-LSKNet layer) and a depthwise separable multi-layer perceptron layer (DwMlp layer) to enhance the feature expression ability of the model.
[0102] It should be noted that the process of passing the signal features after expanding the number of channels into the RP-LSKNet layer includes:
[0103] The signal features after expanding the number of channels are concatenated and passed through a pointwise convolution with a kernel size of 1 and a GELU activation function for non-linear integration.
[0104] The non-linearly integrated signal features are passed through a depth convolution composed of n parallel convolutional kernels with a kernel size of 5 to output feature map C.
[0105] Feature map C is passed through a dilated convolution composed of n parallel convolutional kernels with a kernel size of 7 and a dilation rate of 3 to output feature map D, which is used to dynamically expand the receptive field of the convolution.
[0106] Feature map C is passed through a pointwise convolution with a kernel size of 1 to obtain compressed feature map C, where the number of compressed channels is half of the input.
[0107] Feature map D is passed through a pointwise convolution with a kernel size of 1 to obtain compressed feature map D, where the number of compressed channels is half of the input.
[0108] The compressed feature map C and feature map D are concatenated for the first time to obtain the first concatenated feature map.
[0109] The first concatenated feature map is passed through an average pooling layer and a max pooling layer in parallel to output their respective pooling results, and then concatenated for the second time according to their respective pooling results to obtain the second concatenated feature map. Both the average pooling and max pooling are performed on the channel dimension.
[0110] The second concatenated feature map is passed through a depth convolution with a kernel size of 7 and a Sigmoid activation function to generate a heat map.
[0111] The heat map is evenly divided according to the channel dimension, and each part is multiplied by the compressed feature map C and the feature map D respectively, and the results after multiplication are added together to complete the signal integration of different receptive fields at the same spatial position under different granularity information;
[0112] The heat map after integrating the signals of different receptive fields is subjected to pointwise convolution with a convolution kernel of 1 for signal fusion, and the fusion result is multiplied by the signal feature after expanding the number of channels to complete the screening of high-value regions.
[0113] The present embodiment provides the following specific implementation cases:
[0114] The images in the dataset used are all from different angles of magnetic resonance imaging (MRI) scans, including sagittal plane, axial plane and coronal plane, and the dataset contains 5,249 high-quality MRI images of brain tumors with detailed annotations. The diverse images can ensure the comprehensive coverage of the brain anatomical structure, enhance the robustness of the model trained on this dataset, and the image bounding boxes are manually annotated using an image annotation tool (LabelImg) to ensure the high accuracy and reliability of image annotation. The specific division statistics of the Tumor dataset are shown in Table 1
[0115] Table 1, Tumor dataset division table
[0116]
[0117] Some configuration information for the model training of the present invention is shown in Table 2
[0118] Table 2, Partial configuration table for model training
[0119]
[0120] For fairness, the training strategy adopted for the model training of the present invention is consistent with YOLOv8n. The framework used for training is MMYOLO, and MMYOLO has a slight adjustment to the YOLOv8n architecture but does not affect the final performance. The parameter quantity statistics have a certain deviation from the official statistics but do not affect the results. The final effect comparison can be seen in Table 3.
[0121] Table 3, Experimental effect comparison
[0122]
[0123] From the experimental effect comparison, it can be seen that the recognition ability of the model of this patent has been greatly improved compared with YOLOv8n for micro-cellular tumors such as gliomas, and there is still a large room for further improvement.
[0124] According to the nuclear magnetic resonance (NMR) image brain tumor detection method of the present invention, by improving the C2f structural block in the YOLOv8n model, a double-convolution cross-stage division idea is proposed, aiming to enhance the modeling ability of potential high-dimensional feature information, improve the attention to the semantic information of high-value channels, and further extract the semantic information of high-value channels in the subsequent division block. And to avoid NaN in the division, the two divisor branches are first processed by the Sigmoid function to avoid the situation of the divisor being 0; the large-kernel convolution in the large-kernel dynamic selectable convolution structural block is improved by the reparameterization technology of the reparameterized dynamic selectable large-kernel convolution (RP-LSKM) network to improve the information extraction ability; the PVSS module enhances the model state space modeling and the model global representation learning ability at a low computational cost, and the linear self-attention mechanism used avoids the excessive overhead brought by the quadratic computational complexity of self-attention. Thus, the improved YOLOv8n model proposed in this application can not only perform context modeling at different scales in the channel dimension but also perform context modeling in the global space representation. The NMR image brain tumor detection model based on the improved YOLOv8n model can improve the accuracy and efficiency of brain tumor detection.
[0125] Embodiment 2
[0126] As Figure 7 shown, this embodiment provides a nuclear magnetic resonance (NMR) image brain tumor detection system, which is used to implement the detection method as described above and includes:
[0127] An acquisition module, configured to acquire an initial image dataset of NMR brain tumors, perform annotation and divide it into a training set, a validation set, and a test set;
[0128] A first construction module, configured to construct a double-convolution cross-stage division (DivC2f) network and a partial visual state space (PVSS) network, use the double-convolution cross-stage division (DivC2f) network as the backbone network in the YOLOv8 model, use the partial visual state space (PVSS) network as the feature enhancement network after SPPF in the YOLOv8 model, and adjust the backbone network of the YOLOv8 model to lightweight the YOLOv8 model while improving performance; the double-convolution cross-stage division network includes pointwise convolution, batch normalization, SiLU activation function, ReLU activation function, feature segmentation, feature splicing, channel attention layer (SElayer) processing, and a feature interaction dual-stream division (DivOpera) network that divides feature signals pointwise;
[0129] A second construction module, which is used in the YOLOv8 model with improved performance to use the constructed Reparameterized Large Kernel Dynamic Selectable Convolution (RP-LSKM) network as a multi-scale feature interaction network in the neck structure, and at the same time uses the reparameterization technology to further improve the performance of the YOLOv8 model, so as to obtain an improved YOLOv8 model;
[0130] A training module, which is used to iteratively optimize and train the improved YOLOv8 model according to the initial image dataset to obtain a nuclear magnetic resonance (NMR) image brain tumor detection model;
[0131] An output module, which is used to input the initial image dataset into the NMR image brain tumor detection model for detection and output the NMR brain tumor detection result.
[0132] A NMR image brain tumor detection system in an embodiment of the present application can be a device, or a component, an integrated circuit, or a chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a palmtop computer, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), etc. The embodiments of the present application do not make specific limitations.
[0133] A NMR image brain tumor detection system in an embodiment of the present application can be a device with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0134] A NMR image brain tumor detection system provided by an embodiment of the present application can implement Figure 1 each process implemented by the method embodiment of a NMR image brain tumor detection method. To avoid repetition, it will not be elaborated here.
[0135] According to the nuclear magnetic resonance (NMR) image brain tumor detection system of the embodiments of the present invention, by improving the C2f structural block in the YOLOv8n model, a dual-convolution cross-stage division idea is proposed, aiming to enhance the modeling ability of potential high-dimensional feature information, increase the attention to the semantic information of high-value channels, and further extract the semantic information of high-value channels in the subsequent division block. Moreover, to avoid NaN in division, the two divisor branches are first processed by the Sigmoid function to prevent the situation of the divisor being 0. The information extraction ability of the large kernel convolution in the large kernel dynamic selectable convolution structural block is improved by the reparameterization technology of the reparameterized dynamic selectable large kernel convolution (RP-LSKM) network. The PVSS module enhances the model state space modeling and the model global representation learning ability at a relatively low computational cost, and the linear self-attention mechanism used avoids the excessive overhead brought by the quadratic computational complexity of self-attention. Thus, the improved YOLOv8n model proposed in this application can not only perform context modeling at different scales in the channel dimension but also perform context modeling in the global space representation. The NMR image brain tumor detection model based on the improved YOLOv8n model can improve the accuracy and efficiency of brain tumor detection.
[0136] Optionally, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above-mentioned embodiment of a NMR image brain tumor detection method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0137] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, it implements each process of the above-mentioned embodiment of a NMR image brain tumor detection method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0138] Among them, the processor is the processor in the electronic device in the above-mentioned embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc.
[0139] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the invention.
[0140] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example.
[0141] Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The mention of "embodiment" in this article means that the specific features, structures, or characteristics described in connection with the embodiment can be included in at least one embodiment of this application. The occurrence of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art can explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0142] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present invention. The scope of the present invention is defined by the claims and their equivalents.
Claims
1. A method for detecting brain tumors in nuclear magnetic resonance images, characterized in that, Including: S1. Obtain the initial image dataset of nuclear magnetic resonance (NMR) brain tumors, perform annotation, and divide it into a training set, a validation set, and a test set; S2. Construct a double convolutional cross-stage division network and a partial vision state space network. Use the double convolutional cross-stage division network as the backbone network in the YOLOv8 model, and use the partial vision state space network as the feature enhancement network after SPPF in the YOLOv8 model. Then adjust the backbone network of the YOLOv8 model to lightweight the YOLOv8 model while improving its performance. The double convolutional cross-stage division network includes pointwise convolution, batch normalization, SiLU activation function, ReLU activation function, feature segmentation, feature splicing, channel attention layer processing, and a feature interaction dual-stream division network that divides feature signals point by point; S3. In the YOLOv8 model with improved performance, in the neck structure, use the constructed reparameterized dynamic selectable large kernel convolutional network as a multi-scale feature interaction network, and at the same time use reparameterization to further improve the performance of the YOLOv8 model to obtain an improved YOLOv8 model; S4. Use the initial image dataset to perform iterative optimization training on the improved YOLOv8 model to obtain an NMR image brain tumor detection model; S5. Input the initial image dataset into the NMR image brain tumor detection model for detection and output the NMR brain tumor detection result.
2. The nuclear magnetic resonance image brain tumor detection method according to claim 1, characterized in that, In S2, the processing process of the double convolutional cross-stage division network includes: The input feature passes through a separable convolution module composed of a pointwise convolution with a convolution kernel of 1, batch normalization, and a SiLU activation function to evenly divide the input feature into feature map A and feature map B in the channel dimension, and perform channel attention layer processing on feature map A to obtain high-value channel signal features; Construct a triple convolutional module, and pass feature map B through a double-branch parallel structure into the triple convolutional module in each branch to control the range of the kernel function and perform low-dimensional representation and integration of the feature signals while integrating the input features; Process the feature signals of the first branch in the double branch through a ReLU activation function and a Sigmoid function respectively, and perform a mirror operation on the feature signals of the second branch; The second branch uses the output after the ReLU activation function processing of the first branch as the dividend, and the output after the Sigmoid activation function processing of the second branch as the divisor, and divide the results of the two branches to obtain a first output result, and perform a mirror operation on the first branch in the double-stream division network to obtain a second output result; Concatenate the first output result, the second output result, and the high-value channel signal features in the channel dimension to obtain an output channel signal; Perform convolution on the output channel signal and then output.
3. The nuclear magnetic resonance image brain tumor detection method according to claim 2, wherein, In S2, the process of performing channel attention layer processing on feature map A to obtain high-value channel signal features is specifically: Receive feature map A and perform global average pooling processing to compress and integrate it to obtain global features; Pass the global features through a weighted fusion module to interact and perform weighted fusion on the channel signals of the global features; Enter the weighted fusion channel signals into a HardSigmoid activation function to obtain the weights of high-value channel signal features; Multiply the high-value channel signal feature weights by the input features to extract high-value channel signal features.
4. The nuclear magnetic resonance image brain tumor detection method according to claim 1, wherein In S2, the processing process of the partial vision state space network includes: Receive the features processed by the double convolutional cross-stage division network, obtain output features after passing through a convolutional layer with a convolution kernel of 1, and pass the output features through the residual network block in the partial vision state space network to obtain fused multi-scale features, where the residual network block includes a Mamba layer and a multi-layer perceptron layer; Concatenate the multi-scale features with the output features, and then pass through a convolutional layer with a kernel size of 1 to output the enhanced output features.
5. The method for detecting brain tumors in nuclear magnetic resonance images according to claim 3, characterized in that, In S2, the weighted fusion module consists of a pointwise convolution with a kernel size of 1, a batch normalization regularization function, a SiLU activation function, and a pointwise convolution with a kernel size of 1.
6. The method for detecting brain tumors in nuclear magnetic resonance images according to claim 1, characterized in that, In S3, the processing process of the reparameterized dynamic selectable large kernel convolution includes: Receiving the feature map output in S2 into the image patch embedding module to obtain the signal feature after expanding the number of channels; Passing the signal feature after expanding the number of channels through n cascaded random projection local semantic kernel modules to enhance the feature expression ability of the model, where the random projection local semantic kernel module includes a network layer of a large selective kernel based on the receptive field pyramid and a depthwise separable multi-layer perceptron layer.
7. The method for detecting brain tumors in nuclear magnetic resonance images according to claim 6, wherein, The process of passing the signal feature after expanding the number of channels into the network layer of the large selective kernel based on the receptive field pyramid includes: Cascading the signal feature after expanding the number of channels through a pointwise convolution and a GELU activation function for non-linear integration; Passing the non-linearly integrated signal feature through a depthwise convolution composed of n parallel convolution kernels with a size of 5 to output the feature map C; Passing the feature map C through a dilated convolution composed of n parallel convolution kernels with a size of 7 and a dilation rate of 3 to output the feature map D for dynamically expanding the receptive field of the convolution; Compressing the feature map C and the feature map D through a pointwise convolution to obtain the compressed feature map C and the compressed feature map D; Performing the first concatenation on the compressed feature map C and the feature map D to obtain the first concatenated feature map; Passing the first concatenated feature map in parallel through an average pooling layer and a max pooling layer to output their respective pooling results, and performing the second concatenation according to their respective pooling results to obtain the second concatenated feature map; Passing the second concatenated feature map through a depthwise convolution with a kernel size of 7 and a Sigmoid activation function to generate a heat map; Equalizing the heat map in the channel dimension, multiplying each part with the compressed feature map C and the feature map D respectively, and adding the multiplied results to complete the signal integration of different receptive fields at the same spatial position under different granularity information; Fusing the signals of the heat map after integrating the signals of different receptive fields through a pointwise convolution with a kernel size of 1, and multiplying according to the fusion result with the signal feature after expanding the number of channels to complete the screening of high-value regions.
8. A nuclear magnetic resonance image brain tumor detection system, characterized in that, For implementing the detection method according to any one of claims 1 to 7, including: An acquisition module for acquiring the initial image dataset of the nuclear magnetic brain tumor, performing annotation and dividing it into a training set, a validation set, and a test set; The first building block is used to construct a double convolutional cross-stage division network and a partial vision state space network. The double convolutional cross-stage division network is used as the backbone network in the YOLOv8 model, and the partial vision state space network is used as the feature enhancement network after SPPF in the YOLOv8 model. The backbone network of the YOLOv8 model is adjusted to lightweight the YOLOv8 model while improving performance. The double convolutional cross-stage division network includes pointwise convolution, batch normalization, SiLU activation function, ReLU activation function, feature segmentation, feature splicing, channel attention layer processing, and a feature interaction double-flow division network that divides feature signals point by point; A second construction module for, in the YOLOv8 model for improving performance, constructing the reparameterized dynamic selectable large kernel convolution network as a multi-scale feature interaction network in the neck structure, and simultaneously using reparameterization to further improve the performance of the YOLOv8 model to obtain the improved YOLOv8 model; A training module for iteratively optimizing and training the improved YOLOv8 model according to the initial image dataset to obtain a nuclear magnetic resonance (NMR) image brain tumor detection model; An output module for inputting the initial image dataset into the NMR image brain tumor detection model for detection and outputting the NMR brain tumor detection result.
9. A computer device, characterized in that, It includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the NMR image brain tumor detection method according to any one of claims 1 to 7.
10. A computer storage medium, characterized in that, Instructions are stored in the computer storage medium, and when the instructions are executed on a computer, the computer is caused to execute the NMR image brain tumor detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
MRI (Magnetic Resonance Imaging) brain tumor segmentation method based on multi-scale feature fusion of improved U-Net
CN117876399A
Improved YOLOv8-based dangerous behavior detection method for chemical enterprise personnel
CN118038555A