Nuclear magnetic image brain tumor detection method and system
By improving the YOLOv8 model, a double convolutional cross-stage division network and a partial visual state space network are constructed, and a large nuclear convolutional network can be selected in combination with reparameterized dynamics, solving the problem of brain tumor detection accuracy and efficiency in existing MRI technology, and achieving more efficient brain tumor detection results.
Patent Information
- Application Number
- CN202510552928.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The existing MRI technology has problems of inaccuracy and efficiency in brain tumor detection.
A method for detecting brain tumors in nuclear magnetic images is proposed. By improving the YOLOv8 model, a double convolutional cross-stage division (DivC2f) network and a partial visual state space (PVSS) network are constructed. Combined with reparameterized dynamics, large-core convolution (RP-LSKM) network can be selected to improve the performance and detection efficiency of the model.
It improves the accuracy and efficiency of brain tumor detection, enhances the model's attention to high-value channel semantic information, and reduces the computational complexity.
Smart Images

Figure CN120070455A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and in particular to a method and system for detecting brain tumors in nuclear magnetic resonance (MRI) images. Background Art
[0002] Accurate detection of nuclear magnetic resonance (MRI) images has become the gold standard for screening brain diseases due to its superior imaging characteristics. Compared with imaging technologies such as CT, MRI has higher contrast in soft tissue imaging, can clearly show the boundary between tumors and normal tissues, enabling doctors to accurately judge the nature, size and location of tumors. Accurate detection of brain tumors in MRI images is of great significance in improving the early diagnosis rate, optimizing treatment plans, promoting personalized medicine and clinical research.
[0003] Although existing MRI technologies have many advantages in brain tumor detection, there are still problems of low detection accuracy and efficiency. Summary of the Invention
[0004] The present invention aims to at least improve one of the technical problems existing in the prior art. For this purpose, the present invention proposes a method and system for detecting brain tumors in MRI images.
[0005] The technical solution of the present invention is as follows: A method for detecting brain tumors in MRI images, which includes: S1, obtaining an initial image dataset of nuclear magnetic resonance brain tumors, performing annotation and dividing it into a training set, a validation set and a test set; S2, constructing a Dual Convolution Cross Stage Division (DivC2f) network and a Partial Visual State Space (PVSS) network, using the Dual Convolution Cross Stage Division (DivC2f) network as the backbone network in the YOLOv8 model, and using the Partial Visual State Space (PVSS) network as the feature enhancement network after SPPF in the YOLOv8 model, and adjusting the backbone network of the YOLOv8 model to lightweight the YOLOv8 model while improving performance; the Dual Convolution Cross Stage Division network includes pointwise convolution, batch normalization, SiLU activation function, ReLU activation function, feature segmentation, feature splicing, channel attention layer (SElayer) processing, and a Feature Interaction Dual Division (DivOpera) network that performs pointwise division of feature signals; S3, in the YOLOv8 model with improved performance, using the constructed Reparameterized Large Kernel Dynamic Selectable Convolution (RP-LSKM) network as a multi-scale feature interaction network in the neck structure, and at the same time using reparameterization to further improve the performance of the YOLOv8 model to obtain an improved YOLOv8 model; S4. Use the initial image dataset to iteratively optimize and train the improved YOLOv8 model to obtain a nuclear magnetic resonance (NMR) image brain tumor detection model; S5. Input the initial image dataset into the NMR image brain tumor detection model for detection and output the NMR brain tumor detection results.
[0006] In a possible technical solution, further, in S2, the processing process of the double convolutional cross-stage division network includes: The input features pass through a separable convolution module composed of a pointwise convolution with a convolution kernel of 1, batch normalization, and a SiLU activation function to evenly divide the input features into feature map A and feature map B in the channel dimension, and perform channel attention layer (SE layer) processing on feature map A to obtain high-value channel signal features; Construct a triple convolution module, and let feature map B enter the triple convolution modules in their respective branches through double branches in parallel to control the range of the kernel function and perform low-dimensional representation and integration of the feature signals while integrating the input features; Process the feature signals in the first branch of the double branches through a ReLU activation function and a Sigmoid function respectively, and perform a mirror operation on the feature signals in the second branch; The second branch uses the output after the ReLU activation function processing in the first branch as the dividend and the output after the Sigmoid activation function processing in the second branch as the divisor, and divide the results of the two branches to obtain a first output result, and perform a mirror operation on the first branch in the double-flow division network to obtain a second output result; Concatenate the first output result, the second output result, and the high-value channel signal features in the channel dimension to obtain an output channel signal; Output after convolving the output channel signal.
[0007] In a possible technical solution, further, in S2, the specific operation of performing channel attention layer processing on feature map A to obtain high-value channel signal features is as follows: Receive feature map A for global average pooling processing to compress and integrate to obtain global features; Pass the global features through a weighted fusion module to interact and perform weighted fusion on the channel signals of the global features. The weighted fusion module is composed of a pointwise convolution with a convolution kernel of 1, a batch normalization regularization function, a SiLU activation function, and a pointwise convolution with a convolution kernel of 1; Let the weighted fusion channel signals enter the HardSigmoid activation function to obtain the weights of the high-value channel signal features; Multiply the high-value channel signal features by the input features according to the weights of the high-value channel signal features to extract the high-value channel signal features.
[0008] In a possible technical solution, further, the partial vision state space network processing process includes: Receive the features processed by the double convolutional cross-stage division network, obtain the output features after passing through a convolutional layer with a kernel size of 1, and pass the output features through the residual network blocks in the partial vision state space network to obtain the fused multi-scale features, where the residual network blocks include Mamba layers and multi-layer perceptron layers; Concatenate the multi-scale features with the output features, and then pass through a convolutional layer with a kernel size of 1 to output the enhanced output features.
[0009] In a possible technical solution, further, in S3, the processing process of the reparameterized dynamic selectable large kernel convolution includes: Receive the feature map output in S2 to the PatchEmbed module to obtain the signal features after expanding the number of channels. Specifically, in the PatchEmbed module, the number of channels of the input feature map is expanded to the number of hidden layer channels through a pointwise convolutional layer and a depth convolutional layer and integrated to obtain the signal features after expanding the number of channels; Pass the signal features after expanding the number of channels through n cascaded Random Projection Local Semantic Kernel modules (RP-LSKBlock), where RP-LSKBlock contains a network layer of a large selective kernel based on the receptive field pyramid (RP-LSKNet layer) and a depthwise separable multi-layer perceptron layer (DwMlp layer) to enhance the feature expression ability of the model.
[0010] In a possible technical solution, further, the process of passing the signal features after expanding the number of channels into the RP-LSKNet layer includes: Pass the signal features after expanding the number of channels through a pointwise convolution with a kernel size of 1 and a GELU activation function in series for non-linear integration; Pass the non-linearly integrated signal features through a depth convolution composed of n parallel convolutional kernels with a kernel size of 5 to output the feature map C; Pass the feature map C through a dilated convolution composed of n parallel convolutional kernels with a kernel size of 7 and a dilation rate of 3 to output the feature map D for dynamically expanding the receptive field of the convolution; Compress the feature map C and the feature map D through a pointwise convolution to obtain the compressed feature map C and the compressed feature map D, where the number of compressed channels is half of the input; Perform a first concatenation on the compressed feature map C and the feature map D to obtain the feature map after the first concatenation; Pass the feature map after the first concatenation through an average pooling layer and a max pooling layer in parallel to output their respective pooling results, and perform a second Concat (concatenation) based on their respective pooling results to obtain the feature map after the second concatenation, where both average pooling and max pooling are performed in the channel dimension; Pass the feature map after the second concatenation through a depth convolution with a convolution kernel of 7 and a Sigmoid activation function to generate a heat map, Evenly divide the heat map in the channel dimension, multiply each part by the compressed feature map C and the feature map D respectively, and add the results after multiplication to complete the signal integration of different receptive fields at the same spatial position under different granularity information; Pass the heat map after integrating the signals of different receptive fields through a pointwise convolution with a convolution kernel of 1 for signal fusion, and multiply according to the fusion result and the signal feature after expanding the number of channels to complete the screening of high-value regions.
[0011] According to the nuclear magnetic resonance image brain tumor detection method of the present invention, by improving the C2f structural block in the YOLOv8n model, the idea of double convolution cross-stage division is proposed, aiming to improve the modeling ability of potential high-dimensional feature information, increase the attention to the semantic information of high-value channels, and further extract the semantic information of high-value channels in the subsequent division block. And to avoid NaN in division, the two divisor branches are first processed by the Sigmoid function to avoid the situation of the divisor being 0; improve the large kernel convolution in the large kernel dynamic selectable convolution structural block by the reparameterization technology of the reparameterized dynamic selectable large kernel convolution network to improve the information extraction ability; strengthen the model state space modeling and the model global representation learning ability through the PVSS module at a lower computational cost, and the linear self-attention mechanism used avoids the excessive overhead brought by the quadratic computational complexity of self-attention. Thus, the improved YOLOv8n model proposed in this application can not only perform context modeling at different scales in the channel dimension but also perform context modeling in the global spatial representation. The nuclear magnetic resonance image brain tumor detection model based on the improved YOLOv8n model can improve the accuracy and efficiency of brain tumor detection.
[0012] A nuclear magnetic resonance image brain tumor detection system, which is used to implement the detection method as described above, includes: An acquisition module, which is used to acquire the initial image dataset of nuclear magnetic resonance brain tumors, perform annotation and divide it into a training set, a validation set and a test set; The first construction module is used to construct a Dual Convolution Cross Stage Division (DivC2f) network and a Partial Visual State Space (PVSS) network. The Dual Convolution Cross Stage Division (DivC2f) network is used as the backbone network in the YOLOv8 model, and the Partial Visual State Space (PVSS) network is used as the feature enhancement network after SPPF in the YOLOv8 model. The backbone network of the YOLOv8 model is adjusted to lightweight the YOLOv8 model while improving its performance. The Dual Convolution Cross Stage Division network includes pointwise convolution, batch normalization, SiLU activation function, ReLU activation function, feature segmentation, feature splicing, channel attention layer (SElayer) processing, and a Feature Interaction Dual-stream Division (DivOpera) network that performs pointwise division of feature signals The second construction module is used in the YOLOv8 model with improved performance. In the neck structure, the constructed Reparameterized Large Kernel Dynamic Selectable Convolution (RP-LSKM) network is used as a multi-scale feature interaction network, and at the same time, reparameterization is used to further improve the performance of the YOLOv8 model to obtain an improved YOLOv8 model The training module is used to perform iterative optimization training on the improved YOLOv8 model according to the initial image dataset to obtain a nuclear magnetic resonance image brain tumor detection model The output module is used to input the initial image dataset into the nuclear magnetic resonance image brain tumor detection model for detection and output the nuclear magnetic resonance brain tumor detection result
[0013] A computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the nuclear magnetic resonance image brain tumor detection method as described above
[0014] A computer storage medium, in which instructions are stored. When the instructions are executed on a computer, the computer is made to execute the nuclear magnetic resonance image brain tumor detection method as described above
[0015] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention Description of the Drawings
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts
[0017] Figure 1 It is a flowchart of a method for detecting brain tumors in nuclear magnetic resonance (NMR) images according to an embodiment of the present invention; Figure 2 It is a schematic diagram of the DivC2f network of the method for detecting brain tumors in NMR images according to an embodiment of the present invention; Figure 3 It is a schematic diagram of the DivOpera network of the method for detecting brain tumors in NMR images according to an embodiment of the present invention; Figure 4 It is a schematic diagram of the SELayer network of the method for detecting brain tumors in NMR images according to an embodiment of the present invention; Figure 5 It is a schematic diagram of the RP-LSKM network of the method for detecting brain tumors in NMR images according to an embodiment of the present invention; Figure 6 It is a schematic diagram of the PVSS network of the method for detecting brain tumors in NMR images according to an embodiment of the present invention; Figure 7 It is a schematic diagram of a system for detecting brain tumors in NMR images proposed in the second embodiment of the present application. Detailed implementation manners
[0018] The embodiments of the present invention will be described in detail below. The embodiments described with reference to the accompanying drawings are exemplary. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0019] It should be noted that when an element is referred to as "fixed to" another element, it can be directly on the other element or there may also be a middle element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be a middle element at the same time.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific implementation manners and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0021] In the specification and claims of the present application and the above-mentioned drawings, the terms "first", "second", "third", etc. are used to distinguish different objects and are not used to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a series of steps or units are included, or optionally, steps or units not listed are also included, or optionally, other steps or units inherent to these processes, methods, products or devices are also included.
[0022] Only the parts related to this application are shown in the drawings, not all of the content. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be performed in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0023] The terms "component", "module", "system", "unit", etc. used in this specification are used to represent computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a unit can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or distributed between two or more computers. In addition, these units can be executed from various computer-readable media on which various data structures are stored. A unit can communicate, for example, through local and / or remote processes according to a signal having one or more data packets (such as data from a second unit interacting with a local system, a distributed system, and / or a network. For example, the Internet interacting with other systems through a signal).
[0024] Embodiment 1 As Figures 1 to 6 shown, this embodiment provides a method for detecting brain tumors in nuclear magnetic resonance (NMR) images, which includes: S1. Obtain the initial image dataset of NMR brain tumors, perform annotation and divide it into a training set, a validation set, and a test set; S2. Construct a DivC2f network and a PVSS network. Use the DivC2f network as the backbone network in the YOLOv8 model, and use the PVSS network as the feature enhancement network after SPPF in the YOLOv8 model, and adjust the backbone network of the YOLOv8 model to lightweight the YOLOv8 model while improving its performance; the DivC2f network includes pointwise convolution, batch normalization, SiLU activation function, ReLU activation function, feature segmentation, feature splicing, SElayer processing, and a DivOpera network that divides feature signals point by point; S3. In the YOLOv8 model with improved performance, use the constructed RP-LSKM network as a multi-scale feature interaction network in the neck structure, and at the same time use the reparameterization technology to further improve the performance of the YOLOv8 model to obtain an improved YOLOv8 model; S4. Use the initial image dataset to iteratively optimize and train the improved YOLOv8 model to obtain a nuclear magnetic resonance (NMR) image brain tumor detection model; S5. Input the initial image dataset into the NMR image brain tumor detection model for detection and output the NMR brain tumor detection results.
[0025] It should be noted that in this embodiment, S1 also includes format conversion of the divided dataset.
[0026] Taking a publicly available brain tumor NMR image dataset as an example, images are screened from the image dataset to ensure that the screened images cover diverse brain tumor samples such as benign and malignant ones.
[0027] Next, perform format conversion on the label files of the screened images, converting the original image dataset format into YOLO format annotation files to meet the subsequent training and verification requirements for testing. In this embodiment, according to the scale of the image dataset, a sampling scheme based on a 7:2:1 ratio and adopting the cross-validation method (GroupKFold) for maintaining grouped data is used to sample in different category data folders, and the image dataset is divided into a training set, a validation set, and a test set to ensure a reasonable data distribution among the three.
[0028] To obtain a more accurate and efficient training set, data cleaning is performed on the training set and the validation set. Images without annotation boxes or data with duplicate annotation boxes in the YOLO format annotation files are removed, and unreasonable annotation boxes are re-annotated manually through human intervention. No additional processing is required for the test set.
[0029] During the process of optimizing the training set, a variety of image enhancement techniques are used to expand the diversity and richness of the data, specifically including the following: I. The mosaic enhancement method is adopted: Randomly select four images from the image dataset, and perform independent data augmentation operations on each image. Subsequently, splice the four images after the data augmentation operations into one image, thereby effectively increasing the complexity and diversity of the dataset.
[0030] II. The image mixing technique is adopted: Randomly select two images from the image dataset and mix and classify them according to a preset ratio. The classification results obtained are also distributed according to the corresponding ratio to achieve data enhancement.
[0031] III. The copy-paste strategy is adopted: Randomly select two images from the image dataset for augmentation processing, and randomly select a target subset of one of the images and paste it at a random position on the other image.
[0032] IV. Random flipping technology is adopted: randomly flip the selected images in the horizontal or vertical direction from the image dataset to simulate image changes at different angles.
[0033] V. Random scaling technology is adopted: randomly adjust the size of the selected images from the image dataset, which can be reduced or enlarged, so as to simulate objects of different scales.
[0034] VI. Random affine transformation technology is adopted: randomly perform an affine transformation and translation operation on the selected images from the image dataset, including rotation, translation, scaling, and shearing, etc. These operations can simulate various transformations of objects in three-dimensional space, further enhancing the complexity and generalization ability of the dataset.
[0035] It should be noted that in S2, the processing process of the double convolutional cross-stage division network includes: In the DivOpera network: the input features pass through a separable convolution module composed of a pointwise convolution with a convolution kernel of 1, batch normalization, and SiLU activation function, and the input features are evenly divided into feature map A and feature map B in the channel dimension. The feature map A is processed by the SE layer to obtain high-value channel signal features; Construct a three-convolution module, and let the feature map B enter the three-convolution modules in their respective branches through double branches in parallel. The three-convolution module corresponds to Figure 3 the C3 module in it to control the range size of the kernel function and perform low-dimensional representation and integration on the feature signals while integrating the input features; The feature signals in the first branch of the double branches are respectively processed by a ReLU activation function and a Sigmoid function, and the feature signals in the second branch perform a mirror operation; In the double-stream division network, the second branch uses the output of the first branch processed by the ReLU activation function as the dividend, and the output of the second branch processed by the Sigmoid activation function as the divisor, and divides the results of the two branches to obtain the first output result. Perform a mirror operation on the first branch in the double-stream division network to obtain the second output result. In this embodiment, the mirroring means that this branch repeats the steps of the other branch, that is, the two branches perform the same steps; Concatenate the first output result, the second output result, and the high-value channel signal features in the double-stream division network in the channel dimension to obtain the output channel signal; Pass the output channel signal through a pointwise convolution and a depth convolution to perform interactive integration output on the output channel signal to complete the processing process of the input features in the double convolutional cross-stage division network.
[0036] It should be noted that in this embodiment, the three-convolution module includes: The feature map B is passed through a low-dimensional representation module, which includes a pointwise convolution with a convolution kernel of 1, a batch normalization regularization function, a SiLU activation function, a dilated convolution with a convolution kernel of 3 and a dilation rate of 3, a pointwise convolution with a convolution kernel of 1, and a batch normalization regularization function, so as to control the range size of the kernel function while integrating the input features, perform low-dimensional representation on the feature signals, and integrate them.
[0037] It should be noted that in S2, the process of performing SElayer processing on the feature map A to obtain high-value channel signal features is specifically as follows: Receive the feature map A and perform global average pooling processing for compression and integration to obtain global features; Pass the global features through a weighted fusion module to interact and perform weighted fusion on the channel signals of the global features. The weighted fusion module consists of a pointwise convolution with a convolution kernel of 1, a batch normalization regularization function, a SiLU activation function, and a pointwise convolution with a convolution kernel of 1; The channel signals after weighted fusion enter the HardSigmoid activation function to obtain the weights of the high-value channel signal features, where HardSigmoid is similar to the Sigmoid function but has better computational efficiency; Multiply the high-value channel signal feature weights with the input features to extract the high-value channel signal features.
[0038] It should be noted that the processing process of the partial vision state space network includes: Receive the features processed by the double convolution cross-stage division network, obtain the output features after passing through a convolution layer with a convolution kernel of 1, and pass the output features through the residual network block in the partial vision state space network to obtain the fused multi-scale features, where the residual network block includes a Mamba layer and a multi-layer perceptron layer; Concatenate the multi-scale features with the output features, and then pass through a convolution layer with a convolution kernel of 1 to output the enhanced output features.
[0039] It should be noted that in S3, the processing process of the reparameterized dynamic selectable large kernel convolution includes: Receive the feature map output in S2 to the image patch embedding module (PatchEmbed) to obtain the signal features after expanding the number of channels. Specifically, in the image patch embedding module, the number of channels of the input feature map is expanded to the number of hidden layer channels through a pointwise convolution layer and a depth convolution and integrated to obtain the signal features after expanding the number of channels; The signal features after the number of extended channels are passed through n serially connected Random Projection Local Semantic Kernel Modules (RP-LSKBlock), where RP-LSKBlock includes a network layer with a large selective kernel based on a receptive field pyramid (RP-LSKNet layer) and a depthwise separable multi-layer perceptron layer (DwMlp layer) to enhance the feature expression ability of the model.
[0040] It should be noted that the process of the signal features after the number of extended channels entering the RP-LSKNet layer includes: The signal features after the number of extended channels are serially passed through a pointwise convolution with a convolution kernel of 1 and a GELU activation function for non-linear integration; The non-linearly integrated signal features are passed through a depthwise convolution composed of n parallel convolution kernels with a convolution kernel of 5 to output a feature map C; The feature map C is passed through a dilated convolution composed of n parallel convolution kernels with a convolution kernel of 7 and a dilation rate of 3 to output a feature map D for dynamically expanding the receptive field of the convolution; The feature map C is passed through a pointwise convolution with a convolution kernel of 1 to obtain a compressed feature map C, where the number of compressed channels is half of the input; The feature map D is passed through a pointwise convolution with a convolution kernel of 1 to obtain a compressed feature map D, where the number of compressed channels is half of the input; The compressed feature map C and the feature map D are concatenated for the first time (Concat) to obtain a feature map after the first concatenation; The feature map after the first concatenation is passed through an average pooling layer and a max pooling layer in parallel to output their respective pooling results, and the second concatenation (Concat) is performed according to their respective pooling results to obtain a feature map after the second concatenation, where both the average pooling and the max pooling are performed in the channel dimension; The feature map after the second concatenation is passed through a depthwise convolution with a convolution kernel of 7 and a Sigmoid activation function to generate a heat map; According to the heat map, it is evenly divided in the channel dimension, each part is multiplied by the compressed feature map C and the feature map D respectively, and the results after multiplication are added together to complete the signal integration of different receptive fields at the same spatial position under different granularity information; The heat map after the signal integration of different receptive fields is passed through a pointwise convolution with a convolution kernel of 1 for signal fusion, and the result of the fusion is multiplied by the signal features after the number of extended channels to complete the screening of high-value regions.
[0041] This embodiment provides the following specific implementation cases: The images in the dataset used are from different angles of magnetic resonance imaging (MRI) scans, including sagittal, axial, and coronal planes. The dataset contains 5,249 high-quality MRI images of brain tumors with detailed annotations. The diverse images can ensure comprehensive coverage of the brain anatomy, enhancing the robustness of the model trained on this dataset. The image bounding boxes are manually annotated using the image annotation tool (LabelImg) to ensure high accuracy and reliability of the image annotations. For the specific division statistics of the Tumor dataset, see Table 1. Table 1, Tumor Dataset Division Table Some configuration information for the model training of the present invention is shown in Table 2. Table 2, Partial Configuration Table for Model Training For fairness, the training strategy adopted for the model training of the present invention is consistent with YOLOv8n. The framework used for training is MMYOLO, and MMYOLO makes some fine-tuning to the YOLOv8n architecture but does not affect the final performance. There is a certain deviation between the parameter count statistics and the official statistics, but it does not affect the results. For the comparison of the final effects, see Table 3.
[0042] Table 3, Comparison of Experimental Effects From the comparison of the experimental effects, it can be seen that the recognition ability of the model of this patent has been greatly improved compared with YOLOv8n for micro-cellular tumors such as gliomas, and there is still a large room for further improvement.
[0043] According to the nuclear magnetic resonance (NMR) image brain tumor detection method of the present invention, by improving the C2f structural block in the YOLOv8n model, a dual-convolution cross-stage division idea is proposed, aiming to enhance the modeling ability of potential high-dimensional feature information, improve the attention to the semantic information of high-value channels, and further extract the semantic information of high-value channels in the subsequent division block. And to avoid NaN in division, the two divisor branches are first processed by the Sigmoid function to avoid the situation of divisor being 0; the large-kernel convolution in the large-kernel dynamic selectable convolution structural block is improved by the reparameterization technology of the reparameterized dynamic selectable large-kernel convolution (RP-LSKM) network to improve the information extraction ability; the PVSS module enhances the model state space modeling and model global representation learning ability at a lower computational cost, and the linear self-attention mechanism used avoids the excessive overhead brought by the quadratic computational complexity of self-attention. Thus, the improved YOLOv8n model proposed in this application can not only perform context modeling at different scales in the channel dimension but also perform context modeling in the global space representation. The NMR image brain tumor detection model based on the improved YOLOv8n model can improve the accuracy and efficiency of brain tumor detection.
[0044] Embodiment 2 As Figure 7 shown, this embodiment provides a nuclear magnetic resonance (NMR) image brain tumor detection system, which is used to implement the detection method as described above, and includes: An acquisition module, configured to acquire an initial image dataset of nuclear magnetic resonance brain tumors, perform annotation and divide it into a training set, a validation set, and a test set; A first construction module, configured to construct a dual-convolution cross-stage division (DivC2f) network and a partial visual state space (PVSS) network, use the dual-convolution cross-stage division (DivC2f) network as the backbone network in the YOLOv8 model, use the partial visual state space (PVSS) network as the feature enhancement network after SPPF in the YOLOv8 model, and adjust the backbone network of the YOLOv8 model to lightweight the YOLOv8 model while improving performance; the dual-convolution cross-stage division network includes pointwise convolution, batch normalization, SiLU activation function, ReLU activation function, feature segmentation, feature splicing, channel attention layer (SElayer) processing, and a feature interaction dual-stream division (DivOpera) network that divides feature signals pointwise; A second construction module, configured to use the constructed reparameterized dynamic selectable large-kernel convolution (RP-LSKM) network as a multi-scale feature interaction network in the neck structure of the YOLOv8 model with improved performance, and at the same time use the reparameterization technology to further improve the performance of the YOLOv8 model to obtain an improved YOLOv8 model; A training module for iteratively optimizing and training the improved YOLOv8 model according to the initial image dataset to obtain a nuclear magnetic resonance (NMR) image brain tumor detection model; An output module for inputting the initial image dataset into the NMR image brain tumor detection model for detection and outputting the NMR brain tumor detection result.
[0045] A NMR image brain tumor detection system in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a Network Attached Storage (NAS), a personal computer (PC), etc., which are not specifically limited in the embodiments of the present application.
[0046] A NMR image brain tumor detection system in an embodiment of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.
[0047] A NMR image brain tumor detection system provided in an embodiment of the present application can implement Figure 1 each process implemented by a method embodiment of a NMR image brain tumor detection method. To avoid repetition, it will not be elaborated here.
[0048] According to the nuclear magnetic resonance (NMR) image brain tumor detection system of the embodiments of the present invention, by improving the C2f structural block in the YOLOv8n model, a dual-convolution cross-stage division idea is proposed, aiming to enhance the modeling ability of potential high-dimensional feature information, improve the attention to the semantic information of high-value channels, and further extract the semantic information of high-value channels in the subsequent division block. And to avoid NaN in the division, the two divisor branches are first processed by the Sigmoid function to avoid the situation of the divisor being 0. The large kernel convolution in the large kernel dynamic selectable convolution structural block is improved by the reparameterization technology of the reparameterized dynamic selectable large kernel convolution (RP-LSKM) network to improve the information extraction ability. The PVSS module enhances the model state space modeling and the model global representation learning ability at a lower computational cost, and the linear self-attention mechanism used avoids the excessive overhead brought by the quadratic computational complexity of self-attention. Thus, the improved YOLOv8n model proposed in this application can not only perform context modeling at different scales in the channel dimension but also perform context modeling in the global space representation. The NMR image brain tumor detection model based on the improved YOLOv8n model can improve the accuracy and efficiency of brain tumor detection.
[0049] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above-mentioned embodiment of a NMR image brain tumor detection method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0050] An embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, it implements each process of the above-mentioned embodiment of a NMR image brain tumor detection method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0051] Among them, the processor is the processor in the electronic device in the above-mentioned embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc.
[0052] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the invention.
[0053] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example.
[0054] Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. The mention of "embodiment" in this article means that the specific features, structures, or characteristics described in connection with the embodiment can be included in at least one embodiment of this application. The appearance of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. Those skilled in the art can explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0055] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
Claims
1. A method for detecting brain tumors using nuclear magnetic resonance images, characterized in that: include: S1, obtain the initial image dataset of MRI brain tumors, annotate it and divide it into training set, validation set and test set; S2, constructing a dual convolutional cross-stage division network and a partial visual state space network, using the dual convolutional cross-stage division network as the backbone network in the YOLOv8 model, using the partial visual state space network as the feature enhancement network after SPPF in the YOLOv8 model, and adjusting the YOLOv8 model backbone network to lighten the YOLOv8 model while improving performance; the dual convolutional cross-stage division network includes Point-wise convolution, batch normalization, SiLU activation function, ReLU activation function, feature segmentation, feature concatenation, channel attention layer processing, and feature interaction two-stream division network through point-by-point division of feature signals; S3, in the YOLOv8 model with improved performance, the constructed reparameterized dynamically selectable large kernel convolutional network is used as a multi-scale feature interaction network in the neck structure, and the performance of the YOLOv8 model is further improved by reparameterization to obtain an improved YOLOv8 model; S4, using the initial image data set to iteratively optimize and train the improved YOLOv8 model to obtain a brain tumor detection model for MRI images; S5, inputting the initial image data set into the MRI brain tumor detection model for detection, and outputting the MRI brain tumor detection result.
2. The method for detecting brain tumors using nuclear magnetic resonance images according to claim 1, characterized in that: In S2, the processing process of the dual convolution cross-stage division network includes: The input features are processed through a separation convolution module consisting of a point-by-point convolution with a convolution kernel of 1, batch normalization, and SiLU activation function. The input features are evenly divided into feature maps A and B in the channel dimension. Feature map A is processed by a channel attention layer to obtain high-value channel signal features. Construct a three-convolution module, and pass the feature map B through two branches in parallel into the three-convolution module in each branch, so as to control the range of the kernel function and perform low-dimensional representation and integration of the feature signal while integrating the input features; The feature signal of the first branch of the dual branch is processed by a ReLU activation function and a Sigmoid function respectively, and the feature signal of the second branch is mirrored; The second branch uses the output of the first branch after being processed by the ReLU activation function as the dividend, and the output of the second branch after being processed by the Sigmoid activation function as the divisor, divides the results of the two branches to obtain the first output result, and performs a mirror operation on the first branch in the two-stream division network to obtain the second output result; splicing the first output result, the second output result and the high-value channel signal feature in a channel dimension to obtain an output channel signal; The output channel signal is convolved and then outputted.
3. The method for detecting brain tumors using nuclear magnetic resonance images according to claim 2, characterized in that: In S2, the channel attention layer processing is performed on the feature map A to obtain high-value channel signal features, specifically: Receive feature map A and perform global average pooling and compression integration to obtain global features; The global feature is passed through a weighted fusion module to interact and weightedly fuse the channel signals of the global feature; The weighted fused channel signal enters the HardSigmoid activation function to obtain a high-value channel signal feature weight; The high-value channel signal feature weight is multiplied with the input feature to extract the high-value channel signal feature.
4. The method for detecting brain tumors using nuclear magnetic resonance images according to claim 1, characterized in that: In S2, the partial visual state space network processing process includes: Receiving features processed by a double convolutional cross-stage division network, passing through a convolution layer with a convolution kernel of 1 to obtain output features, passing the output features through a residual network block in a partial visual state space network to obtain fused multi-scale features, wherein the residual network block includes a Mamba layer and a multi-layer perceptron layer; The multi-scale features are concatenated with the output features, and then passed through a convolution layer with a convolution kernel of 1 to output the enhanced output features.
5. The method for detecting brain tumors using nuclear magnetic resonance images according to claim 3, characterized in that: In S2, the weighted fusion module consists of a point-by-point convolution with a convolution kernel of 1, a batch normalization regularization function, a SiLU activation function and a point-by-point convolution with a convolution kernel of 1.
6. The method for detecting brain tumors using nuclear magnetic resonance images according to claim 1, characterized in that: In S3, the processing process of the re-parameterized dynamic selectable large kernel convolution includes: Receive the feature map output in S2 to the image block embedding module to obtain the signal features after the number of channels is expanded; The signal features after the expansion of the number of channels are passed through n series-connected random projection local semantic kernel modules to enhance the feature expression ability of the model, wherein the random projection local semantic kernel module includes a network layer with a large selective kernel based on a receptive field pyramid and a depth-separable multi-layer perceptron layer.
7. The method for detecting brain tumors using nuclear magnetic resonance images according to claim 6, characterized in that: The process of entering the signal features after the expansion of the number of channels into the network layer of the large selective kernel based on the receptive field pyramid includes: The signal features after the expansion of the number of channels are connected in series through a point-by-point convolution and a GELU activation function for nonlinear integration; The signal features after nonlinear integration are subjected to a depthwise convolution consisting of n parallel convolution kernels with a value of 5 to output a feature map C; The feature map C is subjected to a dilated convolution consisting of n parallel convolution kernels of 7 and a dilation rate of 3 to output a feature map D for dynamically expanding the receptive field of the convolution; Compressing the feature map C and the feature map D through a point-by-point convolution to obtain a compressed feature map C and a compressed feature map D; Performing a first splicing of the compressed feature map C and the feature map D to obtain a first spliced feature map; Passing the first spliced feature maps through an average pooling layer and a maximum pooling layer in parallel to output respective pooling results, and performing a second splicing according to the respective pooling results to obtain the second spliced feature maps; The feature map after the second splicing is subjected to a depth convolution with a convolution kernel of 7 and a Sigmoid activation function to generate a heat map. According to the heat map, the heat map is evenly divided according to the channel dimension, each part is multiplied with the compressed feature map C and feature map D respectively, and the multiplication results are added to complete the signal integration of different receptive fields at the same spatial position under different granularity information; The heat map after integrating the signals of different receptive fields is subjected to point-by-point convolution with a convolution kernel of 1 for signal fusion, and the fusion result is multiplied with the signal characteristics after the number of expanded channels to complete the screening of high-value areas.
8. A brain tumor detection system using nuclear magnetic resonance images, characterized in that: Used to implement the detection method according to any one of claims 1 to 7, comprising: An acquisition module is used to acquire the initial image dataset of MRI brain tumors, annotate it, and divide it into a training set, a validation set, and a test set; The first building module is used to build a dual convolutional cross-stage division network and a partial visual state space network, use the dual convolutional cross-stage division network as the backbone network in the YOLOv8 model, use the partial visual state space network as the feature enhancement network after SPPF in the YOLOv8 model, and adjust the YOLOv8 model backbone network to lightweight the YOLOv8 model while improving performance; the dual convolutional cross-stage division network includes Point-wise convolution, batch normalization, SiLU activation function, ReLU activation function, feature segmentation, feature concatenation, channel attention layer processing, and feature interaction two-stream division network through point-by-point division of feature signals; A second construction module is used to use the constructed reparameterized dynamically selectable large kernel convolutional network as a multi-scale feature interaction network in the neck structure in the YOLOv8 model with improved performance, and further improve the performance of the YOLOv8 model by reparameterization to obtain an improved YOLOv8 model; A training module, used for iteratively optimizing and training the improved YOLOv8 model according to the initial image data set to obtain a brain tumor detection model for nuclear magnetic resonance images; The output module is used to input the initial image data set into the nuclear magnetic resonance image brain tumor detection model for detection, and output the nuclear magnetic resonance brain tumor detection result.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the method for detecting brain tumors using nuclear magnetic resonance images as described in any one of claims 1 to 7 when executing the computer program.
10. A computer storage medium, characterized in that: The computer storage medium stores instructions, and when the instructions are executed on a computer, the computer executes the method for detecting brain tumors using magnetic resonance imaging according to any one of claims 1 to 7.
Citation Information
Patent Citations
MRI (Magnetic Resonance Imaging) brain tumor segmentation method based on multi-scale feature fusion of improved U-Net
CN117876399A
Improved YOLOv8-based dangerous behavior detection method for chemical enterprise personnel
CN118038555A
Unified framework for multigrid neural network architecture
US20230196750A1
Cited By
Crop leaf disease and pest detection method and system
CN121438115A
A method and system for detecting crop leaf diseases and pests
CN121438115B
Brain tumor classification method and device based on magnetic resonance image, and medium
CN121811115A