A primary brain tumor subtype image intelligent recognition method

CN122780705APending Publication Date: 2026-09-18SOUTHWEST MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611251800.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-18
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0005]尽管上述方法在脑肿瘤分类中取得了一定效果,但仍存在不足:一是,部分深度学习模型采用较深或较复杂的网络结构,参数量和计算量较高,对硬件资源要求较大,不利于在临床场景或资源受限设备中快速部署

Benefits of technology

1、本发明通过轻量级卷积主干模块提取脑部MRI图像的空间结构特征,通过极坐标傅里叶环形频域描述模块提取不同频率环带上的统计特征,并利用频率提示门控增强模块将频域信息转化为通道级调制权重,对空间特征进行自适应增强;采用径向基函数科尔莫戈罗夫-阿诺德分类模块对增强后的图像特征进行非线性映射和类别判别;本发明在保持网络轻量化的同时充分融合了空间域和频域信息,增强了脑肿瘤MRI图像的判别性特征表达能力,降低了不同肿瘤类别之间的混淆,实现了脑肿瘤类别的高效、稳定和准确识别。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122780705A_ABST
    Figure CN122780705A_ABST
Patent Text Reader

Abstract

The application discloses a kind of primary brain tumor subtype image intelligent identification method, comprising: constructing brain MRI image dataset;End-to-end frequency domain prompt radial basis function Kolmogorov-Arnold classification network is constructed, including lightweight convolution main stem feature extraction module, polar coordinate Fourier ring frequency domain description module, frequency prompt gate enhancement module and radial basis function Kolmogorov-Arnold classification module;Network is trained end to end using training set, adopts cross-entropy loss function with label smoothing, and parameter is updated by AdamW optimizer combined with cosine annealing learning rate scheduling strategy;The image to be identified is input into the network trained, and the brain tumor class prediction result is output.The application reduces the complexity of model by lightweight design, enhances the feature discriminability by fusing spatial domain and frequency domain information, and improves the class discrimination ability by using strong nonlinear classifier, to realize the efficient, stable and accurate identification of brain tumor class.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image analysis technology, and in particular to an intelligent method for identifying subtypes of primary brain tumors. Background Technology

[0002] Brain tumors are an important type of central nervous system disease. Common brain tumors include gliomas, meningiomas, and pituitary tumors, among others. Different brain tumors vary significantly in their tissue origin, biological behavior, treatment methods, and prognosis. Accurate identification of brain tumor types is crucial for assisting clinical diagnosis, developing individualized treatment plans, and improving patient outcomes.

[0003] Magnetic resonance imaging (MRI), with its high soft tissue resolution, has become an important imaging tool for the examination and diagnosis of brain tumors. Brain MRI images can display the location, shape, boundaries, signal intensity, and relationship with surrounding tissues of tumors, providing crucial information for brain tumor classification. However, different types of brain tumors may exhibit overlapping phenotypic patterns on MRI images, especially when tumor boundaries are unclear, signal characteristics are similar, lesions are small, or image quality is limited. Relying solely on manual image interpretation can be easily influenced by physician experience and subjective judgment. Existing research on the three-class classification of brain tumors in T1-weighted enhanced MRI images points out the complex differences in imaging manifestations among meningiomas, gliomas, and pituitary adenomas. It proposes using tumor region amplification and ring partitioning features to improve classification performance, demonstrating the significant discriminative value of regional structure and surrounding tissue information in brain tumor MRI classification.

[0004] In recent years, artificial intelligence methods have been widely applied to brain tumor MRI image classification tasks. Convolutional neural networks (CNNs) can automatically learn edge, texture, morphological, and high-level semantic features in medical images, reducing reliance on manual feature design. In studies using CNNs to classify brain tumors from MRI images, CNNs have demonstrated good feature learning capabilities in automatic brain tumor image recognition. Furthermore, deep transfer learning methods have also been used for brain tumor classification. For example, studies employing deep convolutional neural networks combined with data augmentation strategies to classify brain tumors of different grades have proven that deep learning models can improve classification performance in medical image analysis tasks. Multi-scale spatial feature modeling also contributes to improving the performance of brain tumor MRI image analysis.

[0005] While the aforementioned methods have achieved some success in brain tumor classification, they still have shortcomings: First, some deep learning models employ deep or complex network structures, resulting in high parameter and computational costs, and demanding significant hardware resources, hindering rapid deployment in clinical settings or resource-constrained environments. Although methods such as CNNs, capsule networks, and visual Transformers have been widely applied to brain tumor MRI classification, model complexity, generalization ability, data scale, and clinical interpretability remain issues that require further improvement. Second, frequency distributions in medical images can reflect texture variations, edge structures, and tissue differences, but existing brain tumor classification methods primarily focus on spatial domain feature modeling, underutilizing the potential frequency domain information in MRI images. Furthermore, traditional brain tumor classification networks typically use ordinary fully connected layers or simple linear classification heads for category discrimination, limiting their ability to nonlinearly fit complex medical image features. When different types of tumors exhibit similar image phenotypes on MRI images, the model is prone to category confusion. Summary of the Invention

[0006] To address the aforementioned issues, this invention provides an intelligent image recognition method for primary brain tumor subtypes, which can balance lightweight feature extraction, spatial-frequency domain joint representation, and strong nonlinear discrimination capabilities, thereby improving the accuracy, stability, and deployment efficiency of brain tumor MRI image classification.

[0007] This invention provides an intelligent image recognition method for primary brain tumor subtypes, the specific technical solution of which is as follows: Construct a dataset of brain MRI images; The dataset contains four types of sample images: glioma, meningioma, pituitary adenoma, and no tumor. Specifically, the four types of raw images are preprocessed by size normalization, data augmentation, and standardization to obtain sample images, which are then divided into training and test sets according to a set ratio. An end-to-end lightweight frequency domain cue radial basis function Kolmogorov-Arnold classification network is constructed. The classification network includes a lightweight convolutional backbone feature extraction module, a polar coordinate Fourier ring frequency domain description module, a frequency cue gating enhancement module, and a radial basis function Kolmogorov-Arnold classification module. The convolutional backbone feature extraction module is used to extract spatial feature maps from the input brain MRI image; the polar coordinate Fourier ring frequency domain description module is used to receive the input brain MRI image, process it, and output a frequency domain description vector; the frequency cue gating enhancement module is used to generate channel-level gating weights based on the frequency domain description vector, and fuse the gating weights with the spatial feature map to obtain a frequency cue enhanced spatial feature map; the radial basis function Kolmogorov-Arnold classification module is used to perform feature processing on the enhanced spatial feature map, and perform classification and prediction processing to obtain the final prediction result. The classification network was trained end-to-end using the training set images of the brain MRI image dataset. The labeled smooth cross-entropy loss function was used as the training objective. The network parameters were iteratively updated using the AdamW optimizer combined with a cosine annealing learning rate scheduling strategy. The brain MRI image to be identified is input into the trained classification network, which outputs a brain tumor category prediction result.

[0008] Furthermore, the specific processing procedure of the lightweight convolutional backbone feature extraction module is as follows: The input brain MRI image is sequentially subjected to two-dimensional convolution, batch normalization, and SiLU nonlinear activation to obtain a feature map; The input features are transformed by channel dimension through the first pointwise convolution, spatial features are extracted by the 3×3 depthwise convolution, channel dimension is restored by the second pointwise convolution, and the input features are added to the convolutionally transformed features by the residual connection to obtain the output features of the lightweight residual block.

[0009] Furthermore, during the feature extraction process, the output features are used as input to subsequent network layers. After multiple convolution downsampling operations with a stride of 2, the spatial resolution of the feature map is gradually reduced and the number of channels is increased. Finally, the spatial feature map is sent to the frequency cue gating enhancement module.

[0010] Furthermore, the processing procedure of the polar coordinate Fourier ring frequency domain description module is as follows: Convert the input image to a grayscale image; Perform a two-dimensional fast Fourier transform on the grayscale image and center it to shift the low-frequency components to the center of the spectrum; Calculate the logarithmic magnitude spectrum; Constructing in the spectral plane R A polar coordinate annular mask is used to obtain a frequency domain annular description vector by weighted averaging of the spectral amplitude within each frequency annular band. Perform a logarithmic transformation on the frequency domain ring description vector and Normalization yields the frequency domain description vector.

[0011] Furthermore, the number of polar coordinate annular masks R Set it to 12.

[0012] Furthermore, the enhancement process of the frequency cue gating enhancement module is as follows: The frequency domain description vector is sequentially subjected to layer normalization, first linear mapping, GELU activation, second linear mapping, and Sigmoid activation to generate channel-level gated weights; The gating weights are fused with the spatial feature map by channel-by-channel multiplication to obtain a spatial feature map with enhanced frequency cues.

[0013] Furthermore, the specific processing procedure of the radial basis function Kolmogorov-Arnold classification module is as follows: The enhanced spatial feature map is subjected to global average pooling and layer normalization in sequence to obtain normalized image-level feature vectors; The normalized image-level feature vector is input into the radial basis function Kolmogorov-Arnold classifier; Nonlinear mapping is achieved through Gaussian radial basis function response while preserving the basic linear mapping branch, and the brain tumor category prediction score is output. The brain tumor category prediction score is used to obtain the prediction probability of each category through the Softmax function, and the category corresponding to the maximum probability is taken as the final prediction result.

[0014] Furthermore, the classifier has [specific features] on each input dimension. K There are radial basis function centers, and each center point is uniformly distributed within a preset interval.

[0015] Furthermore, the number of radial basis function centers K Set it to 8.

[0016] Furthermore, the labeled smoothed cross-entropy loss function is expressed as follows:

[0017] in, Represents classification loss, This represents the labeled smooth cross-entropy loss function. This represents the original category score vector output by the network. This represents the true category label corresponding to the input image.

[0018] The beneficial effects of this invention are as follows: 1. This invention extracts spatial structural features from brain MRI images using a lightweight convolutional backbone module, extracts statistical features on different frequency bands using a polar coordinate Fourier ring frequency domain description module, and uses a frequency cue gating enhancement module to convert frequency domain information into channel-level modulation weights for adaptive enhancement of spatial features. A radial basis function Kolmogorov-Arnold classification module is used to perform nonlinear mapping and category discrimination on the enhanced image features. This invention fully integrates spatial and frequency domain information while maintaining network lightweightness, enhances the discriminative feature expression ability of brain tumor MRI images, reduces confusion between different tumor categories, and achieves efficient, stable, and accurate identification of brain tumor categories.

[0019] 2. This invention designs a lightweight convolutional backbone feature extraction module, which uses multiple sets of "convolution-batch normalization-SiLU activation" units and lightweight residual blocks to form the backbone network. The lightweight residual blocks adopt a bottleneck structure of pointwise convolution → depthwise convolution → pointwise convolution, and enhance gradient propagation stability and feature reuse capability through residual connections. This module achieves a lightweight network design through depthwise separable convolution and bottleneck residual structure, which significantly reduces the number of model parameters and computational complexity.

[0020] 3. The polar coordinate Fourier ring frequency domain description module of the present invention transforms the frequency domain energy distribution into structured statistical features through the polar coordinate ring frequency domain description method, effectively extracting frequency domain information that reflects texture changes, edge structures and tissue differences, providing frequency cue information for subsequent spatial feature enhancement, and making up for the limitations of simple spatial domain feature modeling.

[0021] 4. The frequency cue gating enhancement module described in this invention achieves effective fusion of frequency domain information and spatial features, transforming the frequency domain description vector into a channel-level modulation signal. It can adaptively enhance discriminative spatial features related to brain tumor structure, edges, and texture, while suppressing class-independent redundant responses. By selectively enhancing spatial features through frequency domain cues, it improves the expressive ability of tumor region-related features and effectively reduces confusion between different tumor categories caused by similar image phenotypes. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of the network framework structure of the present invention.

[0023] Figure 2 This is a schematic diagram of the ROC curve of the method of the present invention.

[0024] Figure 3 This is a schematic diagram of the confusion matrix of the method of the present invention. Detailed Implementation

[0025] The technical solutions in the embodiments of the present invention are clearly and completely described in the following description. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0026] In the description of the embodiments of the present invention, it should be noted that the indicated orientation or positional relationship is based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the product of the invention is conventionally placed during use, or the orientation or positional relationship in which those skilled in the art conventionally understand it during use. This is only for the convenience of describing the present invention and simplifying the description, and is not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of the present invention. Furthermore, the terms "first" and "second" are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0027] In the description of the embodiments of the present invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set" and "connection" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.

[0028] Example 1 Embodiment 1 of the present invention discloses an intelligent image recognition method for primary brain tumor subtypes, such as... Figure 1 As shown, Figure 1 In this diagram, ConvBNAct represents the convolution-batch normalization-activation module, LiteResidualBlock represents a lightweight residual block, LayerNorm represents layer normalization, Linear represents a fully connected layer, GELU represents the Gaussian error linear unit activation function, Sigmoid represents the sigmoid activation function, RBFKANLayer represents the radial basis function Kolmogorov-Arnold network layer, Dropout represents a random deactivation layer, Softmax represents the softness maximum function probability normalization layer, feature vector represents the feature vector, flatten represents the flattening operation, normalized feature vector represents the normalized feature vector, g represents the channel gate weight, F represents the spatial feature map, and Fgated represents the enhanced feature map. Argmax represents the maximum index selection, and RBF-KAN represents the radial basis function Kolmogorov-Arnold network.

[0029] The specific method is as follows: Construct a dataset of brain MRI images; The dataset contains four types of sample images: glioma, meningioma, pituitary adenoma, and no tumor. Specifically, the four types of raw images were preprocessed by size normalization, data augmentation, and standardization to obtain sample images, which were then divided into training and test sets in an 8:2 ratio. MRI imaging data for four types of tumors—glioma, meningioma, pituitary adenoma, and tumor-free tumors—can be obtained from a magnetic resonance imaging (MRI) scanner.

[0030] An end-to-end lightweight frequency domain cue radial basis function Kolmogorov-Arnold classification network is constructed. The classification network includes a lightweight convolutional backbone feature extraction module, a polar coordinate Fourier ring frequency domain description module, a frequency cue gating enhancement module, and a radial basis function Kolmogorov-Arnold classification module. The convolutional backbone feature extraction module is used to extract spatial feature maps from the input brain MRI image; In this embodiment, the input brain MRI image is recorded as:

[0031] in, This represents the input brain MRI image tensor; Represents the real number field, meaning the pixel values ​​in the image are real numbers; and These represent the height and width of the input image, respectively, with 3 channels. The lightweight convolutional backbone feature extraction module extracts spatial representation features from the input image, and its output spatial feature map is represented as follows:

[0032] in, This represents a lightweight convolutional backbone network. This represents the extracted spatial feature map.

[0033] The lightweight convolutional backbone consists of multiple sets of convolutional normalized activation units and lightweight residual blocks. Each convolutional normalized activation unit is formed by sequentially connecting a two-dimensional convolutional layer, a batch normalized layer, and a SiLU activation function. The calculation process is as follows:

[0034] in, This represents a two-dimensional convolution operation. This indicates a batch normalization operation. This represents the SiLU nonlinear activation function. This represents the feature map output after convolution, batch normalization, and nonlinear activation.

[0035] The lightweight residual block uses pointwise convolution, depthwise convolution, and pointwise convolution to form a bottleneck structure. The computation process is represented as follows:

[0036] in, and They represent Pointwise convolution, express Depthwise convolutions and residual connections are used to combine input features Features after convolution transformation The features are added together to enhance gradient propagation stability and feature reuse capability, ultimately yielding the output features of the lightweight residual module.

[0037] During feature extraction, the output features of lightweight residual blocks As input to subsequent network layers, the feature map undergoes multiple convolutional downsampling operations with a stride of 2, progressively reducing the spatial resolution of the feature map and increasing the number of channels. This expands the receptive field and yields higher-level semantic features, ultimately resulting in the spatial feature map. The final spatial feature map It is then sent to the subsequent frequency prompt gating enhancement module.

[0038] The polar coordinate Fourier ring frequency domain description module is used to receive and process the input brain MRI image, and output a frequency domain description vector; specifically, it first converts the input image into a grayscale image:

[0039] in, Indicates the number of image channels. Indicates the first Images of each channel, This represents a grayscale image.

[0040] Subsequently, a two-dimensional Fourier transform is performed on the grayscale image. And shift the low-frequency components to the center of the spectrum:

[0041] in, This represents a two-dimensional Fast Fourier Transform. This indicates a spectrum centering operation.

[0042] Next, the amplitude response of the Fourier spectrum is calculated, and a logarithmic transform is performed:

[0043] in, This represents the logarithmic magnitude spectrum.

[0044] To obtain statistical characteristics at different frequency radii, a model is constructed in the spectral plane. A polar coordinate annular mask:

[0045] in, Indicates the number of ring bands. For the first A ring-shaped mask. In this embodiment, It can be set to 12. Each annular mask corresponds to a frequency radius interval and has been normalized. The frequency domain statistical characteristics of each frequency band can be expressed as:

[0046] in, Indicates the first Frequency domain statistical characteristic values ​​of each ring frequency band; Represents the logarithmic magnitude spectrum on the frequency coordinate. The amplitude at that point; Indicates the first A ring-shaped mask in frequency coordinates Normalized weights at each location; and These represent the height and width of the spectrum, respectively. Because... Only in the The nth annular region has non-zero weights, therefore this formula is equivalent to applying the nth... The frequency response characteristics of a frequency band are obtained by weighted averaging of the spectral amplitudes within each frequency band. This yields the frequency domain ring description vector:

[0047] Subsequently, a logarithmic transformation is performed on the frequency domain ring description vector and Normalization:

[0048] in, This represents the final frequency domain description vector. This represents the frequency domain annular energy vector obtained through statistical analysis of the annular region in polar coordinates. Indicates to Perform an element-wise logarithmic transformation. express Norms are used to calculate the Euclidean length of a vector after a logarithmic transformation. To prevent the denominator from being too small or from causing division by zero, a very small constant is used.

[0049] This frequency domain description vector can characterize the energy distribution of an image in different frequency bands, providing frequency cue information for subsequent spatial feature enhancement.

[0050] The frequency cue gating enhancement module is used to generate channel-level gating weights based on the frequency domain description vector, and fuse the gating weights with the spatial feature map to obtain a frequency cue enhanced spatial feature map; In this embodiment, the spatial feature map output by the lightweight convolutional backbone is assumed to be:

[0051] in, Indicates the number of feature channels. and Let these represent the feature map height and width, respectively. Let the frequency domain description vector output by the polar Fourier ring frequency domain description module be:

[0052] in, Represents the frequency domain description vector For one 3D real-valued vector, This represents the number of ring bands. The frequency cue gating enhancement module first performs layer normalization and nonlinear mapping on the frequency domain description vector to obtain the channel-level gating weights:

[0053] in, Representation layer normalization, and They represent the linear mapping parameters, This represents the GELU activation function. This represents the Sigmoid function. This represents the generated channel-level gating weights.

[0054] Subsequently, the gating weights are extended to the spatial dimension, and frequency cue enhancement is applied to the spatial feature map:

[0055] Or equivalently represented as:

[0056] in, This indicates channel-by-channel multiplication. This represents the spatial feature map after frequency cues enhancement.

[0057] In this way, the frequency domain description vector is transformed into a channel-level modulated signal to enhance the discriminative spatial features related to brain tumor structure, edges, and texture, while suppressing class-independent redundant responses.

[0058] The radial basis function Kolmogorov-Arnold classification module is used to perform feature processing on the enhanced spatial feature map, and to perform classification and prediction processing to obtain the final prediction result; Specifically, after obtaining the spatial feature map enhanced with frequency cues... Then, image-level feature vectors are first obtained through global average pooling:

[0059] in, This indicates a global average pooling operation. Indicated by enhanced feature maps The resulting image-level feature vectors.

[0060] Subsequently, layer normalization is performed on the image-level feature vectors:

[0061] in, Presentation layer normalization operation; Indicates to The normalized image-level feature vector obtained after layer normalization is then input into the radial basis function Kolmogorov-Arnold classifier, which outputs the class prediction score.

[0062] in, This represents the Kolmogorov-Arnold classifier with radial basis functions. This represents the original class score vector output by the classifier.

[0063] And the predicted probabilities for each category are obtained through the flexible maximum function:

[0064] in, , Indicates the number of categories. Indicates that the sample belongs to the first The predicted probability of the class. The final class prediction result is:

[0065] in, This indicates the type of brain tumor ultimately predicted by the model. Indicates that the sample belongs to the first The predicted probabilities of each category, For category indexing, This indicates that the category index corresponding to the maximum predicted probability among all candidate categories is selected and used as the final classification result of the model.

[0066] The classifier consists of a Kolmogorov-Arnold radial basis function network layer, a Gaussian error linear unit activation layer, a random deactivation layer, and an output radial basis function Kolmogorov-Arnold layer.

[0067] For the Kolmogorov-Arnold layer of radial basis functions, let the input features be:

[0068] in, This represents the feature vector input to the Kolmogorov-Arnold layer of radial basis functions. Indicates the input feature dimension. Indicates the first Feature values ​​on each input dimension.

[0069] Set on each input dimension Radial basis function centers:

[0070] in, This indicates the number of radial basis function centers. Indicates the first Each radial basis function center. In this embodiment, Set to 8, and each center point can be evenly distributed within the preset interval.

[0071] No. The input dimension at the th ... The radial basis function responses at each center are:

[0072] in, Representing input features Compared to the first Radial basis function centers The response value; This is a learnable radial basis function width control parameter used to adjust the width of the radial basis function response range. When The closer to the center hour, The larger the value, the more... Far from the center As time progresses, the response value gradually decreases. The output of the radial basis function Kolmogorov-Arnold layer is expressed as:

[0073] in, Indicates the first The response values ​​of each output node; Indicates the dimension of the input features; This represents the number of radial basis function centers set on each input dimension; Represents the input vector The One eigenvalue; Indicates the first The input feature at the th ... The response values ​​at the center of each radial basis function; This represents the corresponding learnable radial basis function weights; The branch of the basic linear mapping is represented in the th case. Output on each output node; Indicates the first The bias term of each output node.

[0074] The classification network was trained end-to-end using the training set images from the aforementioned brain MRI image dataset.

[0075] Let the true category label be The network prediction output is The labeled smooth cross-entropy loss function is used as the training objective:

[0076] in, Indicates classification loss; This represents the labeled, smoothed cross-entropy loss function; This represents the original category score vector output by the network; This represents the true class label corresponding to the input image. Label smoothing is used to reduce the model's overconfidence in a single class, thereby enhancing the model's generalization ability.

[0077] In this embodiment, the network parameters are updated using the AdamW optimizer, and the learning rate is dynamically adjusted using a cosine annealing learning rate scheduling strategy.

[0078] in, Represents the learnable parameters in the network. This represents the gradient of the classification loss function with respect to the network parameters. This represents the current learning rate. AdamW This indicates that the AdamW optimizer optimizes parameters based on adaptive parameter updates calculated from gradients, combined with decoupled weight decay. It continuously minimizes the classification loss. The model gradually learns discriminative features that can distinguish different types of brain tumors.

[0079] Through the above training process, the network can simultaneously learn the spatial structural features and frequency domain statistical features of brain MRI images, and use the radial basis function Kolmogorov-Arnold classifier to achieve nonlinear category discrimination.

[0080] The brain MRI image to be identified is input into the trained classification network, which outputs a brain tumor category prediction result.

[0081] This embodiment uses an independent test set to evaluate the trained network model. The test set contains 1600 brain MRI images, with 400 images for each of the four categories: pituitary adenoma, no tumor, meningioma, and glioma. After the model training is complete, the optimal model parameters that achieve the highest test accuracy are loaded, and the final performance is evaluated on the test set.

[0082] In this embodiment, accuracy (ACC), precision (PRE), sensitivity (SEN), specificity (SPE), F1 score, and area under the curve (AUC) are used as evaluation metrics. The calculation formulas for each metric are as follows:

[0083] In this model, true positive (TP) refers to the number of positive samples correctly classified by the model; true negative (TN) refers to the number of negative samples correctly classified by the model; false positive (FP) refers to the number of samples incorrectly classified as positive by the model; and false negative (FN) refers to the number of samples incorrectly classified as negative by the model. The F1 score is used to comprehensively evaluate the model's precision and recall, while AUC measures the model's overall ability to distinguish between different categories of samples.

[0084] also, Figure 2 The figure shows the ROC curve results of the method of this invention in a four-class classification task of brain MRI images. As can be seen from the figure, the ROC curves of each category are close to the upper left corner, indicating that the model has a good ability to distinguish pituitary tumors, tumor-free tumors, meningiomas, and gliomas.

[0085] Figure 3 This is the confusion matrix result of the method of this invention in a four-class classification task of brain MRI images. The color depth of each cell in the confusion matrix is ​​positively correlated with the corresponding value; the darker the color, the more samples in that category, and the lighter the color, the fewer samples. As can be seen from the figure, the model prediction results show obvious diagonal clustering characteristics in all four categories, and the number of misclassifications in the off-diagonal areas is small, indicating that the model has good classification consistency.

[0086] In summary, the polar coordinate Fourier ring frequency domain description module proposed in this invention can effectively extract frequency domain statistical features from brain MRI images. The frequency cue gating module helps to strengthen the feature expression related to brain tumor classification, and the radial basis function Kolmogorov-Arnold classification module further improves the model's nonlinear classification ability, thereby improving the reliability of brain tumor type identification. Test set validation results show that the method of this invention achieves high results in accuracy, precision, sensitivity, specificity, F1 score, and AUC, as shown in Table 1 below. Table 1: Performance Evaluation Results of Indicators

[0087] This method can accurately distinguish between four types of brain MRI images: pituitary adenoma, meningioma, glioma, and non-tumor images. It has good intelligent brain tumor identification capabilities and auxiliary diagnostic application value.

[0088] This invention is not limited to the specific embodiments described above. The invention extends to any new feature or combination disclosed in this specification, as well as any new method or process step or combination disclosed herein.

Claims

1. A method for intelligent image recognition of primary brain tumor subtypes, characterized in that, include: Construct a dataset of brain MRI images; An end-to-end frequency domain cue radial basis function Kolmogorov-Arnold classification network is constructed, including a convolutional backbone feature extraction module, a polar coordinate Fourier ring frequency domain description module, a frequency cue gating enhancement module, and a radial basis function Kolmogorov-Arnold classification module; The convolutional backbone feature extraction module is used to extract spatial feature maps from the input brain MRI image; the polar coordinate Fourier ring frequency domain description module is used to receive the input brain MRI image, process it, and output a frequency domain description vector; the frequency cue gating enhancement module is used to generate gating weights based on the frequency domain description vector, and fuse the gating weights with the spatial feature map to obtain a frequency cue enhanced spatial feature map; the radial basis function Kolmogorov-Arnold classification module is used to perform feature processing on the enhanced spatial feature map, and perform classification and prediction processing to obtain the final prediction result. The classification network was trained end-to-end using the training set images of the brain MRI image dataset. The labeled smooth cross-entropy loss function was used as the training objective. The network parameters were iteratively updated using the AdamW optimizer combined with a cosine annealing learning rate scheduling strategy. The brain MRI image to be identified is input into the trained classification network, which outputs a brain tumor category prediction result.

2. The intelligent image recognition method for primary brain tumor subtypes according to claim 1, characterized in that, The specific processing procedure of the convolutional backbone feature extraction module is as follows: The input brain MRI image is sequentially subjected to two-dimensional convolution, batch normalization, and SiLU nonlinear activation to obtain a feature map; The input features are transformed by channel dimension through the first pointwise convolution, then spatial features are extracted by the 3×3 depthwise convolution, and finally the channel dimension is restored by the second pointwise convolution. The input features are added to the convolutionally transformed features through residual connections to obtain the output features.

3. The intelligent image recognition method for primary brain tumor subtypes according to claim 2, characterized in that, During the feature extraction process, the output features are used as input to subsequent network layers. After multiple convolution downsampling operations with a stride of 2, the spatial resolution of the feature map is gradually reduced and the number of channels is increased. Finally, the spatial feature map is sent to the frequency cue gating enhancement module.

4. The intelligent image recognition method for primary brain tumor subtypes according to claim 1, characterized in that, The processing procedure of the polar coordinate Fourier ring frequency domain description module is as follows: Convert the input image to a grayscale image; Perform a two-dimensional fast Fourier transform on the grayscale image and then center it; Calculate the logarithmic magnitude spectrum; Constructing in the spectral plane R A polar coordinate annular mask is used to obtain a frequency domain annular description vector by weighted averaging of the spectral amplitude within each frequency annular band. Perform a logarithmic transformation on the frequency domain ring description vector and Normalization yields the frequency domain description vector.

5. The intelligent image recognition method for primary brain tumor subtypes according to claim 4, characterized in that, Number of polar coordinate ring masks R Set it to 12.

6. The intelligent image recognition method for primary brain tumor subtypes according to claim 1, characterized in that, The enhancement process of the frequency cue gating enhancement module is as follows: The frequency domain description vector is sequentially subjected to layer normalization, first linear mapping, GELU activation, second linear mapping, and Sigmoid activation to generate gated weights; The gating weights are fused with the spatial feature map by channel-by-channel multiplication to obtain a spatial feature map with enhanced frequency cues.

7. The intelligent image recognition method for primary brain tumor subtypes according to claim 1, characterized in that, The specific processing procedure of the radial basis function Kolmogorov-Arnold classification module is as follows: The enhanced spatial feature map is subjected to global average pooling and layer normalization in sequence to obtain normalized image-level feature vectors; The normalized image-level feature vector is input into the radial basis function Kolmogorov-Arnold classifier; Nonlinear mapping is achieved through Gaussian radial basis function response while preserving the basic linear mapping branch, and the brain tumor category prediction score is output. The brain tumor category prediction score is used to obtain the prediction probability of each category through the Softmax function, and the category corresponding to the maximum probability is taken as the final prediction result.

8. The intelligent image recognition method for primary brain tumor subtypes according to claim 7, characterized in that, The classifier has [specific features] on each input dimension. K There are radial basis function centers, and each center point is uniformly distributed within a preset interval.

9. The intelligent image recognition method for primary brain tumor subtypes according to claim 8, characterized in that, Number of radial basis function centers K Set it to 8.

10. The intelligent image recognition method for primary brain tumor subtypes according to claim 1, characterized in that, The labeled smooth cross-entropy loss function is expressed as follows: in, Represents classification loss, This represents the labeled smooth cross-entropy loss function. This represents the original category score vector output by the network. This represents the true category label corresponding to the input image.