A multimodal-based tumor image classification method for the amine region and a terminal device

Through the multimodal amine region tumor image classification method, discrete wavelet transformation and multi-graph convolution network are used to extract and classify tumor images in the intracranial saddle region, which solves the problem of diagnosis in the prior art and realizes efficient automatic tumor identification and classification.

CN114266924BActive Publication Date: 2025-07-11SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111593776.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-23
Publication Date
2025-07-11
Estimated Expiration
2041-12-23

AI Technical Summary

Technical Problem

The prior art cannot realize the automatic and accurate classification of tumor images in the intracranial saddle area. The excessively thick layer of MRI imaging makes it time-consuming and labor-intensive for clinicians to read the film, and the diagnosis is prone to errors.

Method used

The multimodal amine region tumor image classification method is adopted, and by obtaining T1-weighted MRI, T2-weighted MRI and enhanced T1-weighted MRI images, the low-frequency and high-frequency features are extracted using a 3D Resnet network with discrete wavelet transformation, and a feature similarity map is constructed with the Pearson correlation coefficient. The multi-graph convolution network is used for feature optimization, and the multi-head attention mechanism is used for feature fusion, and finally the full connection layer is used for classification.

Benefits of technology

It improves the accuracy of automatic classification of intracranial saddle tumors, reduces diagnosis time, and provides an effective means for clinicians to differential diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114266924B_ABST
    Figure CN114266924B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-modal based tumor image classification method for the amine region and a terminal device. Among them, the method includes the steps of: acquiring MRI images of three modalities of patients with intracranial amine region tumors; using a 3D Resnet network with discrete wavelet transform to extract features from the MRI images of the three modalities to obtain low-frequency features and high-frequency features of the three modalities; based on the low-frequency features and high-frequency features of the three modalities and the clinical information of patients with intracranial amine region tumors, constructing a low-frequency feature similarity map and a high-frequency feature similarity map through Pearson correlation coefficients; using a multi-graph convolutional network to optimize the features of the low-frequency feature similarity map and the high-frequency feature similarity map to obtain optimized features of the three modalities; using a multi-head attention mechanism to fuse the optimized features of the three modalities, and using a fully connected layer to classify the tumor images of the amine region. The multi-graph convolutional network based on discrete wavelets provided by the present invention can achieve accurate classification of sellar region tumor images automatically.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image classification, and particularly to a multimodal-based method for classifying amine region tumor images and a terminal device. Background Art

[0002] Intracranial sellar region tumors are a type of tumors that grow in the tissues near the sellar region. Among them, pituitary tumors, craniopharyngiomas, germ cell tumors, etc. are mostly located in the intrasellar and suprasellar regions, and meningiomas are mostly located beside the sella. Patients with intracranial sellar region tumors have no obvious discomfort in the early stage. As the tumor grows, patients will present different clinical symptoms such as headache, nausea, and vomiting, seriously affecting the patients' life and health. Clearly defining the tumor location and its relationship with surrounding tissues is crucial for the treatment of this disease. The development of MRI imaging technology has provided an effective means of discrimination for clinicians. However, due to the extremely small volume of the sellar region in the brain and the complex relationship of its surrounding tissues, even experienced clinicians may make misdiagnoses. In addition, the excessive number of layers of MRI images makes it time-consuming and laborious for clinicians to read the images.

[0003] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0004] In view of the above deficiencies of the existing technology, the purpose of the present invention is to provide a multimodal-based method for classifying amine region tumor images and a terminal device, aiming to solve the problem that the existing technology cannot automatically and accurately classify amine region tumor images.

[0005] The technical solution of the present invention is as follows:

[0006] A multimodal-based method for classifying amine region tumor images, which includes the steps of:

[0007] Obtaining three-modal MRI images of patients with intracranial amine region tumors, where the three-modal MRI images are T1-weighted MRI images, T2-weighted MRI images, and enhanced T1-weighted MRI images respectively;

[0008] Using a 3D Resnet network with discrete wavelet transform to extract features from the three-modal MRI images, obtaining low-frequency features and high-frequency features of the three modalities;

[0009] Based on the low-frequency features, high-frequency features of the three modalities and the clinical information of patients with intracranial amine region tumors, constructing a low-frequency feature similarity map and a high-frequency feature similarity map through Pearson correlation coefficients;

[0010] Using a multi-graph convolutional network to optimize the features of the low-frequency feature similarity map and the high-frequency feature similarity map, obtaining optimized features of the three modalities;

[0011] The optimized features of the three modalities are fused using the multi-head attention mechanism, and the amine region tumor images are classified using a fully connected layer.

[0012] In the multi-modal based amine region tumor image classification method, the MRI images of the three modalities of the intracranial amine region tumor patients are the MRI images obtained by segmenting the head using an edge detection method.

[0013] In the multi-modal based amine region tumor image classification method, the steps of obtaining the low-frequency features and high-frequency features of the three modalities by using a 3D Resnet network with discrete wavelet transform added to extract features from the MRI images of the three modalities include:

[0014] The MRI images of the three modalities are decomposed using discrete wavelet transform to obtain 8 frequency signals corresponding to the MRI image of each modality, and the 8 frequency signals include 1 low-frequency signal and 7 high-frequency signals;

[0015] 3D Resnet is used to extract features from the 8 frequency signals to obtain the low-frequency features and high-frequency features of the three modalities. The 3D Resnet includes a 3D convolutional block of 7x7x7 arranged in sequence, a batch normalization layer, a RELU layer, a first channel attention module, a first spatial attention module, 4 residual blocks, a second channel attention module, a second spatial attention module, and an average pooling layer.

[0016] In the multi-modal based amine region tumor image classification method, in the step of decomposing the MRI images of the three modalities using discrete wavelet transform to obtain 8 frequency signals corresponding to the MRI image of each modality, the formula of the discrete wavelet transform is: where a represents the scaling factor, b represents the translation factor, ψ a,b (t) represents the wavelet basis, and W f (a, b) represents the spectral signal.

[0017] In the multi-modal based amine region tumor image classification method, the steps of constructing a low-frequency feature similarity map and a high-frequency feature similarity map based on the low-frequency features and high-frequency features of the three modalities and the clinical information of the intracranial amine region tumor patients include:

[0018] The relationship between the intracranial amine region tumor and the clinical information is obtained through the adjacency matrix formula where, and represent the fused feature vectors of intracranial amine region tumor patients m and intracranial amine region tumor patient n, b m and b n represent the gender information of intracranial amine region tumor patients m and n, gm and g n represent the age information of patients m and n with intracranial amine region tumors, and S represents the similarity measure between feature vectors. The calculation method is as follows: Among them, represents the mean value of, represents the mean value of, ⊙ represents dot product, and C() represents the relationship between non-image information.

[0019] For the multi-modal amine region tumor image classification method, among them, the steps of using a multi-graph convolutional network to optimize the low-frequency feature similarity map and the high-frequency feature similarity map to obtain optimized features in three modalities include:

[0020] Use a graph wavelet transform convolutional network to perform convolution on the low-frequency feature similarity map and the high-frequency feature similarity map to obtain optimized features in three modalities. The graph wavelet convolutional network is described as Among them, K represents the convolutional kernel, and ⊙ represents element-wise multiplication. represents the conversion of the signal x from the spatial domain to the frequency domain, x = ψ s x' represents the conversion of the signal in the frequency domain to the spatial domain.

[0021] For the multi-modal amine region tumor image classification method, among them, the steps of using a multi-head attention mechanism to fuse the optimized features in three modalities include:

[0022] The multi-head attention mechanism is composed of three self-attention mechanisms. Through the three self-attention mechanisms, the features of each modality are weighted by the features of the other two modalities. The three self-attention mechanisms are described as:

[0023] Among them, is the feature of the T1-weighted MRI image optimized by the multi-graph convolutional network. is the feature of the enhanced T1-weighted MRI image optimized by the multi-graph convolutional network. represents the transpose of, is the feature of the T2-weighted MRI image optimized by the multi-graph convolutional network. represents the transpose of, and dk represents the dimension of the feature;

[0024] Fuse the optimized features in three modalities through a fusion mechanism. The formula of the fusion mechanism is: Among them, Conat means concatenating the features after three self-attention mechanisms in the feature dimension to finally obtain the fused features.

[0025] The multimodal-based amine region tumor image classification method, wherein the classification results of the amine region tumor images include sellar germinoma, non-functional pituitary adenoma, craniopharyngioma, and meningioma.

[0026] A storage medium, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the multimodal-based amine region tumor image classification method as described in the present invention.

[0027] A terminal device, including: a processor, a memory, and a communication bus; a computer-readable program executable by the processor is stored on the memory;

[0028] The communication bus realizes the connection and communication between the processor and the memory;

[0029] When the processor executes the computer-readable program, the steps in the multimodal-based amine region tumor image classification method as described in the present invention are realized.

[0030] Beneficial effects: The present invention proposes to use a 3D Resnet network with discrete wavelet transform to effectively extract features from T1-weighted MRI images, T2-weighted MRI images, and enhanced T1-weighted MRI images, construct high-frequency feature similarity maps and low-frequency feature similarity maps for the extracted high-frequency features and low-frequency features respectively, and finally use a multi-graph convolutional network to perform accurate classification tasks. The present invention verifies the effectiveness of the discrete wavelet-based multi-graph convolutional network designed by the present invention for the automatic differential diagnosis of sellar tumors on the sellar tumor dataset provided by the cooperative research unit, providing an effective means for clinicians to perform differential diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a flowchart of a multimodal-based amine region tumor image classification method of the present invention.

[0032] Figure 2 It is a schematic diagram of the automatic differential diagnosis of sellar tumor images by the multi-graph convolutional network based on wavelet transform of the present invention.

[0033] Figure 3 They are MRI image slices of three modalities and their corresponding cropped slices.

[0034] Figure 4 It is a schematic diagram of the 3D Resnet with attention mechanism for feature extraction.

[0035] Figure 5It is a schematic diagram of multi-modal fusion based on the multi-head attention mechanism.

[0036] Figure 6 It is a schematic diagram of a terminal device. Specific implementation manners

[0037] The present invention provides a method and a terminal device for classifying amine-region tumor images based on multi-modalities. To make the objectives, technical solutions and effects of the present invention clearer and more definite, the present invention is further described in detail below. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0038] Please refer to Figure 1 , Figure 1 It is a flowchart of a preferred embodiment of a method for classifying amine-region tumor images based on multi-modalities provided by the present invention. As shown in the figure, it includes the following steps:

[0039] S10. Obtain three-modal MRI images of patients with intracranial amine-region tumors. The three-modal MRI images are T1-weighted MRI images, T2-weighted MRI images, and enhanced T1-weighted MRI images respectively;

[0040] S20. Use a 3D Resnet network with discrete wavelet transform to extract features from the three-modal MRI images, and obtain low-frequency features and high-frequency features of the three modalities;

[0041] S30. Based on the low-frequency features, high-frequency features of the three modalities and the clinical information of patients with intracranial amine-region tumors, construct a low-frequency feature similarity map and a high-frequency feature similarity map through the Pearson correlation coefficient;

[0042] S40. Use a multi-graph convolutional network to optimize the features of the low-frequency feature similarity map and the high-frequency feature similarity map, and obtain optimized features of the three modalities;

[0043] S50. Use the multi-head attention mechanism to fuse the optimized features of the three modalities, and use a fully connected layer to classify the amine-region tumor images.

[0044] Specifically, due to the limited research on existing automatic classification methods for intracranial sellar region tumors, there is an urgent need for an automatic classification method for sellar region tumors to assist clinicians in differentiating images. Therefore, the present invention proposes a method: First, use a 3D Resnet network with discrete wavelet transform to effectively extract features from T1-weighted MRI images, T2-weighted MRI images, and enhanced T1-weighted MRI images. Since a pure convolutional neural network cannot effectively extract the high-frequency and low-frequency components of images, wavelet transform is added to extract the low-frequency and high-frequency features of the images, construct a low-frequency feature similarity map and a high-frequency feature similarity map for the extracted low-frequency and high-frequency features, and use a multi-graph convolutional network to differentiate and classify sellar region tumor images. As Figure 2 shown, the method of the present invention mainly includes two tasks: (1) Feature extraction: Use a 3D Resnet network with wavelet transform to extract high-frequency and low-frequency features. The low-frequency features retain the main energy information of the MRI image, and the high-frequency features retain the detailed information of the MRI image; (2) Classification task: Use the Pearson correlation coefficient to construct 1 low-frequency feature similarity map and 7 high-frequency feature similarity maps for the extracted 1 low-frequency feature and 7 high-frequency features, and add clinical information to detect whether it can assist in classification. Use a multi-graph convolutional network to optimize the extracted features, and finally use the multi-head attention mechanism to fuse the optimized features of the three modalities, so that the features of each modality are assisted by the features of the other two modalities, so that the fused features reach the best effect. Finally, use a fully connected layer for accurate differential classification.

[0045] In the present invention, due to the too small volume of the sellar region in the brain, it is very difficult to classify sellar region tumors. Therefore, it is necessary to refer to different MRI images for accurate differentiation. In this embodiment, three modalities of MRI images are used for differentiation to improve the differentiation accuracy. The three modalities of MRI images are T1-weighted MRI images, T2-weighted MRI images, and enhanced T1-weighted MRI images.

[0046] In some embodiments, the three modalities of MRI images of the patient with intracranial sellar region tumors are MRI images obtained by using the edge detection method to segment the head. Specifically, due to scanning differences, the position of the head is not unified, and the sellar region is small. Extracting features from the overall MRI image will cause interference from other parts to the feature extraction of the tumor region. Therefore, in this embodiment, the edge detection method is used to segment the head. In addition, the sellar region is distributed in the middle of the brain, so as to achieve the cropping of the largest sellar region, which can not only reduce the interference of other regions, but also reduce the weight parameters of the network and avoid overfitting. This preprocessing part is completed before inputting into the network.

[0047] In some specific embodiments, 3D MRI images of sellar region tumors from Peking Union Medical College Hospital in Beijing are obtained. The purpose of the experiment is to perform a classification task on four types of sellar region tumor images. The dataset contains a total of 394 samples. Each sample includes T1-weighted images, T2-weighted images, and enhanced T1-weighted images. Among them, there are 115 cases of germinoma in the sellar region, 199 cases of non-functional pituitary adenomas, 90 cases of craniopharyngiomas, and 33 cases of meningiomas. Due to the scanning differences of each sample, in this embodiment, edge detection and cropping methods are used to preprocess the sample images. The size of the initial image is 512×512×slices, and the size of the processed image is 300×300×slices. The slices of T1-weighted images and enhanced T1-weighted images are both 8, while the slices of T2-weighted images are 20. Therefore, the resize function is used to set the slices of T2-weighted images to 8. As Figure 3 shown, it is the display of the multi-modal data slices in this article and the display after preprocessing.

[0048] In some embodiments, discrete wavelet transform is used to decompose the MRI images of the three modalities, and 8 frequency signals corresponding to the MRI images of each modality are obtained. The 8 frequency signals include 1 low-frequency signal and 7 high-frequency signals; for 3D MRI images, they can be regarded as three-dimensional discrete data, which are processed according to one-dimensional discrete signals. One-dimensional discrete signal wavelets decompose the signal through a scaling factor and a translation factor, and have good scalability and translation invariance compared with traditional Fourier transforms. They can not only reduce the amount of calculation, but also obtain good filtering performance. The formula for the discrete wavelet transform is: where a represents the scaling factor, b represents the translation factor, and ψ a,b (t) represents the wavelet basis, and W f (a, b) represents the spectral signal. Convolution operations are performed using a low-pass filter matrix L and a high-pass filter matrix H for filtering, as follows: f(t) l =Lf(t), f(t) h =Hf(t) where f(t) l is the low-frequency signal, and f(t) h is the high-frequency signal. For the 3D MRI signal X, there will be 8 filters, such as LLL, LLH, LHL, LHH, HLL, HLH, HHL, HHH. These 8 matrices are orthogonal to each other. Among them, LLL is a low-pass filter, and the remaining 7 are high-pass filters in different directions.

[0049] In this embodiment, 3D Resnet is used to extract features from the 8 frequency signals, obtaining low-frequency features and high-frequency features of three modalities. The 3D Resnet includes a 3D convolutional block of 7x7x7 arranged in sequence, a batch normalization layer, a RELU layer, a first channel attention module, a first spatial attention module, 4 residual blocks, a second channel attention module, a second spatial attention module, and an average pooling layer. As Figure 4 shown, in this embodiment, first, a 3D convolutional block of 7x7x7 is used to perform convolution on the MRI images of three modalities, and the batch normalization layer and the RELU layer are adopted. Subsequently, a channel attention module and a spatial attention module are added. We use the network of 3D Resnet18. After the attention module, 4 residual blocks are added, then a channel attention module and a spatial attention module are added, and then average pooling is added to expand the receptive field of the network, so as to extract the effective low-frequency and high-frequency features of each modality.

[0050] In this embodiment, the channel attention module considers the combined output of average pooling and max pooling. It can generate a channel attention feature map by utilizing the relationship between channels. The attention feature map is multiplied by the input feature map to sparsify the features, enabling the network to focus on more useful features. The channel attention module can be described by the formula: F c =σ(conv(avgpool(F in )) + conv(maxpool(F in ))), where F in represents the input feature, F c represents the output channel attention feature, avgpool represents average pooling, maxpool represents max pooling, conv represents 3D convolution, and σ represents the activation function.

[0051] In this embodiment, the spatial attention module combines max pooling and average pooling on the channels. It generates a spatial attention map by utilizing the spatial mutual relationship of the features. The attention map is multiplied by the feature map generated by the channel attention. The spatial attention module can be described by the formula: F s =σ(conv([avgpool(F in );maxpool(F in )])), where F s represents the output spatial attention feature, the features after average pooling and max pooling are concatenated, and 3D convolution is used to perform convolution on it

[0052] In some embodiments, based on the low-frequency features and high-frequency features of the three modalities and the clinical information of patients with intracranial amine region tumors, a low-frequency feature similarity map and a high-frequency feature similarity map are constructed through Pearson correlation coefficients. Specifically, the Pearson correlation coefficient is used to explore the relationship between the fused sample features. In this embodiment, the similarity and difference of the relationship between subjects are added on the basis of the extracted fused features, which can more effectively assist in diagnosis. In addition, in order to explore the relationship between sellar region tumors and clinical information, clinical information is also added on this basis to explore whether this information is beneficial to image classification. In this embodiment, the relationship between intracranial amine region tumors and clinical information is obtained through the adjacency matrix formula where and represent the fused feature vectors of patient m and patient n with intracranial amine region tumors, b m and b n represent the gender information of patient m and patient n with intracranial amine region tumors, g m and g n represent the age information of patient m and patient n with intracranial amine region tumors, S represents the similarity measure between feature vectors, and the calculation method is as follows: where represents the mean value of , represents the mean value of , ⊙ represents dot product, and C() represents the relationship between non-image information.

[0053] In some embodiments, traditional graph convolutional networks use Fourier transform to convert signals in the spatial domain to the frequency domain. However, there are some deficiencies in using Fourier transform. The computational complexity of Laplacian matrix decomposition is relatively large, and the multiplication of the dense matrix after Laplacian matrix decomposition and the signal makes the graph Fourier transform not efficient enough. Graph convolution based on Fourier transform does not only perform convolution on a single node. Although the graph Fourier transform can solve this problem by using Chebyshev polynomials, the scalability flexibility of Chebyshev polynomials is not good enough, and it is impossible to confirm the number of terms of its polynomial to approximate graph convolution. Therefore, in this embodiment, a multi-graph convolutional network is used to optimize the features of the low-frequency feature similarity map and the high-frequency feature similarity map to obtain optimized features of the three modalities.

[0054] Specifically, a graph wavelet transform convolutional network is used to implement convolution on the graph. The graph convolution of wavelet transform does not require Laplacian decomposition of the adjacency matrix. Using the wavelet basis ψ s =(ψ s1 , ψ s2 ,..., ψ sn ), where the wavelet ψ siDenote that a signal walks from node \(i\) on the graph, and \(s\) represents the dilation scale parameter. \(\psi\) s can be expressed as \(\psi\) s \( = UG\) s where \(U\) is the Laplacian eigenvector, and \(G\) T , where \(U\) is the Laplacian eigenvector, and \(G\) s \( = diag(g(s\lambda_1),..., g(s\lambda\) n )) is a dilation matrix, Therefore, the graph wavelet convolution can be described as where \(K\) represents the convolution kernel, and \(\odot\) represents element-wise multiplication, represents the transformation of the signal \(x\) from the spatial domain to the frequency domain, \(x = \psi\) s \(x'\) represents the transformation of the signal in the frequency domain to the spatial domain. Due to the sparsity of the wavelet basis, the corresponding network also becomes sparse. In addition, the dilation flexibility of the wavelet basis makes the graph wavelet convolution highly localized in the node domain, which also makes the performance of the graph wavelet convolution better than that of the traditional Fourier transform-based graph convolution. For each modality, in this embodiment, 1 low-frequency graph and 7 high-frequency graphs are constructed, and these graphs are used to optimize the features. Since the graph wavelet convolution can greatly reduce the computational amount, 8 graph convolutions all adopt single-layer graph wavelet convolutions. After single-layer graph wavelet convolution, 8 optimized features will be obtained. The 8 optimized features are concatenated, and each modality is concatenated in the corresponding order.

[0055] In some embodiments, as Figure 5 shown, the multi-head attention mechanism is used to fuse the optimized features of the three modalities. The multi-head attention mechanism is composed of three self-attention mechanisms. Through the three self-attention mechanisms, the features of each modality are weighted by the features of the other two modalities. The three self-attention mechanisms are described as:

[0056] where, is the feature of the T1-weighted MRI image optimized by the multi-graph convolution network, is the feature of the enhanced T1-weighted MRI image optimized by the multi-graph convolution network, represents transpose of, is the feature of the T2-weighted MRI image optimized by the multi-graph convolution network, represents transpose of, \(d\) k represents the dimension of the feature;

[0057] The optimized features of the three modalities are fused through the fusion mechanism. The formula of the fusion mechanism is: Among them, Conat means concatenating the features passing through three self-attention mechanisms in the feature dimension to finally obtain the fused features.

[0058] In some embodiments, the fused features are input into a fully connected layer for sellar region tumor image classification, and the sellar region tumor image classification results include germinoma of the sellar region, non-functional pituitary adenoma, craniopharyngioma, and meningioma.

[0059] In some embodiments, the multi-graph convolutional network based on wavelet transform provided by the present invention is carried out on a Linux server. This experiment performs feature extraction through 5-fold cross-validation and uses the same 5-fold cross-validation for differential diagnosis. In this embodiment, the model is trained using the PyTorch framework on a single TTTAN RTX GPU with 24GB of memory. During the training process, the learning rate is set to 10, the number of iterations is set to 300, and an early stopping mechanism is set to avoid overfitting, and the batch size is set to 4.

[0060] In some embodiments, a storage medium is further provided, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the multi-modal sellar region tumor image classification method as described in the present invention.

[0061] In some embodiments, the present application further provides a terminal device, as Figure 6 shown, which includes at least one processor 20; a display screen 21; and a memory 22, and may further include a communication interface 23 and a bus 24. Among them, the processor 20, the display screen 21, the memory 22, and the communication interface 23 can complete mutual communication through the bus 24. The display screen 21 is set to display a user guidance interface preset in the initial setting mode. The communication interface 23 can transmit information. The processor 20 can call the logical instructions in the memory 22 to execute the method in the above embodiments.

[0062] In addition, when the logical instructions in the above-mentioned memory 22 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.

[0063] The memory 22, as a computer-readable storage medium, can be set to store software programs and computer-executable programs, such as the program instructions or modules corresponding to the methods in the embodiments of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions, or modules stored in the memory 22, that is, implements the methods in the above embodiments.

[0064] The memory 22 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory. For example, various media that can store program codes such as USB flash drives, external hard drives, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs may also be transient storage media.

[0065] In addition, the specific processes of loading and executing multiple instructions in the storage medium and the terminal device have been described in detail in the above method and will not be elaborated here one by one.

[0066] In summary, the present invention proposes to use a 3D Resnet network with discrete wavelet transform to effectively extract features from T1-weighted MRI images, T2-weighted MRI images, and enhanced T1-weighted MRI images, construct high-frequency feature similarity maps and low-frequency feature similarity maps for the extracted high-frequency features and low-frequency features respectively, and finally use a multi-map convolutional network to perform accurate classification tasks. The present invention verifies the effectiveness of the multi-map convolutional network based on discrete wavelets designed by the present invention for the automatic differential diagnosis of sellar tumors on the sellar tumor dataset provided by the cooperative research unit, providing an effective means for clinicians to perform differential diagnosis.

[0067] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A multimodal-based tumor image classification method for the amine region, characterized in that, Including the steps of: Obtaining MRI images of three modalities of patients with intracranial amine region tumors, and the three-modal MRI images are T1-weighted MRI images, T2-weighted MRI images, and enhanced T1-weighted MRI images respectively; Using a 3D Resnet network with discrete wavelet transform to extract features from the three-modal MRI images, obtaining low-frequency features and high-frequency features of the three modalities; Based on the low-frequency features, high-frequency features of the three modalities and the clinical information of patients with intracranial amine region tumors, constructing a low-frequency feature similarity map and a high-frequency feature similarity map through the Pearson correlation coefficient; Using a multi-graph convolutional network to optimize the features of the low-frequency feature similarity map and the high-frequency feature similarity map, obtaining optimized features of the three modalities; Using a multi-head attention mechanism to fuse the optimized features of the three modalities, and using a fully connected layer for amine region tumor image classification.

2. The multimodality-based amine region tumor image classification method according to claim 1, wherein The three-modal MRI images of the patients with intracranial amine region tumors are MRI images obtained by segmenting the head using an edge detection method.

3. The multimodal-based amine region tumor image classification method according to claim 1, wherein The steps of using a 3D Resnet network with discrete wavelet transform to extract features from the three-modal MRI images, obtaining low-frequency features and high-frequency features of the three modalities include: Using discrete wavelet transform to decompose the three-modal MRI images, obtaining 8 frequency signals corresponding to the MRI images of each modality, and the 8 frequency signals include 1 low-frequency signal and 7 high-frequency signals; Using 3D Resnet to extract features from the 8 frequency signals, obtaining low-frequency features and high-frequency features of the three modalities, and the 3D Resnet includes a 7x7x7 3D convolutional block, a batch normalization layer, a RELU layer, a first channel attention module, a first spatial attention module, 4 residual blocks, a second channel attention module, a second spatial attention module, and an average pooling layer arranged in sequence.

4. The multi-modal based amine region tumor image classification method according to claim 3, wherein In the step of using discrete wavelet transform to decompose the three-modal MRI images, obtaining 8 frequency signals corresponding to the MRI images of each modality, the formula of the discrete wavelet transform is: Among them, a represents the scaling factor, b represents the translation factor, and ψ a,b (t) represents the wavelet basis, and W f (a, b) represents the spectral signal.

5. The multimodal-based amine region tumor image classification method according to claim 1, wherein The steps of constructing a low-frequency feature similarity map and a high-frequency feature similarity map through the Pearson correlation coefficient based on the low-frequency features, high-frequency features of the three modalities and the clinical information of patients with intracranial amine region tumors include: Through the adjacency matrix formula Obtain the relationship between the intracranial amine region tumor and clinical information, where and represent the fused feature vectors of patient m and patient n with intracranial amine region tumors, b m and b n represent the gender information of patient m and patient n with intracranial amine region tumors, g m and g n represent the age information of patient m and patient n with intracranial amine region tumors, S represents the similarity measure between feature vectors, and the calculation method is as follows: Where represents the mean value of represents the mean value of represents the dot product, and C( ) represents the relationship between non-image information 6. The multimodal-based amine region tumor image classification method according to claim 1, characterized in that The steps of using a multi-graph convolutional network to optimize the features of the low-frequency feature similarity map and the high-frequency feature similarity map, obtaining optimized features of the three modalities include: Use the graph wavelet transform convolutional network to perform the convolution of the low-frequency feature similarity graph and the high-frequency feature similarity graph, and obtain the optimized features of the three modalities. The graph wavelet convolutional network is described as where K represents the convolution kernel, denotes element-wise multiplication, denotes the conversion of the signal x from the spatial domain to the frequency domain, x = ψ s x' represents the conversion of the signal in the frequency domain to the spatial domain.

7. The multimodal-based amine region tumor image classification method according to claim 1, wherein The steps of using a multi-head attention mechanism to fuse the optimized features of the three modalities include: The multi-head attention mechanism is composed of three self-attention mechanisms, and through the three self-attention mechanisms, the features of each modality are weighted by the features of the other two modalities, and the three self-attention mechanisms are described as: Among them, is the feature of the T1-weighted MRI image optimized by the multi-image convolutional network, is the feature of the enhanced T1-weighted MRI image optimized by the multi-image convolutional network, represents the transpose of, is the feature of the T2-weighted MRI image optimized by the multi-image convolutional network, represents the transpose of, d k represents the dimension of the feature; Fusing the optimized features of the three modalities through a fusion mechanism, and the formula of the fusion mechanism is: Among them, Conat means concatenating the features that have passed through three self-attention mechanisms in the feature dimension to finally obtain the fused features.

8. The multimodal-based amine region tumor image classification method according to claim 1, wherein The classification results of the amine region tumor images include sellar germinoma, non-functional pituitary adenoma, craniopharyngioma, and meningioma.

9. A storage medium, characterized in that, The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the multi-modal amine region tumor image classification method according to any one of claims 1-8.

10. A terminal device, characterized in that, Including: A processor, a memory, and a communication bus; A computer-readable program that can be executed by the processor is stored on the memory; The communication bus realizes the connection and communication between the processor and the memory; When the processor executes the computer-readable program, it realizes the steps in the multi-modal amine region tumor image classification method according to any one of claims 1-8.

Citation Information

Patent Citations

  • MRI tumor optimal segmentation method and system based on multi-modal image fusion

    CN111612754A

  • Tumor image focus area prediction analysis method and system and terminal equipment

    CN112801168A