Gastrointestinal tract disease classification method and system based on convolution and state space model

By adopting a method based on convolution and state space model in medical image classification, texture features and adaptive contrast enhancement are extracted, feature extraction is combined with convolution and state space units, and the final feature map is generated through patch merging, the problem of high computational complexity in the existing technology is solved, and efficient image classification of gastrointestinal diseases is achieved.

CN119992200AActive Publication Date: 2025-05-13SICHUAN UNIV

Patent Information

Application Number
CN202510087812.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

When processing gastrointestinal disease images, the existing medical image classification methods have high computational complexity, resulting in low processing efficiency, making it difficult to effectively capture global information and process long-distance dependence.

Method used

Using a method based on convolution and state space model, the image is segmented into multiple patches by extracting texture features and adaptive contrast enhancement, and the convolution unit and state space unit are combined in the fusion module of the deep learning model to perform local and global feature extraction, and finally the final feature map is generated through patch merging for classification.

Benefits of technology

It realizes that the global information and local features of medical images are effectively captured while maintaining low computational complexity, and improves the accuracy and efficiency of image classification of gastrointestinal diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992200A_ABST
    Figure CN119992200A_ABST
Patent Text Reader

Abstract

The invention discloses a gastrointestinal disease classification method and system based on a convolution and state space model, and the method comprises the steps: extracting the texture features of a gastrointestinal disease medical image, determining a lesion region in the medical image according to the texture features, carrying out the adaptive contrast enhancement of the lesion region, carrying out the preprocessing, and carrying out the recognition of the lesion region. Inputting a patch embedding module of the trained deep learning model, and segmenting the medical image into a plurality of patches; the patches are sequentially input into the two fusion modules, a convolution unit and a state space unit of each fusion module respectively extract local features and global features of the patches, and the local features and the global features are fused; merging the fusion features of the plurality of patches by using a patch merging module, and inputting the merged features into two fusion modules of a deep learning model to generate fusion features of the merged features; and repeating the previous step for preset times to obtain a final feature map, and inputting the final feature map into a classifier of the deep learning model to obtain a classification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image classification technology, and in particular to a gastrointestinal disease classification method and system based on convolution and state space model. Background Art

[0002] In recent years, with the rise of deep learning technology, especially the application of convolutional neural networks (CNN) in computer vision, the task of medical image classification has made significant progress. CNN can automatically extract features from images and perform multi-level feature learning, making the classification of medical images more accurate. However, CNN still has shortcomings in capturing global information, handling long-distance dependencies, and robustness. Specifically, when processing medical images of gastrointestinal diseases (such as ulcers, polyps, tumors, etc.), although CNN can identify local lesions in the image, it often ignores the influence of the global context, which may lead to misjudgment of some minor lesions.

[0003] The Transformer model was originally applied in the field of natural language processing and has also begun to gain widespread attention in the field of image analysis. Transformer models long-distance dependencies between features globally through the self-attention mechanism, thus performing well on certain tasks. The Transformer architecture provides stronger capabilities than CNN when processing medical images, especially in modeling global information. However, the computational complexity of the Transformer grows quadratically with the increase in input data (O(N 2 )), where N represents the dimension of the input data or the size of the image. This quadratic computational complexity causes huge computational pressure and resource consumption when processing large-scale images, especially in medical image analysis, which limits its practical application.

[0004] In order to solve these problems, the state space model (SSM) has been introduced into the field of deep learning in recent years. SSM can effectively handle long-distance dependencies in images and improve classification accuracy by dynamically modeling the global information in images. Research on combining SSM with deep neural networks (such as CNN) has gradually emerged, forming a new generation of image analysis methods. These methods significantly improve the accuracy and robustness of image classification by combining the local feature extraction capabilities of CNN and the global information modeling capabilities of SSM.

[0005] Although some research has made some progress in this regard, the existing fusion methods still have problems such as high computational complexity and low efficiency. Therefore, how to design an efficient image classification method that can capture global information while maintaining low computational complexity is still an important research topic in the field of medical image analysis. Summary of the invention

[0006] In view of the above-mentioned deficiencies in the prior art, the gastrointestinal disease classification method and system based on convolution and state space model provided by the present invention solves the problem of low processing efficiency due to high computational complexity of the prior art methods when performing medical image classification.

[0007] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:

[0008] In a first aspect, a gastrointestinal disease classification method based on convolution and state space model is provided, which comprises the steps of:

[0009] S1. Extracting texture features of medical images of gastrointestinal diseases, determining the lesion area in the medical image according to the texture features, and performing adaptive contrast enhancement on the lesion area;

[0010] S2, preprocessing the medical image after contrast enhancement of the lesion area, and then inputting it into the patch embedding module of the trained deep learning model to segment the medical image into multiple patches;

[0011] S3, the patches are sequentially input into two fusion modules of the deep learning model, the convolution unit and the state space unit of each fusion module extract local features and global features of the patches respectively, and then the local features and global features of the patches are fused to obtain fused features;

[0012] S4, using the patch merging module of the deep learning model to merge the fusion features of multiple patches, and inputting the merged features into two fusion modules of the deep learning model to generate fusion features of the merged features;

[0013] S5. Repeat step S4 for a preset number of times to obtain a final feature map, and then input the final feature map of the medical image into the classifier of the deep learning model to obtain a classification result.

[0014] Furthermore, step S1 further comprises:

[0015] S11. Use Gabor filter to extract texture features of medical images of gastrointestinal diseases:

[0016]

[0017] Where G(x,y) is the texture feature of the pixel with spatial coordinates (x,y); exp is the exponential function; σ is the scale parameter of the Gabor filter; f is the frequency of the Gabor filter; θ is the direction angle of the Gabor filter;

[0018] S12, determining whether the texture feature G(x, y) of the pixel point with spatial coordinates (x, y) is greater than a preset threshold, if so, marking it as a diseased area, otherwise marking it as a normal area;

[0019] S13, calculate the fitness threshold of the pixel points corresponding to the lesion area:

[0020] threshold(x,y)=μ(x,y)+λ·σ(x,y)

[0021] Wherein, threshold(x,y) is the fitness threshold of the pixel with spatial coordinates (x,y); μ(x,y) and σ(x,y) are the mean and standard deviation of the grayscale values ​​of all pixels in the neighborhood centered at (x,y) in the medical image; λ is the sensitivity adjustment parameter;

[0022] S14, determining whether the gray value of the pixel with spatial coordinates (x, y) is greater than its corresponding fitness threshold, if so, proceeding to step S15, otherwise, not performing contrast enhancement on the pixel;

[0023] S15, enhancing the contrast of the pixels:

[0024]

[0025] Among them, I(x,y) and I′(x,y) are the grayscale values ​​of the pixel with spatial coordinates (x,y) before and after enhancement; α is the intensity of contrast enhancement.

[0026] Further, the method of segmenting the medical image into a plurality of patches includes:

[0027] Divide the medical image into patches of size P×P, each patch represents a local area in the medical image;

[0028] A convolution operation is used to transform each patch into a vector of a set dimension. The final embedding is represented as a tensor containing the features of all patches. This tensor X patch_embedding The shape is:

[0029]

[0030] Where B is the batch size, is the number of patches in a medical image; D is the dimension of each patch after mapping to the feature space; H is the height of the medical image; W is the medical image; P is the side length of the patch.

[0031] Furthermore, step S3 further includes:

[0032] S31. Input the patch into the fusion module. The convolution unit of the fusion module uses multiple convolution operations to gradually extract the spatial features of the image, and obtains the local features through batch normalization and nonlinear activation function:

[0033] Y=X p +ReLU(Batch Norm(ReLU(Batch Norm(X p *K1 (1) +b1 (1) ))*K2 (2) +b2 (2) )*K3 (3) +b3 (2) ))

[0034] Among them, Y is the local feature of the patch; ReLU() is the activation function; Batch Norm() is the batch normalization operation; is the first 1×1 convolution kernel, b1 (1) K1 (1) The bias term of is the number of channels after dimensionality reduction; C in is the number of channels of the input tensor; is a 2×2 convolution kernel, b2 (2) For K2 (2) The bias term, C out is the number of channels of the output tensor; is the second 1×1 convolution kernel, b3 (3) For K3 (3) Bias term; * indicates multiplication; X p is the patch obtained after transposition;

[0035] S32, the state space unit of the fusion module extracts global features from the input patch, and then splices the global features and local features corresponding to the same patch to obtain spliced ​​features;

[0036] S33, use the channel attention module of the fusion module to calculate the channel attention of the splicing features to obtain the channel attention map:

[0037] M c (F)=σ′(MLP(AvgPool(F))+MLP(MaxPool(F)))

[0038] Among them, F is the splicing feature; M c (F) is the channel attention map of the spliced ​​feature F; σ′() is the Sigmoid activation function; Avg Pool() is the average pooling operation; MLP() is the multi-layer perceptron; Max Pool() is the maximum pooling operation;

[0039] S34, multiply the concatenated feature by its corresponding channel attention map element by element, and then input it into the spatial attention module of the fusion module for spatial attention calculation to obtain the spatial attention map:

[0040] M s (F′)=σ′(f 7×7 ([Avg Pool(F′);Max Pool(F′)]))

[0041] Among them, F′ is the feature map obtained by element-by-element multiplication; M s (F′) is the spatial attention map corresponding to the feature map F′; f 7 ×7 It is a 7×7 convolution operation;

[0042] S35, multiplying the spatial attention map and the splicing feature element by element to obtain a first fusion feature;

[0043] S36: Input the first fusion feature into another fusion module, and repeat steps S31 to S35 to obtain the final fusion feature of each patch.

[0044] Furthermore, the method for merging the fusion features of multiple patches includes:

[0045] The four segmented regions are merged according to the channel dimension by splicing the features of the four regions together; the feature dimension of each patch is increased fourfold; then, each new feature map obtained after the merger is normalized, and then the normalized features are reduced in dimension using a linear transformation to obtain the merged features.

[0046] Furthermore, the loss function of the deep learning model is expressed as:

[0047]

[0048] in, is the standard cross entropy loss; is the focal loss; is the structural similarity loss; is the boundary loss; λ1, λ2, λ3 and λ4 are and The weight coefficient of .

[0049] Further, and The expressions are:

[0050]

[0051] Among them, α′ is a hyperparameter for adjusting category balance; pt is the prediction probability of the deep learning model for the correct category; γ is the parameter for adjusting the difficult and easy samples; SIMM is the structural similarity index; I and They are the original images during training and the predicted images of the deep learning model; and are the kth original image I during training. k And the corresponding prediction image of the deep learning model The boundary gradient of , 1≤k≤U, U is the total number of medical images in the dataset during training; |·| is the absolute value symbol.

[0052] In a second aspect, a gastrointestinal disease classification system based on convolution and state space model is provided, comprising:

[0053] The image filtering and enhancement module is used to extract the texture features of the medical images of gastrointestinal diseases, determine the lesion area in the medical image according to the texture features, and perform adaptive contrast enhancement on the lesion area;

[0054] A preprocessing module, used for preprocessing the medical image after contrast enhancement of the lesion area;

[0055] A patch generation module is used to input the preprocessed image into the patch embedding module of the trained deep learning model to segment the medical image into multiple patches;

[0056] The first fusion feature generation module is used to input the patches into the two fusion modules of the deep learning model in sequence, and the convolution unit and the state space unit of each fusion module extract local features and global features of the patches respectively, and then fuse the local features and global features of the patches to obtain fusion features;

[0057] A feature merging module, used to merge the fusion features of multiple patches using a patch merging module of a deep learning model to obtain a merged feature;

[0058] A second fusion feature generating module, used for inputting the merged features into the two fusion modules of the deep learning model to generate a fusion feature of the merged features;

[0059] A judgment module is used to judge whether the executed feature merging module and the second fusion feature generating module are the last group. If so, the output of the last second fusion feature generating module is output as the final feature map, otherwise, the next group of feature merging modules and the second fusion feature generating module is entered;

[0060] The classification module is used to input the final feature map of the medical image into the classifier of the deep learning model to obtain the classification result.

[0061] The beneficial effects of the present invention are as follows: the present scheme can identify the lesion area through filtering processing, so as to perform targeted enhancement and thus improve the sensitivity of the model to the lesion area and the classification accuracy; the fusion module can fuse features from different modules to enhance the performance of the image classification task, thereby helping the model to understand a wider range of image information; patch merging is a key step in the present scheme for integrating patches that have undergone feature fusion processing, which reduces the spatial resolution of the feature map while increasing the feature dimension by merging the features of adjacent patches, thereby providing a richer and more compact feature representation for subsequent network processing, and the patch merging operation ensures that the multi-scale features of the image can be effectively fused while maintaining computational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 Flowchart of the gastrointestinal disease classification method based on convolution and state-space models.

[0063] Figure 2 This is a partial principle block diagram of the gastrointestinal disease classification method based on convolution and state-space models.

[0064] Figure 3 This is the principle block diagram of the convolution unit.

[0065] Figure 4 This is the principle block diagram of the state space unit.

[0066] Figure 5 This is the principle block diagram of the fusion module. DETAILED DESCRIPTION

[0067] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.

[0068] refer to Figure 1 , Figure 1 A flow chart of a gastrointestinal disease classification method based on convolution and state space model is shown, as Figure 1 and Figure 2 As shown, the method S includes steps S1 to S5.

[0069] In step S1, texture features of a gastrointestinal disease medical image are extracted, a lesion area in the medical image is determined according to the texture features, and adaptive contrast enhancement is performed on the lesion area;

[0070] In one embodiment of the present invention, step S1 further comprises:

[0071] S11. Use Gabor filter to extract texture features of medical images of gastrointestinal diseases:

[0072]

[0073] Among them, G(x,y) is the texture feature of the pixel point with spatial coordinates (x,y); exp is the exponential function; σ is the scale parameter of the Gabor filter; f is the frequency of the Gabor filter, which is used to control the frequency characteristics of the texture and can capture the frequency information of detail changes; θ is the direction angle of the Gabor filter, which is used to control the direction information of the texture and help identify lesion features in different directions in gastrointestinal images.

[0074] S12, determining whether the texture feature G(x, y) of the pixel point with spatial coordinates (x, y) is greater than a preset threshold, if so, marking it as a diseased area, otherwise marking it as a normal area;

[0075] Since lesion areas usually have different texture patterns, these patterns can be highlighted through the output of the Gabor filter, thereby helping the model to identify lesion areas and highlight the differences between lesion areas and normal tissues.

[0076] S13, calculate the fitness threshold of the pixel points corresponding to the lesion area:

[0077] threshold(x,y)=μ(x,y)+λ·σ(x,y)

[0078] Wherein, threshold(x, y) is the fitness threshold of the pixel with spatial coordinates (x, y); μ(x, y) and σ(x, y) are the mean and standard deviation of the grayscale values ​​of all pixels in the neighborhood centered at (x, y) in the medical image; λ is the sensitivity adjustment parameter;

[0079] S14, determining whether the gray value of the pixel with spatial coordinates (x, y) is greater than its corresponding fitness threshold, if so, proceeding to step S15, otherwise, not performing contrast enhancement on the pixel;

[0080] The calculation of the fitness threshold of this scheme is based on the mean and standard deviation of the local area, and has good adaptability, so as to improve the contrast of the pixels in the lesion area in a targeted manner, so that the classifier can more easily identify and distinguish the lesion area.

[0081] S15, enhancing the contrast of the pixels:

[0082]

[0083] Wherein, I(x, y) and I′(x, y) are the grayscale values ​​of the pixel with spatial coordinates (x, y) before and after enhancement; α is the intensity of contrast enhancement.

[0084] By enhancing the contrast, the diseased areas (such as tumors, ulcers, polyps, etc.) can be made more obvious in the image, so that the subsequent classification algorithm can identify these areas more accurately.

[0085] In step S2, the medical image after contrast enhancement of the lesion area is preprocessed, and then input into the patch embedding module of the trained deep learning model to segment the medical image into multiple patches; the preprocessing here is to adjust the size of the contrast enhanced medical image so that the input image matches the size requirement of the deep learning model.

[0086] In one embodiment of the present invention, a method for segmenting a medical image into a plurality of patches comprises:

[0087] The medical image is divided into small blocks (i.e. patches) of size (P×P). Each patch represents a local area in the image, and the size of the patch is controlled by the parameter patch_size. For example, when the patch size is set to (P=4), each patch is a small area of ​​(4×4).

[0088] Next, a convolution operation is used to convert each patch into a fixed-size feature vector. This convolution operation maps each (P×P) region in the image to a new feature space, and the output feature vector can be used to describe the content of the region.

[0089] Specifically, the convolution operation scans the entire image and divides the image into non-overlapping small blocks according to the set stride (i.e. the distance of each movement). Each small block is mapped to a new feature vector, and the dimension of the feature vector is determined by the setting of the convolution operation.

[0090] After the convolution operation, each patch in the image is represented as a vector of fixed dimension, and the final embedding is represented as a tensor containing the features of all patches. patch_embedding The shape is:

[0091]

[0092] Where B is the batch size, is the number of patches in a medical image; D is the dimension of each patch after mapping to the feature space; H is the height of the medical image; W is the medical image; P is the side length of the patch.

[0093] In step S3, the patches are sequentially input into two fusion modules of the deep learning model. The convolution unit and state space unit of each fusion module extract local features and global features of the patches respectively, and then the local features and global features of the patches are fused to obtain fused features.

[0094] In one embodiment of the present invention, step S3 further comprises:

[0095] S31, the patch is input into the fusion module, the convolution unit of the fusion module (the principle block diagram of the convolution unit can be referred to Figure 3 )Use multiple convolution operations to gradually extract the spatial features of the image, and obtain local features through batch normalization and nonlinear activation functions:

[0096] Y=X p +ReLU(Batch Norm(ReLU(Batch Norm(X p *K1 (1) +b1 (1) ))*K2 (2) +b2 (2) )*K3 (3) +b3 (2) ))

[0097] Among them, Y is the local feature of the patch; ReLU() is the activation function; Batch Norm() is the batch normalization operation; is the first 1×1 convolution kernel, b1 (1) K1 (1) The bias term of is the number of channels after dimensionality reduction; C in is the number of channels of the input tensor; is a 2×2 convolution kernel, b2 (2) For K2 (2) The bias term, C out is the number of channels of the output tensor; is the second 1×1 convolution kernel, b3 (3) For K3 (3) Bias term; * indicates multiplication; X p is the patch obtained after transposition;

[0098] S32, the state space unit of the fusion module extracts global features from the input patch, and then splices the global features and local features corresponding to the same patch to obtain spliced ​​features;

[0099] The State Space Model Unit is a basic module used in this scheme to extract the global features of the input image. It aims to extract the global features of the image and enhance the representation ability of the features through a dynamic state transition model. In this example, the State Space Models (SSM) is a model that efficiently processes sequence data based on continuous linear ordinary differential equations (ODEs); ODEs describe the process of transforming a one-dimensional input function or sequence into a state-dependent variable. To the intermediate hidden state Then to output 's mapping.

[0100] The state matrices A and B are converted into discrete parameters using the zero-order hold discretization rule and

[0101]

[0102] in, is the state matrix; and are projection parameters; t is time; exp is an exponential function used to calculate the conversion from continuous to discrete; I′ is the unit matrix.

[0103] After discretization, linear recursion or global convolution is used to calculate the global features. The expression of linear recursion is:

[0104]

[0105] y(t)=(Ch'(t)

[0106] Where h′(t) is X p (t) is the hidden state corresponding to x; h′(t) is the intermediate parameter; y(t) is the hidden state corresponding to x p (t) corresponds to the global feature;

[0107] The expression for global convolution calculation is:

[0108]

[0109] in, is the structured convolution kernel; L is X p Length; is the structured convolution kernel.

[0110] refer to Figure 4 ,The state space unit extracts global features including the following steps:

[0111] 1. First, the input features are layer normalized to ensure the stability of the input; 2. Then, the input features are divided into two branches, the first branch is processed by a linear layer and an activation function, and the second branch is processed by a depthwise separable convolution and enters the 2D selective scanning (SS2D) module after activation; 3. In the SS2D module, the input image is expanded along four different directions to generate multiple sequences; these sequences are processed by the S6 block and restored to output images of the same size as the input; 4. Finally, the features are mixed through a linear layer and element-by-element multiplication is performed with the output of the first branch to finally fuse the features of the two branches.

[0112] S33, use the channel attention module of the fusion module to calculate the channel attention of the splicing features to obtain the channel attention map:

[0113] M c (F)=σ′(MLP(Avg Pool(F))+MLP(Max Pool(F)))

[0114] Among them, F is the splicing feature; M c (F) is the channel attention map of the spliced ​​feature F; σ′() is the Sigmoid activation function; Avg Pool() is the average pooling operation; MLP() is the multi-layer perceptron; Max Pool() is the maximum pooling operation;

[0115] S34, multiply the concatenated feature by its corresponding channel attention map element by element, and then input it into the spatial attention module of the fusion module for spatial attention calculation to obtain the spatial attention map:

[0116] M s (F′)=σ′(f 7×7 ([Avg Pool(F′);Max Pool(F′)]))

[0117] Among them, F′ is the feature map obtained by element-by-element multiplication; M s (F′) is the spatial attention map corresponding to the feature map F′; f 7 ×7 It is a 7×7 convolution operation;

[0118] S35, multiply the spatial attention map and the splicing feature element by element to obtain the first fusion feature; the principle block diagram of the fusion module processing process corresponding to steps S31 to S35 can be referred to Figure 5 .

[0119] S36: Input the first fusion feature into another fusion module, and repeat steps S31 to S35 to obtain the final fusion feature of each patch.

[0120] In step S4, the fusion features of multiple patches are merged using a patch merging module of the deep learning model, and the merged features are input into two fusion modules of the deep learning model to generate fusion features of the merged features;

[0121] During implementation, the method for merging the fusion features of multiple patches preferably includes:

[0122] The four segmented regions are merged according to the channel dimension by splicing the features of the four regions together. Through this merging, the features of each region are enhanced, and the merged feature map becomes larger in the dimension of each patch. At this point, the merged feature map contains the information of the original four patches, and the feature dimension of each patch is increased by four times. Each new feature map obtained after the merger is normalized, and then the normalized features are reduced in dimension using a linear transformation to obtain the merged features.

[0123] S5. Repeat step S4 for a preset number of times to obtain a final feature map, and then input the final feature map of the medical image into the classifier of the deep learning model to obtain the classification result:

[0124] Y = Softmax(W′F final +b′)

[0125] in, The output of the classifier; W′ and b′ are the weight parameter and bias term of the classifier respectively; F final is the final feature map; N′ is the number of categories;

[0126] In the classification of gastrointestinal disease images, the lesion areas usually have relatively small and complex textures, which makes the traditional loss function unable to fully capture the key features of these areas. Therefore, to address this problem, this solution proposes a weighted combination of focal loss, structural similarity loss (SSIMLoss) and boundary loss based on the standard cross entropy loss function to optimize the classification accuracy, especially in the identification of lesion areas.

[0127] During implementation, the loss function of the deep learning model is preferably expressed as:

[0128]

[0129] in, is the standard cross entropy loss; is the focal loss, used to alleviate category imbalance

[0130] problems, especially improving the model's ability to identify rare lesions by focusing on hard-to-classify areas; Structural similarity loss is used to optimize the structural learning of the lesion area and enhance the structural similarity of the image, especially in complex gastrointestinal images, to help the model accurately model the tiny structures in the lesion area; is the boundary loss, which is specifically used to improve the accuracy of the boundary of the lesion area, reduce the misjudgment of the edge, and optimize the segmentation accuracy of the lesion area; λ1, λ2, λ3 and λ4 are and The weight coefficient of .

[0131] in, and The expressions are:

[0132]

[0133] Among them, α′ is a hyperparameter for adjusting category balance; p t is the prediction probability of the deep learning model for the correct category; γ is the parameter for adjusting the difficult and easy samples; SIMM is the structural similarity index; I and They are the original images during training and the predicted images of the deep learning model; and are the kth original image I during training. k And the corresponding prediction image of the deep learning model The boundary gradient of , 1≤k≤U, U is the total number of medical images in the dataset during training; |·| is the absolute value symbol.

[0134] The loss function provided by this solution has the following advantages:

[0135] Dealing with category imbalance: Focal Loss is used to increase the focus on rare lesion areas, avoid the dominant role of common categories in the training process, and improve the classification accuracy of rare lesion areas.

[0136] Detailed structural learning: Through SSIM Loss, the model is able to learn the complex structure of the lesion area in gastrointestinal disease images, especially in the tiny lesion area, ensuring that the lesion details are not ignored.

[0137] Accurate boundary recognition: Through boundary loss, the recognition of the boundary of the lesion area is optimized, the fuzzy judgment of the model on the lesion area is reduced, and the segmentation accuracy of the lesion area is improved.

[0138] Through the above loss function design, this scheme can significantly improve the performance of the model in the gastrointestinal disease image classification task, especially in the recognition and boundary segmentation of subtle lesions, ensuring that the model can more accurately identify and classify gastrointestinal diseases.

[0139] This solution also provides a gastrointestinal disease classification system based on convolution and state space model, which includes:

[0140] The image filtering and enhancement module is used to extract the texture features of the medical images of gastrointestinal diseases, determine the lesion area in the medical image according to the texture features, and perform adaptive contrast enhancement on the lesion area;

[0141] A preprocessing module, used for preprocessing the medical image after contrast enhancement of the lesion area;

[0142] A patch generation module is used to input the preprocessed image into the patch embedding module of the trained deep learning model to segment the medical image into multiple patches;

[0143] The first fusion feature generation module is used to input the patches into the two fusion modules of the deep learning model in sequence, and the convolution unit and the state space unit of each fusion module extract local features and global features of the patches respectively, and then fuse the local features and global features of the patches to obtain fusion features;

[0144] A feature merging module, used to merge the fusion features of multiple patches using a patch merging module of a deep learning model to obtain a merged feature;

[0145] A second fusion feature generating module, used for inputting the merged features into the two fusion modules of the deep learning model to generate a fusion feature of the merged features;

[0146] A judgment module is used to judge whether the executed feature merging module and the second fusion feature generating module are the last group. If so, the output of the last second fusion feature generating module is output as the final feature map, otherwise, the next group of feature merging modules and the second fusion feature generating module is entered;

[0147] The classification module is used to input the final feature map of the medical image into the classifier of the deep learning model to obtain the classification result.

[0148] In summary, the gastrointestinal disease classification method of this scheme can not only capture the global information of medical images, but also achieve efficient image classification while maintaining low computational complexity.

Claims

1. A gastrointestinal disease classification method based on convolution and state space model, characterized in that: Includes steps: S1. Extracting texture features of medical images of gastrointestinal diseases, determining the lesion area in the medical image according to the texture features, and performing adaptive contrast enhancement on the lesion area; S2, preprocessing the medical image after contrast enhancement of the lesion area, and then inputting it into the patch embedding module of the trained deep learning model to segment the medical image into multiple patches; S3, the patches are sequentially input into two fusion modules of the deep learning model, the convolution unit and the state space unit of each fusion module extract local features and global features of the patches respectively, and then the local features and global features of the patches are fused to obtain fused features; S4, using the patch merging module of the deep learning model to merge the fusion features of multiple patches, and inputting the merged features into two fusion modules of the deep learning model to generate fusion features of the merged features; S5. Repeat step S4 for a preset number of times to obtain a final feature map, and then input the final feature map of the medical image into the classifier of the deep learning model to obtain a classification result.

2. The gastrointestinal disease classification method based on convolution and state space model according to claim 1, characterized in that: Step S1 further comprises: S11. Use Gabor filter to extract texture features of medical images of gastrointestinal diseases: Where G(x,y) is the texture feature of the pixel with spatial coordinates (x,y); exp is the exponential function; σ is the scale parameter of the Gabor filter; f is the frequency of the Gabor filter; θ is the direction angle of the Gabor filter; S12, determining whether the texture feature G(x, y) of the pixel point with spatial coordinates (x, y) is greater than a preset threshold, if so, marking it as a diseased area, otherwise marking it as a normal area; S13, calculate the fitness threshold of the pixel points corresponding to the lesion area: threshold(x,y)=μ(x,y)+λσ(x,y) Wherein, threshold(x,y) is the fitness threshold of the pixel with spatial coordinates (x,y); μ(x,y) and σ(x,y) are the mean and standard deviation of the grayscale values ​​of all pixels in the neighborhood centered at (x,y) in the medical image; λ is the sensitivity adjustment parameter; S14, determine whether the gray value of the pixel with spatial coordinates (x, y) is greater than its corresponding fitness threshold, if so, proceed to step S15, otherwise do not perform contrast enhancement on the pixel; S15, enhancing the contrast of the pixels: Among them, I(x,y) and I′(x,y) are the grayscale values ​​of the pixel with spatial coordinates (x,y) before and after enhancement; α is the intensity of contrast enhancement.

3. The gastrointestinal disease classification method based on convolution and state space model according to claim 1, characterized in that: Methods for segmenting medical images into multiple patches include: Divide the medical image into patches of size P×P, each patch represents a local area in the medical image; A convolution operation is used to transform each patch into a vector of a set dimension. The final embedding is represented as a tensor containing the features of all patches. This tensor X patch_embedding The shape is: Where B is the batch size, is the number of patches in a medical image; D is the dimension of each patch after mapping to the feature space; H is the height of the medical image; W is the medical image; P is the side length of the patch.

4. The gastrointestinal disease classification method based on convolution and state space model according to claim 1, characterized in that: Step S3 further comprises: S31. Input the patch into the fusion module. The convolution unit of the fusion module uses multiple convolution operations to gradually extract the spatial features of the image, and obtains the local features through batch normalization and nonlinear activation function: Y=X p +ReLU(Batch Norm(ReLU(Batch Norm(X p *K1(1)+b1(1)))*K2(2)+b2(2))*K3(3)+b3(2))) Among them, Y is the local feature of the patch; ReLU() is the activation function; Batch Norm() is the batch normalization operation; is the first 1×1 convolution kernel, b1 (1) K1 (1) The bias term of is the number of channels after dimensionality reduction; C in is the number of channels of the input tensor; is a 2×2 convolution kernel, b2 (2) For K2 (2) The bias term, C out is the number of channels of the output tensor; is the second 1×1 convolution kernel, b3 (3) For K3 (3) Bias term; * indicates multiplication; X p is the patch obtained after transposition; S32, the state space unit of the fusion module extracts global features from the input patch, and then splices the global features and local features corresponding to the same patch to obtain spliced ​​features; S33, use the channel attention module of the fusion module to calculate the channel attention of the splicing features to obtain the channel attention map: M c (F)=σ′(MLP(Avg Pool(F))+MLP(Max Pool(F))) Among them, F is the splicing feature; M c (F) is the channel attention map of the concatenated feature F; σ′() is the Sigmoid activation function; AvgPool() is the average pooling operation; MLP() is the multi-layer perceptron; Max Pool() is the maximum pooling operation; S34, multiply the concatenated feature by its corresponding channel attention map element by element, and then input it into the spatial attention module of the fusion module for spatial attention calculation to obtain the spatial attention map: M s (F′)=σ′(f 7×7 ([Avg Pool(F′);Max Pool(F′)])) Among them, F′ is the feature map obtained by element-by-element multiplication; M s (F′) is the spatial attention map corresponding to the feature map F′; f 7×7 It is a 7×7 convolution operation; S35, multiplying the spatial attention map and the splicing feature element by element to obtain a first fusion feature; S36: Input the first fusion feature into another fusion module, and repeat steps S31 to S35 to obtain the final fusion feature of each patch.

5. The gastrointestinal disease classification method based on convolution and state space model according to claim 1, characterized in that: Methods for merging fusion features of multiple patches include: The four segmented regions are merged according to the channel dimension. The merging method is to splice the features of the four regions together, and the feature dimension of each patch is increased four times; then, each new feature map obtained after the merger is normalized, and then the normalized features are reduced in dimension using linear transformation to obtain the merged features.

6. The gastrointestinal disease classification method based on convolution and state space model according to any one of claims 1 to 5, characterized in that: The loss function of the deep learning model is expressed as: in, is the standard cross entropy loss; is the focal loss; is the structural similarity loss; is the boundary loss; λ1, λ2, λ3 and λ4 are and The weight coefficient of .

7. The gastrointestinal disease classification method based on convolution and state space model according to claim 6, characterized in that: and The expressions are: Among them, α′ is a hyperparameter for adjusting category balance; p t is the prediction probability of the deep learning model for the correct category; γ is the parameter for adjusting the difficult and easy samples; SIMM is the structural similarity index; I and They are the original images during training and the predicted images of the deep learning model; and are the kth original image I during training. k And the corresponding prediction image of the deep learning model The boundary gradient of , 1≤k≤U, U is the total number of medical images in the dataset during training; |·| is the absolute value symbol.

8. A gastrointestinal disease classification system based on convolution and state space model, characterized in that: include: The image filtering and enhancement module is used to extract the texture features of the medical images of gastrointestinal diseases, determine the lesion area in the medical image according to the texture features, and perform adaptive contrast enhancement on the lesion area; A preprocessing module, used for preprocessing the medical image after contrast enhancement of the lesion area; A patch generation module is used to input the preprocessed image into the patch embedding module of the trained deep learning model to segment the medical image into multiple patches; The first fusion feature generation module is used to input the patches into the two fusion modules of the deep learning model in sequence, and the convolution unit and the state space unit of each fusion module extract local features and global features of the patches respectively, and then fuse the local features and global features of the patches to obtain fusion features; A feature merging module, used to merge the fusion features of multiple patches using a patch merging module of a deep learning model to obtain a merged feature; A second fusion feature generation module, used for inputting the merged features into the two fusion modules of the deep learning model to generate a fusion feature of the merged features; A judgment module is used to judge whether the executed feature merging module and the second fusion feature generating module are the last group. If so, the output of the last second fusion feature generating module is output as the final feature map, otherwise, the next group of feature merging modules and the second fusion feature generating module is entered; The classification module is used to input the final feature map of the medical image into the classifier of the deep learning model to obtain the classification result.

Citation Information

Patent Citations

  • Colonoscope polyp image detection method based on Mamba and YOLOv8

    CN118762009A

  • Method and system for segmenting kidney region in dynamic kidney development image based on mixed attention branches

    CN118918126A

Cited By

  • Battery top cover defect detection system and method based on image analysis

    CN120235864A