An Intestinal Disease Classification Method Based on Dual-Model Dynamic Feature Fusion

By building a network based on ResNet and Transformer, combining the two-way collaborative deep fusion module and the dynamic channel attention module, the problem of insufficient feature extraction capability in intestinal disease image diagnosis is solved, and higher feature extraction capability and intestinal disease classification accuracy are achieved.

CN118982705BActive Publication Date: 2025-05-27ZHEJIANG AIDA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411047937.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-01
Publication Date
2025-05-27
Estimated Expiration
2044-08-01

AI Technical Summary

Technical Problem

In intestinal disease image diagnosis, the ability of a single model when performing regional feature extraction may be limited, resulting in small target areas being easily overlooked, thereby reducing the accuracy of disease detection.

Method used

A network based on ResNet and Transformer is constructed, combining the two-way collaborative deep fusion module and the dynamic channel attention module to be used to scale collaborative fusion feature maps. The final output feature map is processed with a dynamic gated feedforward network to obtain intestinal disease classification results.

Benefits of technology

Through the two-way collaborative deep fusion module, the advantages of ResNet and Transformer are fully utilized to improve feature extraction capabilities and classification performance; the dynamic channel attention module realizes feature channel weight allocation to adapt to different types of intestinal images; the dynamic gated feedforward network provides effective feature selection and information filtering strategies to improve the generalization ability and classification accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118982705B_ABST
    Figure CN118982705B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for classifying intestinal diseases based on dual-model dynamic feature fusion. A network based on ResNet and Transformer is constructed to extract features from intestinal lesion images and obtain feature maps of different scales. A bidirectional collaborative depth fusion module is set up in conjunction with ResNet and Transformer to fuse the feature maps in a scale-by-scale collaborative manner. Finally, the output feature map is fused by a dynamic channel attention module, and the fused feature map is output to a dynamic gated feed-forward network to obtain the classification result of intestinal diseases. The present invention improves the feature extraction ability and classification performance of the entire network, can better adapt to different types of intestinal images, and gives accurate diagnostic results; the network can adaptively focus on the features that are most critical to the current task, thereby improving the generalization ability of the model and the accuracy of classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of general image data processing or generation, and particularly to a method for classifying intestinal diseases based on dual-model dynamic feature fusion for image analysis and classification in the field of medical image processing. Background Art

[0002] Early screening and early diagnosis and treatment are the keys to preventing colorectal cancer. Among them, early detection through endoscopic examination is an effective means to reduce cancer mortality. Traditional endoscopic examinations mainly rely on hand-held endoscopes. Although hand-held endoscopes have the advantages of high flexibility and convenient real-time observation, there are still some limitations. First, since hand-held endoscopes need to enter the body, they may cause discomfort or pain to patients, resulting in a low acceptance rate of such examinations by patients. Second, due to improper use by operators, problems such as intestinal wall damage may occur, causing secondary harm to patients. Therefore, painless and non-invasive wireless capsule endoscopes have become a more ideal choice. Wireless capsule endoscopes (WCE) can perform non-invasive evaluation of the digestive system without sedation, and can effectively diagnose abnormalities in gastrointestinal tissues. Patients are examined through the natural digestion process, which not only avoids the discomfort and pain that may be caused by traditional endoscopes, but also reduces the risk of trauma to patients.

[0003] However, although deep learning technology has been widely used in computer-aided diagnosis systems and has shown remarkable effects in multiple fields, it still faces unique challenges in the diagnosis process of intestinal image diseases. Specifically, the scale changes of different lesions in intestinal disease images are relatively large, and the ability of a single model to extract regional features may be limited, resulting in small target regions being easily overlooked, thereby reducing the accuracy of disease detection. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a method for classifying intestinal diseases based on dual-model dynamic feature fusion.

[0005] The technical solution adopted by the present invention is a method for classifying intestinal diseases based on dual-model dynamic feature fusion. The method constructs a network based on ResNet and Transformer to extract features from intestinal lesion images and obtain feature maps of different scales. A bidirectional collaborative depth fusion module is set up in cooperation with ResNet and Transformer to fuse the feature maps at different scales. Finally, the output feature map is fused with a dynamic channel attention module, and the fused feature map is output to a dynamic gated feed-forward network to obtain the classification result of intestinal diseases.

[0006] Preferably, the method includes the following steps:

[0007] S1 Establish a sample data set based on intestinal lesion images;

[0008] S2 Construct a collaborative network based on ResNet and Transformer and improve it with a bidirectional collaborative depth fusion module and a dynamic channel attention module. A dynamic gating feed-forward network is set after the network to obtain an improved network;

[0009] S3 Input the data of the sample data set into the improved network for training until it is stable;

[0010] S4 Input the intestinal lesion image to be classified into the trained improved network to obtain the corresponding intestinal disease classification result.

[0011] Preferably, in S2, the improved network includes a correspondingly set ResNet model and Transformer model. The ResNet model includes several Resnet layers, and the Transformer model includes self-attention mechanism blocks corresponding one by one to the Resnet layers;

[0012] The output of the previous group of corresponding Resnet layers and self-attention mechanism blocks is output to the corresponding bidirectional collaborative depth fusion module. After the output of the bidirectional collaborative depth fusion module is added to the output of the previous group of corresponding Resnet layers and self-attention mechanism blocks, it is respectively input into the self-attention mechanism blocks and Resnet layers of the next group;

[0013] The output of the last group of corresponding Resnet layers and self-attention mechanism blocks is output to the dynamic channel attention module, and the output of the dynamic channel attention module is output to the dynamic gating feed-forward network.

[0014] Preferably, the bidirectional collaborative depth fusion module includes a C2T module for fusing the feature map output by the ResNet model into the feature map output by the corresponding Transformer model and a T2C module for fusing the feature map output by the Transformer model into the feature map output by the ResNet model.

[0015] Preferably, the C2T module includes a sequentially arranged channel downsampling layer, an attention mechanism layer, a channel upsampling layer, and an addition module;

[0016] The feature map output by the ResNet model is sequentially input into the attention mechanism layer and the channel upsampling layer after passing through the channel downsampling layer. The feature map output by the ResNet model, the feature map output by the Transformer model, and the output of the channel upsampling layer are output after passing through the addition module.

[0017] Preferably, the T2C module includes a channel ascending layer, a channel descending layer, and an addition module arranged in sequence. Between the channel ascending layer and the channel descending layer, and between the channel descending layer and the addition module, there are a convolutional layer, a normalization layer, and an activation function layer arranged in sequence;

[0018] The feature map output by the Transformer model passes through the channel ascending layer, the channel descending layer, and the corresponding convolutional layer, normalization layer, and activation function layer of the two, and then is added to the feature map output by the Resnet model in the addition module and output.

[0019] Preferably, the dynamic channel attention module includes a global average pooling module, two fully connected layers, one ReLU activation function, and a multiplication module arranged in sequence;

[0020] Two feature maps are input into the dynamic channel attention module. The global average pooling module compresses each feature map of c channels, H×W into a vector of c channels 1×1. Through two fully connected layers and one ReLU activation function, the weight corresponding to each channel is obtained. The multiplication module multiplies each original feature channel by its weight to obtain a weighted feature map.

[0021] Preferably, the weighted feature map is sequentially input into a convolutional layer and a pooling layer and then output to the dynamic gated feed-forward network.

[0022] Preferably, the dynamic gated feed-forward network adjusts the channel weights of the feature map with a dynamic gating mechanism and inputs the adjusted feature map into a feed-forward neural network to obtain the final classification result.

[0023] Preferably, the dynamic gating mechanism is to generate dynamic weights by calculating weight parameters and offset parameters and combining with the sigmoid activation function. After the dynamic weights are multiplied element-wise with the input feature map, the obtained result is dimension-reduced by an average pooling layer, non-linearly transformed with the ReLU activation function, and finally input into the Softmax layer to generate a classification probability distribution.

[0024] The present invention provides a method for classifying intestinal diseases based on dual-model dynamic feature fusion, constructs a network based on ResNet and Transformer, is used to extract features from intestinal lesion images and obtain feature maps of different scales, cooperates with ResNet and Transformer to set a bidirectional collaborative depth fusion module for fusing feature maps at different scales, and finally the output feature map is fused by a dynamic channel attention module, and the fused feature map is output to the dynamic gated feed-forward network to obtain the classification result of intestinal diseases.

[0025] The beneficial effects of the present invention are as follows:

[0026] (1) The output of the Resnet and Transformer models is bidirectionally and synergistically deeply fused by a bidirectional collaborative deep fusion module. Through this fusion strategy, each layer of the module can make full use of the advantages of Resnet and Transformer, thereby improving the feature extraction ability and classification performance of the entire network;

[0027] (2) A dynamic channel attention module utilizes the dynamic channel attention mechanism to assign weights and fuse features for each feature channel in the output feature maps of the last layers of Resnet and Transformer, enabling the model to have the ability to dynamically select between the local features of ResNet and the global features of Transformer. This flexibility allows the model to better adapt to different types of intestinal images, thus giving accurate diagnostic results;

[0028] (3) A dynamic gated feed-forward network integrates the dynamic gating mechanism and the feed-forward neural network to provide an effective feature selection and information filtering strategy for the model, enabling the network to adaptively focus on the features most critical to the current task, thereby improving the generalization ability and classification accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a flowchart for classifying intestinal diseases according to the present invention;

[0030] Figure 2 It is a network structure diagram of the present invention;

[0031] Figure 3 It is a structure diagram of the bidirectional collaborative deep fusion module of the present invention, where (a) is the structure diagram of the C2T module and (b) is the structure diagram of the T2C module;

[0032] Figure 4 It is a structure diagram of the dynamic channel attention module of the present invention;

[0033] Figure 5 It is a structure diagram of the dynamic gated feed-forward network of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0034] The following further describes the present invention in detail with reference to the embodiments, but the protection scope of the present invention is not limited thereto.

[0035] The present invention relates to a method for classifying intestinal diseases based on dual-model dynamic feature fusion. The method constructs a network based on ResNet and Transformer to extract features from intestinal lesion images and obtain feature maps of different scales. A Bidirectional Synergic Fusion (BiSF) module is set up in cooperation with ResNet and Transformer to synergistically fuse the feature maps at different scales. Finally, the output feature map is fused by a Dynamic Weight (DW) module, and the fused feature map is output to a dynamic gated feedforward network to obtain the classification result of intestinal diseases.

[0036] As Figure 1 shown, the inventive concept of the present invention lies in establishing an image data set containing various intestinal lesions, using the networks based on ResNet and Transformer to respectively perform multi-stage feature extraction on the intestinal lesion images to obtain feature maps of different scales. At the same scale, the feature maps output by the two models are deeply fused through the bidirectional synergic depth fusion module, and the fused features are downsampled and then fused with the features of the next layer. The feature map output by the last layer is subjected to dynamic weight feature fusion by the dynamic channel attention module, and the fused feature map is output. Finally, the dynamic gated feedforward network is used to classify the output features of the lesions to obtain the final classification result of intestinal diseases.

[0037] The method includes the following steps:

[0038] S1 Establish a sample data set based on intestinal lesion images;

[0039] S2 Construct a collaborative network based on ResNet and Transformer and improve it with a bidirectional synergic depth fusion module and a dynamic channel attention module, and set a dynamic gated feedforward network behind the network to obtain an improved network;

[0040] S3 Input the data of the sample data set into the improved network for training until it is stable;

[0041] S4 Input the intestinal lesion image to be classified into the trained improved network to obtain the corresponding classification result of intestinal diseases.

[0042] The following is an explanatory description in combination with specific methods.

[0043] (1) Establish a sample data set based on intestinal lesion images;

[0044] In the implementation process of the present invention, this sample data set comes from the image set captured by a wireless capsule endoscope, and these images are all intestinal lesion images that have been carefully analyzed and accurately diagnosed as having diseases by professional doctors.

[0045] In this embodiment, the image resolution is 240×240, with a total of 8000 images.

[0046] (2) Construct a collaborative network based on ResNet and Transformer and improve it with a bidirectional collaborative depth fusion module and a dynamic channel attention module. A dynamic gated feed-forward network is set after the network to obtain an improved network;

[0047] Here, ResNet is stacked by multiple residual blocks. Inside each residual block, the input is directly added to the output through a skip connection, allowing the gradient to flow directly through multiple layers. It is good at capturing local detail features (local key regions), while Transformer adopts an architecture based on the self-attention mechanism. Through the multi-head self-attention mechanism, position encoding, and hierarchical feed-forward network, it realizes the efficient processing and feature extraction of sequence data. It is good at understanding global context information (global dependencies), thus endowing the overall model with local and global feature extraction capabilities. The improved network can accurately extract the key features of the lesions from intestinal images in stages through methods such as convolutional layers, residual connections, and attention mechanisms, generating a series of feature maps of different sizes to comprehensively capture the lesions at different scales, that is, a multi-stage feature extraction process;

[0048] Among them, the bidirectional collaborative depth fusion module refers to a module with flexible fusion characteristics, which can fuse the outputs of the ResNet and Transformer network models, thereby improving the overall performance of the network. The dynamic channel attention module calculates the weight values on each feature channel through learning, enabling the network to give priority to more important feature channels. The dynamic gated feed-forward network provides effective feature selection and information filtering strategies, improving the generalization ability of the model and the accuracy of classification.

[0049] Such as Figure 2 shown, the improved network includes a corresponding ResNet model and a Transformer model. The ResNet model includes several ResNet layers, and the Transformer model includes self-attention mechanism blocks corresponding one by one to the ResNet layers;

[0050] The outputs of the previous group of corresponding ResNet layers and self-attention mechanism blocks are sent to the corresponding bidirectional collaborative depth fusion module. After the output of the bidirectional collaborative depth fusion module is added to the outputs of the previous group of corresponding ResNet layers and self-attention mechanism blocks, they are respectively input into the self-attention mechanism blocks and ResNet layers of the next group;

[0051] The outputs of the last group of corresponding ResNet layers and self-attention mechanism blocks are sent to the dynamic channel attention module, and the output of the dynamic channel attention module is sent to the dynamic gated feed-forward network.

[0052] The two-way collaborative deep fusion module includes a C2T module for fusing the feature map F1 output by the Resnet model into the corresponding feature map output by the Transformer model, and a T2C module for fusing the feature map F2 output by the Transformer model into the feature map output by the Resnet model.

[0053] In the present invention, through this two-way fusion strategy, the C2T module and the T2C module can make full use of the advantages of Resnet and Transformer at each layer, thereby improving the feature extraction ability and classification performance of the entire network.

[0054] As Figure 3 shown in (a) of [], the C2T module includes a channel reduction layer, an attention mechanism layer, a channel increase layer, and an addition module arranged in sequence;

[0055] The feature map output by the Resnet model is sequentially input into the attention mechanism layer and the channel increase layer after passing through the channel reduction layer; the feature map output by the Resnet model, the feature map output by the Transformer model, and the output of the channel increase layer are output after passing through the addition module.

[0056] In this embodiment, the specific implementation manner of the C2T module is as follows:

[0057] (2-1-1) Reduce the number of channels of the Resnet output feature map F C by half through 1×1 convolution to obtain the feature map F C1 with reduced number of channels, satisfying

[0058] F C1 = Conv(F C )

[0059] where Conv(·) is a convolution operation;

[0060] (2-1-2) Use the attention mechanism layer for global information extraction, satisfying

[0061]

[0062] where is a scaling factor;

[0063] (2-1-3) Restore the number of channels through 1×1 convolution, add it to the Resnet input, and add it to the feature map output by the current layer Transformer.

[0064] As Figure 3As shown in Figure (b), the T2C module includes a channel upsampling layer, a channel downsampling layer, and an addition module arranged in sequence. A convolutional layer, a normalization layer, and an activation function layer are sequentially provided between the channel upsampling layer and the channel downsampling layer, and between the channel downsampling layer and the addition module;

[0065] The feature map output by the Transformer model passes through the channel upsampling layer, the channel downsampling layer, and the corresponding convolutional layer, normalization layer, and activation function layer of both, and then is added to the feature map output by the Resnet model in the addition module and output.

[0066] In this embodiment, the specific implementation manner of the T2C module is as follows:

[0067] (2-2-1) Double the number of channels of the Transformer output feature F T through 1×1 convolution to obtain the feature map F T1 with increased number of channels, satisfying

[0068] F T1 = Conv(F T )

[0069] where Conv(·) is the convolution operation;

[0070] (2-2-2) Pass through the convolutional layer, the normalization layer, and the ReLU activation function to obtain

[0071] F T2 = Conv(F T1 )

[0072] F T3 = BN(F T2 )

[0073] O = ReLU(F T3 )

[0074] where BN(·) is the average pooling method, which normalizes the feature map F T2 after the convolution operation to obtain F T3 , and O is the ReLU output, performing the activation function operation on F T3 ;

[0075] (2-2-3) Restore the number of channels through 1×1 convolution, then pass through the convolutional layer, the normalization layer, and the Relu activation function, and add it to the feature map output by the current layer Resnet.

[0076] In an embodiment of the present invention, for the three dimensions H×W×c of the feature map, downsampling of the feature map refers to reducing the size of the feature map through convolution and pooling operations to reduce the computational amount and storage space of the data, while retaining important feature information. Channel increase and channel decrease refer to modifying c of the feature map, that is, adjusting the number of channels, so as to achieve the purpose of reducing the computational cost. Here, the process of adjusting the number of channels is generally to change the number of convolutional kernels, that is, "changing / adjusting the number of channels". An increase in the number of channels is called channel increase, and vice versa is called channel decrease.

[0077] As Figure 4 shown, the dynamic channel attention module includes a global average pooling module, two fully connected layers, one ReLU activation function, and a multiplication module arranged in sequence;

[0078] Two feature maps are input into the dynamic channel attention module. The global average pooling module compresses each c-channel, H×W feature map into a c-channel, 1×1 vector. Through two fully connected layers and one ReLU activation function, the weight corresponding to each channel is obtained. The multiplication module multiplies each original feature channel by its weight to obtain a weighted feature map.

[0079] The weighted feature map is sequentially input into a convolutional layer and a pooling layer and then output to the dynamic gated feedforward network.

[0080] In the present invention, the gating signal generates specific weights to adjust the weights of the channels, thereby realizing "dynamic gating", which actually involves convolution calculation and backpropagation operations.

[0081] In this embodiment, the specific implementation manner of the dynamic channel attention module is as follows:

[0082] (2-3-1) Through global average pooling, each c-channel, H×W feature map is compressed into a c-channel, 1×1 vector. The global average pooling result z of the c-th channel c satisfies

[0083]

[0084] where u c (i,j) represents the value of the c-th channel at the position (i,j), and H×W respectively represent the height and width of the feature map;

[0085] (2-3-2) Through two fully connected layers and one ReLU activation function, the weight s of each channel is obtained, satisfying

[0086] s = σ(W 2 δ(W 1 z))

[0087] where W 1 and W2 where \(s\) is the weight of the two - layer fully - connected layer, \(\delta\) is the ReLU activation function, and \(\sigma\) is the sigmoid activation function;

[0088] (2 - 3 - 3) Multiply the feature \(u\) on each original feature channel \(c\) c by its weight \(s\) c to obtain the weighted feature map satisfying

[0089]

[0090] The dynamic gated feed - forward network adjusts the channel weights of the feature map with a dynamic gating mechanism and inputs the adjusted feature map into the feed - forward neural network to obtain the final classification result.

[0091] As Figure 5 shown, the dynamic gating mechanism is as follows: by calculating the weight parameter and the offset parameter, combined with the sigmoid activation function to generate the dynamic weight. After the dynamic weight is multiplied element - by - element with the input feature map, the obtained result is processed by an average pooling layer for dimensionality reduction, and then undergoes a non - linear transformation with the ReLU activation function, and finally is input into the Softmax layer to generate the classification probability distribution.

[0092] In this embodiment, the specific implementation of the dynamic gated feed - forward network is as follows:

[0093] (2 - 4 - 1) Generation of the gating signal \(G\),

[0094] \(G=\sigma(W\) g \(F + b\) g )

[0095] where \(F\) is the feature map, \(W\) g and \(b\) g are the weight and bias of the gating layer respectively, and \(\sigma\) is the Sigmoid activation function;

[0096] (2 - 4 - 2) Feature map gating satisfies

[0097] \(F\) gated \(=\ G\odot F\)

[0098] where \(\odot\) represents element - by - element multiplication, and \(F\) gated is the adjusted feature map;

[0099] (2 - 4 - 3) Input the feature map adjusted by the gating mechanism into the feed - forward neural network to obtain the final classification result, satisfying

[0100] \(v = AvgPool(F\) gated )

[0101] \(O = ReLU(W\) f \(v + b\)f )

[0102] P = Softmax(W c O + b c )

[0103] where AvgPool(·) is the average pooling method that converts the gated feature map F gated into the feature vector v, O is the ReLU output, W f and b f are the weights and biases of the FFN, and W c and b c are the weights and biases of the classification layer, and P represents the classification probability distribution.

[0104] In the present invention, the dynamic gating process can adaptively adjust the channel weights of the feature map, selectively enhance or suppress the features in the feature map, thereby improving the feature expression ability and classification performance of the network and optimizing the accuracy of intestinal disease classification.

[0105] (3) Input the data of the sample data set into the improved network for training until it is stable;

[0106] (4) Input the intestinal lesion image to be classified into the trained improved network to obtain the corresponding intestinal disease classification result.

[0107] The present invention also relates to a computer-readable storage medium in application, on which a program for classifying intestinal diseases based on dual-model dynamic feature fusion is stored, and when the program is executed by a processor, the above-mentioned method for classifying intestinal diseases based on dual-model dynamic feature fusion is implemented.

[0108] The present invention also relates to a computer device in application, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned method for classifying intestinal diseases based on dual-model dynamic feature fusion is implemented.

[0109] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0110] The present invention is described with reference to flowchart illustrations and / or block diagram illustrations of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowchart illustrations and / or block diagram illustrations, and combinations of flows and / or blocks in the flowchart illustrations and / or block diagram illustrations, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing apparatus create means for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.

[0111] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.

[0112] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.

[0113] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.

[0114] It is apparent that those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A method for intestinal disease classification based on dual-model dynamic feature fusion, characterized by: The method constructs a network based on ResNet and Transformer to extract features from intestinal lesion images and obtain feature maps of different scales, and sets a bidirectional collaborative deep fusion module with ResNet and Transformer to collaboratively fuse feature maps at different scales; The bidirectional collaborative deep fusion module includes a C2T module for fusing the feature map output by the Resnet model into the feature map output by the corresponding Transformer model and a T2C module for fusing the feature map output by the Transformer model into the feature map output by the Resnet model; The C2T module includes a channel descent layer, an attention mechanism layer, a channel ascending layer, and an addition module that are arranged in sequence; the feature map output by the Resnet model is sequentially input into the attention mechanism layer and the channel ascending layer after passing through the channel descent layer; the feature map output by the Resnet model, the feature map output by the Transformer model, and the output of the channel ascending layer are output after passing through the addition module; The T2C module includes a channel ascending layer, a channel descending layer and an addition module which are arranged in sequence, and a convolution layer, a normalization layer and an activation function layer are arranged in sequence between the channel ascending layer and the channel descending layer, and between the channel descending layer and the addition module respectively; the feature map output by the Transformer model passes through the channel ascending layer, the channel descending layer and the corresponding convolution layer, normalization layer and activation function layer, and then is added to the feature map output by the Resnet model in the addition module and output; The final output feature map is fused using a dynamic channel attention module, and the fused feature map is output to a dynamic gated feedforward network to obtain the intestinal disease classification results.

2. According to claim 1, a method for classifying intestinal diseases based on dual-model dynamic feature fusion, characterized in that: The method comprises the following steps: S1 establishes a sample dataset based on intestinal lesion images; S2 builds a collaborative network based on ResNet and Transformer and improves it with a bidirectional collaborative deep fusion module and a dynamic channel attention module. A dynamic gated feedforward network is set after the network to obtain an improved network. S3: inputting data of the sample data set into the improved network for training until it is stable; S4 inputs the intestinal lesion image to be classified into the trained improved network to obtain the corresponding intestinal disease classification result.

3. The intestinal disease classification method based on dual-model dynamic feature fusion according to claim 2 is characterized in that: In S2, the improved network includes a correspondingly set ResNet model and a Transformer model, the ResNet model includes a plurality of Resnet layers, and the Transformer model includes a self-attention mechanism block corresponding to the Resnet layer one by one; The outputs of the Resnet layer and self-attention mechanism block corresponding to the first group are sent to the corresponding bidirectional collaborative deep fusion module. The outputs of the bidirectional collaborative deep fusion module are added to the outputs of the Resnet layer and self-attention mechanism block corresponding to the first group, and then are input to the self-attention mechanism block and Resnet layer corresponding to the second group respectively. The last set of corresponding Resnet layers and self-attention mechanism blocks output to the dynamic channel attention module, which outputs to the dynamic gated feedforward network.

4. A method for classifying intestinal diseases based on dual-model dynamic feature fusion according to any one of claims 1 to 3, characterized in that: The dynamic channel attention module includes a global average pooling module, two fully connected layers, a ReLU activation function, and a multiplication module arranged in sequence; The two feature maps are input into the dynamic channel attention module. The global average pooling module compresses each c-channel, H×W feature map into a c-channel 1×1 vector. Through two fully connected layers and a ReLU activation function, the weight corresponding to each channel is obtained. The multiplication module multiplies each original feature channel with its weight to obtain the weighted feature map.

5. The intestinal disease classification method based on dual-model dynamic feature fusion according to claim 4 is characterized in that: The weighted feature map is sequentially input into the convolution layer and the pooling layer and then output to the dynamic gated feedforward network.

6. The intestinal disease classification method based on dual-model dynamic feature fusion according to claim 1, characterized in that: The dynamic gated feedforward network adjusts the channel weights of the feature map with a dynamic gating mechanism, and inputs the adjusted feature map into the feedforward neural network to obtain the final classification result.

7. The intestinal disease classification method based on dual-model dynamic feature fusion according to claim 6, characterized in that: The dynamic gating mechanism is to generate dynamic weights by calculating weight parameters and offset parameters in combination with the sigmoid activation function. After the dynamic weights are multiplied element by element with the input feature map, the obtained result is processed by average pooling layer for dimensionality reduction, and after nonlinear transformation with ReLU activation function, it is finally input into Softmax layer to generate classification probability distribution.