Rare disease image classification method and device based on small sample learning and storage medium

Through a small sample learning method, a feature backtracking fusion encoder and a multi-level prototype reconstruction network are used, and combined with a classifier based on Euclidean distance, the problem of difficult identification of lesion features in gastrointestinal diseases and large differences in class are solved, and high-precision image classification is achieved.

CN120032170AActive Publication Date: 2025-05-23ANHUI UNIV +1

Patent Information

Application Number
CN202510117339.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-23
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

The prior art is difficult to identify lesion characteristics in the image classification of gastrointestinal diseases, and the differences in distribution and size between similar lesions lead to inaccurate classification.

Method used

Using a small sample learning method, image features are extracted through feature backtracking and fusion encoder, and the semantic correlation between samples is captured using multi-level prototype reconstruction network. Combined with a classifier based on Euclidean distance, a calibration class center suitable for the current query sample is generated and classification results are output.

Benefits of technology

It effectively solves the problem that the lesion characteristics are difficult to identify lesions due to low contrast between the lesion and the intestinal wall, narrows the intra-class differences, and significantly improves the accuracy of image classification of gastrointestinal diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032170A_ABST
    Figure CN120032170A_ABST
Patent Text Reader

Abstract

According to the rare disease image classification method and device based on small sample learning and the medium, the feature backtracking fusion encoder is ingeniously constructed, and the attention mask is generated for the current layer by using the next-layer features, so that useless noise information in the low-layer features is reduced, and the classification efficiency is improved. Therefore, the space detail information in the low-level features is better fused into the high-level features, and the classification performance of the model on gastrointestinal tract disease areas is effectively improved. Second, the multi-level prototype reconstruction network further captures semantic relevance between the support set and query set samples to enhance the differentiated region on the support image representation, generating a calibration class center for each query sample that is suitable for the current query sample. A classifier based on the Euclidean distance outputs a classification result of the query sample, model optimization is guided by using a cross entropy function, and the precision of the classification result is ensured. Finally, the rare disease image classification model based on small sample learning can output the classification result of each image through strict training and testing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image classification, and in particular to a rare disease image classification method, device and storage medium based on small sample learning. Background Art

[0002] Endoscopy is the gold standard for gastrointestinal disease examination. It can help clinicians accurately identify gastrointestinal diseases in early screening for gastrointestinal cancer, thereby developing more precise treatment plans to reduce patient mortality. Therefore, accurate identification of gastrointestinal diseases, including rare diseases, is crucial to promote early diagnosis of gastrointestinal cancer patients and improve patient survival rates.

[0003] In recent years, with the rapid development of artificial intelligence technology represented by deep learning, some intelligent methods for endoscopic-assisted diagnosis have been initially explored in the gastrointestinal tract, proving the feasibility of artificial intelligence to promote endoscopic examinations. However, these technologies still face challenges in clinical applications. In medical endoscopic images, gastrointestinal diseases usually present colors and textures similar to the surrounding intestinal wall, and the boundaries and morphological features of the lesions are not obvious enough, making it difficult to accurately identify the lesion features. At the same time, due to the randomness of the shooting position and posture of the medical endoscope equipment, the distribution and size of similar lesions in gastrointestinal images are significantly different. Although the existing deep learning model can achieve end-to-end fully automatic gastrointestinal image classification, its success depends largely on sufficient training with a large amount of labeled data, which is difficult to meet for the diagnosis of rare diseases in the gastrointestinal tract.

[0004] In summary, developing a rare disease image classification method based on small sample learning has great practical significance and important scientific research value for the development of medical artificial intelligence. However, current technology still faces challenges in accuracy and has not yet been widely used in clinical practice. Summary of the invention

[0005] The present invention proposes a rare disease image classification method, device and storage medium based on small sample learning, which can solve the technical problem of how to solve the problem of difficulty in identifying lesion characteristics due to low contrast between lesions and intestinal walls and the inaccurate classification of gastrointestinal diseases including rare diseases due to large intra-class differences caused by differences in distribution and size between similar lesions.

[0006] To achieve the above object, the present invention adopts the following technical solutions: A rare disease image classification method based on small sample learning includes the following steps: S1. Collect an endoscopic image dataset containing rare diseases in the gastrointestinal tract and divide the dataset into a support set and a query set. S2, constructing a feature backtracking fusion encoder for extracting support image features in the support set and query image features in the query set, and passing these feature information to a multi-level prototype reconstruction network; S3, using a multi-level prototype reconstruction network to capture the semantic relevance between support set and query set samples, and using semantic relevance to enhance the distinguishing regions on the support image representation, generating the corresponding support set class center for each query sample; S4. Construct a classifier based on Euclidean distance to calculate the Euclidean distance between the query image features and all the support set features output by the multi-level prototype reconstruction network, and use the opposite of the Euclidean distance as the classification score, thus forming a rare disease image classification model based on small sample learning; S5. Using the collected training samples to train the rare disease image classification model based on small sample learning, and using the validation set samples to verify and select the optimal network parameters, so as to form the final rare disease image classification model based on small sample learning; S6. When applied, endoscopic images of the gastrointestinal tract containing rare diseases are collected and input into the above-mentioned rare disease image classification model based on small sample learning, and the classification results of each query sample are output after calculation.

[0007] Furthermore, the step S1 comprises: S11. Collect a dataset of endoscopic images of the gastrointestinal tract containing rare diseases and divide it into a training set , validation set and test set , and there is no overlapping category between the three, in the training set , validation set and test set It is further divided into support set and query set; S12, using bilinear interpolation method to adjust the resolution of all images to 224×224; S13, perform data augmentation by randomly flipping the images and their labels in the training dataset horizontally, rotating them 90 degrees clockwise, and rotating them 180 degrees clockwise; S14. Using the support image, support label, and query image as inputs of a neural network, and using the query label as supervisory information output by the rare disease image classification model based on small sample learning. Furthermore, the step S2 comprises: S21, construct a feature backtracking fusion encoder, the input support set image and the query set image share the feature backtracking fusion encoder, and obtain the support image feature and the query image features ; S22, the feature backtracking fusion encoder consists of a ResNet18 encoder with initialized pre-trained weights, a backtracking attention learning module, and a multi-scale fusion module; S221, initialize the ResNet18 encoder with pre-trained weights to input the support set image and query set image , extract the feature representation of each image respectively; The ResNet18 encoder consists of 1 convolutional layer, 1 maximum pooling layer and 4 residual blocks, where each residual block contains 2 convolutional layers and 1 residual connection; The ResNet18 encoder converts the 224×224 raw data into 7×7 feature encoding through a series of convolution and pooling operations; Assume the encoder input is First, a 7×7 convolutional layer is used to extract the 112×112 low-level features of the image, and a 56×56 intermediate layer feature is obtained through a maximum pooling layer. , then In the 4 residual blocks of the input series, the intermediate layer features of the 4 residual blocks output are recorded as ,in Indicates The output features of the residual blocks and , the residual block is expressed as ; S222, the backtracking attention learning module receives and The output features of the residual block and , and outputs a two-dimensional attention mask , where C is the number of channels of the feature map, H is the height of the feature map, W is the width of the feature map, and at this time middle , that is, for the features of each layer of residual blocks, in addition to the features And the last layer of features , which uses the features of the next layer of residual blocks to calculate a two-dimensional attention mask , and is used to refine the current residual block features; The retrospective attention learning module first uses convolutional layers with three convolution kernel sizes of 3×3, 5×5, and 7×7 to Perform convolution operations to obtain information features with different receptive fields ,in represents the convolution kernel size, Indicates The residual block features are convolved with a kernel size of k×k, and then two consecutive point-by-point convolutions are performed to ensure better nonlinear learning ability and capture the complex correlation between spatial positions; the first 1×1 convolution layer is to convert the features The channel from Compress to , is a reduction ratio, set to , then through Layer and After the layer, the second 1×1 convolutional layer is used to convert the channel from Converted to 1, the feature is then upsampled by a factor of 2 and passed through The function gets a two-dimensional attention mask , and finally the three two-dimensional attention masks The average obtained ; The above process is expressed as: ; ; in yes function, It is the upsampling process; finally, for the features of each layer of residual blocks, in addition to the features And the last layer of features , the refined features are ; S223, the multi-scale fusion module combines the features of each layer of residual blocks, except for the features And the last layer of features , the refined features are As input, we use the refined low-level features Layer-by-layer fusion is used to supplement low-level features at different levels, and the final features after the fusion of low-level features and high-level features are output. ; The multi-scale fusion module first transforms the input features The features of three different receptive fields are obtained through three pooling operations of 3×3, 5×5 and 7×7. , and then sum the corresponding elements of the three to get the feature ,in , and then use global average pooling Get a channel-level embedding vector , and after three fully connected layers and normalization operations, three weight vectors are obtained , multiply it by the corresponding three different receptive field features Get each layer feature , and finally merge layer by layer. As input features, we get , and so on to get the final feature .

[0008] Furthermore, the step S3 comprises: S31, multi-level prototype reconstruction network to feature back-fusion encoder to pass the final feature As input, it includes supporting image features and query image features ; S32, the multi-level prototype reconstruction network consists of a contrast level relationship module, a saliency level relationship module and an attention level relationship module in parallel S321, the comparison level relationship module first compares the query image features Perform global average pooling to get the embedding vector , and then use the subtraction operation to make the embedding vector of the query image and all supported image features , and compare each position to output the relationship between the instance-level objects of the two ; The output characteristic relationship formula is expressed as ,in Indicates a broadcast operation. Repeated expansion to , Indicates absolute value operation, and finally ; S322, the saliency level relationship module first performs query image features Perform global average pooling to get the embedding vector , then As a convolution kernel, it supports image features in a deep separation convolution manner. Perform a convolution operation to output the feature relationship between the two ; The output characteristic relationship formula is expressed as ,in represents the depth separation convolution operation, and finally ; S323, the attention level relationship module first obtains the two embedded features of the support image and the query image through 1×1 convolution, and then calculates the spatial attention weight matrix of the two according to the matrix multiplication between the two embedded features , then query the image features With the attention weight matrix Perform matrix multiplication to obtain spatially aware query image features , and extract spatially aware query image features through subtraction and supporting image features Local semantic relationship between , output Expressed as ,at last ; S324, the multi-level prototype reconstruction network will finally and , , Perform channel-level concatenation and compress the channel through 1×1 convolution , get the reconstructed supporting image features ,at last .

[0009] Furthermore, the step S4 comprises: S41, Euclidean distance-based classifier reconstructs the supporting image features transmitted in the multi-level prototype network and query image features As input, output the score of the query image belonging to each category vector; S411, the classifier based on Euclidean distance will first support image features Flatten and average the embedding vectors of the same category to get the prototype center of the category in Euclidean space , similarly, query image features Flattened to embedding vector ; S412, Euclidean distance-based classifier calculates query image embedding vector With prototype centers for each category to obtain the query image Belongs to each category Score , select the category with the highest score as the prediction result of the query image category.

[0010] Furthermore, the step S5 comprises: S51. Build a rare disease image classification model based on small sample learning, use the Adam optimizer, and set the initial learning rate to , the training epoch is set to 400, and each epoch consists of 200 episodes; S52, during the training process, using the training set The rare disease image classification model based on small sample learning is trained and optimized using the cross entropy loss function, which is expressed as: ,in Indicates the number of supported image categories; S53. Train the model through the loss function, adjust the parameters through back propagation, and select the model parameters with the highest accuracy in the validation set for saving.

[0011] Furthermore, the step S5 further includes: S54. Select evaluation indicators including accuracy to evaluate model performance.

[0012] On the other hand, the present invention further discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.

[0013] On the other hand, the present invention further discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.

[0014] It can be seen from the above technical scheme that the rare disease image classification method and system based on small sample learning of the present invention is intended to accurately classify various diseases in the gastrointestinal tract, including rare diseases. This method first cleverly constructs a feature backtracking fusion encoder and uses the next layer of features to generate an attention mask for the current layer to reduce useless noise information in low-level features, thereby better integrating spatial detail information in low-level features into high-level features, effectively improving the model's classification performance for gastrointestinal disease areas. Secondly, the multi-level prototype reconstruction network further captures the semantic relevance between the support set and query set samples to enhance the distinguishing areas on the support image representation, and generates a calibration class center suitable for the current query sample for each query sample. At the same time, the classifier based on Euclidean distance outputs the classification result of the query sample, and uses the cross entropy function to guide model optimization to ensure the accuracy of the classification result. Finally, the rare disease image classification model based on small sample learning is able to output the classification result of each image after rigorous training and testing. The present invention not only effectively solves the problem of difficulty in identifying lesion characteristics due to low contrast between lesions and intestinal walls, but also reduces intra-class differences and significantly improves classification accuracy, providing strong technical support for the diagnosis and treatment of gastrointestinal diseases including rare diseases. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a schematic diagram of the steps of the classification method according to an embodiment of the present invention; Figure 2 It is a structural diagram of a classification model according to an embodiment of the present invention; Figure 3 It is a schematic diagram of the structure of a feature backtracking fusion encoder according to an embodiment of the present invention; Figure 4 A schematic diagram of a multi-level prototype reconstruction network structure according to an embodiment of the present invention; Figure 5 It is a schematic diagram of the application effect of an embodiment of the present invention. DETAILED DESCRIPTION

[0016] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.

[0017] like Figure 1 As shown, the rare disease image classification method based on small sample learning described in this embodiment includes the following steps: Step 1: Collect endoscopic image datasets containing rare diseases in the gastrointestinal tract, which contain endoscopic images of different gastrointestinal tissues and corresponding tissue and lesion category annotations, divide the datasets into training set, validation set and test set, and further divide the training samples, validation samples and test samples into support set and query set respectively; Specifically, the step 1 includes: Step 1.1: Collect a dataset of endoscopic images of the gastrointestinal tract containing rare diseases and divide it into a training set , validation set and test set , and there is no overlapping category between the three, in the training set , validation set and test set It is further divided into support set and query set; Step 1.2: Use bilinear interpolation to adjust the resolution of all images to 224×224; Step 1.3: Perform data augmentation by randomly flipping the images and their labels in the training dataset horizontally, rotating them 90 degrees clockwise, and rotating them 180 degrees clockwise. Step 1.4: Use the support image, support label, and query image as inputs of the neural network, and use the query label as supervisory information output by the rare disease image classification model based on small sample learning.

[0018] Step 2: Figure 2 As shown, a feature backtracking fusion encoder is constructed to extract support image features in the support set, extract query image features in the query set, and pass these feature information to a multi-stage prototype reconstruction network; The step S2 comprises: Step 2.1: Construct a feature back-tracing fusion encoder. The input support set image and the query set image share the feature back-tracing fusion encoder to obtain the support image feature. and the query image features ; Step 2.2, the feature backtracking fusion encoder consists of a ResNet18 encoder with initialized pre-trained weights, a backtracking attention learning module, and a multi-scale fusion module. The backtracking attention learning module refines shallow features and cleverly fuses shallow features with spatial details and deep information with rich semantic context information layer by layer, so as to use feature information at different levels to solve the problem of gastrointestinal diseases with low contrast between the intestines and difficult to identify lesions; Step 2.2.1: Initialize the ResNet18 encoder with pre-trained weights and input the support set image and query set image , extract the feature representation of each image respectively. The ResNet18 encoder consists of 1 convolution layer, 1 maximum pooling layer and 4 residual blocks, where each residual block contains 2 convolution layers and 1 residual connection. After a series of convolution and pooling operations, the ResNet18 encoder can convert the 224×224 raw data into 7×7 feature encoding. Assume that the encoder input is First, a 7×7 convolutional layer is used to extract the 112×112 low-level features of the image, and a 56×56 intermediate layer feature is obtained through a maximum pooling layer. , then In the 4 residual blocks of the input series, the intermediate layer features of the 4 residual blocks output are recorded as ,in Indicates The output features of the residual blocks and , the residual block is expressed as ; Step 2.2.2: The backtracking attention learning module receives the and The output features of the residual block and , and outputs a two-dimensional attention mask , where C is the number of channels of the feature map, H is the height of the feature map, W is the width of the feature map, and at this time middle , that is, for the features of each layer of residual blocks (except for the features And the last layer of features ), which uses the features of the next layer of residual blocks to calculate a two-dimensional attention mask , and is used to refine the features of the current residual block. The retrospective attention learning module first uses convolutional layers with three convolution kernel sizes of 3×3, 5×5, and 7×7 to Perform convolution operations to obtain information features with different receptive fields ,in represents the convolution kernel size, Indicates The features of the residual block are convolved with a kernel size of k×k, and then two consecutive point-by-point convolutions are performed to ensure better nonlinear learning ability and capture the complex correlations between spatial positions. The first 1×1 convolution layer is to convert the features The channel from Compress to ( is a reduction ratio, set to ), then through Layer and After the layer, the second 1×1 convolutional layer is used to convert the channel from Converted to 1, the feature is then upsampled by a factor of 2 and passed through The function gets a two-dimensional attention mask , and finally the three two-dimensional attention masks The average obtained The above process can be expressed as: ; ; in yes function, is the upsampling process. Finally, the features of each layer of residual blocks (except for the features And the last layer of features ) The refined features are . Step 2.2.3: The multi-scale fusion module combines the features of each layer of residual blocks (except for the features And the last layer of features ) The refined features are As input, we use the refined low-level features Layer-by-layer fusion is used to supplement low-level features at different levels, and the final features after the fusion of low-level features and high-level features are output. The multi-scale fusion module first transforms the input features The features of three different receptive fields are obtained through three pooling operations of 3×3, 5×5 and 7×7. , and then sum the corresponding elements of the three to get the feature ,in , and then use global average pooling Get a channel-level embedding vector , and after three fully connected layers and normalization operations, three weight vectors are obtained , multiply it by the corresponding three different receptive field features Get each layer feature , and finally merge layer by layer. As input features, we get , and so on to get the final feature .

[0019] Step 3: Figure 3 As shown, a multi-level prototype reconstruction network is constructed and used to capture the semantic correlation between the support set and the query set samples, and the semantic correlation is used to enhance the distinguishing regions on the support image representation, and the corresponding support set class center is generated for each query sample; The step S3 specifically includes: Step 3.1: Multi-level prototype reconstruction network uses feature back-fusion to transfer the final features of the encoder As input, it includes supporting image features and query image features ; Step 3.2, the multi-level prototype reconstruction network is composed of a contrast level relationship module, a saliency level relationship module and an attention level relationship module in parallel; Step 3.2.1: The contrast-level relation module first compares the query image features Perform global average pooling to get the embedding vector , and then use the subtraction operation to make the embedding vector of the query image and all supported image features , and compare each position to output the relationship between the instance-level objects of the two , to better represent the global features such as the contour boundary of the query object. The output feature relationship formula is expressed as ,in Indicates a broadcast operation. Repeated expansion to , Indicates absolute value operation, and finally .

[0020] Step 3.2.2: The salient level relationship module first analyzes the query image features Perform global average pooling to get the embedding vector , then As a convolution kernel, it supports image features in a deep separation convolution manner. Perform a convolution operation to output the feature relationship between the two , in order to better highlight the main area features of the query object. The output feature relationship formula is expressed as ,in represents the depth separation convolution operation, and finally .

[0021] Step 3.2.3: The attention-level relationship module first obtains the two embedded features of the support image and the query image through 1×1 convolution, and then calculates the spatial attention weight matrix of the two by matrix multiplication between the two embedded features. , then query the image features With the attention weight matrix Perform matrix multiplication to obtain spatially aware query image features , and extract spatially aware query image features through subtraction and supporting image features Local semantic relationship between , output Expressed as ,at last ; Step 3.2.4, the multi-level prototype reconstruction network finally and , , Perform channel-level concatenation and compress the channel through 1×1 convolution , get the reconstructed supporting image features ,Right now , thereby making full use of the complementary advantages to comprehensively extract the semantic information generated by the three different relationship modules and enhance the query feature area owned by the supporting features to solve the problem of large differences within the gastrointestinal image class.

[0022] Step 4: Figure 4 As shown, a classifier based on Euclidean distance is constructed to calculate the Euclidean distance between the query image features and all the support set features output by the multi-level prototype reconstruction network, and the opposite of the Euclidean distance is used as the classification score; The step S4 specifically includes: Step 4.1: The Euclidean distance-based classifier reconstructs the supporting image features transmitted in the multi-level prototype network and query image features As input, output the score of the query image belonging to each category vector; Step 4.1.1: The classifier based on Euclidean distance first supports image features Flatten and average the embedding vectors of the same category to get the prototype center of the category in Euclidean space , similarly, query image features Flattened to embedding vector ; Step 4.1.2: Calculate the query image embedding vector based on the Euclidean distance classifier With prototype centers for each category to obtain the query image Belongs to each category Score , select the category with the highest score as the prediction result of the query image category.

[0023] Step 5: Use the collected training samples to train the gastrointestinal disease classification model composed of the feature backtracking fusion encoder, the multi-level prototype reconstruction network and the classifier based on Euclidean distance, and use the validation set samples to verify and select the optimal network parameters to form the final model; The step S5 specifically includes: Step 5.1: Build a rare disease image classification model based on small sample learning, use the Adam optimizer, and set the initial learning rate to , the training epoch is set to 400, and each epoch consists of 200 episodes.

[0024] Step 5.2: During the training process, use the training set The rare disease image classification model based on small sample learning is trained and optimized using the cross entropy loss function, which is expressed as: ,in Indicates the number of supported image categories.

[0025] Step 5.3: Train the model through the loss function, adjust the parameters through back propagation, and select the model parameters with the highest accuracy in the validation set for saving; Step 5.4: Select evaluation indicators including accuracy to evaluate model performance.

[0026] Step 6: When applied, collect endoscopic images of the gastrointestinal tract containing rare diseases, and input them into the rare disease image classification model based on small sample learning, and output the classification results of each query sample after calculation.

[0027] like Figure 5 As shown in the figure, the input image is an image taken by an endoscope in the gastrointestinal tract, the label is the image classification result manually annotated by professional gastrointestinal disease experts, and the model output result is the classification accuracy result automatically output by the rare disease image classification model based on small sample learning. The average classification accuracy of the five diseases reached 83.56%.

[0028] In summary, the embodiment of the present invention cleverly constructs a feature backtracking fusion encoder and uses the next layer of features to generate an attention mask for the current layer to reduce useless noise information in the low-level features, thereby better integrating the spatial detail information in the low-level features into the high-level features, and effectively improving the classification performance of the model for gastrointestinal disease areas including rare diseases. Secondly, the introduction of the multi-level prototype reconstruction network further captures the semantic relevance between the support set and the query set samples to enhance the distinguishing areas on the support image representation, and generates a calibration class center suitable for the current query sample for each query sample. At the same time, the classifier based on the Euclidean distance outputs the classification result of the query sample, and the cross entropy function is used to guide the model optimization to ensure the accuracy of the classification result. Finally, the rare disease image classification model based on small sample learning has shown excellent classification performance after rigorous training and testing. It not only solves the problem of difficulty in identifying lesion features due to low contrast between lesions and intestinal walls, but also reduces intra-class differences, providing strong technical support for the diagnosis and treatment of gastrointestinal diseases including rare diseases.

[0029] On the other hand, the present invention further discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.

[0030] On the other hand, the present invention further discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.

[0031] In another embodiment provided in the present application, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute any of the rare disease image classification methods based on small sample learning in the above embodiments.

[0032] It is understandable that the system, device and storage medium provided in the embodiments of the present invention correspond to the method provided in the embodiments of the present invention, and the explanation, examples and beneficial effects of the relevant contents can refer to the corresponding parts in the above methods.

[0033] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a server or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive Solid State Disk (SSD)), etc.

[0034] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0035] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0036] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A rare disease image classification method based on small sample learning, characterized in that: The following steps are included: S1. Collect an endoscopic image dataset containing rare diseases in the gastrointestinal tract and divide the dataset into a support set and a query set. S2, constructing a feature backtracking fusion encoder for extracting support image features in the support set and query image features in the query set, and passing these feature information to a multi-level prototype reconstruction network; S3, using a multi-level prototype reconstruction network to capture the semantic relevance between support set and query set samples, and using semantic relevance to enhance the distinguishing regions on the support image representation, generating the corresponding support set class center for each query sample; S4. Construct a classifier based on Euclidean distance to calculate the Euclidean distance between the query image features and all the support set features output by the multi-level prototype reconstruction network, and use the opposite of the Euclidean distance as the classification score, thus forming a rare disease image classification model based on small sample learning; S5. Using the collected training samples to train the rare disease image classification model based on small sample learning, and using the validation set samples to verify and select the optimal network parameters, so as to form the final rare disease image classification model based on small sample learning; S6. When applied, endoscopic images of the gastrointestinal tract containing rare diseases are collected and input into the rare disease image classification model based on small sample learning, and the classification results of each query sample are output after calculation.

2. The rare disease image classification method based on small sample learning according to claim 1, characterized in that: The step S1 comprises: S11. Collect a dataset of endoscopic images of the gastrointestinal tract containing rare diseases and divide it into a training set , validation set and test set , and there is no overlapping category between the three, in the training set , validation set and test set It is further divided into support set and query set; S12, using bilinear interpolation method to adjust the resolution of all images to 224×224; S13, perform data augmentation by randomly flipping the images and their labels in the training dataset horizontally, rotating them 90 degrees clockwise, and rotating them 180 degrees clockwise; S14. Using the support image, support label, and query image as inputs of a neural network, and using the query label as supervisory information output by the rare disease image classification model based on small sample learning.

3. The rare disease image classification method based on small sample learning according to claim 1, characterized in that: The step S2 comprises: S21, construct a feature backtracking fusion encoder, the input support set image and the query set image share the feature backtracking fusion encoder, and obtain the support image feature and the query image features ; S22, the feature backtracking fusion encoder consists of a ResNet18 encoder with initialized pre-trained weights, a backtracking attention learning module, and a multi-scale fusion module; S221, initialize the ResNet18 encoder with pre-trained weights to input the support set image and query set image , extract the feature representation of each image respectively; The ResNet18 encoder consists of 1 convolutional layer, 1 maximum pooling layer and 4 residual blocks, where each residual block contains 2 convolutional layers and 1 residual connection; The ResNet18 encoder converts the 224×224 raw data into 7×7 feature encoding through a series of convolution and pooling operations; Assume the encoder input is First, a 7×7 convolutional layer is used to extract the 112×112 low-level features of the image, and a 56×56 intermediate layer feature is obtained through a maximum pooling layer. , then In the 4 residual blocks of the input series, the intermediate layer features of the 4 residual blocks output are recorded as ,in Indicates The output features of the residual blocks and , the residual block is expressed as ; S222, the backtracking attention learning module receives and The output features of the residual block and , and outputs a two-dimensional attention mask , where C is the number of channels of the feature map, H is the height of the feature map, W is the width of the feature map, and at this time middle , that is, for the features of each layer of residual blocks, in addition to the features And the last layer of features , which uses the features of the next layer of residual blocks to calculate a two-dimensional attention mask , and is used to refine the current residual block features; The retrospective attention learning module first uses convolutional layers with three convolution kernel sizes of 3×3, 5×5, and 7×7 to Perform convolution operations to obtain information features with different receptive fields ,in represents the convolution kernel size, Indicates The residual block features are convolved with a kernel size of k×k, and then two consecutive point-by-point convolutions are performed to ensure better nonlinear learning ability and capture the complex correlation between spatial positions; the first 1×1 convolution layer is to convert the features The channel from Compress to , is a reduction ratio, set to , then through Layer and After the layer, the second 1×1 convolutional layer is used to convert the channel from Converted to 1, the feature is then upsampled by a factor of 2 and passed through The function gets a two-dimensional attention mask , and finally the three two-dimensional attention masks The average obtained ; The above process is expressed as: ; ; in yes function, It is the upsampling process; finally, for the features of each layer of residual blocks, in addition to the features And the last layer of features , the refined features are ; S223, the multi-scale fusion module combines the features of each layer of residual blocks, except for the features And the last layer of features , the refined features are As input, we use the refined low-level features Layer-by-layer fusion is used to supplement low-level features at different levels, and the final features after the fusion of low-level features and high-level features are output. ; The multi-scale fusion module first transforms the input features The features of three different receptive fields are obtained through three pooling operations of 3×3, 5×5 and 7×7. , and then sum the corresponding elements of the three to get the feature ,in , and then use global average pooling Get a channel-level embedding vector , and after three fully connected layers and normalization operations, three weight vectors are obtained , multiply it by the corresponding three different receptive field features Get each layer feature , and finally merge layer by layer. As input features, we get , and so on to get the final feature .

4. The rare disease image classification method based on small sample learning according to claim 3 is characterized by: The step S3 comprises: S31, multi-level prototype reconstruction network to feature back-fusion encoder to pass the final feature As input, it includes supporting image features and query image features ; S32, the multi-level prototype reconstruction network consists of a contrast level relationship module, a saliency level relationship module and an attention level relationship module in parallel S321, the comparison level relationship module first compares the query image features Perform global average pooling to get the embedding vector , and then use the subtraction operation to make the embedding vector of the query image and all supported image features , and compare each position to output the relationship between the instance-level objects of the two ; The output characteristic relationship formula is expressed as ,in Indicates a broadcast operation. Repeated expansion to , Indicates absolute value operation, and finally ; S322, the saliency level relationship module first performs query image features Perform global average pooling to get the embedding vector , then As a convolution kernel, it supports image features in a deep separation convolution manner. Perform a convolution operation to output the feature relationship between the two ; The output characteristic relationship formula is expressed as ,in represents the depth separation convolution operation, and finally ; S323, the attention level relationship module first obtains the two embedded features of the support image and the query image through 1×1 convolution, and then calculates the spatial attention weight matrix of the two according to the matrix multiplication between the two embedded features , then query the image features With the attention weight matrix Perform matrix multiplication to obtain spatially aware query image features , and extract spatially aware query image features through subtraction and supporting image features Local semantic relationship between , output Expressed as ,at last ; S324, the multi-level prototype reconstruction network will finally and , , Perform channel-level concatenation and compress the channel through 1×1 convolution , get the reconstructed supporting image features ,at last .

5. The rare disease image classification method based on small sample learning according to claim 4 is characterized in that: The step S4 comprises: S41, Euclidean distance-based classifier reconstructs the supporting image features transmitted in the multi-level prototype network and query image features As input, output the score of the query image belonging to each category vector; S411, the classifier based on Euclidean distance will first support image features Flatten and average the embedding vectors of the same category to get the prototype center of the category in Euclidean space , similarly, query image features Flattened to embedding vector ; S412, Euclidean distance-based classifier calculates query image embedding vector With prototype centers for each category to obtain the query image Belongs to each category Score , select the category with the highest score as the prediction result of the query image category.

6. The rare disease image classification method based on small sample learning according to claim 5, characterized in that: The step S5 comprises: S51. Build a rare disease image classification model based on small sample learning, use the Adam optimizer, and set the initial learning rate to , the training epoch is set to 400, and each epoch consists of 200 episodes; S52, during the training process, using the training set The rare disease image classification model based on small sample learning is trained and optimized using the cross entropy loss function, which is expressed as: ,in Indicates the number of supported image categories; S53. Train the model through the loss function, adjust the parameters through back propagation, and select the model parameters with the highest accuracy in the validation set for saving.

7. The rare disease image classification method based on small sample learning according to claim 6 is characterized by: The step S5 further comprises: S55. Select evaluation indicators including accuracy to evaluate model performance.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the computer program is executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Intestinal polyp segmentation method and system based on small sample learning

    CN115049603A

  • Small sample fine-grained image classification method and system

    CN116824274A

  • Fine-grained small sample classification method based on task specific channel reconstruction network

    CN116843970A

  • Instance-level small sample learning method based on inter-class distance expansion

    CN118015382A

  • Rare fundus disease automatic classification method and system based on OCT image

    CN118411573A

Cited By

  • Endoscope image data enhancement method and system based on adversarial network

    CN121304471A

  • A hierarchical latent state prototype generation method for low-resource biomedical image classification

    CN122637034A