Fundus Image Classification Method and System

By introducing the second branch block into ResNet-50, the processing of fundus images by the ResNet-50 model is optimized, and the problem of low classification accuracy of fundus images is solved, achieving higher classification accuracy and better feature extraction effects.

CN119399546BActive Publication Date: 2025-06-20江苏富翰医疗产业发展有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411576083.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-06
Publication Date
2025-06-20
Estimated Expiration
2044-11-06

AI Technical Summary

Technical Problem

The classification accuracy of fundus images is low, and the existing ResNet-50 model fails to effectively capture the subtle structural changes and specific features in OCT images.

Method used

By modifying the residual block of ResNet-50, a second branch block is introduced, including a separable convolutional layer, a fully connected layer, a star operation layer and a separable convolutional layer, optimized noise processing and feature extraction.

Benefits of technology

The classification accuracy of fundus images is improved, and through more effective feature extraction and denoising processing, the network's ability to capture OCT image details is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399546B_ABST
    Figure CN119399546B_ABST
Patent Text Reader

Abstract

This application relates to the field of computer vision technology, and provides a fundus image classification method and system. The method includes obtaining a first image to be classified, and then constructing a classification network. The classification network includes at least a first module and a second module. The second module includes at least one first branch block and a second branch block. The first image to be classified is input into the first module to output a first intermediate image. Then, the second intermediate image is input into the second branch block to output a first feature map. The second intermediate image is the image output by the first branch block. The second branch block includes a first separable convolutional layer, a first fully connected layer, a star operation layer, a second fully connected layer, and a second separable convolutional layer. Based on the first feature map, a first classification result is output through the classification network. By changing the second module, the second module optimizes noise and can extract and process the detailed information of the image to be classified, so as to improve the classification accuracy of fundus images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular, to a fundus image classification method and system. Background Art

[0002] ResNet-50 (Residual Network with 50 layers) is a deep convolutional neural network used to solve the problems of gradient disappearance and gradient explosion in the training of deep networks. ResNet-50 is divided into five stages, and each stage contains several residual blocks. Each residual block uses a skip connection to directly add the input to the output, and through the convolutional layer structure, feature extraction and dimensionality reduction are achieved.

[0003] The second residual module of ResNet-50 is located in the third stage of the network and contains four residual blocks. Each block consists of a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer. Each layer is equipped with batch normalization and a ReLU activation function. These blocks combine the input and output through skip connections to achieve feature extraction and non-linear mapping within the ResNet-50 framework, providing a basis for the four-classification task of OCT images.

[0004] However, for the four-classification task of OCT images, no optimization is performed for the characteristics of fundus images, resulting in a low classification accuracy of fundus images. Summary of the Invention

[0005] This application provides a fundus image classification method and system to solve the problem of low classification accuracy of fundus images.

[0006] This application provides a fundus image classification method, including:

[0007] Obtain a first image to be classified, where the first image to be classified is an OCT image;

[0008] Construct a classification network, where the classification network includes at least a first module and a second module, and the second module includes at least one first branch block and a second branch block;

[0009] Input the first image to be classified into the first module to output a first intermediate image;

[0010] Input a second intermediate image into the second branch block to output a first feature map. The second intermediate image is the image obtained by inputting the first intermediate image into the first branch block, and the output of the first branch block. The second branch block includes a first depthwise separable convolutional layer, a first fully connected layer, a star operation layer, a second fully connected layer, and a second depthwise separable convolutional layer. The first feature map is the feature map obtained by fusing the first intermediate image and the output of the second depthwise separable convolutional layer;

[0011] Based on the first feature map, output a first classification result through the classification network.

[0012] In some feasible embodiments, the method further includes:

[0013] Input a second intermediate image into the first separable convolutional layer to output a second feature map;

[0014] Input the second feature map into two of the first fully connected layers respectively to output a first vector and a second vector;

[0015] Input the first vector and the second vector into the star operation layer to output a third vector, and the star operation layer is used to perform an element-wise multiplication operation on the first vector and the second vector;

[0016] Input the third vector into the second fully connected layer to output a fourth vector;

[0017] Input the fourth vector into the second separable convolutional layer to output a third feature map, and the third feature map is the output of the second separable convolutional layer.

[0018] In some feasible embodiments, the inputting the first processed image into the second branch block to output a first feature map includes:

[0019] According to the dimension of the third feature map, transform the dimension of the second intermediate image so that the dimensions of the third feature map and the second intermediate image are consistent;

[0020] Perform an addition operation on the transformed second intermediate image and the third feature map to output a first feature map.

[0021] In some feasible embodiments, the inputting the first vector and the second vector into the star operation layer to output a third vector includes:

[0022] Obtain a first linear transformation matrix and a second linear transformation matrix;

[0023] Based on the number of rows and columns of the first linear transformation matrix, generate a first transposed matrix, and based on the number of rows and columns of the second linear transformation matrix, generate a second transposed matrix;

[0024] Input the first vector into the first transposed matrix to generate a first linear result, and input the second vector into the second transposed matrix to generate a second linear result;

[0025] Perform an element-wise multiplication on the first linear result and the second linear result to output a third vector.

[0026] In some feasible embodiments, the second module further includes a third branch block, and the third branch block has the same structure as the second branch block;

[0027] Input the first feature map into the third branch block to output a first processed image.

[0028] In some feasible embodiments, the classification network further includes a third module, a fourth module, and a fifth module;

[0029] The method further includes:

[0030] Input the first processed image into the third module to output a second processed image, and the third module includes at least three of the second branch blocks;

[0031] Input the second processed image into the fourth module to output a third processed image, and the third module includes at least four of the second branch blocks;

[0032] Input the third processed image into the fifth module to output a fourth processed image, and the third module includes at least five of the second branch blocks.

[0033] In some feasible embodiments, the classification network further includes a first auxiliary module, a second auxiliary module, and a third auxiliary module;

[0034] Outputting a first classification result based on the first feature map through the classification network includes:

[0035] Input the first image to be classified, the first processed image, and the second processed image into the first auxiliary module to perform fused features through the first auxiliary module to generate a first fused image;

[0036] Input the first fused image, the first processed image, the second processed image, and the third processed image into the second auxiliary module to perform fused features through the second auxiliary module to generate a second fused image;

[0037] Input the second fused image, the first processed image, the second processed image, the third processed image, and the fourth processed image into the third auxiliary module to perform fused features through the third auxiliary module to generate a first classification result.

[0038] In some feasible embodiments, after outputting the first classification result based on the first feature map through the classification network, it further includes:

[0039] Obtain a second image to be classified;

[0040] Input the second image to be classified into the classification network for inference calculation through the classification network, and output a second classification result. The classification network further includes a third module, a fourth module, and a fifth module.

[0041] In some feasible embodiments, before inputting the second intermediate image into the second branch block to output a first feature map, it includes:

[0042] Input the first image to be classified into a convolutional layer to output a third intermediate image. The first module includes a convolutional layer and a pooling layer;

[0043] Input the third intermediate image into the pooling layer to output a first intermediate image;

[0044] Input the first intermediate image into the first branch block to output a second intermediate image.

[0045] In a second aspect, the present application provides a fundus image classification system, including:

[0046] An acquisition unit configured to acquire a first image to be classified, where the first image to be classified is an OCT image;

[0047] A construction unit configured to construct a classification network, where the classification network at least includes a first module and a second module, and the second module at least includes one first branch block and a second branch block;

[0048] It is further configured to input the second intermediate image into the second branch block to output a first feature map. The second intermediate image is the output of the first branch block. The second branch block includes a first depthwise separable convolutional layer, a first fully connected layer, a star operation layer, a second fully connected layer, and a second depthwise separable convolutional layer. The first feature map is a feature map obtained by fusing the output of the first processed image and the second depthwise separable convolutional layer;

[0049] And, based on the first feature map, output a first classification result through the classification network.

[0050] As can be seen from the above technical solutions, the present application provides a fundus image classification method and system. The method includes obtaining a first image to be classified, and then constructing a classification network. The first image to be classified is an OCT image. The classification network includes at least a first module and a second module. The second module includes at least one first branch block and a second branch block. The first image to be classified is input into the first module to output a first intermediate image. Then, the second intermediate image is input into the second branch block to output a first feature map. The second intermediate image is the image obtained by inputting the first intermediate image into the first branch block and output by the first branch block. The second branch block includes a first depthwise separable convolutional layer, a first fully connected layer, a star operation layer, a second fully connected layer, and a second depthwise separable convolutional layer. The first feature map is the feature map obtained by fusing the output of the first intermediate image and the second depthwise separable convolutional layer. Based on the first feature map, a first classification result is output through the classification network. By changing the second module, the second module optimizes the noise and can extract and process the detailed information of the image to be classified, so as to improve the classification accuracy of the fundus image. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions of the present application, the drawings required for the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0052] Figure 1 It is a schematic structural diagram of the ResNet-50 network provided by the embodiment of the present application;

[0053] Figure 2 It is a schematic flowchart of the fundus image classification method provided by the embodiment of the present application;

[0054] Figure 3 It is a schematic structural diagram of the classification network provided by the embodiment of the present application;

[0055] Figure 4 It is a schematic structural diagram of the first branch block provided by the embodiment of the present application;

[0056] Figure 5 It is a schematic diagram of the first structure of the second branch block provided by the embodiment of the present application;

[0057] Figure 6 It is a schematic diagram of the second structure of the second branch block provided by the embodiment of the present application;

[0058] Figure 7 It is a schematic diagram of the third structure of the second branch block provided by the embodiment of the present application;

[0059] Figure 8Schematic diagram of the auxiliary network structure provided by the embodiments of the present application;

[0060] Figure 9 Schematic diagram of the structure of the fundus image classification system provided by the embodiments of the present application. Detailed implementation manners

[0061] The embodiments will be described in detail below, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following embodiments do not represent all implementation manners consistent with the present application. They are only examples of the systems and methods consistent with some aspects of the present application detailed in the claims.

[0062] The four-class classification task of ophthalmic OCT (Optical Coherence Tomography) images refers to classifying OCT images into four different categories, which are related to retinal diseases. Among them, the four categories include: Choroidal Neovascularization (CNV), Diabetic Macular Edema (DME), Age-related Macular Degeneration (AMD), and NORMAL (normal retina).

[0063] CNV is the newly formed blood vessels abnormally formed in the choroid, which is related to various retinal diseases, such as AMD, etc. In OCT images, CNV appears as a neovascular membrane and related subretinal fluid.

[0064] DME is a complication of diabetic retinopathy, which is manifested as thickening of the retina in the macular area and accumulation of intraretinal fluid. In OCT images, the intraretinal fluid related to the thickening of the retina in DME can be clearly observed.

[0065] AMD is an age-related change in the structure of the macular area, and abnormal metabolites of pigment epithelial cells are deposited on the retina to form drusen. In OCT images, dry AMD is manifested as atrophy in the macular area, hyperreflectivity under the RPE (retinal pigment epithelium), and choroidoretinal atrophy, while wet AMD is manifested as a choroidal neovascular membrane in the macular area, retinal pigment epithelial detachment, and subretinal fluid.

[0066] NORMAL represents the normal retinal structure without any retinal diseases. In OCT images, the normal retina has a preserved foveal contour and no retinal fluid or edema.

[0067] In practical applications, to improve the accuracy and efficiency of the four-classification task for OCT images, technologies such as deep learning are adopted to build classification models. These models can automatically learn the feature representations in OCT images and optimize the classification performance through a large amount of training data.

[0068] For example, ResNet50, which is a deep convolutional neural network used to address the problems of vanishing gradients and exploding gradients in the training of deep networks, as Figure 1 shown, the ResNet50 network is divided into five stages, namely STAGE0, STAGE1, STAGE2, STAGE3, and STAGE4. Each stage contains several residual blocks and has a total of 50 layers. Each residual block uses a skip connection to add the input to the output, and through the convolutional layer structure of 1×1, 3×3, and 1×1, feature extraction and dimensionality reduction are achieved.

[0069] The second residual module of ResNet-50 is located in STAGE3 of the network and contains four residual blocks. Each residual block is composed of a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer, and each layer is followed by batch normalization and the ReLU activation function. These blocks directly add the input to the output through skip connections, which are used to achieve effective feature extraction and non-linear mapping within the ResNet50 framework and can provide a basis for the four-classification task of OCT images.

[0070] However, the depth of ResNet-50 does not match its complexity. Although ResNet-50 has sufficient depth to extract high-level features in images, for OCT images, its complexity, such as detailed textures, tiny lesions, etc., may require a deeper or more refined network structure to capture. The subtle structural changes in OCT images are important for distinguishing different lesion types, but the standard ResNet-50 may not fully pay attention to these subtle differences.

[0071] Secondly, ResNet-50 increases the depth of the network by stacking convolutional layers and residual connections, but this limits the receptive field size of some layers in the network. For the association between the global structure and local details in OCT images, a larger receptive field may be required to capture. Insufficient receptive fields may cause the network to be unable to effectively distinguish lesion types with similar global structures but significant differences in local details.

[0072] Lesion types in ophthalmic OCT images, such as choroidal neovascularization, diabetic macular edema, and cystoid macular edema, often have specific morphological features, and these features may not be prominent enough in the general feature representation of ResNet-50. The standard ResNet-50 may not fully learn these specific features during the training process, resulting in a decrease in classification accuracy.

[0073] Although ResNet-50 performs well in many tasks, its high computational complexity and large number of parameters may lead to computational resource bottlenecks when dealing with large-scale or high-resolution medical images. For the four-class classification task of OCT images, if the dataset is large or the image resolution is high, ResNet-50 may experience a decrease in processing speed due to excessive computational resource requirements.

[0074] ResNet-50 is designed for image recognition tasks and is not optimized for the characteristics of ophthalmic images. For example, problems such as noise, artifacts, and uneven brightness in ophthalmic images may require specific preprocessing steps or network structures to overcome. In addition, the types of lesions in ophthalmic images are diverse and complex, and domain-specific prior knowledge or network structures may need to be introduced to improve classification performance.

[0075] To solve the above problems, the present application provides a fundus image classification method and system. The method modifies the residual blocks of ResNet-50 for the OCT four-class classification task. Among them, the convolutional layers of the modified residual blocks reduce the amount of computation and the number of parameters, improving the computational efficiency of the network. When processing images, it has both the deep feature extraction ability of ResNet and can utilize the efficient denoising function of the second branch block, thereby improving the effect and accuracy of image processing. The denoising effect is made more obvious by optimizing the noise points through the star operation layer. In addition, through the feature transformation of multiple fully connected layers, the modified residual blocks can better extract and process the detailed information in the images, improving the accuracy of fundus image classification.

[0076] As Figure 2 shown, the fundus image classification method provided by some embodiments of the present application includes the following steps:

[0077] S100: Obtain a first image to be classified.

[0078] The first image to be classified is an OCT image that needs to be classified. OCT imaging is a technique that uses optical principles to obtain cross-sectional images of biological tissues.

[0079] OCT images can be retinal images, optic nerve images, corneal images, or other ocular structure images. OCT technology is used to obtain detailed images of the retina to evaluate the layered structure, blood vessel distribution, and possible pathological changes of the retina, such as retinal detachment, macular degeneration, etc. OCT is also used to observe cross-sectional images of the optic nerve to diagnose optic nerve diseases. OCT technology can also be used to evaluate the structure of the cornea, for example, in evaluating corneal thickness, corneal diseases, etc. OCT can also be used to observe other structures of the eye, such as the sclera, anterior chamber, etc.

[0080] Input the first image to be classified into the ResNet-50 model after modifying the residual block. The model will automatically extract features from the image using a deep learning architecture and predict the category to which the first image to be classified belongs. In the four-class classification task of OCT images, these four categories may be defined based on specific diseases, structural features, or pathological changes. For example, normal, choroidal neovascularization, diabetic macular edema, and cystoid macular edema.

[0081] S200: Build a classification network.

[0082] The classification network is modified based on the ResNet-50 network. As Figure 1 , Figure 3 shown, the ResNet-50 network includes five stages, namely STAGE0, STAGE1, STAGE2, STAGE3, and STAGE4. Among them, STAGE0 corresponds to Figure 3 the first module in Figure 3 , STAGE1 corresponds to Figure 3 the second module in Figure 3 , STAGE2 corresponds to Figure 3 the third module in

[0083] The structure of the first module of the classification network is the same as that of the ResNet-50 network, both including a convolutional layer and a pooling layer. Among them, the convolutional layer is 7×7, and it may also include a BN (Batch Normalization) layer, a RELU (Rectified Linear Unit) function, etc.

[0084] The second module includes a first branch block and two second branch blocks. For the convenience of description, the two second branch blocks are defined as the second branch block and the third branch block; the third module includes a first branch block and three second branch blocks, the fourth module includes a first branch block and four second branch blocks, and the fifth module includes a first branch block and five second branch blocks.

[0085] As Figure 4 described, the first branch block includes a 1×1 convolutional layer, a 3×3 convolutional layer, and a ReLU activation function; the input of the first branch block comes from the output of the first module.

[0086] The input features first pass through the first 1×1 convolutional layer, which can reduce the dimension, decrease the computational amount, and change the number of channels of the feature map. Meanwhile, a ReLU activation function can be included after this layer to normalize and non-linearly transform the features. The features processed by the first 1×1 convolutional layer enter the 3×3 convolutional layer for further feature learning to extract higher-level feature representations. Similarly, a ReLU activation function can also be included after this layer. The features processed by the 3×3 convolutional layer then pass through the second 1×1 convolutional layer to restore the number of channels to match the input number of channels or adjust according to the needs of the residual connection.

[0087] The features that pass through the 1×1 convolutional layer and then the ReLU activation function are added to the output of the second 1×1 convolutional layer to obtain a new feature map, whose number of channels is four times that of the initial input, and the width and height are reduced by a quarter.

[0088] The first branch block enables effective information transmission even when the network depth increases, avoiding the problem of gradient vanishing.

[0089] As Figure 5 shown, the second branch block includes a first depthwise separable convolution (DW-Conv) layer, a first fully connected layer (FC), a star operation layer, a second fully connected layer, and a second depthwise separable convolution layer; the structure of the third branch block is the same as that of the second branch block; the input of the second branch block is the output of the first branch block, and the input of the third branch block is the output of the second branch block, that is to say, the output of the third branch block is the output of the second module.

[0090] The first depthwise separable convolution layer and the second depthwise separable convolution layer are depthwise separable convolution layers. DW Conv consists of depthwise convolution and pointwise convolution.

[0091] Depthwise convolution performs convolution operations on each channel of the input feature map separately, and each channel is only convolved by one convolution kernel. The number of convolution kernels used is the same as the number of channels of the input feature map. Therefore, the number of output feature maps is also the same as the input number of channels. It can effectively reduce the number of parameters and provide better local feature extraction ability.

[0092] Pointwise convolution, namely 1×1 convolution, uses a 1×1 convolutional kernel to perform an inter-channel linear combination of the output of channel-wise convolution. Through the cross-channel 1×1 convolution, features from different channels are fused and combined to generate a new feature map. Pointwise convolution increases or decreases the number of channels while keeping the size of the feature map unchanged, and provides a richer feature representation.

[0093] Depthwise separable convolution reduces the number of parameters by decomposing the convolution operation. Compared with standard convolution, the number of parameters can be reduced to 1 / 10 to 1 / 4. And due to the reduction in the number of parameters, the computational complexity of depthwise separable convolution is reduced, making the network more computationally efficient. Although the number of parameters is reduced, depthwise separable convolution can still maintain good model expressiveness by fusing features from different channels through the pointwise convolution layer, improving the accuracy and generalization ability of the model.

[0094] The first fully connected layer and the second fully connected layer have the same structure. In Figure 5 a fully connected layer can learn and extract global features from the input data, which are then used for tasks such as classification and regression. In a convolutional neural network, the fully connected layer is located after the convolutional layer and the pooling layer, and is used to integrate the local features extracted by the convolutional layer to form a global feature representation. In a classification task, the fully connected layer can act as a classifier, mapping the learned features to the final class labels. The number of neurons in the last fully connected layer is equal to the number of classes, and the output value of each neuron represents the probability that the input data belongs to the corresponding class. The fully connected layer can also change the dimension of the data by adjusting the number of neurons to adapt to different network structures and task requirements.

[0095] The star operation layer is used to perform the star operation (Star Operation), which fuses features from different subspaces through element-wise multiplication. In a single layer of a neural network, the star operation represents fusing the features of two linear transformations through element-wise multiplication, and is written as:

[0096] ;

[0097] where and are two linear transformation matrices, T is the transpose, and X is the input feature map.

[0098] Using the star operation in a d-dimensional space, the following implicit dimensional feature space can be obtained:

[0099] ;

[0100] Thus, while significantly amplifying the feature dimension, no additional computational overhead is generated within a single layer.

[0101] S300: Inputting the first image to be classified into the first module to output a first intermediate image.

[0102] For a first image to be classified, the first module is first inputted. In some embodiments, the first image to be classified is inputted into a convolution layer to output a third intermediate image. The first module includes a convolution layer and a pooling layer. The third intermediate image is inputted into the pooling layer to output a first intermediate image. In other words, the output of the first module is the first intermediate image.

[0103] The convolution layer performs a sliding window operation on the first image to be classified through the convolution kernel to extract local features. For OCT images, the convolution layer can extract features such as retinal layers and vascular structures.

[0104] The pooling layer is used to reduce the dimension of the feature map, i.e., height and width, while retaining important information. For OCT images, the pooling layer further refines the features and removes redundant information. The pooling layer reduces the size of the feature map to make the output first intermediate image more compact and convenient for subsequent processing. While reducing the size, it also retains the key information in the image and removes unimportant details, allowing the classification network to focus on features that are useful for classification or classification tasks.

[0105] For the first intermediate image output by the first module, the first intermediate image is input to the first branch block, a second intermediate image is output, and the second intermediate image is used as input of the second branch block.

[0106] Based on the above, we can see the structure of the first branch block, where, through different convolutional layers, the classification network can further refine the features in the first intermediate image and extract higher-level abstract features. The residual connection allows the classification network to learn residual mapping instead of direct input-to-output mapping, which helps to alleviate the gradient vanishing problem.

[0107] By further extracting and refining features through the first branch block and generating a second intermediate image, the OCT image is transformed into a series of more abstract but more representative features, which are used in subsequent classification or categorization tasks.

[0108] S400: Input the second intermediate image to the second branch block to output a first feature map.

[0109] like Figure 5As shown, in some embodiments, the second intermediate image is input into the first separable convolutional layer to output a second feature map; the second feature map is respectively input into two of the first fully-connected layers to output a first vector and a second vector; the first vector and the second vector are input into the star operation layer to output a third vector, and the star operation layer is used to perform an element-wise multiplication operation on the first vector and the second vector; the third vector is input into the second fully-connected layer to output a fourth vector; the fourth vector is input into the second separable convolutional layer to output a third feature map, which is the output of the second separable convolutional layer, and an addition operation is performed on the second intermediate image and the third feature map to output a first feature map.

[0110] First, the second intermediate image is a feature map. The separable convolutional layer can perform spatial convolution, i.e., depth convolution, on each channel of the feature map, and then perform 1×1 convolution to combine channel information, which can reduce the number of parameters and computational complexity while maintaining the feature extraction ability. After being processed by the first separable convolutional layer, the details of the retinal structure, such as the vascular network, retinal layers, etc., will be further highlighted.

[0111] The output second feature map is respectively input into two parallel first fully-connected layers. The fully-connected layer is used to flatten the feature map into a one-dimensional vector and extract global features through the transformation of the weight matrix. The two fully-connected layers are used to extract features from different angles or levels for more complex operations in the future.

[0112] Performing an element-wise multiplication operation on the first vector and the second vector in the star operation layer enables the classification network to learn the element-level interaction information between the two vectors to capture more complex feature combinations. In some embodiments, a first linear transformation matrix and a second linear transformation matrix are obtained; based on the number of rows and columns of the first linear transformation matrix, a first transpose matrix is generated, and based on the number of rows and columns of the second linear transformation matrix, a second transpose matrix is generated; the first vector is input into the first transpose matrix to generate a first linear result, and the second vector is input into the second transpose matrix to generate a second linear result; an element-wise multiplication is performed on the first linear result and the second linear result to output a third vector. Optimization is performed for star point noise to make the denoising effect more significant.

[0113] The third vector is then input into the second fully-connected layer to further extract or transform features. The fourth vector is reshaped and input into the second separable convolutional layer to convert the feature map output by the fully-connected layer back to the spatial dimension for further convolutional operations or fusion with other feature maps. Through the feature transformation of multiple fully-connected layers, the second branch block can better extract and process the detailed information in the image, resulting in a high-quality and delicate processed image.

[0114] After being processed by the second separable convolutional layer, the output third feature map is the original OCT image, which is the representation of the first image to be classified after a series of transformations and feature extractions, containing key information for subsequent classification, sorting, or other tasks.

[0115] By fusing the features of the second intermediate image and the third feature map, the classification network can utilize both low-level details and high-level abstract features while maintaining the low-level details. The first feature map contains both the detailed information and high-level semantic information of the image, which helps the classification network capture more important information in the OCT image more accurately, such as blood vessel structures, retinal layers, etc.

[0116] It can be understood that for feature fusion, the dimensions of the two need to be kept consistent. If they are inconsistent, based on the third feature map, the dimension of the second intermediate image is transformed, and then the addition operation is performed.

[0117] For the second branch block, there are also other implementation methods, such as Figure 6 As shown, the second feature map is input into the first fully-connected layer to output a first vector; the first vector and the second feature map are input into the star operation layer to output a third vector, and the star operation layer is used to perform an element-wise multiplication operation on the first vector and the second feature map; the third vector is input into the second fully-connected layer to output a fourth vector; the fourth vector is input into the second separable convolutional layer to output a third feature map, which is the output of the second separable convolutional layer, and an addition operation is performed on the second intermediate image and the third feature map to output a first feature map.

[0118] Such as Figure 7 As shown, in some embodiments, the second feature map is input into the first fully-connected layer to output a first vector; two first vectors are respectively input into the star operation layer to output a third vector, and the star operation layer is used to perform an element-wise multiplication operation on the two first vectors; the third vector is input into the second fully-connected layer to output a fourth vector; the fourth vector is input into the second separable convolutional layer to output a third feature map, which is the output of the second separable convolutional layer, and an addition operation is performed on the second intermediate image and the third feature map to output a first feature map.

[0119] Figure 6 、 Figure 7 The structure in Figure 7 can make the key features in the second feature map more prominent by changing the connections and outputs of the first fully connected layer and the inputs of the star operation layer, which helps to improve the classification performance of the network. It helps to increase the attention of the classification network to key features, thereby improving the classification performance, and can also achieve the same effect as the first structure in the second branch block.

[0120] S500: Based on the first feature map, output the first classification result through the classification network.

[0121] In some embodiments, input the first feature map into the third branch block to output the first processed image. Processing it again through the second branch block can further enhance the model's ability to understand the features of OCT images, help capture deeper retinal structure information, and improve the model's performance through non-linear transformation and feature fusion mechanisms.

[0122] Then input the first processed image into the third module to output the second processed image, input the second processed image into the fourth module to output the third processed image, and input the third processed image into the fifth module to output the fourth processed image. That is to say, the first processed image is the output of the second module, use the output of the second module as the input of the third module, use the output of the third module as the input of the fourth module, and use the output of the fourth module as the input of the fifth module.

[0123] The features of the fourth processed image have been highly abstracted, containing higher-level retinal structure features; and after multiple downsamplings, the size of the feature map is relatively small, but the number of channels increases, which helps to capture more complex patterns. After multiple feature extractions and refinements, the detailed information is removed, and only the important features are retained, which are important for distinguishing different OCT image categories.

[0124] Moreover, the fourth processed image undergoes multiple downsamplings and feature extractions, reducing the computational complexity, which helps to improve the running speed and efficiency of the model. Through operations such as the star operation layer, the features are effectively fused, emphasizing the important features, which helps the model to classify more accurately. After multiple non-linear transformations, it can learn more complex feature representations and capture the subtle differences in OCT images. After multiple processes, it is more robust to noise and quality changes in the image, which helps to improve the accuracy of classification.

[0125] To highlight the important regions, in some embodiments, an attention mechanism is also included. Through multiple processes and fusions, the fourth processed image can highlight the important regions through the attention mechanism, which is important for identifying specific retinal structures.

[0126] The fourth processed image contains highly abstract features and is suitable for the four-classification task of OCT images. The fourth processed image can be used as the final feature representation and input into a classifier for classification.

[0127] To further improve the classification accuracy of the classification network, in some embodiments, the classification network further includes a first auxiliary module, a second auxiliary module, and a third auxiliary module. The first auxiliary module, the second auxiliary module, and the third auxiliary module are used to fuse the inputs to generate a new feature representation.

[0128] As Figure 8 shown, in some embodiments, the first image to be classified, the first processed image, and the second processed image are input into the first auxiliary module to perform fused features through the first auxiliary module to generate a first fused image; the first fused image, the first processed image, the second processed image, and the third processed image are input into the second auxiliary module to perform fused features through the second auxiliary module to generate a second fused image; the second fused image, the first processed image, the second processed image, the third processed image, and the fourth processed image are input into the third auxiliary module to perform fused features through the third auxiliary module to generate a third fused image; the third fused image is fused with the fourth processed image to generate a first classification result.

[0129] The auxiliary module is used to enhance the feature extraction ability and performance of the classification network. In the first auxiliary module, the inputs of the classification network, the outputs of the first module and the second module, namely the first image to be classified, the first processed image, and the second processed image, are simultaneously passed to the first auxiliary module. The first auxiliary module is used to fuse these inputs to generate a new feature representation. By combining the features of the input data, the first module, and the second module, the first auxiliary module can provide a richer feature representation to enhance the perception ability and processing ability of the classification network, ensuring that the classification network can comprehensively utilize the initial input and the features extracted in the early stage, and improving the diversity and information content of the feature representation.

[0130] In the second auxiliary module, the outputs of the first module, the second module, the third module, and the first auxiliary module, namely the first fused image, the first processed image, the second processed image, and the third processed image, are simultaneously passed to the second auxiliary module. The second auxiliary module further fuses these data to extract deeper features, enabling the classification network to capture more details and global information, which helps to improve the feature extraction effect. By combining features at more levels, the second auxiliary module can perform feature fusion at different scales and different abstraction levels, thereby enhancing the classification ability of the classification network for complex patterns and features. This multi-level feature fusion strategy significantly improves the expression ability of the classification network.

[0131] In the third auxiliary module, the outputs of the first module, the second module, the third module, the fourth module, and the second auxiliary module, i.e., the outputs of the second fusion image, the first processed image, the second processed image, the third processed image, and the fourth processed image, are simultaneously transmitted to the third auxiliary module. The third auxiliary module fuses this data to further refine and enhance the features.

[0132] The output of the third auxiliary module and the output of the fourth module are transmitted to the output layer. By fusing the outputs of multiple modules with the features of the branch module, the classification network can have stronger expressiveness and accuracy when dealing with complex tasks. This multi-path fusion strategy can not only ensure that the features extracted by each module are fully utilized, but also improve the quality and accuracy of the final output through layer-by-layer fusion and refinement. Moreover, by introducing the auxiliary module, the overall performance of the network is enhanced.

[0133] The first classification result is a vector, where each element corresponds to the predicted probability of a category, and the category corresponding to the highest probability represents the result of the classification by the classification network.

[0134] It can be understood that after the classification network is trained, an inference stage is required, which is a process of using the trained classification network to predict the second image to be classified.

[0135] Although the classification network includes a backbone network and an auxiliary network, where the backbone network is the first module, the second module, the third module, the fourth module, and the fifth module, and the auxiliary network includes the first auxiliary module, the second auxiliary module, and the third auxiliary module. In the training stage, the backbone network and the auxiliary network are jointly trained. However, to improve the inference efficiency, in the inference stage, the auxiliary network will be disabled, and only the backbone network is used for inference, thus simplifying the calculation and ensuring the efficiency and speed of inference. Through the strategy of separating the training and inference stages, not only the learning ability during the training process is improved, but also the efficiency and practicality of the inference stage are guaranteed.

[0136] Through the classification accuracy test, the accuracy of the OCT image four-classification task performed by Resnet-50 is 91.45%, the accuracy of the OCT image four-classification task performed by Resnet-50 and the auxiliary network is 92.36%, the accuracy of the OCT image four-classification task performed by the backbone network provided by the embodiments of the present application is 94.61%, and the accuracy of the OCT image four-classification task performed by the backbone network and the auxiliary network provided by the embodiments of the present application is 95.00%.

[0137] Based on the above-mentioned fundus image classification method, as Figure 9 shown, some embodiments of the present application further provide a fundus image classification system, including:

[0138] An acquisition unit 100 is configured to acquire a first image to be classified, where the first image to be classified is an OCT image;

[0139] A construction unit 200 is configured to construct a classification network, where the classification network at least includes a first module and a second module, and the second module at least includes a first branch block and a second branch block;

[0140] The construction unit 200 is further configured to input a second intermediate image into the second branch block to output a first feature map, where the second intermediate image is the output of the first branch block, and the second branch block includes a first separable convolutional layer, a first fully connected layer, a star operation layer, a second fully connected layer, and a second separable convolutional layer, and the first feature map is a feature map obtained by fusing the first processed image and the output of the second separable convolutional layer;

[0141] And, based on the first feature map, output a first classification result through the classification network.

[0142] The present application provides a fundus image classification method and system. The method acquires a first image to be classified and then constructs a classification network. Wherein, the first image to be classified is an OCT image; the classification network at least includes a first module and a second module, and the second module at least includes a first branch block and a second branch block; input the first image to be classified into the first module to output a first intermediate image; then input the second intermediate image into the second branch block to output a first feature map, where the second intermediate image is the image obtained by inputting the first intermediate image into the first branch block and output by the first branch block, and the second branch block includes a first separable convolutional layer, a first fully connected layer, a star operation layer, a second fully connected layer, and a second separable convolutional layer, and the first feature map is a feature map obtained by performing fusion on the first intermediate image and the output of the second separable convolutional layer; based on the first feature map, output a first classification result through the classification network. By changing the second module, the second module optimizes the noise and can extract and process the detailed information of the image to be classified, so as to improve the classification accuracy of the fundus image.

[0143] For the similarity parts between the embodiments provided in the present application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of the present application, and do not constitute a limitation on the protection scope of the present application. For those skilled in the art, any other implementation manner extended based on the solution of the present application without creative efforts belongs to the protection scope of the present application.

Claims

1. A fundus image classification method, characterized in that: include: Acquire a first image to be classified, where the first image to be classified is an OCT image; Constructing a classification network, wherein the classification network includes at least a first module and a second module, wherein the second module includes at least a first branch block and a second branch block; Inputting the first image to be classified into the first module to output a first intermediate image; Inputting a second intermediate image into the second branch block to output a first feature map, wherein the second intermediate image is an image output by the first branch block after the first intermediate image is input into the first branch block, the second branch block comprises a first separable convolutional layer, a first fully connected layer, a star operation layer, a second fully connected layer, and a second separable convolutional layer, and the first feature map is a feature map obtained by fusing the first intermediate image and the output of the second separable convolutional layer; Based on the first feature graph, outputting a first classification result through the classification network, wherein, based on the first feature graph, outputting a first classification result through the classification network includes: Inputting the first image to be classified, the first processed image and the second processed image into a first auxiliary module, so as to generate a first fused image by performing fusion features through the first auxiliary module, wherein the first processed image is the output of the second module, and the classification network further comprises a first auxiliary module, a second auxiliary module and a third auxiliary module; inputting the first fused image, the first processed image, the second processed image and the third processed image into the second auxiliary module so as to perform fusion features through the second auxiliary module to generate a second fused image; inputting the second fused image, the first processed image, the second processed image, the third processed image and the fourth processed image into the third auxiliary module so as to perform fusion features through the third auxiliary module to generate a third fused image; The third fused image is fused with the fourth processed image to generate a first classification result.

2. The fundus image classification method according to claim 1, characterized in that: The method further comprises: Inputting the second intermediate image into the first separable convolutional layer to output a second feature map; Inputting the second feature map into two of the first fully connected layers respectively to output a first vector and a second vector; Input the first vector and the second vector into the star operation layer to output a third vector, wherein the star operation layer is used to perform an element-by-element multiplication operation on the first vector and the second vector; Inputting the third vector into the second fully connected layer to output a fourth vector; The fourth vector is input to the second separable convolutional layer to output a third feature map, where the third feature map is the output of the second separable convolutional layer.

3. The fundus image classification method according to claim 2, characterized in that: The step of inputting the second intermediate image to the second branch block to output a first feature map includes: transforming the dimension of the second intermediate image according to the dimension of the third feature map so that the dimensions of the third feature map and the second intermediate image are consistent; An addition operation is performed on the transformed second intermediate image and the third feature map to output a first feature map.

4. The fundus image classification method according to claim 2, characterized in that: The step of inputting the first vector and the second vector into the star operation layer to output a third vector comprises: Obtain a first linear transformation matrix and a second linear transformation matrix; Generate a first transposed matrix based on the number of rows and columns of the first linear transformation matrix, and generate a second transposed matrix based on the number of rows and columns of the second linear transformation matrix; Inputting the first vector into the first transposed matrix to generate a first linear result, and inputting the second vector into the second transposed matrix to generate a second linear result; An element-by-element multiplication is performed on the first linear result and the second linear result to output a third vector.

5. The fundus image classification method according to claim 2, characterized in that: The second module further includes a third branch block, and the structure of the third branch block is the same as that of the second branch block; The first feature map is input to the third branch block to output a first processed image.

6. The fundus image classification method according to claim 5, characterized in that: The classification network also includes a third module, a fourth module and a fifth module; The method further comprises: Inputting the first processed image to the third module to output a second processed image, wherein the third module includes at least three of the second branch blocks; Inputting the second processed image to the fourth module to output a third processed image, wherein the third module includes at least four of the second branch blocks; The third processed image is input to the fifth module to output a fourth processed image, and the third module includes at least five of the second branch blocks.

7. The fundus image classification method according to claim 5, characterized in that: After outputting the first classification result through the classification network based on the first feature graph, the method further includes: Acquire a second image to be classified; The second image to be classified is input into the classification network to perform inference calculation through the classification network and output a second classification result. The classification network also includes a third module, a fourth module and a fifth module.

8. The fundus image classification method according to claim 1, characterized in that: The step of inputting the second intermediate image to the second branch block to output the first feature map comprises: Inputting the first image to be classified into a convolutional layer to output a third intermediate image, wherein the first module includes a convolutional layer and a pooling layer; Inputting the third intermediate image to the pooling layer to output the first intermediate image; The first intermediate image is input to the first branch block to output a second intermediate image.

9. A fundus image classification system, characterized in that: The method for classifying fundus images according to any one of claims 1 to 8 comprises: An acquisition unit, the acquisition unit is used to acquire a first image to be classified, the first image to be classified is an OCT image; A construction unit, the construction unit is used to construct a classification network, the classification network includes at least a first module and a second module, the second module includes at least a first branch block and a second branch block; The method is also used for inputting a second intermediate image into the second branch block to output a first feature map, wherein the second intermediate image is the output of the first branch block, the second branch block includes a first separable convolution layer, a first fully connected layer, a star operation layer, a second fully connected layer, and a second separable convolution layer, and the first feature map is a feature map obtained by fusion of the first processed image and the output of the second separable convolution layer; And, based on the first feature map, outputting a first classification result through the classification network, wherein, based on the first feature map, outputting a first classification result through the classification network includes: Inputting the first image to be classified, the first processed image and the second processed image into a first auxiliary module, so as to generate a first fused image by performing fusion features through the first auxiliary module, wherein the first processed image is the output of the second module, and the classification network further comprises a first auxiliary module, a second auxiliary module and a third auxiliary module; inputting the first fused image, the first processed image, the second processed image and the third processed image into the second auxiliary module so as to perform fusion features through the second auxiliary module to generate a second fused image; inputting the second fused image, the first processed image, the second processed image, the third processed image and the fourth processed image into the third auxiliary module so as to perform fusion features through the third auxiliary module to generate a third fused image; The third fused image is fused with the fourth processed image to generate a first classification result.

Citation Information

Patent Citations

  • Light-weight wind generating set surface defect detection method

    CN118735900A