Image classification method

By designing a multi-module convolutional neural network in the medical image classification task, gradually extracting and enhancing features, problems such as overfitting and high computational complexity in the existing technology are solved, and the accuracy and efficiency of the classification task are improved.

CN119251591BActive Publication Date: 2025-06-20江苏富翰医疗产业发展有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411576128.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-06
Publication Date
2025-06-20
Estimated Expiration
2044-11-06

AI Technical Summary

Technical Problem

In medical imaging classification tasks, the existing technology has problems such as overfitting, high computational complexity, feature redundancy, local feature limitations and high memory usage, resulting in low accuracy and efficiency of classification tasks.

Method used

An image classification method is proposed to gradually increase the complexity of feature extraction and enhancement, including the multi-module convolutional neural network structure, and to gradually extract and enhance features using multiple unit modules to reduce the risk of overfitting and capture more global information.

Benefits of technology

While maintaining the complexity of the classification network, it reduces the risk of overfitting, improves the accuracy and efficiency of classification tasks, effectively processes images and improves the robustness and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119251591B_ABST
    Figure CN119251591B_ABST
Patent Text Reader

Abstract

This application relates to the field of image processing technology. This application provides an image classification method, which includes: obtaining an image to be classified and a classification network, where the classification network includes a first module, a second module, a third module, and a fourth module; then inputting the preprocessed image to be classified into the first module to output a first feature map, where a first unit is used to extract features and a second unit is used to enhance feature representation; inputting the first feature map into the second module to output a second feature map; inputting the second feature map into the third module to output a third feature map; and inputting the third feature map into the fourth module to output a classification result. By increasing the complexity of feature extraction and enhancement, this method can reduce the risk of overfitting while maintaining the complexity of the classification network. By adding deeper feature extraction, more global information can be captured, thereby reducing the limitations of local features, and the accuracy and efficiency of the classification task can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly to an image classification method. Background Art

[0002] In medical image classification tasks, Convolutional Neural Network (CNN) is applied due to its feature extraction ability. By combining different types of CNN architectures, the accuracy and generalization ability of the model can be further improved.

[0003] Different types of CNN architectures can be used to perform different tasks. For example, combine Faster Net, CSP (Channel-Spatial Pyramid), and PCBS (Patch-based Convolutional Boosting System). Faster Net provides efficient feature extraction ability, CSP strengthens multi-scale feature extraction, and PCBS improves the classification performance of local features.

[0004] However, in the combination process, there are problems such as overfitting, high computational complexity, feature redundancy, limitations of local features, and high memory occupancy, resulting in low accuracy and efficiency of the classification task. Summary of the Invention

[0005] This application provides an image classification method to solve the problem of low accuracy and efficiency of the classification task.

[0006] This application provides an image classification method, including:

[0007] Obtain the image to be classified and a classification network, where the classification network includes a first module, a second module, a third module, and a fourth module;

[0008] Input the preprocessed image to be classified into the first module to output a first feature map. The first module includes a first unit and at least two second units. The first unit is used to extract features, and the second unit is used to enhance feature representation;

[0009] Input the first feature map into the second module to output a second feature map. The second module includes a first unit and at least three second units;

[0010] Input the second feature map into the third module to output a third feature map. The third module includes a first unit and at least four second units;

[0011] Input the third feature map into the fourth module to output a classification result. The fourth module includes a first unit and at least three second units.

[0012] In some feasible embodiments, the inputting the preprocessed image to be classified into the first module to output a first feature map includes:

[0013] Input the preprocessed image to be classified into the first unit to output a first intermediate feature map, where the number of channels of the first intermediate feature map is a first multiple of the number of the image to be classified;

[0014] Input the first intermediate feature map into the second unit to output a second intermediate feature map;

[0015] Input the second intermediate feature map into the second unit to output a first feature map.

[0016] In some feasible embodiments, the first unit includes a first-size convolution and a second-size convolution;

[0017] The inputting the preprocessed image to be classified into the first unit to output a first intermediate feature map includes:

[0018] Input the preprocessed image to be classified into two first-size convolutions respectively to output a first convolution map and a second convolution map;

[0019] Input the first convolution map into the second-size convolution to output a third convolution map;

[0020] Input the third convolution map into the first-size convolution to output a fourth convolution map;

[0021] Pass the second convolution map and the fourth convolution map through an activation function to output a first intermediate feature map.

[0022] In some feasible embodiments, the second unit includes a segmentation layer, a first feature fusion block, and a first convolution enhancement block;

[0023] The inputting the first intermediate feature map into the second unit to output a second intermediate feature map includes:

[0024] Input the first intermediate feature map into the segmentation layer to output a first segmentation map and a second segmentation map;

[0025] Input the second segmentation map into the first feature fusion block to output a first fusion map;

[0026] Input the first fusion map into the first convolution enhancement block to output a first enhancement map;

[0027] Perform a splicing operation on the second segmentation map and the first enhanced map to output a second intermediate feature map.

[0028] In some feasible embodiments, the second unit further includes a second feature fusion block and a second convolutional enhancement block;

[0029] The step of inputting the first intermediate feature map into the second unit to output a second intermediate feature map includes:

[0030] Input the first enhanced map into the second feature fusion block to output a second fusion map;

[0031] Input the second fusion map into the second convolutional enhancement block to output a second enhanced map;

[0032] Perform a splicing operation on the second enhanced map, the second segmentation map, and the first enhanced map to output a second intermediate feature map.

[0033] In some feasible embodiments, the first feature fusion block includes a first convolutional layer, a fast block, and a second convolutional layer;

[0034] The step of inputting the second segmentation map into the first feature fusion block to output a first fusion map includes:

[0035] Input the second segmentation map into the two first convolutional layers respectively to output a fifth convolutional map and a sixth convolutional map;

[0036] Input the sixth convolutional map into the fast block to output a first local feature map;

[0037] Perform a splicing operation on the fifth convolutional map and the first local feature map to output a first spliced map;

[0038] Input the first spliced map into the second convolutional layer to output a first fusion map.

[0039] In some feasible embodiments, the first convolutional layer includes a two-dimensional convolutional layer and a first batch normalization layer;

[0040] The step of inputting the second segmentation map into the two first convolutional layers respectively to output a fifth convolutional map and a sixth convolutional map includes:

[0041] Input the second segmentation map into the two-dimensional convolutional layer to output a seventh convolutional map;

[0042] Input the seventh convolutional map into the first batch normalization layer to output a first transformed feature map;

[0043] Input the first transformed feature map through an activation function to output a fifth convolutional map or a sixth convolutional map.

[0044] In some feasible embodiments, the first convolutional enhancement block includes a third convolutional layer and a second batch normalization layer;

[0045] The step of inputting the first fusion map into the first convolutional enhancement block to output a first enhanced map includes:

[0046] Input the first fusion map into the third convolutional layer to output an eighth convolutional map;

[0047] Input the eighth convolutional map into the second batch normalization layer to output a second transformed feature map;

[0048] Input the second transformed feature map through an activation function to output a first enhanced map.

[0049] In some feasible embodiments, the third convolutional layer includes a two-dimensional convolutional layer;

[0050] The step of inputting the first fusion map into the third convolutional layer to output an eighth convolutional map includes:

[0051] Perform a segmentation operation on the first fusion map to generate a first segmentation map and a second segmentation map. The number of channels of the first segmentation map is one-fourth of that of the first fusion map, and the number of channels of the second segmentation map is three-fourths of that of the first fusion map;

[0052] Input the first segmentation map into the two-dimensional convolutional layer to output a ninth convolutional map;

[0053] Perform a splicing operation on the ninth convolutional map and the second segmentation map to output an eighth convolutional map.

[0054] In some feasible embodiments, the classification network further includes: a fourth convolutional layer and a max pooling layer;

[0055] Before inputting the preprocessed image to be classified into the first module to output a first feature map, it includes:

[0056] Input the image to be classified into the fourth convolutional layer to output a fifth convolutional map;

[0057] Input the fifth convolutional map into the max pooling layer to output the preprocessed image to be classified.

[0058] As can be seen from the above technical solutions, the present application provides an image classification method, including: obtaining an image to be classified and a classification network, where the classification network includes a first module, a second module, a third module, and a fourth module; then inputting the preprocessed image to be classified into the first module to output a first feature map, the first module includes a first unit and at least two second units, the first unit is used to extract features, and the second unit is used to enhance feature representation; inputting the first feature map into the second module to output a second feature map, the second module includes a first unit and at least three second units; inputting the second feature map into the third module to output a third feature map, the third module includes a first unit and at least four second units; inputting the third feature map into the fourth module to output a classification result, the fourth module includes a first unit and at least three second units. By gradually increasing the complexity of feature extraction and enhancement, the method can reduce the risk of overfitting while maintaining the complexity of the classification network. By adding deeper feature extraction units, more global information can be captured, thereby reducing the limitations of local features. It can effectively process images and improve the accuracy and efficiency of the classification task. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0060] Figure 1 Schematic flowchart of the image classification method provided by the embodiment of the present application;

[0061] Figure 2 Schematic structural diagram of the classification network provided by the embodiment of the present application;

[0062] Figure 3 Schematic flowchart of the processing of the image to be classified provided by the embodiment of the present application;

[0063] Figure 4 Schematic structural diagram of the first unit provided by the embodiment of the present application;

[0064] Figure 5 Schematic flowchart of the processing of the second unit provided by the embodiment of the present application;

[0065] Figure 6 Schematic structural diagram of the first feature fusion block provided by the embodiment of the present application;

[0066] Figure 7 Schematic structural diagram of the first convolutional layer provided by the embodiment of the present application;

[0067] Figure 8 Schematic diagram of the first convolutional enhancement block provided by the embodiment of the present application;

[0068] Figure 9 Schematic diagram of the third convolutional layer structure provided by the embodiment of the present application. Detailed implementation manners

[0069] The embodiments will be described in detail below, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following embodiments do not represent all implementation manners consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application detailed in the claims.

[0070] Medical images refer to images of the internal structure of the human body obtained through imaging techniques, such as X-ray, CT, MRI, ultrasound, etc. Medical image classification tasks are mainly used for disease diagnosis, lesion localization, condition assessment, assisting surgical planning, pathological analysis, etc.

[0071] In medical image data, diseases or lesions are manifested as local patterns or structures in the images, and meaningful feature representations can be automatically learned from the original pixel values. Through a series of convolutional layers and pooling layers, the CNN can capture local and global features in the image. Different CNN architectures, such as Res Net, Inception, VGG, etc., have different characteristics, and models can be constructed by combining these architectures. For example, Res Net solves the problem of gradient disappearance in deep networks by introducing residual connections, improving the training stability of the model.

[0072] By fusing different CNN architectures, their advantages can be fully utilized to improve the overall performance of the model. For example, the residual connections of Res Net can be combined with the multi-scale characteristics of the Inception module. Another example is to combine FasterNet, CSP, and PCBS.

[0073] Faster Net is a convolutional neural network that optimizes the depth and parameter configuration of convolutional layers. Different from other deep networks, Faster Net reduces the network depth and improves convolutional operations to reduce the computational complexity. This enables Faster Net to achieve fast and accurate feature extraction at a lower computational cost. Through a streamlined network architecture and an efficient calculation method, the speed and performance of image processing are optimized.

[0074] The CSP network enhances the feature extraction ability through a channel spatial pyramid structure. This structure divides the image feature extraction into different scales. By performing feature extraction at multiple spatial levels and combining channel features and spatial features, the model's ability to capture information at different scales is improved. The pyramid structure of CSP can effectively capture the details and local features in the image. And by processing the features at multiple levels, the understanding of the image structure and details is enhanced.

[0075] PCBS adopts a convolutional enhancement system based on image block division. It divides the input image into several small blocks and performs independent convolutional operations on each block, which can capture local features more meticulously. When processing high-resolution images, it can effectively enhance the analysis ability of local details. By enhancing local features, PCBS improves the overall classification performance, making the model perform better in the recognition and classification of fine-grained features. By deeply analyzing the local blocks of the image, the model's ability to process complex images is improved.

[0076] In some embodiments, the input medical image is preprocessed, for example, normalized, cropped, etc., to ensure that the image is suitable for model input. The preprocessed image is input into Faster Net to extract preliminary features. The feature map output by Faster Net is input into the CSP module. The CSP module fuses features through different channel and spatial pyramid structures to capture information at different scales. The feature map output by CSP is input into the PCBS module. The PCBS module enhances the local feature representation by independently processing small blocks of the image and reduces feature redundancy. The feature map enhanced by PCBS is input into the fully connected layer or other classifiers to output the final classification result.

[0077] However, due to the relatively small medical image dataset and the combination of different modules, problems such as overfitting of the model are likely to occur. Using multi-layer convolutional networks and complex feature fusion techniques will increase the computational burden. There may be duplicate features when using multi-scale feature extraction. Only focusing on local features may cause the model to be unable to capture global information. Using large convolutional kernels and deep networks will lead to an increase in memory consumption, etc., resulting in low accuracy and efficiency of the classification task.

[0078] As Figure 1 shown, to solve the above problems, an embodiment of the present application provides an image classification method, including:

[0079] S100: Obtain the image to be classified and the classification network.

[0080] In a medical image classification task, the image to be classified can be medical imaging data, and the imaging data can be from different imaging techniques, including but not limited to X-ray images, Computed Tomography (CT) images, Magnetic Resonance Imaging (MRI) images, ultrasound images, digital pathology images, Optical Coherence Tomography (OCT) images, and Fluorescein Angiography (FA) images.

[0081] Among them, X-ray images are used to show bone structures, lung conditions, etc. Computed tomography images provide cross-sectional images. Magnetic resonance imaging images provide detailed images of soft tissues. Ultrasound images are used to show fetal development, heart conditions, abdominal organs, etc. Digital pathology images are used for microscopic image analysis of tissue samples. Optical coherence tomography images are used to show ophthalmic imaging. Fluorescein angiography images are also used to show eye images, such as diabetic retinopathy.

[0082] Before inputting the image to be classified into the classification network, in some embodiments, preprocessing is performed on the image to be classified, and the preprocessing includes one or more of image normalization, size adjustment, cropping and padding, data augmentation, and channel conversion. The preprocessing enables the classification network to be effectively trained and classified.

[0083] Image normalization is to scale the pixel values of the image to a specific range to improve the convergence speed of the classification network. Size adjustment is to adjust the image to a unified size to meet the input requirements of the classification network. Cropping and padding are to remove irrelevant regions or make the image reach the required size. Data augmentation increases the diversity of the dataset through operations such as rotation, flipping, and scaling to improve the generalization ability of the classification network. For grayscale images, they can be converted into single-channel images; for color images, their original RGB format is maintained.

[0084] In some embodiments, the image to be classified is a color image with a height and width of 224 pixels, that is, the image to be classified is a color image with a resolution of 224×224 pixels.

[0085] After preprocessing, the image to be classified is processed into a color image with a resolution of 224×224 pixels and input into the classification network. Before input, a pre-trained classification network is first obtained. For the training of the classification network, it at least includes steps such as obtaining a dataset, defining a model, training, parameter adjustment, and verification testing. The specific implementation methods are known to those skilled in the art and will not be elaborated here.

[0086] Such asFigure 2 , Figure 3 As shown, the pre-trained classification network is used for medical image classification. The classification network includes a first module, a second module, a third module, and a fourth module. The first module includes a first unit and at least two second units. The second module includes a first unit and at least three second units. The third module includes a first unit and at least four second units. The fourth module includes a first unit and at least three second units. Among them, the second module and the fourth module have the same structure.

[0087] To facilitate the extraction and classification of the image to be classified, in some embodiments, the classification network further includes: a fourth convolutional layer and a max pooling layer; input the image to be classified into the fourth convolutional layer to output a fifth convolutional map; input the fifth convolutional map into the max pooling layer to output the pre-processed image to be classified.

[0088] After pre-processing, the size of the image to be classified changes from (3, 244, 244) to (64, 112, 112) through the convolutional layer, and then changes to (64, 56, 56) through the max pooling layer, that is, the size of the pre-processed image to be classified is (64, 56, 56).

[0089] S200: Input the pre-processed image to be classified into the first module to output a first feature map.

[0090] Since the first module includes a first unit and two second units, in some embodiments, first input the pre-processed image to be classified into the first unit to output a first intermediate feature map. Among them, as Figure 4 shown, the first unit includes multiple convolutional layers (CONV), that is, the first-size convolution and the second-size convolution.

[0091] Taking the first-size convolution as a 1×1 convolution and the second-size convolution as a 3×3 convolution as an example, input the pre-processed image to be classified into two 1×1 CONVs respectively to obtain a first convolutional map and a second convolutional map. Then input the first convolutional map into a 3×3 CONV to obtain a third convolutional map. Then input the third convolutional map into a 1×1 CONV to obtain a fourth convolutional map. Process the second convolutional map and the fourth convolutional map through the RELU function to obtain a first intermediate feature map. After being processed by the first unit, the size changes from (64, 56, 56) to (256, 56, 56), that is, the size of the first intermediate feature map is (256, 56, 56).

[0092] Then input the first intermediate feature map into the second unit to output a second intermediate feature map, as Figure 5As shown, the second unit includes a segmentation layer, a first feature fusion block, and a first convolutional enhancement block. In some embodiments, the first intermediate feature map is input into the segmentation layer to output a first segmentation map and a second segmentation map; the second segmentation map is input into the first feature fusion block to output a first fusion map; the first fusion map is input into the first convolutional enhancement block to output a first enhancement map; a concatenation (concat) operation is performed on the second segmentation map and the first enhancement map to output a second intermediate feature map.

[0093] The segmentation layer divides the input image into a first segmentation map and a second segmentation map. The segmentation layer evenly divides the first intermediate feature map along the second dimension into two blocks, which are the first segmentation map and the second segmentation map. The segmentation layer is used for parallel processing of features to capture information from different aspects. The second segmentation map is input into the first feature fusion block, and the output first fusion map contains features from the second segmentation map; the first fusion map is input into the first convolutional enhancement block, and the output first enhancement map enhances the features in the first fusion map. The second segmentation map and the first enhancement map are concatenated together to form the second intermediate feature map. The concatenation operation combines two different feature representation forms, thereby providing richer information.

[0094] Among them, as Figure 6 shown, the first feature fusion block includes a first convolutional layer, a fast block, and a second convolutional layer; the first feature fusion block is a Faster CSP block constructed by combining Faster Block and CSP, that is, the fast block. Faster Block can accelerate the training and inference processes of the classification network, while the CSP network can optimize the gradient flow and parameter usage efficiency. The FasterCSP block can enhance the feature extraction ability of the network, while reducing the computational burden and improving the efficiency and performance of the classification network. Through the combination of parallel convolution and fast feature extraction, the first feature fusion block can generate diverse feature representations. This diversity helps the network to better generalize in subsequent classification tasks.

[0095] In some embodiments, the second segmentation map is first input into two first convolutional layers respectively to output a fifth convolutional map and a sixth convolutional map. The sixth convolutional map is input into the fast block to output a first local feature map. A concatenation operation is performed on the fifth convolutional map and the first local feature map to output a first concatenated map. The first concatenated map is input into the second convolutional layer to output a first fusion map.

[0096] The Faster CSP block divides the input layer into two parts. Among them, one part undergoes CBS convolution, that is, the first convolutional layer, and the other part first undergoes CBS convolution and then passes through the Faster Block. After splicing the data, CBS convolution is performed to obtain the output of this block structure, and the input and output scales are the same. The first local feature map of the output contains the features from the sixth convolutional map, the first spliced map of the output contains the combination of two different features, and the first fused map of the output contains the features from the first spliced map.

[0097] Among them, the first convolutional layer and the second convolutional layer have the same structure, that is, CBS convolution, as Figure 7 shown. The CBS convolution includes a two-dimensional convolutional layer and the first batch normalization (BN) layer. For the first convolutional layer, in some embodiments, the second split map is input into the two-dimensional convolutional layer, the seventh convolutional map is output, and then the seventh convolutional map is input into the first batch normalization layer to output the first transformed feature map. The first transformed feature map is passed through an activation function to output the fifth convolutional map or the sixth convolutional map.

[0098] For the second convolutional layer, in some embodiments, the first spliced map is input into the two-dimensional convolutional layer, the output of the two-dimensional convolutional layer is input into the first batch normalization layer, and the output of the first batch normalization layer is passed through an activation function to output the first fused map.

[0099] The first fused map is input into the first convolutional enhancement block to output the first enhanced map. Among them, as Figure 8 shown, the first convolutional enhancement block includes a third convolutional layer and a second batch normalization layer. In some embodiments, the first fused map is input into the third convolutional layer to output the eighth convolutional map, the eighth convolutional map is input into the second batch normalization layer to output the second transformed feature map, and the second transformed feature map is passed through an activation function to output the first enhanced map. The normalization layer helps to stabilize the training process, prevent the problem of gradient disappearance or explosion, and helps to accelerate convergence. The activation function introduces a non-linear mapping, enabling the classification network to fit more complex functional relationships.

[0100] Among them, as Figure 9 shown, the third convolutional layer includes a two-dimensional convolutional layer. In some embodiments, the first fused map is subjected to a split operation to generate a first split map and a second split map. The number of channels of the first split map is one-fourth of that of the first fused map, and the number of channels of the second split map is three-fourths of that of the first fused map. The first split map is input into the two-dimensional convolutional layer to output the ninth convolutional map. A splicing operation is performed on the ninth convolutional map and the second split map to output the eighth convolutional map.

[0101] By segmenting and separately processing the first fusion graph, the classification network can better capture different feature levels. The small segmented features, i.e., the first segmented graph, are further enhanced through convolution, while the larger features, i.e., the second segmented graph, remain unchanged or are slightly processed. By concatenating the ninth convolutional graph and the second segmented graph, the classification network can integrate the information of the two parts to obtain a more comprehensive feature representation.

[0102] While maintaining high computational efficiency and strong feature extraction capabilities, it further enhances the model's ability to process incomplete inputs. By combining the high computational efficiency of the Faster Block and the partial convolutional characteristics of the PCBS block, i.e., the first convolutional enhancement block, and introducing the cross-stage partial connection of CSP Net and the ELAN (Enhanced Local Aggregation Network) structure, the PFaster CSPELAN block can optimize feature fusion and gradient flow while ensuring computational efficiency, improving the robustness and adaptability of the model.

[0103] Continue to refer to Figure 5 , in some embodiments, the second unit further includes a second feature fusion block and a second convolutional enhancement block. Among them, the second feature fusion block has the same structure as the first feature fusion block, and the second convolutional enhancement block has the same structure as the first convolutional enhancement block. Repeatedly processing features using the same structure helps alleviate the problem of gradient disappearance in deep networks.

[0104] After outputting the first enhanced graph, input the first enhanced graph into the second feature fusion block to output a second fusion graph; input the second fusion graph into the second convolutional enhancement block to output a second enhanced graph; perform a concatenation operation on the second enhanced graph, the second segmented graph, and the first enhanced graph to output a second intermediate feature graph.

[0105] By feeding the first enhanced graph into the feature fusion block and the convolutional enhancement block again, the features can be further refined and enhanced, increasing the expressive power of the classification network. The concatenation operation combines feature graphs at different stages, i.e., the first enhanced graph, the second enhanced graph, and the second segmented graph, can fuse multi-scale information, and can improve the training stability of the classification network through skip connections and other means. Although the segmentation operation is introduced, only partial feature graphs are convolved, which helps to maintain computational efficiency to a certain extent. At the same time, the concatenation operation also avoids the redundancy that may be brought by completely independent feature processing paths.

[0106] In summary, the first unit is used to extract and enhance key features from the original input and output a preliminarily enhanced feature graph. The second unit further processes these features on this basis, and through repeated feature fusion and enhancement operations, outputs a more refined and comprehensive feature representation.

[0107] The first module performs preliminary feature extraction and enhancement on the input data. Two second units can further fuse and enhance these features to provide more detailed information.

[0108] S300: Input the first feature map into the second module to output a second feature map.

[0109] For the second module, the output of the first module, i.e., the first feature map, is further processed. Since the second module includes a first unit and three second units, where the first unit and the second units of the second module have the same structure as those of the first module, the second module continues to perform feature refinement and enhancement on the basis of the first module.

[0110] Among them, the processing procedures of the first unit and the second units are the same as those of the first unit and the second units in the first module. Through more second units, deeper feature information can be captured.

[0111] Exemplarily, as Figure 3 shown, the first feature map is the output of the first module. Input the first feature map into the first unit of the second module. The first unit adjusts the size of the first feature map (256, 56, 56) to (512, 28, 28), then input the output of the first unit into the first second unit, input the output of the first second unit into the second second unit, input the output of the second second unit into the third second unit, and the output of the third second unit is the second feature map.

[0112] S400: Input the second feature map into the third module to output a third feature map.

[0113] Similarly to the second module, the input of the third module is the output of the second module. The third module further deepens the learning level of features through more second units, improving the expression ability of the classification network. On the basis of the first module and the second module, the third module can provide a more comprehensive and refined feature representation.

[0114] Exemplarily, as Figure 3 shown, the second feature map is the output of the second module. Input the second feature map into the first unit of the third module. The first unit adjusts the size of the second feature map (512, 28, 28) to (1024, 14, 14), then input the output of the first unit into the first second unit, input the output of the first second unit into the second second unit, input the output of the second second unit into the third second unit, input the output of the third second unit into the fourth second unit, and the output of the fourth second unit is the third feature map.

[0115] S500: Input the third feature map into the fourth module to output a classification result.

[0116] Similarly to the second and third modules, the input of the fourth module is the output of the third module. Based on the output of the third module, the fourth module further optimizes the feature representation. The classification network can capture more hierarchical features, from low-level textures to high-level semantic information, and prepares the final output for the results of classification, regression, or other tasks.

[0117] Exemplarily, as Figure 3 shown, the third feature map is the output of the third module. The third feature map is input into the first unit of the fourth module. The first unit adjusts the size of the third feature map from (1024, 14, 14) to (2048, 7, 7), and then the output of the first unit is input into the first second unit, the output of the first second unit is input into the second second unit, the output of the second second unit is input into the third second unit, and the output of the third second unit is the classification result.

[0118] According to the first, second, third, and fourth modules, as the module hierarchy deepens, the size of the feature map gradually decreases while the number of channels gradually increases. That is to say, the classification network becomes more meticulous in processing features and can capture higher-level feature information.

[0119] The number of channels increases from 256 to 512, 1024, and 2048. As the classification network deepens, it can capture more feature types and details. The spatial resolution of the feature map decreases from (56, 56) to (28, 28), then to (14, 14), and finally to (7, 7), indicating that as the classification network deepens, the focus of the feature map shifts from local details to more abstract concepts. The final reduction of the feature map size to (7, 7) helps the classification network perform feature integration and optimization in preparation for the classification task.

[0120] In some embodiments, the classification network further includes a global pooling layer and a fully connected layer. The global pooling layer is connected to the fourth module, and the fully connected layer is connected to the output of the global pooling layer to output the classification result, which represents the probability estimation of the classification network for the input image belonging to a certain category. For example, a vector where each element corresponds to a category and represents the probability corresponding to that category.

[0121] Exemplarily, if the image to be classified is a pathological section image of breast tissue containing a tumor, the classification network may output a probability vector, such as [0.1, 0.9], indicating that the probability of the image belonging to "benign" is 10% and the probability of belonging to "malignant" is 90%.

[0122] In some embodiments, the medical image classification task can also be performed by Resnet50 with an accuracy of 0.9145, while the classification network provided in this application has an accuracy of 0.9382.

[0123] As can be seen from the above technical solutions, the present application provides an image classification method, including: obtaining an image to be classified and a classification network, where the classification network includes a first module, a second module, a third module, and a fourth module; then inputting the preprocessed image to be classified into the first module to output a first feature map, where the first module includes a first unit and at least two second units, the first unit is used to extract features, and the second unit is used to enhance feature representation; inputting the first feature map into the second module to output a second feature map, where the second module includes a first unit and at least three second units; inputting the second feature map into the third module to output a third feature map, where the third module includes a first unit and at least four second units; inputting the third feature map into the fourth module to output a classification result, where the fourth module includes a first unit and at least three second units. By gradually increasing the complexity of feature extraction and enhancement, the method can reduce the risk of overfitting while maintaining the complexity of the classification network. By adding deeper feature extraction units, more global information can be captured, thereby reducing the limitations of local features. It can effectively process images and improve the accuracy and efficiency of the classification task.

[0124] For the similar parts between the embodiments provided in the present application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of the present application and do not constitute a limitation on the protection scope of the present application. For those skilled in the art, any other implementation manners extended based on the solution of the present application without creative efforts belong to the protection scope of the present application.

Claims

1. An image classification method, characterized in that: include: Acquire an image to be classified and a classification network, wherein the classification network includes a first module, a second module, a third module and a fourth module; Inputting the preprocessed image to be classified into the first module to output a first feature map, the first module includes a first unit and at least two second units, the first unit is used to extract features, the second unit is used to enhance feature representation, and the second unit includes a segmentation layer, a first feature fusion block and a first convolution enhancement block; The step of inputting the preprocessed image to be classified into the first module to output a first feature map comprises: inputting the preprocessed image to be classified into the first unit to output a first intermediate feature map, wherein the number of channels of the first intermediate feature map is a multiple of the first number of the image to be classified; inputting the first intermediate feature map into the second unit to output a second intermediate feature map; The step of inputting the first intermediate feature map to the second unit to output a second intermediate feature map comprises: Inputting the first intermediate feature map to the segmentation layer to output a first segmentation map and a second segmentation map; Inputting the second segmentation map to the first feature fusion block to output a first fusion map, the first feature fusion block includes a first convolution layer, a fast block and a second convolution layer, the first convolution layer and the second convolution layer are CBS convolutions, and the fast block is constructed by combining Faster Block and CSP; The step of inputting the second segmentation map into the first feature fusion block to output a first fusion map includes: Inputting the second segmentation map into two of the first convolutional layers respectively to output a fifth convolutional map and a sixth convolutional map; Inputting the sixth convolutional map into the fast block to output a first local feature map; Performing a splicing operation on the fifth convolutional map and the first local feature map to output a first spliced ​​map; Inputting the first spliced ​​image into the second convolutional layer to output a first fused image; Inputting the first fusion map to the first convolution enhancement block to output a first enhanced map; Performing a splicing operation on the second segmentation map and the first enhancement map to output a second intermediate feature map; Inputting the first feature map into the second module to output a second feature map, the second module comprising a first unit and at least three second units; Inputting the second feature map into the third module to output a third feature map, the third module comprising a first unit and at least four second units; The third feature map is input to the fourth module to output a classification result, and the fourth module includes a first unit and at least three second units.

2. The image classification method according to claim 1, characterized in that: The step of inputting the preprocessed image to be classified into the first module to output a first feature map also includes: inputting the second intermediate feature map into the second unit to output a first feature map.

3. The image classification method according to claim 2, characterized in that: The first unit includes a first-size convolution and a second-size convolution; The step of inputting the preprocessed image to be classified into the first unit to output a first intermediate feature map comprises: Inputting the preprocessed image to be classified into two convolutions of the first size respectively to output a first convolution map and a second convolution map; Inputting the first convolution map into the second-size convolution to output a third convolution map; Inputting the third convolution map into the first-size convolution to output a fourth convolution map; The second convolution map and the fourth convolution map are passed through an activation function to output a first intermediate feature map.

4. The image classification method according to claim 1, characterized in that: The second unit also includes a second feature fusion block and a second convolution enhancement block; The step of inputting the first intermediate feature map to the second unit to output a second intermediate feature map comprises: Inputting the first enhanced image into the second feature fusion block to output a second fusion image; Inputting the second fused image into the second convolution enhancement block to output a second enhanced image; A concatenation operation is performed on the second enhanced image, the second segmentation image, and the first enhanced image to output a second intermediate feature map.

5. The image classification method according to claim 1, characterized in that: The first convolutional layer includes a two-dimensional convolutional layer and a first batch of normalization layers; The step of inputting the second segmentation map to two of the first convolutional layers respectively to output a fifth convolutional map and a sixth convolutional map comprises: Inputting the second segmentation map into the two-dimensional convolutional layer to output a seventh convolutional map; Inputting the seventh convolutional map into the first batch of normalization layers to output a first transformed feature map; The first transformed feature map passes through an activation function to output a fifth convolution map or a sixth convolution map.

6. The image classification method according to claim 1, characterized in that: The first convolution enhancement block includes a third convolution layer and a second batch of normalization layers; The step of inputting the first fusion map to the first convolution enhancement block to output a first enhanced map comprises: Inputting the first fusion map to the third convolutional layer to output an eighth convolutional map; Inputting the eighth convolutional map into the second batch normalization layer to output a second transformed feature map; The second transformed feature map is passed through an activation function to output a first enhanced map.

7. The image classification method according to claim 6, characterized in that: The third convolutional layer includes a two-dimensional convolutional layer; The step of inputting the first fusion graph into the third convolutional layer to output an eighth convolutional graph comprises: Performing a segmentation operation on the first fusion image to generate a first segmentation image and a second segmentation image, wherein the number of channels of the first segmentation image is one quarter of that of the first fusion image, and the number of channels of the second segmentation image is three quarters of that of the first fusion image; Inputting the first segmentation map into the two-dimensional convolutional layer to output a ninth convolutional map; A splicing operation is performed on the ninth convolution map and the second segmentation map to output an eighth convolution map.

8. The image classification method according to claim 1, characterized in that: The classification network also includes: a fourth convolutional layer and a maximum pooling layer; The step of inputting the preprocessed image to be classified into the first module to output the first feature map comprises: Inputting the image to be classified into the fourth convolutional layer to output a fifth convolutional map; The fifth convolutional map is input into the maximum pooling layer to output the preprocessed image to be classified.

Citation Information

Patent Citations

  • Sight line estimation method and system

    CN118506430A

  • Line-of-sight estimation method

    CN118762394A