Image classification method and device, electronic equipment and storage medium
By preprocessing the image and using semantic segmentation models for pixel-level classification, the problem of loss of feature information and contextual relationships in the image classification model in the prior art is solved, and higher image classification accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510275319.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-05-30
AI Technical Summary
When processing the image characteristics of images, the existing image classification model loses the feature information in each channel and the context feature relationship between the channels, resulting in insufficient accuracy and robustness of the classification.
By preprocessing the initial image, the region contrast of the region of interest is improved, and pixel-level classification is used to use the pre-trained semantic segmentation model to fully learn the discernible features, and the feature map is restored to the resolution scale of the original input through the feature decoder and convolution module.
The accuracy and robustness of image classification are improved, the interference of non-correlated image features on feature extraction in the region of interest is avoided, and image-level classification results are obtained that are more accurate than direct image classification.
Smart Images

Figure CN120071016A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and particularly relates to an image classification method, apparatus, electronic device, and storage medium. Background Art
[0002] With the emergence of big data and the reduction of computing hardware costs, it provides basic support for the wide application of artificial intelligence technology, computer vision technology, etc. in the field of images. Currently, using image classification models based on image big data and convolutional neural networks to identify certain image categories has achieved better accuracy and efficiency than manual diagnostic discrimination.
[0003] However, the imaging features of images are relatively complex, such as their morphological features, volume features, spatial distribution features, etc. Existing classification models can be divided into two parts: a feature learning part and a feature-based classification part. When using existing classification models for image classification, the convolutional part and the fully connected part finally compress the features learned into a single value by channel, and finally classify based on the compressed features. Although the computational complexity is reduced, on the one hand, the rich feature information contained in the feature maps of each channel is lost, and on the other hand, the contextual feature relationship between channels is lost. Therefore, the accuracy and robustness of classification are ultimately affected. Summary of the Invention
[0004] The present invention provides an image classification method, apparatus, electronic device, and storage medium, which can improve the accuracy and robustness of image classification when performing image classification.
[0005] In a first aspect, an image classification method is provided, including: Performing image preprocessing on an initial image to obtain a target image corresponding to the initial image, where at least one region of interest is included in the target image, and at least one imaging feature of a preset category is included in each region of interest; Inputting the target image and region of interest information into a pre-trained semantic segmentation model to obtain a pixel-level classification result for each region of interest, where the semantic segmentation model includes a feature encoder, a feature decoder, and a convolutional module; Generating an image-level classification result for the target image according to the pixel-level classification result; Inputting the target image and region of interest information into a pre-trained semantic segmentation model to obtain a pixel-level classification result for each region of interest, including: Inputting the target image and region of interest information into the feature encoder to obtain encoded feature images output by each of the multiple first network layers of the feature encoder, where the number of channels configured for each first network layer increases layer by layer; The encoded feature image is decoded by using a feature decoder to obtain decoded feature images output by each of multiple second network layers of the feature decoder, where the number of channels configured for each second network layer decreases layer by layer; Based on the decoded feature image, a convolution module is used to classify the pixel points in each region of interest into a preset category to obtain a pixel-level classification result corresponding to each region of interest.
[0006] In a second aspect, an image classification device is provided, including: A processing module configured to perform image preprocessing on an initial image to obtain a target image corresponding to the initial image, where the target image includes at least one region of interest, and each region of interest includes at least one imaging feature of a preset category; A classification module configured to input the target image and region-of-interest information into a pre-trained semantic segmentation model to obtain a pixel-level classification result for each region of interest, where the semantic segmentation model includes a feature encoder, a feature decoder, and a convolution module; A generation module configured to generate an image-level classification result of the target image according to the pixel-level classification result; The classification module is specifically configured to: Input the target image and region-of-interest information into the feature encoder to obtain encoded feature images output by each of multiple first network layers of the feature encoder, where the number of channels configured for each first network layer increases layer by layer; The encoded feature image is decoded by using a feature decoder to obtain decoded feature images output by each of multiple second network layers of the feature decoder, where the number of channels configured for each second network layer decreases layer by layer; Based on the decoded feature image, a convolution module is used to classify the pixel points in each region of interest into a preset category to obtain a pixel-level classification result corresponding to each region of interest.
[0007] In a third aspect, an electronic device is provided, including: a processor and a memory, where the memory is configured to store a computer program, and the processor is configured to call and run the computer program stored in the memory to execute the method according to the first aspect or its various implementation manners.
[0008] In a fourth aspect, a computer-readable storage medium is provided, configured to store a computer program, where the computer program causes a computer to execute the method according to the first aspect or its various implementation manners.
[0009] Through the technical solution provided by the present invention, by performing image preprocessing on the initial image, while ensuring the integrity of image features, the regional contrast of the region of interest can be improved, and interference from non-related image features to feature extraction within the region of interest can be avoided. When using the semantic segmentation model for pixel-level classification of the target image, the changes in the channel dimension in the feature encoder can be utilized to fully learn discriminative features. At the same time, the learned feature map is restored to the resolution scale of the original input through the feature decoder and the convolutional module, thereby realizing semantic classification of each pixel point in the region of interest. Classifying and recognizing different regions of interest simultaneously can further fully learn the imaging features of the same preset category and the contextual feature relationship between the imaging features of other preset categories. In addition, by performing pixel-level classification on the target image and converting the pixel-level classification result into an image-level classification result, the aggregation of the feature map into a single feature can be avoided, thereby discarding the resolution scale, and then obtaining a more accurate image-level classification result than directly performing image classification.
[0010] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. Other features and advantages of the present invention will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1 Schematic diagram of the model structure of an existing classification model provided by an embodiment of the present application; Figure 2 Schematic diagram of an application scenario provided by an embodiment of the present application; Figure 3 Schematic diagram of the flow of an image classification method provided by an embodiment of the present application; Figure 4 Schematic diagram of the training process of a semantic segmentation model provided by an embodiment of the present application; Figure 5 Schematic diagram of the calculation process of converting the pixel-level classification result of the validation set into an image-level classification result provided by an embodiment of the present application; Figure 6 Schematic diagram of the flow of an image classification method provided by another embodiment of the present application; Figure 7 Schematic diagram of the network structure of a semantic segmentation model provided by an embodiment of the present application; Figure 8 Schematic diagram of the calculation process of a projection excitation layer provided by an embodiment of the present application; Figure 9 Schematic diagram of the calculation process of converting the pixel-level classification result of a test set into an image-level classification result provided by an embodiment of the present application; Figure 10 Schematic diagram of the structure of a region of interest segmentation device provided by an embodiment of the present invention; Figure 11 Schematic diagram of the structure of a region of interest segmentation device provided by another embodiment of the present invention; Figure 12 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0013] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0014] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0015] With the emergence of big data and the reduction of the cost of computing hardware, it provides a basic support for the wide application of artificial intelligence technology, computer vision technology, etc. in the field of images. At present, based on image big data and image classification models based on convolutional neural networks, the identification of certain image categories has achieved better accuracy and efficiency than manual identification and diagnosis.
[0016] However, the current implementation solutions have the following deficiencies: I. Data level 1. Adopt the scheme of unlabeled data. The training and validation set data only have classification labels, and the regions of interest are not labeled. For images with significant differences in imaging signs (including density, transparency, morphology, volume, distribution, etc.), good results can be obtained. However, for images with highly similar imaging signs, the image classification results obtained by the model are almost unusable.
[0017] 2. Adopt the scheme of labeled data. The training and validation set data not only have classification labels, but also the regions of interest are labeled. It can be divided into two cases: (1) Only collect training images with a single category. That is, the collected training images are special, and the model trained in this case is also special. Good results can be obtained for the prediction results of the validation set, but the prediction effect for a general test set will be relatively poor or even unusable. Because general test images often combine multiple regions of interest, and the imaging signs of these regions of interest often have similarities.
[0018] (2) Collect general training images, but only label the main regions of interest. For disease discrimination, for example, in cases of lung infections, multiple infections or lung infection diseases are often combined, and clinically a main diagnosis and multiple secondary diagnoses are often given. Discrimination is often based on the main diagnosis. Since only the regions of interest associated with the main diagnosis are labeled, and the regions of interest associated with the main diagnosis and those associated with multiple secondary diagnoses have great similarities in imaging signs, the learned model may produce relatively high false positive results when performing predictive analysis.
[0019] II. Model level In a specific application scenario, a classification model based on deep learning usually requires the input shape to be . represents the batch size, that is, the number of training images input each time; represents the number of channels of the training image. For training samples with region of interest labels, the number of channels is 2, that is, the training image image and its corresponding region of interest annotation file mask are combined to form a two-channel input; , and respectively represent the depth (i.e., the number of layers or slices), height and width of the training image, which are called the resolution scale or spatial size.
[0020] The architecture of the classification model is as Figure 1As shown in the figure, the classification model can be divided into two parts: the feature learning part (i.e., the convolutional part) and the feature-based classification part (i.e., the fully connected part). The feature learning part consists of multiple convolutional layers (including feature activation layers and pooling layers), which are used to learn the features of the labeled objects (objects of interest) from the input training set. Through multiple convolutional layers and changes in the channel dimension, discriminative features can be fully learned. To reduce the computational complexity and obtain an abstract representation of the object of interest, the local features are pooled through the pooling layer, while the resolution scale of the feature map is reduced. Before feeding the features into the fully connected part for classification, the resolution scale of the feature maps of each channel is changed to 1*1*1, that is, the feature maps of each channel are pooled into a single feature. Assume that the output shape of the convolutional part is (1, 32, 1, 1, 1), that is, the number of output channels is 32.
[0021] The fully connected part first flattens the input (1, 32, 1, 1, 1) of the convolutional part into (1, 32), that is, completely removes the resolution scale (i.e., the spatial dimension size) of the input features. The batch size and the features of each channel form regular tabular data. The fully connected part performs a weighted operation on each feature by channel based on the single features of multiple channels of the input of the convolutional part, and further compresses the output channels according to the number of categories, that is, the number of output channels is equal to the number of categories. Finally, the outputs of each channel are converted into a probability distribution through the softmax operation, and the channel index corresponding to the maximum probability is the predicted category value. Assume a three-class classification (the class values are 0, 1, and 2 respectively), then the output shape of the fully connected part is (1, 3), that is, for each sample, the prediction result will have 3 values, corresponding to the probability of each class respectively. Usually, the position index (starting from 0) with a larger probability is taken as the final predicted category.
[0022] However, the imaging features of images are relatively complex, such as their morphological features, volume features, spatial distribution features, etc. When using the existing classification model for image classification, the convolutional part and the fully connected part finally compress the learned features into a single value by channel, and finally classify based on the compressed features. Although the computational complexity is reduced, on the one hand, the rich feature information contained in the feature maps of each channel is lost, and on the other hand, the contextual feature relationship between channels is lost. Therefore, the accuracy and robustness of classification are ultimately affected.
[0023] In view of the above deficiencies of the classification model, the inventive concept of the present invention is: For the training set and the validation set, all regions of interest of different categories (categories concerned by the classification target) existing in the same sample image are labeled to fully learn the imaging features of the regions of interest of the same category and the contextual feature relationship between the imaging features of the regions of interest of other categories.
[0024] A multi-class semantic segmentation model is adopted to achieve more refined pixel-level classification (in general, in a 2D graphic, each basic unit constituting the graphic is called a pixel; in 3D medical image data, each basic unit constituting the image should be called a voxel, but it is still habitually called a pixel). Image classification classifies the entire input image (i.e., a 2D or 3D image) as a unit. Different from image classification, semantic segmentation is used to segment an image into regions belonging to different semantic classes, and the annotation and prediction of its semantic regions are at the pixel level, that is, semantic segmentation can identify and understand the content of each pixel in the image. The resolution scale of the feature map learned in the feature learning stage (encoder stage) of semantic segmentation is generally lower than the original input resolution scale, but it will not, like in the classification model, aggregate the feature map into a single feature and thus completely discard the resolution scale. At the same time, semantic segmentation will restore the learned feature map to the original input resolution scale through the decoder stage, so as to achieve semantic classification of each pixel point in the original input.
[0025] Based on the pixel-level classification results obtained by semantic segmentation, through post-processing analysis, it is abstracted into image-level classification results, obtaining more accurate image-level classification results than directly performing image classification.
[0026] It should be noted that image classification, as an important application in the field of artificial intelligence, plays a key role in many fields. The technical solutions in this application can be applied to image processing in various fields, such as the fields of autonomous driving, smart home, satellite technology, medical assistance, and educational assistance, etc., and no specific limitations are made here.
[0027] It should be understood that the technical solutions of the present invention can be applied to the following scenarios, but are not limited to: In some implementable ways, Figure 2 This is an application scenario diagram provided by an embodiment of the present invention. As Figure 2 shown, this application scenario may include an electronic device 110 and a network device 120. The electronic device 110 can establish a connection with the network device 120 through a wired network or a wireless network.
[0028] Exemplarily, the electronic device 110 can be a desktop computer, a laptop computer, a tablet computer, etc., but is not limited thereto. The network device 120 can be a terminal device or a server, but is not limited thereto. In an embodiment of the present invention, the electronic device 110 can send a request message to the network device 120, and this request message can be used to request to obtain the image-level classification result of the target image. Further, the electronic device 110 can receive a response message sent by the network device 120, and this response message includes the image-level classification result of the target image.
[0029] In addition,Figure 2 Exemplarily, an electronic device 110 and a network device 120 are given. In fact, other numbers of electronic devices and network devices may be included, and the present invention places no limitation thereon.
[0030] In some other implementable manners, the technical solution of the present invention may also be executed by the above-mentioned electronic device 110, or the technical solution of the present invention may also be executed by the above-mentioned network device 120, and the present invention places no limitation thereon.
[0031] After introducing the application scenarios of the embodiments of the present invention, the technical solution of the present invention will be elaborated in detail below: Figure 3 is a flowchart of an image classification method provided for an embodiment of the present invention. This method may be executed by the electronic device 110 as shown in Figure 2 but is not limited thereto. As shown in Figure 3 , this method may include the following steps: Step 310: Perform image preprocessing on an initial image to obtain a target image corresponding to the initial image.
[0032] Among them, the initial image is an original image sequence to be subjected to image classification processing, and multiple continuously acquired images may be included in the original sequence image. The image preprocessing may include data verification processing and data preprocessing. The data verification processing may include, for example: availability verification, integrity verification, repeatability verification, etc. Through the data verification processing, the correctness and consistency of the image sequence can be ensured. The availability verification is used to confirm whether the image data exists and can be accessed and used, and to check whether the image file is damaged, such as data loss caused by transmission errors, storage medium failures or other reasons; the integrity verification is used to verify whether the image data is complete and whether there is any missing or tampering; the repeatability verification is used to ensure that each frame in the image sequence is unique, so as to avoid repeated calculation or analysis of the same content during processing. The data preprocessing may include, for example, foreground cropping processing and normalization processing. Through the foreground cropping processing, the locking of the range where the features are located, the positioning of specific features and the extraction of information can be accelerated, and the image classification efficiency can be improved. The purpose of performing normalization processing on the initial image is to make the pixel values of each pixel point in the initial image be normalized and distributed between 0 and 1, so as to improve the contrast of the region of interest, that is, to make the region of interest more obvious relative to other regions, which is convenient for subsequent extraction of image features. Specifically, a pixel value range may be set first, and the pixel value range may be set differently according to different regions of interest, and specific limitations are not provided herein. Through the image preprocessing, the high availability of the data can be ensured, and the generalization ability and robustness of the method can be improved.
[0033] After performing image preprocessing on the initial image to obtain the target image, the target image may contain at least one region of interest (ROI), and each ROI contains at least one radiological feature of a preset category. Among them, the region of interest (ROI) usually refers to the part of the target image that is particularly important for image classification decisions. The purpose of determining the ROI is to improve the accuracy and efficiency of image classification. By focusing on the most informative parts of the image, the performance of the model can be optimized, unnecessary computational workload can be reduced, and the interpretability of the classification results can be improved. For example, in object recognition on a lane, the lane area in the image can be set as the ROI; in animal classification, the main body part of the animal in the image can be set as the ROI; in scene classification, the main structures and elements of the scene in the image can be set as the ROI. Radiological features can include morphological features, volume features, and spatial distribution features, etc. Morphological features involve observable attributes such as the external shape, size, color, and texture of an object or phenomenon. For example, in medical images, the shape, size, and surface texture of an organ are all important morphological features; volume features refer to the size of the space occupied by an object or phenomenon; spatial distribution features describe the distribution status and pattern of an object or phenomenon in the geographical space. For example, in ecology, features such as the distribution density, aggregation, and spatial heterogeneity of species; in urban planning, the form and layout of urban buildings can correspond to spatial distribution features.
[0034] Step 320: Input the target image and the ROI information into the pre-trained semantic segmentation model to obtain the pixel-level classification result of each ROI.
[0035] Among them, the semantic segmentation model includes a feature encoder, a feature decoder, and a convolution module; the semantic segmentation model can be any deep learning network model trained for the prediction task of pixel point classification. For example, it can be a Bi-directional Long Short-Term Memory (BiLSTM) model with an attention mechanism, an Att-BiLstm (Attention-Based Bidirectional Long Short-Term Memory) model, and a 3D U-Net model, or it can also be a network model with some network layer modifications based on the above-mentioned network models, and no specific limitation is made here. In the following embodiments of the present disclosure, taking the 3D U-Net model as the basic backbone network and adding a projection excitation module on this basis to obtain the semantic segmentation model as an example to illustrate the technical solutions in the present disclosure, but it does not constitute a specific limitation. By adding a projection excitation module to the model, the learning of the target of interest can be strengthened.
[0036] When pre-training a semantic segmentation model, an iterative training is performed on the semantic segmentation model using a sample image that annotates the region of interest and the classification label of each pixel point in the region of interest until the semantic segmentation model reaches a convergence state, that is, the corresponding dice coefficient is greater than a preset threshold, and it is determined that the pre-training of the semantic segmentation model is completed. When performing iterative training on the semantic segmentation model using the sample image (i.e., training data), data verification processing and data preprocessing can be performed on the sample image first. Among them, the data verification processing can include availability verification, integrity verification, repeatability verification, and consistency verification. The implementation processes of the availability verification, integrity verification, and repeatability verification are the same as those in step 310 of the embodiment and will not be elaborated here. The consistency verification is used to verify whether the annotation data of the sample image is consistent with the actual situation of the sample image. Through the consistency verification, the training accuracy of the semantic segmentation model can be ensured, and it can be avoided from being interfered by mislabeled data. The data preprocessing can include foreground cropping processing, normalization processing, establishing a data dictionary, data loading, annotation file processing, resampling, random cropping, and data augmentation processing, etc. The implementation processes of the foreground cropping processing and normalization processing are the same as those in step 310 of the embodiment and will not be elaborated here. When establishing the data dictionary, for each sample image, a mapping relationship between it and the corresponding annotation training label can be created and encapsulated as a list of dictionaries in the format of a Python language dictionary, which is convenient for verifying the consistency and accuracy of the annotation through the list of dictionaries to ensure that the annotation of each sample image is correct; when performing data loading, for each sample image and its corresponding annotation file, a channel dimension can be added in front of its spatial dimension. Adding the channel dimension can be used to represent whether each pixel point or region belongs to multiple categories; when the sample image is a CT image, resampling is used to resample the CT image and the annotation to unify the anatomical coordinates and voxel spacing so that the semantic segmentation model can learn consistent region-of-interest target features. The anatomical coordinate system is a continuous three-dimensional space composed of three planes (transverse plane, coronal plane, sagittal plane), and the anatomical coordinates are used to describe the position of the standard human body anatomically; random cropping is used to randomly crop one or more fixed-size input blocks (crops) from the corresponding positions of the sample image and the annotation. On the one hand, it is for batch training of the sample image. On the other hand, through this random cropping, the region of interest is located at different positions in space, weakening the sensitivity of the semantic segmentation model to the target position to improve the generalization ability of the training model. In each iteration cycle during the model training process, the spatial positions of the random cropping are almost different; data augmentation is used to perform image augmentation on the sample image to ensure the data volume of the training samples. When performing data augmentation, random affine transformations (rotation, scaling, etc.) can be performed on the input blocks after random cropping, also to improve the generalization ability of the training model.
[0037] After performing the above image preprocessing on the sample image, such as Figure 4As shown, 20% of the samples in the sample images can be further randomly divided into the validation set, and 80% into the training set. The training set is used to continuously optimize and train the semantic segmentation model (Task 1), and the validation set is used to verify the training accuracy of the semantic segmentation model during the training process (Task 2) until the semantic segmentation model is trained. During the execution of Task 2, the pixel-level classification prediction results of the validation set can be converted into image-level classification prediction results to achieve the classification function (Task 3). After the semantic segmentation model reaches a certain accuracy through the validation of the validation set, the semantic segmentation model can be used to perform pixel-level classification of the regions of interest in the test set (i.e., the target images) (Task 4), and the pixel-level classification results of the test set are converted into image-level classification results (Task 5).
[0038] When converting the pixel-level classification prediction results of the validation set into image-level classification prediction results, since each sample image in the validation set includes a corresponding manual annotation file, based on the pixel-level classification prediction results and the corresponding manual annotation file, calculate the dice metric for each region of interest on each sample image, that is, quantitatively evaluate the pixel-level classification prediction results. Table 1 shows two examples of dice metrics: Table 1: Two examples of dice metrics
[0039] Table 1 shows examples of dice metrics for semantic segmentation with three classifications (excluding the background class). "... / masks / ly_010.nii.gz" is the manual annotation, and "... / vals / ly_010.nii.gz" is the corresponding pixel-level classification prediction result. By comparing and calculating these two, the corresponding dice metric is obtained. The dice metric values for each category are listed in the dice metric.
[0040] The algorithm process of converting the pixel-level classification prediction results of the validation set into image-level classification prediction results is as Figure 5As shown, for each region of interest, the corresponding dice metric can be traversed. In the dice metric, the corresponding maximum dice value, i.e., max_dice, can be filtered. When max_dice is not 0, the class value corresponding to max_dice can be used as the classification prediction result of the region of interest. In addition, according to the classification prediction result and the true class of the sample image, the number of samples correctly predicted for each class can be counted, and then the accuracy of the semantic segmentation model can be evaluated to obtain multiple classification metrics. The multiple classification metrics can include the accuracy recall1, recall2, and recall3 of each class, the overall accuracy acc, and the overall average accuracy avg_acc. After the classification metrics show that the semantic segmentation model meets certain accuracy requirements, it can be determined that the training of the semantic segmentation model is completed, and then the network weights obtained by training the semantic segmentation model can be used to perform pixel-level classification of the regions of interest on the test set (i.e., the target image).
[0041] Correspondingly, when pre-training the semantic segmentation model, the steps of the embodiment may include: determining the sample image after image preprocessing, the sample region of interest information corresponding to the sample image, and the preset training label corresponding to the sample image, where the preset training label at least includes the classification label of each pixel point in the region of interest; inputting the sample image, the sample region of interest information, and the preset training label into the semantic segmentation model, and using the feature encoder, the feature decoder, and the convolutional module to perform prediction training on pixel point classification of the semantic segmentation model; where, during the prediction training of pixel point classification, the sample image and the sample region of interest information are used as input features, and the preset training label is used as the training label, and the model parameters in the semantic segmentation model are iteratively updated based on the loss function and the optimizer until the dice coefficient of the semantic segmentation model is greater than the preset threshold, and it is determined that the training of the semantic segmentation model is completed. Wherein, the preset threshold is a value between 0 and 1, and the specific value can be set according to the actual application scenario and will not be specifically limited here.
[0042] For the embodiments of the present disclosure, the step of inputting the target image and the region of interest information into the pre-trained semantic segmentation model in step 320 to obtain the pixel-level classification result of each region of interest may include the following steps: Step 320-1: Input the target image and the region of interest information into the feature encoder to obtain the encoded feature images output by each of the multiple first network layers of the feature encoder, where the number of channels configured for each first network layer increases layer by layer.
[0043] Among them, the feature encoder includes multiple first network layers. For each of the multiple first network layers, different numbers of channels can be configured respectively, and the number of channels increases layer by layer among the multiple network layers. For example, when the feature encoder includes 4 first network layers, the number of channels corresponding to each first network layer can be 16, 32, 64, and 128 in sequence. Each first network layer can gradually fuse deeper features and search for effective regions in the target image to obtain an encoded feature image. Specifically, the more channels configured for the network layer, the smaller the actual image size obtained, the higher the feature depth of the initial encoded features, and it is easier to extract more hidden features (such as gray-scale features, position features, etc.).
[0044] For the embodiments of the present disclosure, the target image can be input into the feature encoder, and the first network layers with gradually increasing numbers of channels in the feature encoder are used to perform encoded feature recognition in sequence. For the first first network layer, the target image can be used as the input feature to obtain an encoded feature image at the feature resolution of this layer; for each first network layer after the first first network layer, the encoded feature image output by the previous first network layer can be used as the input feature to further obtain an encoded feature image at a larger feature resolution of this layer. Repeat the above process until the last first network layer outputs an encoded feature image.
[0045] Step 320-2: Use the feature decoder to perform decoding processing on the encoded feature image to obtain the decoded feature images output by each of the multiple second network layers of the feature decoder, where the number of channels corresponding to each second network layer decreases layer by layer.
[0046] Among them, the feature decoder can also include multiple second network layers. For the multiple second network layers in the feature decoder, different numbers of channels can be configured respectively, and the number of channels decreases layer by layer among the multiple second network layers. For example, when the feature decoder includes 4 second network layers, the number of channels corresponding to each second network layer can be 256, 128, 64, and 32 in sequence.
[0047] For the embodiments of the present disclosure, the feature decoder can fuse the encoded feature image and the decoded feature images output by each second network layer to obtain the decoded feature image at a smaller feature resolution of this layer. Repeat the above process until the last second network layer outputs the decoded feature image. In this way, the resolution is restored layer by layer to the same as that of the input image, which can achieve an end-to-end effect, reduce unnecessary workload, and improve the speed of pixel classification.
[0048] Step 320-3: Use the convolution module to classify the pixel points in each region of interest based on the decoded feature image to obtain the pixel-level classification result corresponding to each region of interest.
[0049] For the embodiments of the present disclosure, the convolutional module can be configured with the same number of channels as the first first network layer, such as 16, so that the output of the semantic segmentation model has the same spatial resolution scale as the input. Specifically, the convolutional module can classify the pixel points in each region of interest into preset categories based on the encoded feature image output by the first first network layer in the feature encoder and the decoded feature image output by the last second network layer in the feature decoder, to obtain the pixel-level classification result corresponding to each region of interest. Specifically, the convolutional module can output the probability values corresponding to each preset classification for each pixel point, and then determine the preset classification with the largest corresponding probability value as the pixel-level classification result of the pixel point.
[0050] Step 330, generate an image-level classification result of the target image according to the pixel-level classification result.
[0051] For the embodiments of the present disclosure, by first performing pixel-level classification on the target image, an image-level classification result of the target image is generated based on the pixel-level classification result. By reducing the classification prediction granularity of the model, the error of image classification can be reduced, and a more accurate image-level classification result than directly performing image classification can be obtained.
[0052] In summary, according to the image classification method provided by the present invention, through image preprocessing of the initial image, while ensuring the integrity of image features, the regional contrast of the region of interest can be improved, and interference of non-related image features on feature extraction within the region of interest can be avoided. When using the semantic segmentation model to perform pixel-level classification of the target image, the change in the channel dimension in the feature encoder can be utilized to fully learn discriminative features. At the same time, the learned feature map is restored to the resolution scale of the original input through the feature decoder and the convolutional module, so as to realize semantic classification of each pixel point in the region of interest. Classifying and recognizing different regions of interest simultaneously can fully learn the imaging features of the same preset category and the contextual feature relationship between the imaging features of other preset categories. In addition, by performing pixel-level classification on the target image and converting the pixel-level classification result into an image-level classification result, the aggregation of the feature map into a single feature can be avoided, thereby discarding the resolution scale, and a more accurate image-level classification result than directly performing image classification can be obtained.
[0053] Based on Figure 3 the embodiments shown, as a refinement and extension of the above embodiments, in order to fully illustrate the specific implementation process of the method in this embodiment, this embodiment provides a specific method as Figure 6 shown. Figure 6 Based on Figure 3 the embodiments shown, the steps of the embodiments are further defined. As shown in 6, the method includes the following steps: Step 610: Perform image preprocessing on the initial image to obtain a target image corresponding to the initial image.
[0054] For the embodiments of the present disclosure, the specific implementation process can refer to the relevant descriptions in step 310 of the embodiments, which will not be elaborated here.
[0055] Step 620: Input the target image and the region of interest information into the feature encoder to obtain the encoded feature images output by each of the multiple first network layers of the feature encoder.
[0056] Among them, the feature encoder includes N first network layers, and the number of channels configured for each first network layer is different, and the number of channels increases layer by layer among the N first network layers. The N first network layers respectively include a first convolutional layer, a first projection excitation layer, and a downsampling layer. In the following embodiment steps of the present disclosure, taking N = 4 as an example, the technical solutions in the present disclosure are described, but it does not constitute a specific limitation.
[0057] Exemplarily, as Figure 7 shown, when N = 4, the feature encoder includes 4 first network layers (i.e., convolution module 1, convolution module 2, convolution module 3, and convolution module 4). Convolution module 1, convolution module 2, convolution module 3, and convolution module 4 can be equivalent to downsampling modules, and the corresponding feature map channel numbers are 16, 32, 64, and 128 respectively. Each first network layer respectively includes two first convolutional layers (Conv + BN + ReLU), a first projection excitation layer (Project&Excite), and a downsampling layer (Max Pool). Each time the target image passes through a downsampling module in the feature encoder, the size of the feature map becomes smaller and the output channel number increases.
[0058] For the embodiments of the present disclosure, the embodiment steps may include: for any one of the N first network layers, using the first convolutional layer configured therein, sequentially perform convolution processing, batch normalization processing, and activation processing on the first input feature of the current first network layer to obtain a first output feature image; use the first projection excitation layer to perform projection calibration processing on the first output feature image to obtain a first calibration feature image corresponding to the first output feature image; use the downsampling layer to perform a max pooling operation on the first calibration feature image to obtain the encoded feature image output by the current first network layer; among them, when the current first network layer is the first first network layer, the first input feature is the target image and the region of interest information; when the current first network layer is any one of the N first network layers except the first first network layer, the first input feature is the target encoded feature image, and the target encoded feature image is the encoded feature image output by the previous first network layer corresponding to the current first network layer.
[0059] For example, as Figure 7 shown, when N takes the value of 4, the target image can be input into the feature encoder to obtain the encoded feature images output by each of the 4 first network layers; among them, for any one of the 4 first network layers, the first input feature is convolved, batch-normalized, and activated using the first convolutional layer configured therein to obtain the first output feature image; then, the first projection excitation layer is used to perform projection calibration processing on the first output feature image to obtain the first calibration feature image corresponding to the first output feature image; the downsampling layer is used to perform max pooling operation on the first calibration feature image to obtain the encoded feature image output by the current first network layer. Specifically, when the current first network layer is the 1st first network layer (convolution module 1), the first input feature is the target image and the region of interest information, and the region of interest information can specifically be the position information of each region of interest in the target image; when the current first network layer is the 2nd first network layer (convolution module 2), the corresponding first input feature is the encoded feature image output by the previous first network layer (convolution module 1) corresponding to the current first network layer (convolution module 2); when the current first network layer is the 3rd first network layer (convolution module 3), the corresponding first input feature is the encoded feature image output by the previous first network layer (convolution module 2) corresponding to the current first network layer (convolution module 3); when the current first network layer is the 4th first network layer (convolution module 4), the corresponding first input feature is the encoded feature image output by the previous first network layer (convolution module 3) corresponding to the current first network layer (convolution module 4).
[0060] In a specific application scenario, since the features of different layers and different channels have different degrees of importance, and the conventional downsampling layer operation cannot distinguish the finest-grained features and the coarsest-grained features spatially, a projection excitation layer can be added to each convolutional module corresponding to the first network layer to re-verify the output features of each convolutional module through projection excitation operations, thereby retaining relatively important image regions.
[0061] The calculation process of the projection excitation layer is as Figure 8 shown. Assume that the output feature image of the mth convolutional module is , and its shape is , represents the number of channels, respectively represent 's depth, height, and width. The projection excitation layer hopes to find a mapping to generate a calibrated weight for re-calibrating the output feature image , and this weight gradually biases towards important regions (i.e., regions of interest) while ignoring uninformative regions.
[0062] Correspondingly, the step of performing projection calibration processing on the first output feature image by using the first projection excitation layer in step 620 to obtain a first calibration feature image corresponding to the first output feature image may include the following steps: Step 620-1: Perform projections in the depth, height, and width directions on each channel image in the first output feature image to obtain a first depth projection feature, a first height projection feature, and a first width projection feature.
[0063] For each channel image in the output feature image perform projections in the depth, height, and width directions respectively to obtain a first depth projection feature , a first height projection feature and a first width projection feature , and the formulas are as follows:
[0064]
[0065]
[0066] wherein, , , respectively represent weight parameters in the depth, height, and width directions, and they are initialized as random values in a distribution with a mean of 0 and a variance of 1.
[0067] Step 620-2: Calculate a spatially feature image of the first output feature image after projection compression based on the first depth projection feature, the first height projection feature, and the first width projection feature.
[0068] Furthermore, the first depth projection feature , the first height projection feature and the first width projection feature can be respectively restored to the original output feature size based on the broadcast mechanism, and then fused to obtain a final spatially feature image , and the formula is as follows:
[0069] Step 620-3: Recombine multiple spatially feature images corresponding to multiple channel images in the first output feature image to obtain a first calibration weight.
[0070] After obtaining the spatially feature compressed by projection through mapping , next, further recombine multiple channels through an activation operation to obtain a first calibration weight :
[0071] and represents the convolution kernel, * represents the convolution operation, and f represents the ReLU non-linear activation function. Using to perform convolution to reduce the number of channels to , while using to perform convolution to restore the number of channels to . Through this reduction and increase of channels, multiple channels are reorganized to emphasize the information channels and suppress redundant channels.
[0072] Step 620-4: Calibrate the first output feature image using the first calibration weight to obtain the first calibrated feature image.
[0073] After calculating the first calibration weight , the first output feature image can be calibrated using the first calibration weight, and the obtained first calibrated feature image is:
[0074] Step 630: Decode the encoded feature image using the feature decoder to obtain the decoded feature images output by each of the multiple second network layers of the feature decoder.
[0075] Among them, the feature decoder includes N second network layers, and the channel numbers configured for each second network layer are different, and the channel numbers decrease layer by layer among the N second network layers. The N second network layers respectively include a second convolutional layer, a second projection excitation layer, and an upsampling layer. In the following embodiment steps of the present disclosure, taking N = 4 as an example to illustrate the technical solution in the present disclosure, but it does not constitute a specific limitation.
[0076] Exemplarily, such as Figure 7As shown, when N takes the value of 4, the feature decoder includes 4 second network layers (i.e., convolution module 5, convolution module 6, convolution module 7, and convolution module 8). Convolution module 5, convolution module 6, convolution module 7, and convolution module 8 can be equivalent to upsampling modules, and their corresponding feature map channel numbers are 256, 128, 64, and 32 respectively. Each second network layer respectively includes two second convolutional layers (Conv+BN+ReLU), a second projection excitation layer (Project&Excite), and an upsampling layer (Upsampling). After passing through each upsampling module, the size of the feature map becomes larger and the number of output channels decreases. The output connection of convolution module 4 and convolution module 5 serves as the input of convolution module 6. The output connection of convolution module 3 and convolution module 6 serves as the input of convolution module 7. The output connection of convolution module 2 and convolution module 7 serves as the input of convolution module 8. The output connection of convolution module 1 and convolution module 8 serves as the input of convolution module 9.
[0077] For the embodiments of the present disclosure, the embodiment steps may include: for any one of the N second network layers, using the second convolutional layer configured therein, sequentially performing convolution processing, batch normalization processing, and activation processing on the second input feature of the current second network layer to obtain a second output feature image; using the second projection excitation layer to perform projection calibration processing on the second output feature image to obtain a second calibration feature image corresponding to the second output feature image; using the upsampling layer to perform upsampling processing on the second calibration feature image to obtain a decoded feature image output by the current second network layer; wherein, when the current second network layer is the first second network layer, the second input feature is the encoded feature image output by the last first network layer among the N first network layers; when the current second network layer is any one of the N second network layers except the first second network layer, the second input feature includes the decoded feature image output by the previous second network layer corresponding to the current second network layer, and the encoded feature image output by the first network layer with the same channel number configured corresponding to the current second network layer.
[0078] For example, as Figure 7As shown, when N takes the value of 4, for any one of the four second network layers, the second input feature can be subjected to image segmentation processing to obtain a decoded feature image corresponding to the configured number of channels of the current second network layer. Among them, when the current second network layer is the first second network layer (convolution module 5), the second input feature includes the encoded feature image output by the last first network layer (convolution module 4) among the N first network layers; when the current second network layer is the second second network layer (convolution module 6), the second input feature includes the decoded feature image output by the previous second network layer (convolution module 5) corresponding to the current second network layer, and the encoded feature image output by the first network layer (convolution module 4) with the same number of channels as the configuration corresponding to the current second network layer (convolution module 6); when the current second network layer is the third second network layer (convolution module 7), the second input feature includes the decoded feature image output by the previous second network layer (convolution module 6) corresponding to the current second network layer, and the encoded feature image output by the first network layer (convolution module 3) with the same number of channels as the configuration corresponding to the current second network layer (convolution module 7); when the current second network layer is the fourth second network layer (convolution module 8), the second input feature includes the encoded feature image output by the previous second network layer (convolution module 7) corresponding to the current second network layer, and the encoded feature image output by the first network layer (convolution module 2) with the same number of channels as the configuration corresponding to the current second network layer (convolution module 8).
[0079] Correspondingly, the process of using the second projection excitation layer to perform projection calibration processing on the second output feature image to obtain the second calibration feature image corresponding to the second output feature image may include the following steps. The specific implementation process can refer to the relevant description in step 620 of the embodiment and will not be elaborated here: Step 630-1: Perform projections on each channel image in the second output feature image in the depth, height, and width directions to obtain a second depth projection feature, a second height projection feature, and a second width projection feature.
[0080] Step 630-2: Calculate the spatially feature image of the second output feature image after projection compression based on the second depth projection feature, the second height projection feature, and the second width projection feature.
[0081] Step 630-3: Recombine the multiple spatially feature images corresponding to the multiple channel images in the second output feature image to obtain a second calibration weight.
[0082] Step 630-4: Use the second calibration weight to perform calibration processing on the second output feature image to obtain the second calibration feature image.
[0083] Step 640: Based on the decoded feature image, use the convolution module to classify the pixel points in each region of interest into a preset category, and obtain the pixel-level classification result corresponding to each region of interest.
[0084] Among them, the convolution module includes a third convolutional layer (Conv + BN + ReLU), a third projection excitation layer (Project&Excite), and a fourth convolutional layer (Conv + ReLU). As Figure 7 shown, the convolution module can be convolution module 9, and the outputs of convolution module 1 and convolution module 8 are connected as the input of convolution module 9.
[0085] For the embodiments of the present disclosure, the embodiment steps may include: performing convolution processing on the third input feature in the third convolutional layer to obtain a third output feature image, where the third input feature includes the encoded feature image output by the first first network layer in the feature encoder and the decoded feature image output by the last second network layer in the feature decoder; using the third projection excitation layer to perform projection calibration processing on the third output feature image to obtain a third calibrated feature image corresponding to the third output feature image; in the fourth convolutional layer, based on the third calibrated feature image, classify the pixel points in each region of interest into a preset category respectively, and obtain the pixel-level classification result corresponding to each region of interest.
[0086] Correspondingly, the step of using the third projection excitation layer to perform projection calibration processing on the third output feature image to obtain a third calibrated feature image corresponding to the third output feature image in step 640 may include the following steps, and the specific implementation process may refer to the relevant description in embodiment step 620 and will not be elaborated here: Step 640-1: Perform projections in the depth, height, and width directions on each channel image in the third output feature image to obtain a third depth projection feature, a third height projection feature, and a third width projection feature.
[0087] Step 640-2: Based on the third depth projection feature, the third height projection feature, and the third width projection feature, calculate the spatially compressed feature image of the third output feature image after projection.
[0088] Step 640-3: Recombine the multiple spatially compressed feature images corresponding to the multiple channel images in the third output feature image to obtain a third calibration weight.
[0089] Step 640-4: Use the third calibration weight to perform calibration processing on the third output feature image to obtain a third calibrated feature image.
[0090] Step 650: Generate an image-level classification result of the target image according to the pixel-level classification result.
[0091] AsFigure 9 As shown, when converting the pixel-level classification result of the test set (target image) into an image-level classification result, the number of pixel points for each preset category within each region of interest can be counted, that is, the maximum value n_max. When n_max is not 0, the category value corresponding to n_max can be used as the classification prediction result for the corresponding region of interest. Alternatively, after counting the number of pixel points for each preset category within each region of interest, the pixel points can also be sorted according to the number of pixel points to determine the main classification result and the secondary classification result within the corresponding region of interest. The method for counting the number of pixel points corresponding to each preset category is as follows: Use logical operations to set the voxels equal to a certain category value to True, and the remaining voxels to False, then convert the boolean values to floating-point values, convert True to 1.0, False to 0.0, and finally count the number of values equal to 1.0 or calculate their sum, that is, obtain the number of voxels of this category.
[0092] In addition, according to the classification prediction result and the true category of the target image, the number of samples with correct predictions for each category can be counted, and then the accuracy of the semantic segmentation model can be evaluated to obtain multiple classification metrics. The multiple classification metrics can include the accuracy recall1, recall2, and recall3 for each category, the overall accuracy acc, and the overall average accuracy avg_acc.
[0093] Correspondingly, for the embodiments of the present disclosure, the embodiment steps may include: counting the number of pixel points corresponding to each preset category according to the pixel-level classification result of each region of interest; determining the image-level classification result corresponding to each region of interest in the target image based on the number of pixel points corresponding to each preset category.
[0094] In summary, the technical solution in this application can, while ensuring the integrity of image features, improve the regional contrast of the region of interest and avoid interference from non-related image features on feature extraction within the region of interest by performing image preprocessing on the initial image. When using the semantic segmentation model for pixel-level classification of the target image, the changes in the channel dimension in the feature encoder can be utilized to fully learn discriminative features, and at the same time, the learned feature map can be restored to the resolution scale of the original input through the feature decoder and the convolutional module, thereby realizing semantic classification of each pixel point in the region of interest. Classifying and recognizing different regions of interest simultaneously can further fully learn the imaging features of the same preset category and the context feature relationship between the imaging features of other preset categories. In addition, by performing pixel-level classification on the target image and converting the pixel-level classification result into an image-level classification result, it is possible to avoid aggregating the feature map into a single feature and thus discarding the resolution scale, and further obtain a more accurate image-level classification result than directly performing image classification.
[0095] Based on the above Figure 3 , Figure 6 specific description of the provided image classification method, as Figure 10 shown, Figure 10 is a block diagram of an image classification device shown according to an exemplary embodiment. As Figure 10 shown, the device includes: A processing module 1010, which can be used to perform image preprocessing on the initial image to obtain a target image corresponding to the initial image, where the target image contains at least one region of interest, and each region of interest contains at least one radiological feature of a preset category; A classification module 1020, which can be used to input the target image and region of interest information into a pre-trained semantic segmentation model to obtain a pixel-level classification result for each region of interest. The semantic segmentation model includes a feature encoder, a feature decoder, and a convolution module; A generation module 1030, which can be used to generate an image-level classification result of the target image based on the pixel-level classification result; The classification module 1020 is specifically used for: Input the target image and region of interest information into the feature encoder to obtain the encoded feature images output by each of the multiple first network layers of the feature encoder, where the number of channels configured for each first network layer increases layer by layer; Use the feature decoder to perform decoding processing on the encoded feature images to obtain the decoded feature images output by each of the multiple second network layers of the feature decoder, where the number of channels configured for each second network layer decreases layer by layer; Use the convolution module to classify the pixel points in each region of interest based on the decoded feature images to obtain the pixel-level classification result corresponding to each region of interest.
[0096] In some embodiments of the present application, the feature encoder includes N first network layers, the number of channels configured for each first network layer is different, and the number of channels increases layer by layer in the N first network layers. The N first network layers respectively include a first convolutional layer, a first projection excitation layer, and a downsampling layer; When inputting the target image and the region of interest information into the feature encoder to obtain the encoded feature images output by each of the multiple first network layers of the feature encoder, the classification module 1020 can be used for any one of the N first network layers. By using the first convolutional layer configured therein, the first input feature of the current first network layer is sequentially subjected to convolution processing, batch normalization processing, and activation processing to obtain a first output feature image; the first output feature image is subjected to projection calibration processing by using the first projection excitation layer to obtain a first calibrated feature image corresponding to the first output feature image; the first calibrated feature image is subjected to a max-pooling operation by using the downsampling layer to obtain the encoded feature image output by the current first network layer; wherein, when the current first network layer is the first first network layer, the first input feature is the target image and the region of interest information; when the current first network layer is any one of the N first network layers except the first first network layer, the first input feature is the target encoded feature image, and the target encoded feature image is the encoded feature image output by the previous first network layer corresponding to the current first network layer.
[0097] In some embodiments of the present application, when the first output feature image is subjected to projection calibration processing by using the first projection excitation layer to obtain a first calibrated feature image corresponding to the first output feature image, the classification module 1020 can be used to perform projections on each channel image in the first output feature image in the depth, height, and width directions respectively to obtain a first depth projection feature, a first height projection feature, and a first width projection feature; based on the first depth projection feature, the first height projection feature, and the first width projection feature, calculate a spatially compressed feature image of the first output feature image after projection; recombine the multiple spatially compressed feature images corresponding to the multiple channel images in the first output feature image to obtain a first calibration weight; use the first calibration weight to perform calibration processing on the first output feature image to obtain the first calibrated feature image.
[0098] In some embodiments of the present application, the feature decoder includes N second network layers, the number of channels configured for each second network layer is different, and the number of channels decreases layer by layer among the N second network layers. Each of the N second network layers includes a second convolutional layer, a second projection excitation layer, and an upsampling layer; When decoding the encoded feature image using the feature decoder to obtain the decoded feature images output by each of the multiple second network layers of the feature decoder, the classification module 1020 can be used for any one of the N second network layers. Using the second convolutional layer configured therein, the second input features of the current second network layer are sequentially subjected to convolutional processing, batch normalization processing, and activation processing to obtain a second output feature image; using the second projection excitation layer to perform projection calibration processing on the second output feature image to obtain a second calibrated feature image corresponding to the second output feature image; using the upsampling layer to perform upsampling processing on the second calibrated feature image to obtain the decoded feature image output by the current second network layer; wherein, when the current second network layer is the first second network layer, the second input feature is the encoded feature image output by the last first network layer among the N first network layers; when the current second network layer is any one of the N second network layers except the first second network layer, the second input feature includes the decoded feature image output by the previous second network layer corresponding to the current second network layer, and the encoded feature image output by the first network layer configured with the same number of channels as the current second network layer.
[0099] In some embodiments of the present application, the convolutional module includes a third convolutional layer, a third projection excitation layer, and a fourth convolutional layer; When using the convolutional module to classify the pixel points in each region of interest based on the decoded feature image to obtain the pixel-level classification result corresponding to each region of interest, the classification module 1020 can be used to perform convolutional processing on the third input feature in the third convolutional layer to obtain a third output feature image, where the third input feature includes the encoded feature image output by the first first network layer in the feature encoder and the decoded feature image output by the last second network layer in the feature decoder; using the third projection excitation layer to perform projection calibration processing on the third output feature image to obtain a third calibrated feature image corresponding to the third output feature image; in the fourth convolutional layer, classifying the pixel points in each region of interest based on the third calibrated feature image respectively to obtain the pixel-level classification result corresponding to each region of interest.
[0100] In some embodiments of the present application, the generation module 1030 can be used to count the number of pixel points corresponding to each preset category according to the pixel-level classification result of each region of interest; based on the number of pixel points corresponding to each preset category, determine the image-level classification result corresponding to each region of interest in the target image.
[0101] In some embodiments of the present application, as Figure 11 shown, the device further includes: a training module 1040; The training module 1040 can be used to determine a sample image after image preprocessing, sample region-of-interest information corresponding to the sample image, and a preset training label corresponding to the sample image when training a semantic segmentation model. The preset training label includes at least the classification label of each pixel point in the region of interest. Input the sample image, the sample region-of-interest information, and the preset training label into the semantic segmentation model, and use a feature encoder, a feature decoder, and a convolutional module to perform predictive training on pixel classification of the semantic segmentation model. Among them, during the predictive training of pixel classification, the sample image and the sample region-of-interest information are used as input features, and the preset training label is used as the training label. Based on the loss function and the optimizer, the model parameters in the semantic segmentation model are iteratively updated until the dice coefficient of the semantic segmentation model is greater than a preset threshold, and it is determined that the training of the semantic segmentation model is completed.
[0102] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0103] In the embodiments of the present application, by performing image preprocessing on the initial image, the regional contrast of the region of interest can be improved while ensuring the integrity of the image features, and interference from non-related image features to feature extraction within the region of interest can be avoided. When using the semantic segmentation model for pixel-level classification of the target image, the change in the channel dimension in the feature encoder can be utilized to fully learn discriminative features. At the same time, the learned feature map is restored to the resolution scale of the original input through the feature decoder and the convolutional module, thereby realizing semantic classification of each pixel point in the region of interest. Classifying and recognizing different regions of interest simultaneously can further fully learn the imaging features of the same preset category and the contextual feature relationship between the imaging features of other preset categories. In addition, by performing pixel-level classification on the target image and converting the pixel-level classification result into an image-level classification result, the aggregation of the feature map into a single feature can be avoided, thereby discarding the resolution scale, and an image-level classification result more accurate than directly performing image classification can be obtained.
[0104] In the foregoing, the image classification device according to the embodiments of the present invention has been described from the perspective of functional modules. It should be understood that the functional modules can be implemented in the form of hardware, or in the form of instructions in software, or in a combination of hardware and software modules. Specifically, each step of the method embodiment of the image classification method in the embodiments of the present invention can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in the form of software. The steps of the image classification method applied in combination with the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the method embodiment of the above image classification.
[0105] Figure 12 FIG. 4 is a schematic block diagram of an electronic device 1200 according to an embodiment provided by the present invention.
[0106] As Figure 12 shown, the electronic device 1200 may include: A memory 1210 and a processor 1220. The memory 1210 is used to store a computer program and transmit the program code to the processor 1220. In other words, the processor 1220 can call and run the computer program from the memory 1210 to implement the method in the embodiments of the present invention.
[0107] For example, the processor 1220 can be used to execute the above method embodiment according to the instructions in the computer program.
[0108] In some embodiments of the present invention, the processor 1220 may include, but is not limited to: A general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the like.
[0109] In some embodiments of the present invention, the memory 1210 includes, but is not limited to: Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be Read-Only Memory (ROM), Programmable ROM (PROM), Erasable PROM (EPROM), Electrically Erasable PROM (EEPROM), or flash memory. The volatile memory can be Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double DataRate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synch link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0110] In some embodiments of the present invention, the computer program may be divided into one or more modules, which are stored in the memory 1210 and executed by the processor 1220 to complete the method provided by the present invention. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the controller.
[0111] As Figure 12 shown, the electronic device 1200 may further include: A transceiver 1230, which may be connected to the processor 1220 or the memory 1210.
[0112] Among them, the processor 1220 can control the transceiver 1230 to communicate with other devices. Specifically, it can send data to other devices or receive data sent by other devices. The transceiver 1230 may include a transmitter and a receiver. The transceiver 1230 may further include an antenna, and the number of antennas may be one or more.
[0113] It should be understood that the various components in the electronic device are connected through a bus system. Among them, the bus system includes, in addition to the data bus, a power bus, a control bus, and a status signal bus.
[0114] The present invention also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer can execute the methods in the above method embodiments. Or rather, an embodiment provided by the present invention also provides a computer program product containing instructions. When the instructions are executed by a computer, the computer executes the methods in the above method embodiments.
[0115] When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or a data center that integrates one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a Digital Video Disc (DVD)), or a semiconductor medium (such as a Solid State Disk (SSD)), etc.
[0116] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments claimed in this application can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0117] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or modules can be in electrical, mechanical, or other forms.
[0118] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of the present application, the various functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.
[0119] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An image classification method, characterized in that: include: Performing image preprocessing on the initial image to obtain a target image corresponding to the initial image, wherein the target image includes at least one region of interest, and each region of interest includes at least one imaging feature of a preset category; Inputting the target image and the region of interest information into a pre-trained semantic segmentation model to obtain a pixel-level classification result for each region of interest, wherein the semantic segmentation model includes a feature encoder, a feature decoder, and a convolution module; Generating an image-level classification result of the target image according to the pixel-level classification result; The step of inputting the target image and the region of interest information into a pre-trained semantic segmentation model to obtain a pixel-level classification result for each region of interest includes: Inputting the target image and the region of interest information into the feature encoder, obtaining an encoded feature image output by each of the multiple first network layers of the feature encoder, wherein the number of channels configured corresponding to each of the first network layers increases layer by layer; Decoding the encoded feature image using the feature decoder to obtain a decoded feature image output by each of the plurality of second network layers of the feature decoder, wherein the number of channels configured corresponding to each of the second network layers decreases layer by layer; The convolution module is used to classify the pixels in each of the regions of interest into preset categories based on the decoded feature image, so as to obtain a pixel-level classification result corresponding to each of the regions of interest.
2. The method according to claim 1, characterized in that The feature encoder includes N first network layers, each of which has a different number of channels configured accordingly, and the number of channels increases layer by layer in the N first network layers, and the N first network layers respectively include a first convolutional layer, a first projection excitation layer, and a downsampling layer; The step of inputting the target image and the region of interest information into the feature encoder to obtain a coded feature image output by each of the plurality of first network layers of the feature encoder comprises: For any one of the N first network layers, using the first convolutional layer configured therein, sequentially perform convolution processing, batch normalization processing, and activation processing on the first input feature of the current first network layer to obtain a first output feature image; Using the first projection excitation layer to perform projection calibration processing on the first output characteristic image to obtain a first calibration characteristic image corresponding to the first output characteristic image; Performing a maximum pooling operation on the first calibration feature image using the downsampling layer to obtain a coded feature image output by the current first network layer; Among them, when the current first network layer is the first first network layer, the first input feature is the target image and the area of interest information; when the current first network layer is any first network layer among the N first network layers except the first first network layer, the first input feature is a target coding feature image, and the target coding feature image is a coding feature image output by the current first network layer corresponding to the previous first network layer.
3. The method according to claim 2, characterized in that The step of performing projection calibration processing on the first output feature image by using the first projection excitation layer to obtain a first calibration feature image corresponding to the first output feature image includes: Projecting each channel image in the first output feature image in depth, height and width directions respectively to obtain a first depth projection feature, a first height projection feature and a first width projection feature; Calculate a spatial feature image of the first output feature image after projection compression based on the first depth projection feature, the first height projection feature, and the first width projection feature; Recombining a plurality of spatial feature images corresponding to a plurality of channel images in the first output feature image to obtain a first calibration weight; The first output feature image is calibrated using the first calibration weight to obtain a first calibration feature image.
4. The method according to claim 1, characterized in that: The feature decoder includes N second network layers, each of which has a different number of channels configured accordingly, and the number of channels decreases layer by layer in the N second network layers, and the N second network layers respectively include a second convolutional layer, a second projection excitation layer, and an upsampling layer; The decoding process of the encoded feature image by the feature decoder to obtain a decoded feature image output by each of the plurality of second network layers of the feature decoder includes: For any one of the N second network layers, using the second convolutional layer configured therein, sequentially perform convolution processing, batch normalization processing, and activation processing on the second input features of the current second network layer to obtain a second output feature image; Performing projection calibration processing on the second output characteristic image by using the second projection excitation layer to obtain a second calibration characteristic image corresponding to the second output characteristic image; Performing upsampling processing on the second calibration feature image by using the upsampling layer to obtain a decoded feature image output by the current second network layer; Among them, when the current second network layer is the first second network layer, the second input feature is the encoded feature image output by the last first network layer among the N first network layers; when the current second network layer is any second network layer among the N second network layers except the first second network layer, the second input feature includes the decoded feature image output by the previous second network layer corresponding to the current second network layer, and the encoded feature image output by the first network layer configured with the same number of channels as the current second network layer.
5. The method according to claim 1, characterized in that The convolution module includes a third convolution layer, a third projection excitation layer and a fourth convolution layer; Using the convolution module to classify the pixels in each of the regions of interest into preset categories based on the decoded feature image, and obtaining a pixel-level classification result corresponding to each of the regions of interest, including: Performing convolution processing on the third input feature in the third convolution layer to obtain a third output feature image, wherein the third input feature includes the encoded feature image output by the first first network layer in the feature encoder and the decoded feature image output by the last second network layer in the feature decoder; Performing projection calibration processing on the third output characteristic image by using the third projection excitation layer to obtain a third calibration characteristic image corresponding to the third output characteristic image; In the fourth convolutional layer, the pixels in each of the regions of interest are classified into preset categories based on the third calibration feature image to obtain a pixel-level classification result corresponding to each of the regions of interest.
6. The method according to claim 1, characterized in that Generating the image-level classification result of the target image according to the pixel-level classification result includes: According to the pixel-level classification result of each region of interest, counting the number of pixels corresponding to each preset category; Based on the number of pixels corresponding to each of the preset categories, an image-level classification result corresponding to each of the regions of interest in the target image is determined.
7. The method according to claim 1, characterized in that The method also includes a training method for the semantic segmentation model: Determine a sample image after the image preprocessing, sample region of interest information corresponding to the sample image, and a preset training label corresponding to the sample image, wherein the preset training label at least includes a classification label for each pixel in the region of interest; Inputting the sample image, the sample region of interest information and the preset training label into a semantic segmentation model, and performing prediction training of pixel classification on the semantic segmentation model using the feature encoder, the feature decoder and the convolution module; In which, in the prediction training process of pixel classification, the sample image and the sample area of interest information are used as input features, and the preset training label is used as a training label, and the model parameters in the semantic segmentation model are iteratively updated based on the loss function and the optimizer until the dice coefficient of the semantic segmentation model is greater than the preset threshold, and the semantic segmentation model training is judged to be completed.
8. An image classification device, characterized in that: include: A processing module, configured to perform image preprocessing on an initial image to obtain a target image corresponding to the initial image, wherein the target image includes at least one region of interest, and each region of interest includes at least one imaging feature of a preset category; A classification module, used for inputting the target image and the region of interest information into a pre-trained semantic segmentation model to obtain a pixel-level classification result for each region of interest, wherein the semantic segmentation model includes a feature encoder, a feature decoder and a convolution module; A generating module, used for generating an image-level classification result of the target image according to the pixel-level classification result; The classification module is specifically used for: Inputting the target image and the region of interest information into the feature encoder, obtaining an encoded feature image output by each of the multiple first network layers of the feature encoder, wherein the number of channels configured corresponding to each of the first network layers increases layer by layer; Decoding the encoded feature image using the feature decoder to obtain a decoded feature image output by each of the plurality of second network layers of the feature decoder, wherein the number of channels configured corresponding to each of the second network layers decreases layer by layer; The convolution module is used to classify the pixels in each of the regions of interest into preset categories based on the decoded feature image, so as to obtain a pixel-level classification result corresponding to each of the regions of interest.
9. An electronic device, characterized in that: include: A processor and a memory, the memory being used to store a computer program, and the processor being used to call and run the computer program stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Used to store a computer program, wherein the computer program enables a computer to execute the method according to any one of claims 1 to 7.