Method for acquiring image classification model, image classification method, device and medium
By introducing technologies such as scale scaling rate encoding layer and random generation of crop parameters in neural network models, the existing digital pathological image classification methods have solved the problem of high false positive rates and low accuracy due to dependence on the characteristics of the nucleus size, and achieved higher image classification accuracy.
Patent Information
- Application Number
- CN202510467473.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-06-17
AI Technical Summary
The existing digital pathological image classification methods rely on the nucleus size as a key feature, resulting in high false positive rates and low accuracy. Especially when nuclei enlargement occurs in benign tumor cells and normal cells, they are prone to misclassification.
By introducing a scale scaling rate encoding layer into the neural network model, random generation of cropping parameters and sub-image scaling are used, combined with feature extraction and filtering operations, the influence of nuclear size characteristics is suppressed, so that the model relies more on more reliable nuclear boundary shape features such as nuclear atypia and nuclear pleomorphism for classification.
It effectively reduces the false positive rate of pathological image classification, improves the accuracy of image classification, and avoids incorrect classification caused by incorrect reasoning of nucleus size characteristics.
Smart Images

Figure CN120164046A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and particularly to a method for obtaining an image classification model, an image classification method, an apparatus, and a storage medium. Background Art
[0002] Whole Slide Imaging (WSI) refers to the process of scanning an entire slide and converting it into a digital format, which includes scanning a slide under a microscope and converting it into a digital image for viewing, analyzing, and sharing on a computer screen. Applying whole slide imaging in the field of pathology is digital pathology WSI, which is a technology that converts traditional pathological slides into digital images using high-resolution scanning technology. Digital pathological images obtained based on digital pathology WSI are widely used in cancer diagnosis, such as in the detection of cancers like breast cancer, lung cancer, prostate cancer, etc. With the help of automated analysis algorithms, cancer cells in digital pathological images can be quickly identified and classified. With the continuous development of artificial intelligence technology, automated analysis algorithms can be implemented based on deep learning algorithms, thus assisting pathologists in performing more accurate and faster image analysis. For example, automatically identifying cancer cells, lesion areas, and marking immunohistochemical stains in digital pathological images based on Artificial Intelligence (AI), thereby improving the classification efficiency.
[0003] Although the application prospect of image classification of digital pathological images based on artificial intelligence is broad, there are still problems in aspects such as accuracy, stability, and interpretability. Specifically, existing digital pathological image classification based on artificial intelligence uses the nucleus size as a key feature for image classification. For example, cells with larger nuclei are identified as malignant tumor cells.
[0004] However, in some cases, benign tumor cells and normal cells may also show an increase in the nucleus size. At this time, existing image classification methods will classify these clusters of benign tumor cells and normal cells as clusters of malignant tumor cells, resulting in a relatively high false positive rate and reducing the accuracy of image classification.
[0005] For example, in hematoxylin and eosin stained slides (H&E stained slides), the nuclei of malignant tumors are often larger than those of normal tissues, and the nuclear chromatin is denser, resulting in hyperchromasia. However, slightly enlarged nuclei or hyperchromasia may also occur in benign tumors or inflammation. For example, aneurysmal bone cysts, a type of benign bone tumor, sometimes show some characteristics of enlarged nuclei and abnormal chromatin in histology, especially in fibroblasts or endothelial cells within the tumor. These cells may have nuclei slightly larger than normal cells, especially during the active proliferation stage. At the same time, in some cases, the nuclei of normal cells may also enlarge. For example, hepatocytes may show a certain degree of nuclear dilation under certain physiological responses or metabolic stimuli and may exhibit slightly deeper staining in a short period. At this time, if the nuclear size is used as a key feature for image classification, it will lead to a high false positive rate and low accuracy in image classification. Summary of the Invention
[0006] In view of this, the present disclosure provides a method for obtaining an image classification model, an image classification method, a device, and a medium, which can solve the problems of high false positive rate and low accuracy of traditional image classification methods.
[0007] According to one aspect of the present disclosure, there is provided a method for obtaining an image classification model, the method comprising:
[0008] For each training batch, randomly generate cropping parameters within a preset parameter range;
[0009] Crop a pathological image sample based on the cropping parameters to obtain multiple sub-images;
[0010] Scale each sub-image to a target size to obtain a scaled sub-image; the target size is adapted to the input size supported by the feature extraction layer in the neural network model to be trained;
[0011] Input the scaled sub-images of each training batch into the neural network model to obtain the sub-classification results of each scaled sub-image; the neural network model includes a feature extraction layer and a scale scaling rate encoding layer connected to the feature extraction layer. The feature extraction layer is used to extract features from the scaled sub-images to obtain a feature map; the scale scaling rate encoding layer is used to encode the scaling rate of the sub-images to obtain a filtering function containing encoding information, and filter the feature map through the filtering function containing encoding information to obtain a filtered feature map; and determine the sub-classification results based on the filtered feature map.
[0012] Combine the sub-classification results corresponding to each sub-image and the classification label of each sub-image to train the neural network model to obtain the image classification model, which is used to classify the target pathological image to be classified.
[0013] In a possible implementation manner, the scale scaling rate encoding layer encodes the scaling rate of the sub-images to obtain a filtering function containing encoding information, including:
[0014] Determine the scaling rate based on the cropping parameter and the maximum and minimum values of the parameter range;
[0015] Determine the filtering radius of the disk function based on the scaling rate; the disk function is used to perform low-pass filtering on the feature map through a circle with the center point of the feature map as the center and the filtering radius as the radius;
[0016] Determine the disk function based on the filtering radius to obtain the filtering function.
[0017] In a possible implementation manner, the determining the filtering radius of the disk function based on the scaling rate includes:
[0018] Use a preset smoothing function to smooth the scaling rate to obtain a smoothed scaling rate;
[0019] Input the smoothed scaling rate into a pre-created radius determination function to obtain the filtering radius.
[0020] In a possible implementation manner, the scaling rate x is represented by the following formula:
[0021]
[0022] where p represents the cropping parameter, p max represents the maximum value of the parameter range, and p min represents the minimum value of the parameter range;
[0023] The smoothing function is represented by the following formula:
[0024] t = 6x 5 -15x 4 +10x 3
[0025] Wherein, x represents the scaling ratio, and t represents the smoothed scaling ratio;
[0026] The radius determination function is represented by the following formula:
[0027] R = R0 + t(L / 2 - R0);
[0028] Wherein, R represents the filtering radius, R0 represents the preset minimum radius, L represents the size parameter of the feature map, and t represents the smoothed scaling ratio.
[0029] In a possible implementation manner, filtering the feature map through a filtering function containing encoding information to obtain a filtered feature map, including:
[0030] Performing a two-dimensional Fourier transform on the feature map to obtain a feature frequency spectrum map; multiplying the feature frequency spectrum map by a disk function having the filtering radius to obtain a filtered feature frequency spectrum map; performing an inverse two-dimensional Fourier transform on the filtered feature frequency spectrum map to obtain the filtered feature map;
[0031] The disk function having the filtering radius is represented by the following formula:
[0032]
[0033] Wherein, (x0, y0) represents the coordinates of the center point of the feature map, (x, y) represents the independent variable of the disk function, R represents the filtering radius, and S(x, y) represents the value of the disk function when the independent variable is (x, y);
[0034] Or,
[0035] Performing an inverse two-dimensional Fourier transform on the disk function having the filtering radius to obtain a transformed disk function; convolving the feature map with the transformed disk function to obtain the filtered feature map;
[0036] Wherein, the transformed disk function is represented by the following formula:
[0037]
[0038] Among them, (x0, y0) represents the coordinates of the center point of the feature map, (x, y) represents the independent variable of the transformed disc function, R represents the filtering radius, I(x, y) represents the value of the transformed disc function when the independent variable is (x, y), and J1 represents the Bessel function of order 1.
[0039] In a possible implementation manner, the feature tensor output by the feature extraction layer includes multiple channels, and each channel corresponds to a feature map.
[0040] If the scales of the feature tensors include at least two types, the feature maps corresponding to each channel in the feature tensors of each scale are respectively input into the scale scaling rate encoding layer to filter each feature map.
[0041] Or,
[0042] After feature fusion of the feature tensors of various scales, the fused feature tensors corresponding to each scale are obtained; the feature maps corresponding to each channel in the fused feature tensors of each scale are respectively input into the scale scaling rate encoding layer to filter each feature map.
[0043] In a possible implementation manner, the feature extraction layer performs feature extraction based on the convolutional layer in the convolutional neural network to obtain the feature tensor, and the feature tensors of various scales perform feature fusion based on Cross Stage Partial Connection.
[0044] In a possible implementation manner, combining the sub-classification results corresponding to each sub-image and the classification label of each sub-image to train the neural network model to obtain the image classification model includes:
[0045] Based on the difference between the sub-image classification result and the sub-image classification label, iteratively update the model parameters of the neural network model to obtain the image classification model.
[0046] According to another aspect of the present disclosure, an image classification method is provided, and the method includes:
[0047] Traverse and crop the target pathological image to be classified to obtain sub-images to be classified.
[0048] Classify the sub-images to be classified based on a pre-trained image classification model to obtain sub-classification results corresponding to each sub-image to be classified; the image classification model is trained based on the method for obtaining the image classification model described above.
[0049] Vote on each sub-classification result to obtain the classification result of the target pathological image.
[0050] In a possible implementation, the voting on the respective sub-classification results to obtain the classification result of the target pathological image includes:
[0051] Determining the classification with the largest number from the sub-classification results corresponding to the respective sub-images to be classified of the target pathological image, and the classification with the largest number is the classification result of the target pathological image; or, performing weighted summation on the respective sub-classification results corresponding to the respective sub-images to be classified of the target pathological image to obtain the classification with the highest probability, and the classification with the highest probability is the classification result of the target pathological image;
[0052] In a possible implementation, the traversing and cropping of the target pathological image to be classified to obtain the sub-images to be classified includes:
[0053] Traversing and cropping the target pathological image according to the input size supported by the feature extraction layer in the image classification model to obtain the sub-images to be classified.
[0054] According to another aspect of the present disclosure, there is provided a computer device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the above method.
[0055] According to another aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0056] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, or a non-volatile computer-readable storage medium carrying the computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0057] By randomly generating cropping parameters within the preset parameter range for each training batch among multiple training batches; cropping the pathological image samples to be classified based on the cropping parameters to obtain multiple sub-images; scaling each sub-image to the target size to obtain the scaled sub-images; classifying each scaled sub-image based on a neural network model to obtain the sub-classification results of each scaled sub-image; the neural network model includes a feature extraction layer and a scale scaling rate encoding layer connected to the feature extraction layer, the feature extraction layer is used to extract features from the scaled sub-images to obtain feature maps; the scale scaling rate encoding layer is used to encode the scaling rate of the sub-images to obtain a filtering function containing encoded information, and the filtering function containing encoded information filters the feature maps to obtain the filtered feature maps, so as to determine the sub-classification results based on the filtered feature maps; combining the respective sub-classification results and the classification labels of each sub-image to train the neural network model; since the neural network model can suppress the neural network from reasoning through the feature of the nucleus size during the training process, so that the neural network reasons from more reliable nucleus boundary shape features such as nuclear atypia and nuclear pleomorphism, therefore, it can solve the problem that the existing methods for obtaining neural network models in digital pathology often misclassify normal cells and benign tumor cells with enlarged nuclei as malignant tumor cells, resulting in low accuracy of image classification, and can reduce the false positive rate of pathological image classification and improve the classification accuracy.
[0058] Other features and aspects of the present disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The accompanying drawings, which are included in and constitute a part of this specification, illustrate exemplary embodiments, features, and aspects of the present disclosure together with the specification and are used to explain the principles of the present disclosure.
[0060] Figure 1 A flowchart showing a method for obtaining an image classification model according to an embodiment of the present disclosure;
[0061] Figure 2 A schematic diagram showing the model structure of a neural network model according to an embodiment of the present disclosure;
[0062] Figure 3 A schematic diagram showing the calculation process of a scale scaling rate encoding layer according to an embodiment of the present disclosure;
[0063] Figure 4 A flowchart showing an image classification method according to an embodiment of the present disclosure;
[0064] Figure 5 A block diagram showing an apparatus for obtaining an image classification model according to an embodiment of the present disclosure;
[0065] Figure 6 A block diagram showing an image classification device according to an embodiment of the present disclosure;
[0066] Figure 7 A block diagram showing a computer device according to an embodiment of the present disclosure. Detailed implementation manners
[0067] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0068] As used herein, the terms "comprising", "including", "having", or variations thereof are open-ended and include one or more stated features, wholes, elements, steps, components, or functions, but do not exclude the existence or addition of one or more other features, wholes, elements, steps, components, functions, or groups thereof.
[0069] When an element is referred to as being "connected", "coupled", "responsive" or variations thereof to another element, it may be directly connected, coupled, or responsive to the other element, or there may be intervening elements.
[0070] Although the terms first, second, third, etc. may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Thus, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.
[0071] The word "exemplary" as used herein means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" need not be construed as superior to or better than other embodiments.
[0072] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can be implemented without some of these specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0073] Figure 1 A flowchart showing a method for obtaining an image classification model according to an embodiment of the present disclosure. As Figure 1 shown, the method includes:
[0074] Step 101, for each training batch, randomly generate cropping parameters within a preset parameter range.
[0075] Exemplarily, the lower limit value of the parameter range is determined based on the size information of the cells in the pathological image sample. For example, the lower limit value of the parameter range is greater than or equal to the image size of the nucleus diameter. The upper limit value of the parameter range is determined based on the input size supported by the feature extraction layer in the following neural network model. For example, the upper limit value of the parameter range is less than or equal to the input size. Exemplarily, the parameter range is 64 ≤ p ≤ 256, where p represents the cropping parameter.
[0076] The cropping parameter refers to the numerical setting required when performing a cropping operation on a pathological image sample. Exemplarily, the cropping parameter is used to indicate the size of the cropping area. For example, the cropping parameter includes the width and height of the cropping. Exemplarily, if the width and height of the cropping are the same, the cropping parameter is a single value. Correspondingly, the parameter range includes the parameter range corresponding to a single cropping parameter; if the width and height of the cropping are different, the cropping parameter includes a width value and a height value. Correspondingly, the parameter range includes the parameter range corresponding to the width value and the parameter range corresponding to the height value.
[0077] Generate a cropping parameter randomly within a preset parameter range, including: generating a cropping parameter within the parameter range using a random algorithm. Among them, the random algorithm can be a random number generation function or a custom random cropping algorithm. The implementation manner of the random algorithm is not limited in this embodiment.
[0078] Optionally, the same training batch corresponds to the same cropping parameter. For example, batch size = 16, that is, in each training batch, the neural network model processes 16 sub-images. At this time, the 16 sub-images are cropped using the same cropping parameter, that is, the sizes of the 16 sub-images are the same. Or, the same training batch corresponds to different cropping parameters. For example, batch size = 16, that is, in each training batch, the neural network model processes 16 sub-images. At this time, the 16 sub-images are cropped using at least two cropping parameters. For example, 8 of the sub-images are cropped using cropping parameter 1, and the other 8 sub-images are cropped using cropping parameter 2. At this time, the same training batch corresponds to two cropping parameters.
[0079] The cropping parameters corresponding to different training batches are different. For example: If each training batch corresponds to one cropping parameter, then training batch 1 corresponds to cropping parameter 1, and training batch 2 corresponds to cropping parameter 2. If each training batch corresponds to at least two cropping parameters, then there is at least one different cropping parameter between different training batches. For example: Training batch 1 corresponds to cropping parameter 1 and cropping parameter 2, and training batch 2 corresponds to cropping parameter 1 and cropping parameter 3, where cropping parameter 2 and cropping parameter 3 are different; Another example: Training batch 1 corresponds to cropping parameter 1 and cropping parameter 2, and training batch 2 corresponds to cropping parameter 3 and cropping parameter 4, where the two cropping parameters of training batch 1 are different from the two cropping parameters of training batch 2.
[0080] Step 102: Crop the pathological image sample to be classified based on the cropping parameter to obtain multiple sub-images.
[0081] Among them, the pathological image sample can be a digital image obtained by scanning a pathological section through a high-resolution scanning technology. Generally, the size of the pathological image sample is larger than the upper limit value of the parameter range.
[0082] In one example, in one training batch, if the same cropping parameter is used to crop the pathological image sample, then the size of each sub-image is the same. Correspondingly, cropping the pathological image sample based on the cropping parameter to obtain multiple sub-images includes: Cropping the pathological image sample with the cropping parameter every preset pixel in each pathological image sample until the number of sub-images meets the quantity requirement of each training batch. The pathological image samples cropped by different training batches can be exactly the same, or partially the same, or completely different. Among them, the preset pixel makes there be an overlapping area or no overlapping area between adjacent sub-images. The setting method of the preset pixel is not limited in this embodiment. When the number of training batches is large enough, each pathological image sample will be approximately completely traversed.
[0083] Among them, cropping the pathological image sample using the cropping parameter every preset pixel means: for the cropping reference point in the pathological image sample (such as the cropping reference point is initially set as the top-left vertex pixel (0, 0) of the pathological image sample), cropping the pathological image sample based on this cropping reference point according to the cropping parameter to obtain the first sub-image. The top-left pixel position of the first sub-image is (0, 0), and the bottom-right pixel position is (p, p); then, move the cropping reference point by the preset pixel in the width direction or the length direction (for example: move p pixels in the length direction) to obtain the updated cropping reference point (for example: updated to (p, 0)); crop the pathological image sample again based on this updated cropping reference point according to the cropping parameter to obtain the second sub-image. The top-left pixel position of the second sub-image is (p, 0), and the bottom-right pixel position is (2p, p); then move the cropping reference point by the preset pixel in the width direction or the length direction again to obtain the updated cropping reference point (for example: updated to (2p, 0)); crop the pathological image sample again based on this updated cropping reference point according to the cropping parameter to obtain the third sub-image, and so on in a cycle until all the pathological image samples are completely traversed and cropped. Then crop the next pathological image sample until the number requirement of a training batch is reached or until all the pathological image samples are completely cropped, so as to select the sub-images required for a training batch from the obtained sub-images.
[0084] In another example, if for each training batch, at least two cropping parameters are used to crop the pathological image sample, then there are at least two sub-images with different sizes in each training batch. Correspondingly, crop the pathological image sample to be classified based on the cropping parameter to obtain multiple sub-images, including: crop every first preset pixel in each pathological image sample using the first cropping parameter to obtain sub-images of the first cropping parameter size; crop every second preset pixel in each pathological image sample using the second cropping parameter to obtain sub-images of the second cropping parameter size until the number of sub-images meets the requirements of the training batch. Or, for a part of the pathological image samples in the sample set, crop every first preset pixel in this part of the pathological image samples using the first cropping parameter; for another part of the pathological image samples in the sample set, crop every second preset pixel in this other part of the pathological image samples using the second cropping parameter again until the number of sub-images meets the requirements of the training batch; this embodiment does not limit the method of using at least two cropping parameters to crop the pathological image sample.
[0085] Among them, the first cutting parameter is different from the second cutting parameter, and the first preset pixel may be the same as or different from the second preset pixel. For the method of performing cutting using the first cutting parameter every first preset pixel and using the second cutting parameter every second preset pixel, please refer to the above description of "performing cutting using a cutting parameter every preset pixel", and this embodiment will not elaborate here.
[0086] Taking the example where the width and height of the cutting are the same and the preset parameter range is 64 ≤ p ≤ 256, assuming that the cutting parameter p = 128 is selected within this parameter range, and the size of a pathological image sample is 20000×20000, then by cutting the pathological image sample according to p×p×c every p pixels, multiple non-overlapping sub-images with a size of 128×128×c can be obtained. Among them, c represents the number of channels of the pathological image sample. For example, c = 3 indicates that the pathological image sample is a color image.
[0087] Step 103: Scale each sub-image to the target size to obtain the scaled sub-image; the target size is adapted to the input size supported by the feature extraction layer in the neural network model to be trained.
[0088] In this embodiment, during the training process of the model, the feature extraction layer of the neural network model has size limitations on the input image. Therefore, each sub-image needs to be scaled to the size supported by the feature extraction layer of the neural network model. Specifically, each sub-image of p×p×3 is scaled to have the same length and width as the input tensor of the neural network model. For example, if the size of the input tensor of the neural network model is 256×256×3, then the sub-image of 128×128×3 is scaled to 256×256×3, and then the scaled sub-image is input into the neural network model.
[0089] Exemplarily, scaling each sub-image to the target size to obtain the scaled sub-image includes: scaling each sub-image to the target size based on an image interpolation method. Among them, the image interpolation method includes but is not limited to: nearest neighbor interpolation, bilinear interpolation, bicubic interpolation, etc. This embodiment does not limit the implementation manner of the image interpolation method.
[0090] Step 104: Input the scaled sub-images of each training batch into the neural network model to obtain the sub-classification results of each scaled sub-image.
[0091] During the training process, for each training batch, select the sub-image with the size of the cutting parameter corresponding to this training batch. The number of such sub-images is the batch size of this training batch. Input the sub-images of this batch size into the neural network model to perform iterative training on the neural network model.
[0092] For example, when the batch size of each training batch is 16, the cropping parameter of training batch 1 is 128, and the cropping parameter size of training batch 2 is 64. Then, for training batch 1, the training samples input to the neural network model are a set of sub-images with a cropping parameter of 128, that is, this set includes 16 sub-images of 128×128×3; for training batch 2, the training samples input to the neural network model are a set of sub-images with a cropping parameter of 64, that is, this set includes 16 sub-images of 64×64×3.
[0093] The neural network model includes a feature extraction layer and a scale scaling rate encoding layer connected to the feature extraction layer. The feature extraction layer is used to extract features from the scaled sub-images to obtain feature maps; the scale scaling rate encoding layer is used to encode the scaling rate of the sub-images to obtain a filtering function containing encoded information, and filter the feature maps through the filtering function containing encoded information to obtain filtered feature maps; and determine sub-classification results based on the filtered feature maps.
[0094] The feature tensor output by the feature extraction layer is a three-dimensional matrix, and this feature tensor includes multiple channels, with each channel corresponding to a feature map. For example: the feature tensor output by the feature extraction layer is 20×20×1024, indicating that this feature tensor has 1024 channels, and the size of the feature map for each channel is 20×20.
[0095] In one example, as Figure 2 shown, the feature extraction layer is directly connected to the scale scaling rate encoding layer. At this time, the feature maps corresponding to each channel are respectively input to the scale scaling rate encoding layer to filter each feature map. For example: if the feature tensor output by the feature extraction layer is 20×20×1024, then 1024 feature maps of 20×20 are sequentially input to the scale scaling rate encoding layer for filtering.
[0096] Optionally, if the scale of the feature tensor includes at least two types, then the feature maps corresponding to each channel in the feature tensors of each scale are respectively input to the scale scaling rate encoding layer to filter each feature map.
[0097] Among them, the feature tensors of at least two scales refer to the size of the feature maps corresponding to each channel in the feature tensor. For example: if the feature tensor output by the feature extraction layer has three scales, which are: 20×20×1024, 40×40×512, 80×80×256, then the feature maps of the three scales, that is, 1024 feature maps of 20×20, 512 feature maps of 40×40, and 256 feature maps of 80×80, are respectively input to the scale scaling rate encoding layer to filter each feature map.
[0098] In another example, the feature extraction layer is connected to the scale scaling rate encoding layer through a feature fusion layer (Figure 2 (not shown in the figure). At this time, if the scale of the feature tensor includes at least two types, after the feature tensors of various scales are fused, the fused feature tensors corresponding to each scale are obtained; the feature maps corresponding to each channel in the fused feature tensors of each scale are respectively input into the scale scaling rate encoding layer to filter each feature map.
[0099] For example: the feature tensors output by the feature extraction layer are of three scales, which are: 20×20×1024, 40×40×512, 80×80×256. After the feature tensors of the three scales are fused by the feature fusion layer, the fused feature tensors corresponding to each scale are obtained, that is, the fused feature tensors are also: 20×20×1024, 40×40×512, 80×80×256. The 1024 20×20 feature maps, 512 40×40 feature maps, and 256 80×80 feature maps in the fused feature tensors are respectively input into the scale scaling rate encoding layer to filter each feature map.
[0100] Among them, the feature fusion layer is used to fuse the feature tensors of different scales output by the feature extraction layer. For example: through specific fusion algorithms, such as weighted summation, convolution, upsampling, downsampling, splicing, etc., the information of the feature tensors of different scales is fused to obtain the fused feature tensors corresponding to each scale. In this way, the information of the feature tensors of different scales can be integrated to improve the model's ability to understand and process data.
[0101] Exemplarily, referring to Figure 2 the model structure of the neural network model shown, the feature extraction layer is constructed based on the convolutional layer in the convolutional neural network. The model parameters initialized by this convolutional layer can be randomly generated, or the model parameters of the existing convolutional layer for image recognition can be used. This embodiment does not limit the acquisition method of the model parameters initialized by this convolutional layer.
[0102] Among them, the convolutional layer in the convolutional neural network can be constructed based on the DarkNet backbone network, ResNet backbone network, DenseNet backbone network, or MobileNet backbone network, etc. This embodiment does not limit the implementation method of the convolutional layer in the convolutional neural network.
[0103] Exemplarily, the feature fusion layer is constructed based on Cross Stage Partial Connection, that is, feature tensors of various scales are subjected to feature fusion based on Cross Stage Partial Connection. Cross Stage Partial Connection is used to selectively connect and fuse features of different stages or different scales, enabling the neural network model to perform information interaction between features at different levels, retaining both the detailed information in low-level features and combining the semantic information in high-level features, so that the fused feature map contains more comprehensive and richer information, which is beneficial to improving the accuracy and generalization ability of the model.
[0104] In one example, the scale scaling rate encoding layer encodes the scaling rate of the sub-image to obtain a filtering function containing encoded information, including: determining the scaling rate based on the maximum and minimum values of the cropping parameter and the parameter range; determining the filtering radius of the disk function based on the scaling rate; the disk function is used to perform low-pass filtering on the feature map through a circle with the center point of the feature map as the center and the filtering radius as the radius; determining the disk function based on the filtering radius to obtain the filtering function.
[0105] In this embodiment, different cropping parameters can be standardized to [0, 1] by calculating the scaling rate, which is convenient for subsequent dynamic adjustment of the filtering radius.
[0106] Exemplarily, the scaling rate x is represented by the following formula:
[0107]
[0108] where p represents the cropping parameter, p max represents the maximum value of the parameter range, and p min represents the minimum value of the parameter range.
[0109] For example: the parameter range is 64 ≤ p ≤ 256, p = 128. At this time, p max = 256, p min = 64,
[0110]
[0111] Since cell images have the property that rotation does not change the semantic features of cells, and the disk function is applicable to this scenario, in this embodiment, by using the disk function as a filtering function to filter the feature map, it can be adapted to the cell classification scenario. The filtering radius in the disk function is a distance parameter based on which the low-pass filtering operation is performed. The disk function uses the filtering radius to draw a circle with the center point of the feature map as the center, and the radius of this circle is the filtering radius. At this time, the feature values outside the circle are filtered out, thereby achieving low-pass filtering. Since high-frequency feature values (i.e., feature values greater than the filtering radius) are usually high-frequency noises, in this embodiment, by filtering the high-frequency feature values, the model can classify based on features related to cells, improving the classification accuracy. At the same time, filtering out the high-frequency components can also reduce the risk of overfitting. Moreover, the scale encoding function filtering can strengthen the features related to the scale information of the feature map.
[0112] Optionally, determining the filtering radius of the disk function based on the scaling ratio includes: performing smoothing processing on the scaling ratio using a preset smoothing function to obtain the smoothed scaling ratio; inputting the smoothed scaling ratio into a pre-created radius determination function to obtain the filtering radius.
[0113] Exemplarily, the smoothing function is established based on a polynomial equation and can make the scaling ratio within the numerical range of [0, 1]. At this time, the smoothing function can sample more values close to L / 2, that is, obtain disk function samples that can fully cover the feature map, where L represents the size parameter of the feature map.
[0114] For example: the smoothing function is represented by the following formula:
[0115] t = 6x 5 - 15x 4 + 10x 3
[0116] where x represents the scaling ratio and t represents the smoothed scaling ratio.
[0117] In other embodiments, the smoothing function can also be implemented as:
[0118] t = 35x 4 - 84x 5 + 70x 6 - 20x 7 ,
[0119] In other embodiments, the smoothing function can also be selected as the sigmoid function, and this embodiment does not limit the implementation manner of the smoothing function.
[0120] Exemplarily, the radius determination function can make the filtering radius R within the numerical range of [R0, L / 2]. For example: the radius determination function is represented by the following formula:
[0121] R = R0 + t(L / 2 - R0);
[0122] Wherein, R represents the filtering radius. R0 represents a preset minimum radius, and R0 is less than L / 2. For example, R0 = 5. In actual implementation, the value of R0 can also be other values, and this embodiment does not limit the value of R0. L represents the size parameter of the feature map. If the length and width of the feature map are equal, then L is the value of the length or width of the feature map. For example, in the example where the above feature extraction layer includes three scales, the values of L are L = 80, L = 40, and L = 20 respectively, and each L corresponds to a filtering radius, and different L values correspond to different filtering radii. t represents the scaling rate after smoothing.
[0123] In this example, by combining the scaling rate t to determine the filtering radius, the model can obtain the scaling information of the sub-image, so as to perform inference based on more information, further ensuring the accuracy of the model inference.
[0124] Exemplarily, filtering the feature map through a filtering function containing encoding information to obtain the filtered feature map, including one of the following methods:
[0125] The first method: perform a two-dimensional Fourier transform on the feature map to obtain a feature frequency spectrum map; multiply the feature frequency spectrum map by a disk function with a filtering radius to obtain a filtered feature frequency spectrum map; perform an inverse two-dimensional Fourier transform on the filtered feature frequency spectrum map to obtain a filtered feature map.
[0126] Wherein, the disk function S(x, y) with a filtering radius is represented by the following formula:
[0127]
[0128] Wherein, (x0, y0) represents the coordinates of the center point of the feature map, (x, y) represents the independent variable of the disk function, R represents the filtering radius, and S(x, y) represents the value of the disk function when the independent variable is (x, y).
[0129] For example: referring to Figure 3 , for each p×p sub-image, the size of the feature map obtained by extracting the sub-image through the feature extraction layer is L×L. The scale scaling rate encoding layer inputs p into the smoothing function to obtain the scaling rate; the scaling rate t output by the smoothing function and the size L of the feature map are input into the radius determination function to obtain the filtering radius R of the disk function. The disk function determined based on this filtering radius R is multiplied by the two-dimensional Fourier transform of the feature map to obtain a filtered feature map (or the feature map encoded by the scale scaling rate encoding layer).
[0130] The second method: perform an inverse two-dimensional Fourier transform on the disk function with a filtering radius to obtain the transformed disk function; convolve the feature map with the transformed disk function to obtain the filtered feature map.
[0131] Among them, the transformed disk function I(x, y) is represented by the following formula:
[0132]
[0133] Among them, (x0, y0) represents the coordinates of the center point of the feature map, (x, y) represents the independent variable of the transformed disk function, R represents the filtering radius, I(x, y) represents the value of the transformed disk function when the independent variable is (x, y), and J1 represents the Bessel function of order 1.
[0134] As Figure 2 shown, the neural network model further includes a pooling layer and a fully connected layer connected to the scale scaling rate encoding layer. The pooling layer is used to reduce the dimension of the filtered feature map, and the fully connected layer is used to determine the probability vector of each classification based on the filtered feature map.
[0135] Among them, Figure 2 the pooling layer is taken as an example of the global maximum pooling layer for illustration. In actual implementation, the pooling layer can also be a global average pooling layer. This embodiment does not limit the implementation manner of the pooling layer.
[0136] Optionally, the fully connected layer is further connected to a classification layer, and the classification layer is used to determine the sub-classification result based on the probability vector of each classification. Among them, the classification layer can be implemented based on the argmax function, and the argmax function is used to determine the classification with the largest probability vector as the sub-classification result.
[0137] After the scaled sub-images of different training batches are cropped into sub-images according to random cropping parameters, they will be scaled to a unified target size. At this time, for the pathological image samples of cells with larger cell nuclei, if the pathological image samples are cropped according to larger cropping parameters, they will be enlarged or reduced to a lesser extent; while for the pathological image samples of cells with smaller cell nuclei, if the pathological image samples are cropped according to smaller cropping parameters, they will be enlarged to a greater extent. Eventually, the sizes of cells with larger cell nuclei in the scaled sub-images and the sizes of cells with smaller cell nuclei in the scaled sub-images may be basically the same, while their classification labels may be different. Therefore, the neural network model can ignore the influence of the cell nucleus size and learn features other than the cell nucleus size, such as features like the boundary shape of the cell nucleus. Malignant tumor cell clusters usually have nuclear atypia and nuclear pleomorphism, which are manifested as changes in the boundary shape of the cell nucleus in digital pathological images. Therefore, the model learning features such as the boundary shape of the cell nucleus can improve the accuracy and reliability of image classification.
[0138] Step 105: Combine the sub-classification results corresponding to each sub-image and the classification label of each sub-image, and train the neural network model to obtain an image classification model. The image classification model is used to classify the sub-images to be classified of the target pathological image to be classified.
[0139] In this embodiment, each sub-image corresponds to a classification label. Correspondingly, combining the sub-classification results corresponding to each sub-image and the classification label of the sub-image, and training the neural network model to obtain an image classification model includes:
[0140] Based on the difference between the classification result and the classification label, iteratively update the model parameters of the neural network model to obtain an image classification model.
[0141] Among them, for the sub-images in each training batch, the model parameters of the neural network model are updated once, and the neural network model is iteratively updated based on the sub-images in multiple training batches to obtain an image classification model. The iterative update method can be implemented based on the backpropagation algorithm and cross-entropy minimization. This embodiment does not limit the specific implementation method of the iterative update.
[0142] Since each sub-image of each pathological image sample corresponds to a classification label, therefore, by using a loss function to compare the classification label with the classification result of the sub-image, the model parameters of the neural network model can be iteratively updated through methods such as the backpropagation algorithm and cross-entropy minimization until the preset iterative stop condition is reached, and then an image classification model is obtained.
[0143] In summary, for the method for obtaining an image classification model provided in this embodiment, for each training batch, cropping parameters are randomly generated within a preset parameter range; the pathological image sample is cropped based on the cropping parameters to obtain multiple sub-images; each sub-image is scaled to a target size to obtain a scaled sub-image; the scaled sub-images of each training batch are input into a neural network model to obtain a sub-classification result for each scaled sub-image; the neural network model includes a feature extraction layer and a scale scaling rate encoding layer connected to the feature extraction layer. The feature extraction layer is used to extract features from the scaled sub-images to obtain feature maps; the scale scaling rate encoding layer is used to encode the scaling rate of the sub-images to obtain a filtering function containing encoded information, and the feature maps are filtered through the filtering function containing encoded information to obtain filtered feature maps; the sub-classification result is determined based on the filtered feature maps; the neural network model is trained by combining the sub-classification results corresponding to each sub-image and the classification labels of the sub-images to obtain an image classification model; since the neural network model can suppress the neural network from reasoning through the feature of the nucleus size during the training process, so that the neural network reasons from more reliable nucleus boundary shape features such as nuclear atypia and nuclear pleomorphism, therefore, it can solve the problem that existing digital pathological image classification methods often misclassify normal cells and benign tumor cells with enlarged nuclei as malignant tumor cells, resulting in low image classification accuracy, and can reduce the false positive rate of pathological image classification and improve the classification accuracy.
[0144] Figure 4 FIG. 4 shows a flowchart of an image classification method according to an embodiment of the present disclosure. As Figure 1 shown, the method includes:
[0145] Step 401, traverse and crop the target pathological image to be classified to obtain sub-images to be classified.
[0146] In one example, traversing and cropping the target pathological image to be classified to obtain sub-images to be classified includes: traversing and cropping the target pathological image according to the input size supported by the feature extraction layer in the image classification model to obtain sub-images to be classified.
[0147] Among them, the acquisition method of the image classification model refers to the above embodiment, and will not be elaborated herein.
[0148] In this example, after the image classification model is trained, the size of the sub-images to be cropped is the same as the input size supported by the feature extraction layer, that is, there is no scaling operation. Since the image classification model already has the ability to suppress reasoning based on cell size features, therefore, directly using the input size supported by the feature extraction layer to traverse and crop the target pathological image can save the computing resources required for scaling operations, thereby improving the image classification efficiency.
[0149] In other embodiments, the target pathological image can also be traversed and cropped by randomly selecting cropping parameters in the manner of step 101. After that, the sub-images to be classified obtained by cropping are scaled to the input size supported by the feature extraction layer. This embodiment does not limit the cropping method of the target pathological image.
[0150] Step 402: Classify the sub-images to be classified based on a pre-trained image classification model to obtain the sub-classification results corresponding to each sub-image to be classified.
[0151] In one example, if the size of the sub-image to be classified is the same as the input size supported by the feature extraction layer, the sub-image to be classified is directly input into the image classification model to obtain the sub-classification results corresponding to each sub-image to be classified.
[0152] In other embodiments, if the size of the sub-image to be classified is not the same as the input size supported by the feature extraction layer, the sub-image to be classified is scaled to the target size to be consistent with the input size supported by the feature extraction layer, and the scaled sub-image to be classified is input into the image classification model to obtain the sub-classification results corresponding to each sub-image to be classified.
[0153] Step 403: Vote on the respective sub-classification results to obtain the classification result of the target pathological image.
[0154] In one example, voting on the respective sub-classification results to obtain the classification result of the target pathological image includes:
[0155] Determine the classification with the largest number from the sub-classification results corresponding to the respective sub-images to be classified of the target pathological image, and the classification with the largest number is the classification result of the target pathological image; or, perform weighted summation on the respective sub-classification results corresponding to the respective sub-images to be classified in the target pathological image to obtain the classification with the highest probability, and the classification with the highest probability is the classification result of the target pathological image;
[0156] Taking the case of determining the classification with the largest number from the sub-classification results corresponding to the respective sub-images to be classified of the target pathological image, and the classification with the largest number is the classification result of the target pathological image as an example: At this time, the sub-classification result of each sub-image to be classified includes the classification identifier of the cells in the sub-image to be classified. For example, 1 represents a benign tumor, 2 represents a malignant tumor, and 0 represents normal tissue. At this time, if the number of benign tumors in the respective sub-classification results corresponding to the same target pathological image is greater than the number of malignant tumors, the classification result of the target pathological image is a benign tumor; if the number of benign tumors in the respective sub-classification results corresponding to the same target pathological image is less than the number of malignant tumors, the classification result of the target pathological image is a malignant tumor.
[0157] Taking the weighted sum of each sub-classification result corresponding to each sub-image to be classified in the target pathological image to obtain the classification with the highest probability, and the classification with the highest probability is the classification result of the target pathological image as an example: At this time, the sub-classification result of each sub-image to be classified in the same target pathological image includes the probability corresponding to each classification. For example, the sub-classification result of a sub-image to be classified is 60% for benign tumor and 40% for malignant tumor. At this time, the probabilities of benign tumor in each sub-classification result are weighted and summed, the probabilities of malignant tumor in each sub-classification result are weighted and summed, and the classification with the highest probability is used as the classification result of the target pathological image.
[0158] Optionally, the weights corresponding to different sub-images to be classified are the same or different. For example, the weights of different sub-images to be classified are the same and all are 1, or the weights of different sub-images to be classified are different. When the sub-images to be classified are mapped back to the target pathological image, the weights of the sub-images to be classified in the target pathological image where the same classification is in the connected domain are greater than the weights of the sub-images to be classified where the same classification is scattered in the target pathological image. This embodiment does not limit the implementation manner of the weights used in weighted averaging.
[0159] In summary, the image classification method provided in this embodiment traverses and cuts the target pathological image to be classified to obtain sub-images to be classified; classifies the sub-images to be classified based on a pre-trained image classification model to obtain the classification result of each sub-image to be classified; votes on each sub-classification result to obtain the classification result of the target pathological image; since the image classification model can suppress reasoning based on the feature of the nucleus size, so that the image classification model reasons from more reliable nucleus boundary shape features such as nuclear atypia and nuclear pleomorphism. Therefore, it can solve the problem that the existing digital pathological image classification method often misclassifies normal cells with enlarged nuclei and benign tumor cells as malignant tumor cells, resulting in low image classification accuracy, can reduce the false positive rate of pathological image classification, and improve the classification accuracy.
[0160] In addition, since the image classification model already has the ability to suppress reasoning based on cell size features, directly traversing and cutting the target pathological image using the input size supported by the feature extraction layer can save the computing resources required for scaling operations, thereby improving the image classification efficiency.
[0161] Figure 5 The block diagram of the device for obtaining an image classification model according to an embodiment of the present disclosure is shown. As Figure 5 shown, the device includes: a parameter generation module 510, an image cutting module 520, an image scaling module 530, an image classification module 540, and a model training module 550.
[0162] A parameter generation module 510, configured to randomly generate cropping parameters within a preset parameter range for each training batch;
[0163] An image cropping module 520, configured to crop a pathological image sample based on the cropping parameters to obtain multiple sub-images;
[0164] An image scaling module 530, configured to scale each sub-image to a target size to obtain a scaled sub-image; the target size is adapted to the input size supported by a feature extraction layer in a neural network model to be trained;
[0165] An image classification module 540, configured to input the scaled sub-images of each training batch into the neural network model to obtain sub-classification results of each scaled sub-image; the neural network model includes a feature extraction layer and a scale scaling rate encoding layer connected to the feature extraction layer, the feature extraction layer is configured to perform feature extraction on the scaled sub-image to obtain a feature map; the scale scaling rate encoding layer is configured to encode the scaling rate of the sub-image to obtain a filtering function containing encoding information, and filter the feature map through the filtering function containing encoding information to obtain a filtered feature map; and determine the sub-classification result based on the filtered feature map;
[0166] A model training module 550, configured to train the neural network model by combining the sub-classification results corresponding to each sub-image and the classification label of each sub-image to obtain the image classification model, and the image classification model is configured to classify a target pathological image to be classified.
[0167] For relevant details, see the above method embodiments.
[0168] Figure 6 The block diagram of an image classification device according to an embodiment of the present disclosure is shown. As Figure 5 shown, the device includes: an image cropping module 610, an image classification module 620, and a voting classification module 630.
[0169] An image cropping module 610, configured to traverse and crop a target pathological image to be classified to obtain sub-images to be classified;
[0170] An image classification module 620, configured to classify the sub-images to be classified based on a pre-trained image classification model to obtain sub-classification results corresponding to each sub-image to be classified; wherein, the image classification model is trained based on the method for obtaining an image classification model described above;
[0171] A voting classification module 630, configured to vote on each sub-classification result to obtain the classification result of the target pathological image.
[0172] For relevant details, please refer to the method embodiments above.
[0173] In some embodiments, the functions or modules included in the apparatus provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0174] The embodiments of the present disclosure also provide an image classification apparatus, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above method.
[0175] The embodiments of the present disclosure also provide a non-volatile computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.
[0176] The embodiments of the present disclosure also provide a computer program product, including a computer program, or a non-volatile computer-readable storage medium carrying the computer program. When the computer program is executed by a processor, the steps of the above method are implemented.
[0177] Figure 7 It is a block diagram of a computer device 1900 shown according to an exemplary embodiment. For example, the apparatus 1900 can be provided as a server or a terminal device. The computer device can be an image classification apparatus or an apparatus for obtaining an image classification model. The apparatus 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0178] The apparatus 1900 may further include a power supply component 1926 configured to perform power management of the apparatus 1900, a wired or wireless network interface 1950 configured to connect the apparatus 1900 to a network, and an input / output interface 1958 (I / O interface). The apparatus 1900 can operate based on an operating system stored in the memory 1932, such as Windows Server TM , MacOS X TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0179] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, and the computer program instructions can be executed by a processing component 1922 of the device 1900 to complete the above method.
[0180] A computer-readable storage medium can be a tangible device that can hold and store programs / instructions used by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in a groove having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0181] The computer programs (or computer-readable program instructions) described herein can be downloaded from the computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.
[0182] A computer program (or computer program instructions) for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via an Internet service provider through the Internet). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0183] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0184] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0185] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0186] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or by combinations of special purpose hardware and computer instructions.
[0187] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A method for obtaining an image classification model, characterized in that: The method comprises: For each training batch, cropping parameters are randomly generated within the preset parameter range; Cropping the pathological image sample based on the cropping parameters to obtain a plurality of sub-images; Scaling each sub-image to a target size to obtain a scaled sub-image; the target size is adapted to an input size supported by a feature extraction layer in a neural network model to be trained; Input the scaled sub-images of each training batch into the neural network model to obtain the sub-classification result of each scaled sub-image; the neural network model includes a feature extraction layer and a scale scaling rate encoding layer connected to the feature extraction layer, the feature extraction layer is used to extract features from the scaled sub-images to obtain a feature map; the scale scaling rate encoding layer is used to encode the scaling rate of the sub-image to obtain a filter function containing encoding information, and filter the feature map by the filter function containing encoding information to obtain a filtered feature map; and determine the sub-classification result based on the filtered feature map; The neural network model is trained in combination with the sub-classification results corresponding to each sub-image and the classification label of each sub-image to obtain the image classification model, and the image classification model is used to classify the target pathological image to be classified.
2. The method according to claim 1, characterized in that The scale scaling rate coding layer encodes the scaling rate of the sub-image to obtain a filter function containing coding information, including: Determining the scaling ratio based on the cropping parameter and the maximum and minimum values of the parameter range; Determine a filter radius of a disk function based on the scaling rate; the disk function is used to perform a low-pass filter on the feature map through a circle with a center point of the feature map as the center and the filter radius as the radius; A disk function is determined based on the filter radius to obtain the filter function.
3. The method according to claim 2, characterized in that The step of determining the filter radius of the disk function based on the scaling rate comprises: Using a preset smoothing function to smooth the scaling rate to obtain a smoothed scaling rate; The smoothed scaling rate is input into a pre-created radius determination function to obtain the filter radius.
4. The method according to claim 3, characterized in that The scaling factor x is expressed by the following formula: Wherein, p represents the cutting parameter, p max represents the maximum value of the parameter range, p min Indicates the minimum value of the parameter range; The smoothing function is expressed by the following formula: t=6x 5 -15x 4 +10x 3 Wherein, x represents the scaling rate, and t represents the scaling rate after smoothing; The radius determination function is expressed by the following formula: R = R0 + t(L / 2-R0); Among them, R represents the filtering radius, R0 represents the preset minimum radius, L represents the size parameter of the feature map, and t represents the scaling rate after smoothing.
5. The method according to claim 2, characterized in that: The filtering of the feature map by a filter function containing encoding information to obtain a filtered feature map includes: Performing a two-dimensional Fourier transform on the characteristic graph to obtain a characteristic spectrum graph; multiplying the characteristic spectrum graph with a disk function having the filtering radius to obtain a filtered characteristic spectrum graph; performing a two-dimensional inverse Fourier transform on the filtered characteristic spectrum graph to obtain the filtered characteristic graph; The disk function with the filter radius is expressed by the following formula: Wherein, (x0, y0) represents the coordinates of the center point of the feature map, (x, y) represents the independent variable of the disk function, R represents the filter radius, and S(x, y) represents the value of the disk function when the independent variable is (x, y); or, Performing a two-dimensional inverse Fourier transform on the disk function having the filtering radius to obtain a transformed disk function; performing a convolution on the feature map and the transformed disk function to obtain the filtered feature map; The transformed disk function is expressed by the following formula: Among them, (x0, y0) represents the coordinates of the center point of the feature map, (x, y) represents the independent variable of the transformed disk function, R represents the filtering radius, I(x, y) represents the value of the transformed disk function when the independent variable is (x, y), and J1 represents a Bessel function of order 1.
6. The method according to claim 1, characterized in that The feature tensor output by the feature extraction layer includes multiple channels, each channel corresponds to a feature map; If the feature tensor has at least two scales, the feature map corresponding to each channel in the feature tensor of each scale is respectively input to the scale scaling rate encoding layer to filter each feature map; or, After the feature tensors of various scales are fused, the fused feature tensors corresponding to each scale are obtained; The feature map corresponding to each channel in the fused feature tensor of each scale is respectively input into the scale scaling rate encoding layer to filter each feature map.
7. The method according to claim 6, characterized in that The feature extraction layer performs feature extraction based on the convolutional layer in the convolutional neural network to obtain the feature tensor, and feature tensors of various scales perform feature fusion based on the cross-stage partial connection CrossStage Partial Connection.
8. The method according to claim 1, characterized in that The step of training the neural network model by combining the sub-classification results corresponding to each sub-image and the classification label of each sub-image to obtain the image classification model includes: Based on the difference between the classification result and the classification label, the model parameters of the neural network model are iteratively updated to obtain the image classification model.
9. An image classification method, characterized in that: The method comprises: Traverse and crop the target pathological image to be classified to obtain a sub-image to be classified; The sub-images to be classified are classified based on a pre-trained image classification model to obtain a sub-classification result corresponding to each sub-image to be classified; wherein the image classification model is trained based on the method for obtaining an image classification model according to any one of claims 1 to 8; Voting is performed on each sub-classification result to obtain the classification result of the target pathological image.
10. The method according to claim 9, characterized in that The target pathological image to be classified is traversed and cut to obtain sub-images to be classified, including: According to the input size supported by the feature extraction layer in the image classification model, the target pathological image is traversed and cropped to obtain the sub-image to be classified.
11. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 10.
12. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
Citation Information
Cited By
Landslide mass identification method, device, equipment, medium and program product
CN121190872A