Microscopic image bone marrow cell recognition method based on feature mixing

By introducing shallow feature mixing module into the Mask R-CNN network, the problem of degradation of recognition accuracy when processing bone marrow cell images of different color styles is solved, and higher recognition accuracy and generalization capabilities are achieved.

CN120070403AActive Publication Date: 2025-05-30SOUTH CHINA UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510222875.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

When processing bone marrow cell images under a microscope, existing deep learning models are affected by differences in equipment model, dye type and stain quality, resulting in a decrease in recognition accuracy and lack of generalization ability to different color styles.

Method used

Improve the feature extraction module of the Mask R-CNN network, and add shallow feature mixing modules, including the current furthest point sampling module (CFPS), style feature mixing module (MixStyle) and internal channel mixing module (InterMix) to improve the network's recognition and generalization of different microscope models and staining conditions.

Benefits of technology

The feature mixing module generates more color style features outside the distribution, which improves the recognition accuracy and generalization ability of the model, and can more accurately identify bone marrow cells of different color styles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070403A_ABST
    Figure CN120070403A_ABST
Patent Text Reader

Abstract

The invention discloses a microscopic image bone marrow cell recognition method based on feature mixing, which comprises the following steps: acquiring a marked bone marrow cell microscopic image to construct a data set, and dividing the data set into a training set and a test set; performing data enhancement on the images in the training set to obtain an enhanced training set; the enhanced training set is sent to the improved Mask R-CNN network for training, and an optimal model is obtained; and inputting the test set into the optimal model to obtain prediction information, including position frame, category and segmentation mask information of each bone marrow cell, drawing the information on an original image, and marking the predicted bone marrow cell category above the position frame, thereby completing identification of the bone marrow cells. According to the invention, the identification generalization of bone marrow cell images with various color styles is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target recognition of bone marrow blood cells, and in particular to a method for identifying bone marrow cells in microscopic images based on feature mixing. Background Art

[0002] The challenges of automatically identifying bone marrow cells mainly stem from the diversity of cell types and the density of distribution in microscopic images. Although modern deep learning techniques have made remarkable progress in the field of image recognition, when dealing with cell images under a microscope, due to differences in equipment models, types of stains, and staining quality, the variation in color styles may greatly affect the recognition accuracy.

[0003] Existing deep learning models usually perform well on specific training sets but often lack generalization ability. That is, when faced with new color styles different from the training data set, their performance drops sharply. This situation is particularly obvious in bone marrow cell recognition because the differences in image color styles caused by different experimental conditions and equipment settings make it difficult for the model to accurately identify cell types in new samples. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for identifying bone marrow cells in microscopic images based on feature mixing. Batch image feature layer statistical information interaction fusion is performed in the shallow network of the feature extraction module of the improved Mask R-CNN network, thereby improving the recognition generalization ability of the network for color style differences caused by different microscope models, staining conditions, etc.

[0005] To achieve the above purpose, the technical solution provided by the present invention is: A method for identifying bone marrow cells in microscopic images based on feature mixing. This method realizes the target recognition of bone marrow cells based on the improved Mask R-CNN network. The improved Mask R-CNN network improves the feature extraction module of the traditional Mask R-CNN network, that is, a shallow feature mixing module is added to the shallow network of the feature extraction module. The shallow feature mixing module is improved on the basis of the style feature mixing module called MixStyle. The specific improvement is as follows: ① A current farthest point sampling module called CFPS is added before MixStyle. The CFPS calculates the Euclidean distance between the shallow features of pairwise image samples and selects the sample with the largest Euclidean distance from the current sample to provide as input to MixStyle; ② An internal channel mixing module called InterMix is added after MixStyle. The InterMix performs random style feature mixing between the respective feature channels of the sample using MixStyle.

[0006] The specific implementation of the microscopic image bone marrow cell recognition method includes the following steps:

[0007] 1) Obtain a labeled microscopic image of bone marrow cells to construct a dataset, where the annotation of each image includes the position bounding box, category, and segmentation mask of each cell, and then divide the dataset into a training set and a test set;

[0008] 2) Perform data augmentation on the images in the training set to obtain an augmented training set after shape and color transformation;

[0009] 3) Feed the augmented training set into the improved Mask R-CNN network. Generate more out-of-distribution color style features through the shallow feature mixing module of the feature extraction module, and obtain the feature information of different types of bone marrow cells through the feature extraction module. Perform feature multi-scale fusion on the extracted feature information with the Feature Pyramid Network (FPN), then obtain the regions of interest through the Region Proposal Network (RPN), and finally input them into the classifier, bounding box regression module, and mask segmentation module of the improved Mask R-CNN network to obtain the prediction results of bone marrow cells; among them, in backpropagation, use the binary cross-entropy and Smooth-L1 loss functions to calculate the loss values of classification and regression, and after multiple iterations until the loss value is minimized, obtain the optimal model;

[0010] 4) Input the test set into the optimal model to obtain prediction information, including: the position bounding box, category, and segmentation mask information of each bone marrow cell, draw this information on the original image, and label the predicted bone marrow cell category above the position bounding box, thus completing the recognition of bone marrow cells.

[0011] Furthermore, the data augmentation includes: randomly adjusting the contrast of the image, histogram equalization, inverting pixel values, randomly rotating the image, randomly adjusting the brightness of the image, randomly adjusting the sharpness of the image, and flipping the image in the horizontal or vertical direction.

[0012] Furthermore, the shallow feature mixing module includes:

[0013] The current farthest point sampling module, called CFPS: This module calculates the Euclidean distance between the shallow features of pairwise images in the same batch, and then for the current base sample X, selects the sample point C that is farthest from it as the subsequent feature fusion object. CFPS is represented by the following formula:

[0014]

[0015] In the formula, X represents the currently selected base sample, i is the index of the current sample X in the batch, X j represents the jth sample, D(X,X j ) represents the Euclidean distance between X and X jThe Euclidean distance between, argmax represents the sample when the function D(X, X j ) obtains the maximum value;

[0016] The style feature mixing module, called MixStyle: This module regularizes the Mask R-CNN training by perturbing the style information of the source domain instances. By calculating the feature statistics of two instances and mixing them with random convex weights, the purpose of style feature mixing is achieved. Its specific process is represented by the following formula:

[0017] γ mix = λσ(f)+(1 - λ)σ(f′)

[0018] β mix = λμ(f)+(1 - λ)μ(f′)

[0019]

[0020] In the formula, γ mix represents the standard deviation obtained after shallow feature mixing, β mix represents the mean obtained after shallow feature mixing, f and f′ respectively represent two groups of sample feature maps for mixing, μ and σ respectively represent the functions for calculating the channel mean and standard deviation, and λ is a random weight sampled from the beta distribution;

[0021] The internal channel mixing module, called InterMix: This module randomly selects two channels in each feature channel of the sample itself for feature mixing, and the mixing method uses MixStyle to obtain more out-of-distribution color style features.

[0022] Furthermore, the shallow feature mixing module specifically performs the following operations:

[0023] Use CFPS to select the image feature C' with the largest style difference from the batch of images, then perform style mixing on C' and the current image feature X' to obtain the feature M, and then use InterMix to perform internal feature mixing on the internal channel features of the newly generated feature M to generate more implicit color style features; among them, when improving the Mask R-CNN network training, the shallow feature mixing module is inserted between the shallow networks of the feature extraction module of the improved Mask R-CNN network. During the training process, some color style features that the improved Mask R-CNN network has not seen are generated with the style mixing transformation of the shallow feature mixing module, thereby improving the generalization and robustness of the improved Mask R-CNN network.

[0024] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0025] 1. Based on the traditional style transformation method, the present invention proposes to select mixed samples by the current farthest point sampling module (CFPS), which improves the difference and diversity of feature mixing.

[0026] 2. The present invention proposes an internal channel mixing module (InterMix), which further enhances the generated implicit style features.

[0027] 3. Based on the traditional object detection model method, the present invention incorporates the improved style feature mixing module (MixStyle), and proposes an improved Mask R-CNN network based on shallow feature mixing for object detection of unlabeled bone marrow blood cells, which has high recognition accuracy and out-of-distribution generalization. Brief Description of the Drawings

[0028] Figure 1 It is a schematic diagram of the logic flow of the method of the present invention.

[0029] Figure 2 It is a structural diagram of the improved Mask R-CNN network; in the figure, ResBlock is the residual module in the feature extraction network ResNet50, RPN is the region recommendation module, and FPN is the feature pyramid module. Detailed Embodiments

[0030] The present invention will be further described in detail below in conjunction with the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.

[0031] In this embodiment, 20-class cell recognition of pronormoblasts, normoblasts, metanormoblasts, myeloblasts, promyelocytes, abnormal promyelocytes, neutrophilic myelocytes, neutrophilic metamyelocytes, neutrophilic band cells, neutrophilic segmented cells, eosinophilic myelocytes, eosinophilic segmented cells, promonocytes, monocytes, original lymphocytes, lymphocytes, plasma cells, myeloma cells, smear cells, and other cells in bone marrow cells is performed.

[0032] Such as Figure 1As shown in the figure, this embodiment discloses a microscopic image bone marrow cell recognition method based on feature mixing. This method realizes the target recognition of bone marrow cells based on an improved Mask R-CNN network. The improved Mask R-CNN network improves the feature extraction module of the traditional Mask R-CNN network, that is, a shallow feature mixing module is added to the shallow network of the feature extraction module. The shallow feature mixing module is improved on the basis of the style feature mixing module (which can be called MixStyle). The specific improvements are as follows: ① A current farthest point sampling module (which can be called CFPS) is added before MixStyle. The CFPS calculates the Euclidean distance between pairwise image samples and selects the sample with the largest Euclidean distance from the current sample to provide it to MixStyle as input; ② An internal channel mixing module (which can be called InterMix) is added after MixStyle. The InterMix randomly mixes the style features of each feature channel of the sample using MixStyle. The specific steps of this method are as follows:

[0033] 1) Obtain a labeled microscopic image of bone marrow cells to construct a dataset. The annotation of each image includes the position bounding box, category, and segmentation mask of each cell. Then, divide the dataset into a training set and a test set;

[0034] The categories are respectively 20 category attributes including pronormoblast, normoblast, orthochromatic normoblast, myeloblast, promyelocyte, abnormal promyelocyte, neutrophilic myelocyte, neutrophilic metamyelocyte, neutrophilic stab granulocyte, neutrophilic segmented granulocyte, eosinophilic myelocyte, eosinophilic segmented granulocyte, promonocyte, monocyte, original lymphocyte, lymphocyte, plasma cell, myeloma cell, smear cell, and other cells. The annotation contains the position bounding box, category, and segmentation mask of each cell.

[0035] 2) Perform data augmentation on the images in the training set to obtain an augmented training set after shape and color transformation. The data augmentation technique aims to simulate the variations that may be encountered under various actual imaging conditions, thereby enhancing the adaptability of the model to different imaging settings. This step is the key to improving the stability and accuracy of the model in practical applications.

[0036] The data augmentation includes: a. randomly adjusting the contrast of the image; b. histogram equalization; c. inverting the pixel values, i.e., changing the pixel value x to 255 - x; d. randomly rotating by an angle from 0 to 180 degrees and filling the blank areas with 0; e. randomly selecting a value from 0 to 4 and setting to 0 the bit positions of each pixel point in the three RGB channels that are lower than x; f. randomly selecting a value from 0 to 255 and inverting those higher than x; g. randomly adjusting the image brightness; h. randomly adjusting the sharpness of the image; i. with a probability of 0.5, flipping in the horizontal or vertical direction.

[0037] 3) Feed the augmented training set into the improved Mask R-CNN network. Generate more out-of-distribution color style features through the shallow feature mixing module of the feature extraction module, and obtain the feature information of different types of bone marrow cells through the feature extraction module. Perform feature multi-scale fusion on the extracted feature information with the Feature Pyramid Network (FPN) module, then obtain the regions of interest through the Region Proposal Network (RPN) module, and finally input it into the classifier, bounding box regression module, and mask segmentation module of the improved Mask R-CNN network to obtain the prediction results of bone marrow cells; among them, in backpropagation, use binary cross-entropy and Smooth-L1 loss functions to calculate the loss values of classification and regression, and after multiple iterations until the loss value is minimized, obtain the optimal model.

[0038] As Figure 2 shown, the shallow feature mixing module described above includes the following several modules:

[0039] Current Furthest Point Sampling module (which can be called CFPS): This module calculates the Euclidean distance between the shallow features of images in the same batch pairwise. Then, for the current base sample X, select the sample point C that is furthest from it as the subsequent feature fusion object. CFPS is represented by the following formula:

[0040]

[0041] In the formula, X represents the currently selected base sample, i is the index of the current sample X in the batch, X j represents the jth sample, D(X, X j ) represents the Euclidean distance between X and X j , and argmax represents the sample when the function D(X, X j ) reaches the maximum value;

[0042] Style Feature Mixing module (which can be called MixStyle): This module regularizes the Mask R-CNN training by perturbing the style information of source domain instances, calculates the feature statistics of two instances, and mixes them through random convex weights to achieve the purpose of style feature mixing. Its specific process is represented by the following formula:

[0043] γ mix = λσ(f) + (1 - λ)σ(f′)

[0044] β mix = λμ(f) + (1 - λ)μ(f′)

[0045]

[0046] In the formula, γ mix represents the standard deviation obtained after shallow feature mixing, and β mix represents the mean value obtained after shallow feature mixing. f and f′ respectively represent two groups of sample feature maps for mixing, μ and σ respectively represent the functions for calculating the channel mean value and standard deviation, and λ is a random weight sampled from the beta distribution.

[0047] Internal Channel Mixing Module (which can be called InterMix): In the existing style fusion methods in the work, the feature style information of each channel inside the sample features is not considered. However, there are also a lot of pixel difference information hidden between different channels. Making good use of this information helps us generate more potential new color styles and at the same time can limit the color style change beyond the existing range in reality. Therefore, we propose the Internal Channel Mixing Module (which can be called InterMix). This module randomly selects two channels from each feature channel of the sample itself for feature mixing, and the mixing method uses the MixStyle method to obtain more implicit style features.

[0048] The overall process of the shallow feature mixing module is as follows: From the batch of images, use the current farthest point sampling module (which can be called CFPS) to select the image feature C' with the largest style difference, and then mix the style of C' with the current image feature X' to obtain the feature M. Then use the internal channel mixing module (which can be called InterMix) to perform internal feature mixing on the internal channel features of the newly generated feature M, so as to generate more implicit color style features. Among them, when improving the training of the Mask R-CNN network, insert the shallow feature mixing module between the shallow networks of the improved Mask R-CNN network. During the training process, with the style mixing transformation of the shallow feature mixing module, some color style features that the improved Mask R-CNN network has not seen will be generated, so as to improve the generalization and robustness of the improved Mask R-CNN network.

[0049] 4) Input the test set into the optimal model to obtain prediction information, including: the position bounding box, category, and segmentation mask information of each bone marrow cell. Draw this information on the original image, and mark the predicted bone marrow cell category above the position bounding box, thus completing the recognition of bone marrow cells.

[0050] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A method for identifying bone marrow cells in microscopic images based on feature mixing, characterized in that: The method realizes the target recognition of bone marrow cells based on the improved Mask R-CNN network. The improved Mask R-CNN network improves the feature extraction module of the traditional Mask R-CNN network, that is, a shallow feature mixing module is added to the shallow network of the feature extraction module. The shallow feature mixing module is improved on the basis of the style feature mixing module. The style feature mixing module is called MixStyle. The specific improvements are as follows: ① Add the current farthest point sampling module before MixStyle. The current farthest point sampling module is called CFPS. The CFPS calculates the Euclidean distance between the shallow features of the two image samples, and selects the sample with the largest Euclidean distance to the current sample and provides it to MixStyle as input; ② Add an internal channel mixing module after MixStyle. The internal channel mixing module is called InterMix. The InterMix uses MixStyle to randomly mix the style features between the feature channels of the sample; The specific implementation of the microscopic image bone marrow cell identification method includes the following steps: 1) Obtain annotated bone marrow cell microscopic images to construct a dataset, where the annotations of each image include the location border, category, and segmentation mask of each cell, and then divide the dataset into a training set and a test set; 2) Perform data enhancement on the images in the training set to obtain an enhanced training set after shape and color transformation; 3) The enhanced training set is sent to the improved Mask R-CNN network, and more out-of-distribution color style features are generated through the shallow feature mixing module of the feature extraction module, and the feature information of different types of bone marrow cells is obtained through the feature extraction module. The extracted feature information is fused with the feature pyramid module FPN for multi-scale features, and then the region of interest is obtained through the region recommendation module RPN. Finally, it is input into the classifier, bounding box regression module and mask segmentation module of the improved Mask R-CNN network to obtain the prediction results of bone marrow cells; wherein, the binary cross entropy and Smooth-L1 loss function are used in the back propagation to calculate the loss values ​​of classification and regression, and the optimal model is obtained after multiple iterations until the loss value is minimized; 4) Input the test set into the optimal model to obtain prediction information, including: the location border, category and segmentation mask information of each bone marrow cell. Draw this information on the original image and mark the predicted bone marrow cell category above the location border to complete the identification of bone marrow cells.

2. The method for identifying bone marrow cells in microscopic images based on feature mixing according to claim 1, characterized in that: The data enhancement includes: randomly adjusting the contrast of the image, histogram equalization, inverting pixel values, randomly rotating the image, randomly adjusting the brightness of the image, randomly adjusting the sharpness of the image, and flipping the image in the horizontal or vertical direction.

3. The method for identifying bone marrow cells in microscopic images based on feature mixing according to claim 1, characterized in that: The shallow feature mixing module includes: The current farthest point sampling module, called CFPS: This module calculates the Euclidean distance between the shallow features of the same batch of images, and then selects the sample point C farthest from the current base sample X as the subsequent feature fusion object. CFPS is expressed by the following formula: In the formula, X represents the currently selected base sample, i is the index of the current sample X in the batch, and X j represents the jth sample, D(X,X j ) represents X and X j The Euclidean distance between them, argmax means that the function D(X,X j ) The sample at the maximum value; Style feature mixing module, called MixStyle: This module normalizes Mask R-CNN training by perturbing the style information of the source domain instance, calculates the feature statistics of the two instances, and mixes them with random convex weights to achieve the purpose of style feature mixing. The specific process is expressed by the following formula: c mix =λσ(f)+(1-λ)σ(f′) b mix =λμ(f)+(1-λ)μ(f′) In the formula, γ mix represents the standard deviation after shallow feature mixing, β mix represents the average value after shallow feature mixing, f and f′ represent the two sets of sample feature maps used for mixing, μ and σ represent the functions used to calculate the channel average and standard deviation, respectively, and λ is the random weight sampled from the beta distribution; Internal channel mixing module, called InterMix: This module randomly selects two channels from each feature channel of the sample itself for feature mixing, and the mixing method uses MixStyle to obtain more out-of-distribution color style features.

4. The method for identifying bone marrow cells in microscopic images based on feature mixing according to claim 1, characterized in that: The shallow feature mixing module specifically performs the following operations: CFPS is used to select the image feature C' with the largest style difference from the batch images, and then C' is style-mixed with the current image feature X' to obtain feature M, and then InterMix is ​​used to perform internal feature mixing on the internal channel features of the newly generated feature M, thereby generating more implicit color style features; wherein, when training the improved Mask R-CNN network, a shallow feature mixing module is inserted between the shallow networks of the feature extraction module of the improved Mask R-CNN network. During the training process, some color style features that have not been seen by the improved Mask R-CNN network are generated as the style mixing transformation of the shallow feature mixing module occurs, thereby improving the generalization and robustness of the improved Mask R-CNN network.

Citation Information

Patent Citations

  • Multi-modal microscopic image cell segmentation method based on convolutional neural network

    CN116229457A

  • Bone marrow cell detection system based on improved Mask R-CNN

    CN117649657A

  • Systems and methods for automatically classifying cell types in medical images

    US20220335736A1