A method for cell segmentation of M-ROSE images
By constructing a 2D deep-supervised decomposable segmentation network and using an adaptive feedback mechanism to adjust the network layer depth, the problem of uneven pixel segmentation in M-ROSE images was solved, achieving fast and high-precision cell segmentation and improving the efficiency of clinical diagnosis.
Patent Information
- Application Number
- CN202510247841.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-03-04
AI Technical Summary
Existing technologies struggle to achieve fast and accurate cell segmentation in M-ROSE images, especially due to the varying difficulty of pixel segmentation, which leads to an imbalance in computational resources and time for deep learning models, impacting the efficiency of clinicians.
A 2D deep-supervised decomposable segmentation network is constructed. An adaptive feedback mechanism is adopted to adjust the network layer depth according to the difficulty of pixels. The correctly segmented pixel positions are deactivated through the adaptive network layer depth feedback mechanism, so as to achieve accurate pixel segmentation.
It improves the segmentation speed and accuracy of M-ROSE images, reduces the waste of computing resources, lowers the workload of clinicians, and achieves fast and accurate cell segmentation.
Smart Images

Figure CN120182291B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of cell segmentation, and particularly relates to a method for cell segmentation of M-ROSE pictures. BACKGROUND
[0002] Pulmonary infection is a common disease of the respiratory system, which is caused by different pathogens such as bacteria, fungi, viruses, etc. Patients in intensive care units are mostly in the stage of severe pneumonia, and have the characteristics of high morbidity, rapid disease progression, high mortality, poor prognosis, etc. An international study on the prevalence and outcomes of infections in intensive care units showed that infection was the main cause of morbidity and mortality in ICU, and 64% of the infections were pulmonary infections. Therefore, it is urgent in clinical practice to judge the patient's condition change and quickly identify the pathogen to develop a precise treatment plan. However, the most common pathogen detection method in clinical practice is to obtain sputum specimens from patients for smear and culture to identify pathogenic bacteria. This method has the disadvantages of long time consumption and low positive rate of culture. Rapid on-site microbiological evaluation (M-ROSE) obtains alveolar lavage fluid or tracheal aspirate specimens from patients through bronchoscopy for rapid on-site smearing and staining, and makes a preliminary judgment on the infectious pathogens according to the staining and microscopic reading results, and further evaluates the patient's condition and treatment effect through cell background such as cell classification and proportion, so as to guide individualized treatment of patients and improve prognosis.
[0003] The basic steps of M-ROSE include that clinical specimens (alveolar lavage fluid, tracheal aspirate, etc.) are obtained from patients in intensive care units through bedside bronchoscopy. The tracheal aspirate is taken as a typical specimen and is smeared on a glass slide with a suitable thickness. The alveolar lavage fluid can be taken 3-5 ml and placed in a centrifuge. After centrifugation for 10 minutes, the supernatant is discarded, and the remaining specimen is mixed and smeared. After the specimen is dried, staining is performed (1) Diff staining: the glass slide is placed in A liquid (about 30 s), phosphate buffered saline (PBS) for washing, B liquid (about 10 s), and water for washing, and then observed under a microscope after drying. (2) Gram staining: the glass slide is dyed with crystal violet solution for 10 s, washed with water; dyed with iodine solution for 10 s, washed with water; dyed with decolorizing solution (95% ethanol) for about 10 s until the purple color falls off, and then washed with water; dyed with counterstaining solution for 10 s, washed with water, and then observed under a microscope after drying. During the microscopic examination, the whole slide is first observed under a low-power microscope, and then the cell morphology, quantity and pathogenic microorganisms are judged under a high-power microscope.
[0004] The advantages of M-ROSE include:
[0005] (1) Identify infection and non-infection: First, according to the results of identifying the morphology, classification and proportion of cells, judge the quality of the specimen and whether it meets the interpretation standard. The proportion of squamous cells in all cells (excluding red blood cells) in the qualified bronchoalveolar lavage sample under low power is less than 1%, the proportion of columnar epithelial cells is less than 5%, and the proportion of red blood cells is less than 10% (excluding trauma or bleeding related factors). Second, in the qualified specimen, identify infection, colonization and contamination. The proportion of neutrophils in normal bronchoalveolar lavage is <3%. If the proportion of neutrophils is >50%, it often indicates acute lung injury, aspiration pneumonia or purulent infection. If the proportion of neutrophil phagocytosis is >5%, there is also infection, and the pathogenic bacteria are phagocytosed, so that targeted treatment can be carried out.
[0006] (2) Preliminary determination of infectious pathogens: Patients in intensive care units often develop nosocomial infections, and common pathogenic bacteria include Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, Staphylococcus aureus, etc. Gram-positive bacteria appear purple under Gram staining, and Gram-negative bacteria appear red. According to the color, shape and arrangement of the pathogen observed under the microscope, M-ROSE can distinguish Gram-positive bacteria, Gram-negative bacteria and fungi.
[0007] (3) Determine the treatment effect and prognosis
[0008] According to the results of M-ROSE, the pathogenic bacteria of the patient can be preliminarily determined, so that individualized anti-infection treatment plan can be quickly developed to reduce the occurrence of antibiotic abuse. At the same time, M-ROSE has high repeatability and can be performed at different stages of the patient's disease, and the doctor can further determine the patient's condition change according to the results of M-ROSE.
[0009] However, after obtaining the M-ROSE picture of the patient, it is often necessary to rely on professional laboratory experts to read the picture, and the timeliness, rapidity and bedside characteristics of M-ROSE will inevitably lead to this part of the work being completed by the clinical doctors, and the clinical doctors need to spend more time and energy to achieve the same accuracy as the laboratory experts, which undoubtedly increases the workload of the clinical doctors. Therefore, an artificial intelligence technology is needed to replace the clinical doctors to interpret the picture, and through deep learning technology to realize the automatic analysis of the specimen. In order to achieve this effect, the M-ROSE picture must be intelligently segmented first. However, the quality of the M-ROSE picture is usually affected by many factors, such as the sampling process, specimen quality, individual differences of patients, staining level, etc. Therefore, different segmentation requirements exist when segmenting each picture.
[0010] Meanwhile, balancing the speed and accuracy of the network model is a big problem for current segmentation methods. Shallow networks are fast, but difficult to segment difficult signals, and deep networks have improved accuracy, but have huge computational and parameter quantities, and slow operation. Existing deep learning methods often focus on the design of different models, less from the pixel perspective, ignore the difficulty of different pixels, and use the same model for unified processing of all pixels, which seriously overcomputes on a large number of easily segmented pixels, which reduces the operation speed. SUMMARY
[0011] To solve the above technical problems, the present application provides a cell segmentation method for M-ROSE pictures, comprising:
[0012] According to the difficulty of pixel segmentation in the M-ROSE picture, an adaptive feedback mechanism of different difficulty pixels required network layer depth is obtained;
[0013] A 2D deep supervision detachable segmentation network is constructed to output the prediction information of each pixel in the M-ROSE picture for each network layer;
[0014] Based on the adaptive feedback mechanism of network layer depth required by pixels of different difficulty, a pixel difficulty perception adaptive network layer depth feedback mechanism is constructed, the difficulty of pixel segmentation is perceived according to the output result of each network layer, and the neurons at the pixel position correctly segmented by the current layer are inactivated, so as to realize the cell segmentation of the image.
[0015] Preferably, the 2D deep supervision detachable segmentation network adopts an encoding and decoding segmentation method, uses a long-short connection path, combines information of different layer depths, and performs segmentation.
[0016] Preferably, the 2D deep supervision detachable segmentation network is an end-to-end network structure, the input is an image and a segmentation label, and the output is five network prediction results with the same size as the input image, each result corresponding to the output of each network layer.
[0017] Preferably, the 2D deep supervision detachable segmentation network includes 16 convolution network layers, a maximum value pooling layer, and a deconvolution kernel.
[0018] Among them, 11 convolution networks in the convolution network layer use a 3x3 convolution kernel to extract image features, and 5 1x1 convolution layers are used to reduce the final output channel number to 1; the convolution network channel number for feature extraction includes 8, 16 and 32;
[0019] The maximum value pooling layer is used to reduce the input image size by using the maximum value pooling method, reduce the calculation amount, and reduce the mean shift caused by parameter error;
[0020] The deconvolution kernel is a 2*2 deconvolution kernel for changing the down-sampling data into an image of the same size as the input image.
[0021] Preferably, the 2D deep supervision disassembled segmentation network adopts a deep supervision strategy, respectively outputs and calculates a loss function corresponding to the segmentation result of the 5-layer network, and takes all the loss functions as the final loss function.
[0022] Preferably, the process of taking all the loss functions as the final loss function comprises:
[0023] A mixed function composed of the Dice loss and the weighted cross-entropy loss is taken as the segmentation loss function L;
[0024] The corresponding formula expression comprises:
[0025]
[0026] Wherein, the parameter α1 is a variable constant for balancing the loss function L dicek and L cenk ; k represents an arbitrary network layer number, the value is from 1 to 5; j is an index value of an arbitrary pixel V j ; M is the total number of pixels of a 2D image; g j is the input label corresponding to the pixel V j , 1 represents foreground, and 0 represents background; p jk is the foreground prediction probability of the pixel V j output by the k-th network layer, the value is in [0, 1]; and ε is a constant.
[0027] Preferably, the process of perceiving the difficulty of pixel segmentation according to the output result of each network layer comprises:
[0028] A pixel difficulty perception feedback network layer deep mechanism is constructed and fused with the 2D deep supervision disassembled segmentation network to obtain the pixel difficulty perception adaptive network layer deep feedback mechanism.
[0029] Based on the pixel difficulty perception adaptive network layer deep feedback mechanism, the network successively judges whether each pixel has been correctly predicted by each network layer according to the output result of each layer, if the shallow network has realized accurate prediction of a certain pixel, the neuron at the corresponding position is inactivated by using the artificial neuron inactivation mechanism, and the pixel is no longer calculated by the subsequent network layer.
[0030] Preferably, the process of successively judging whether each pixel has been correctly predicted by each network layer according to the output result of each layer comprises:
[0031] Based on the pixel difficulty perception adaptive network layer deep feedback mechanism, the feature map F k1 ,k∈[1,5] and the output result Fk2 To perform the operation, first normalize the F of the output result. k2 The value is used to obtain the predicted probability map of the signal. k ∈[0,1]; Based on the predicted probability of each pixel v, determine whether it has been accurately predicted. If so, deactivate the neuron at the corresponding position of the pixel according to the feedback result.
[0032] Then, based on the inactivation judgment value, a mask image of the corresponding network layer is constructed. k By using the feature map F of each network layer k1 With mask k Multiply to update the feature map F of each network layer k1 Select neurons with an inactivation probability of 0;
[0033] Then, the normalized updated F k1 To ensure that after selective inactivation, the network layers have the same mean and variance;
[0034] Repeat the operation until the value of k changes from 1 to 5, completing the output of all network layers.
[0035] Preferably, the formula for the inactivation judgment value is:
[0036]
[0037] Where, p th This is the probability feedback threshold.
[0038] Compared with the prior art, the present invention has the following advantages and technical effects:
[0039] This invention employs an adaptive feedback mechanism to determine the required network depth for pixels of varying segmentation difficulty in M-ROSE images. A 2D deep-supervised decomposable segmentation network is constructed, outputting the prediction information for each pixel in the M-ROSE image from each network layer. Based on this adaptive feedback mechanism, a pixel difficulty-aware adaptive network depth feedback mechanism is built. This mechanism perceives the segmentation difficulty of each pixel based on the output of each network layer and feeds it back to the 2D deep-supervised decomposable segmentation network. This deactivates neurons at the locations of correctly segmented pixels in the current layer, achieving cell segmentation of the image. Building upon the aforementioned ultra-fast segmentation network, this invention rapidly segments cells within M-ROSE images, improving segmentation speed while balancing segmentation accuracy and precision. Attached Figure Description
[0040] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0041] Figure 1 This is a schematic diagram of the structure of a 2D deep-supervised decomposable segmentation network according to an embodiment of the present invention. Detailed Implementation
[0042] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0043] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0044] like Figure 1 As shown, this embodiment provides a method for cell segmentation of M-ROSE images, including:
[0045] An adaptive feedback mechanism is used to determine the network layer depth required for pixels of different difficulty levels based on the segmentation difficulty of pixels in the M-ROSE image.
[0046] Construct a 2D deep supervised decomposable segmentation network and output the prediction information of each layer of the network for each pixel in the M-ROSE image;
[0047] Based on the adaptive feedback mechanism of the network layer depth required for pixels of different difficulty, a pixel difficulty-aware adaptive network layer depth feedback mechanism is constructed. The difficulty of pixel segmentation is perceived according to the output results of each network layer, and fed back to the 2D deep supervised decomposable segmentation network. The neurons at the pixel positions that have been correctly segmented in the current layer are deactivated, thereby realizing cell segmentation of the image.
[0048] Furthermore, this embodiment constructs a concise 2D deep-supervised decomposable segmentation network (High Speed DeepSupervision Net, abbreviated as HSDS-Net). HSDS-Net employs an encoding and decoding segmentation method, utilizing long and short connection paths and combining information from different layer depths to improve segmentation accuracy. Its structure is as follows: Figure 1 As shown.
[0049] Furthermore, the 2D deep supervised decomposable segmentation network is an end-to-end network structure. The input is an image and segmentation labels, and the output is five network prediction results of the same size as the input image. Each result corresponds to the output of each layer of the network.
[0050] Furthermore, the 2D deep supervised decomposable segmentation network includes 16 convolutional network layers, max pooling layers, and deconvolution kernels;
[0051] Among them, the 11 convolutional network layers in the convolutional network layer use 3×3 convolutional kernels for image feature extraction, and 5 1×1 convolutional layers are used to reduce the final output channel number to 1; the number of channels of the convolutional network for feature extraction includes 8, 16 and 32.
[0052] Max pooling layers are used to reduce the size of the input image by using the max pooling method, thereby reducing the amount of computation and reducing mean shift caused by parameter errors.
[0053] The deconvolution kernel is a 2×2 deconvolution kernel, used to transform the downsampled data into an image the same size as the input image.
[0054] Furthermore, the 2D deep-supervised decomposable segmentation network adopts a deep-supervised strategy, outputting and calculating the loss functions corresponding to the segmentation results of the 5-layer network respectively, and using the sum of all loss functions as the final loss function.
[0055] Furthermore, the process of using all loss functions and sum as the final loss function includes:
[0056] A hybrid function consisting of Dice loss and weighted cross-entropy loss is used as the segmentation loss function L;
[0057] The corresponding formula expressions include:
[0058]
[0059] Wherein, parameter α1 is a variable constant used to balance the loss function L dicek and L cenk In this example, the value is set to 0.5; k represents any number of network layers, ranging from 1 to 5; j represents any pixel V. j The index value; M is the total number of pixels in the 2D image; g j For pixel V j The corresponding input labels are 1 for foreground and 0 for background; p jk For the k-th layer network, the pixel V j The probability of the prospect is in the range [0, 1]; ε is a constant with a value of 1.
[0060] Furthermore, the process of perceiving the difficulty of pixel segmentation based on the output of each network layer includes:
[0061] A pixel difficulty-aware feedback network layer depth mechanism is constructed and fused with a 2D deep-supervised decomposable segmentation network to obtain a pixel difficulty-aware adaptive network layer depth feedback mechanism.
[0062] Based on the pixel difficulty-aware adaptive network layer depth feedback mechanism, the network sequentially judges whether each layer has correctly predicted each pixel based on the output results of each layer. If a shallow layer has achieved accurate prediction of a certain pixel, the neurons at the specific location are deactivated using the artificial neuron deactivation mechanism, and subsequent network layers will no longer calculate that pixel.
[0063] Furthermore, the process by which the network determines whether each layer has correctly predicted each pixel based on the output of each layer includes:
[0064] Based on a pixel-difficulty-aware adaptive deep feedback mechanism, the feature map F of each network layer is processed. k1 ,k∈[1,5] and output result F k2 To perform the operation, first normalize the F of the output result. k2 The value is used to obtain the predicted probability map of the signal. k ∈[0,1]; Based on the predicted probability of each pixel v, determine whether it has been accurately predicted. If so, deactivate the neuron at the corresponding position of the pixel based on the feedback result.
[0065] Then, based on the inactivation judgment value, a mask image of the corresponding network layer is constructed. k By using the feature map F of each network layer k1 With mask k Multiplication updates the feature map F of each network layer k1 Select neurons with an inactivation probability of 0;
[0066] F k1 =F k1 ·mask k
[0067] Then, the normalized updated F k1 This ensures that after selective inactivation, the network layers have the same mean and variance;
[0068] F k1 =F k1 ·count(mask k ) / count_ones(mask k )
[0069] Repeat the operation until the value of k changes from 1 to 5, completing the output of all network layers.
[0070] Furthermore, the formula for the inactivation judgment value is as follows:
[0071]
[0072] Where, p thThis is the probability feedback threshold. To ensure the accuracy of pixel estimation, in this embodiment, p... th The value is greater than 0.95.
[0073] This embodiment, based on the aforementioned ultra-fast segmentation network, performs rapid segmentation of cells within M-ROSE images, improving segmentation speed while balancing segmentation precision and accuracy.
[0074] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for cell segmentation of M-ROSE images, characterized in that, include: An adaptive feedback mechanism is used to determine the network layer depth required for pixels of different difficulty levels based on the segmentation difficulty of pixels in the M-ROSE image. Construct a 2D deep supervised decomposable segmentation network and output the prediction information of each layer of the network for each pixel in the M-ROSE image; Based on the adaptive feedback mechanism of the network layer depth required for pixels of different difficulty, a pixel difficulty-aware adaptive network layer depth feedback mechanism is constructed. The difficulty of pixel segmentation is perceived according to the output results of each network layer, and fed back to the 2D deep supervised decomposable segmentation network. The neurons at the pixel positions that have been correctly segmented in the current layer are deactivated, thereby realizing cell segmentation of the image. The process of perceiving the difficulty of pixel segmentation based on the output of each network layer includes: A pixel difficulty-aware feedback network layer depth mechanism is constructed and fused with the 2D deep supervised decomposable segmentation network to obtain the pixel difficulty-aware adaptive network layer depth feedback mechanism. Based on the pixel difficulty-aware adaptive network layer depth feedback mechanism, the network sequentially determines whether each layer has correctly predicted each pixel based on the output results of each layer. If the shallow network has achieved accurate prediction of a certain pixel, the corresponding neuron is deactivated using the artificial neuron deactivation mechanism, and subsequent network layers will no longer calculate that pixel.
2. The method according to claim 1, characterized in that, The 2D deep-supervised decomposable segmentation network employs an encoding and decoding segmentation method, utilizing long and short connection paths and combining information from different depths for segmentation.
3. The method according to claim 1, characterized in that, The 2D deep supervised decomposable segmentation network is an end-to-end network structure. The input is an image and segmentation labels, and the output is five network prediction results of the same size as the input image. Each result corresponds to the output of each layer of the network.
4. The method according to claim 1, characterized in that, The 2D deep supervised decomposable segmentation network includes 16 convolutional network layers, max pooling layers, and deconvolution kernels; The convolutional network layer consists of 11 layers with 3×3 convolutional kernels for image feature extraction, and 5 layers with 1×1 convolutions to reduce the final output channel number to 1. The number of channels in the feature extraction convolutional network includes 8, 16, and 32. The max pooling layer is used to reduce the size of the input image by using the max pooling method, thereby reducing the amount of computation and reducing the mean shift caused by parameter errors. The deconvolution kernel is a 2×2 deconvolution kernel, used to transform the downsampled data into an image of the same size as the input image.
5. The method according to claim 1, characterized in that, The 2D deep-supervised decomposable segmentation network adopts a deep-supervised strategy, outputting and calculating the loss functions corresponding to the segmentation results of the 5-layer network respectively, and using the sum of all loss functions as the final loss function.
6. The method according to claim 5, characterized in that, The process of using all loss functions as the final loss function includes: A mixture of Dice loss and weighted cross-entropy loss is used as the segmentation loss function. ; The corresponding formula expressions include: Among them, parameters It is a variable constant used to balance the loss function. and ; This represents any number of network layers, with values ranging from 1 to 5. For any pixel The index value; The total number of pixels in a 2D image; For pixels The corresponding input labels are 1 for foreground and 0 for background; For the first Layer network for pixels The probability of the predicted future, with a value in [0, 1]; It is a constant.
7. The method according to claim 1, characterized in that, The process by which the network determines whether each layer has correctly predicted each pixel based on the output of each layer includes: Based on the pixel difficulty-aware adaptive network layer deep feedback mechanism, the feature maps of each network layer are processed respectively. and output results To perform this operation, the output result is first normalized. The value is used to obtain the predicted probability map of the signal. Based on each pixel The predicted probability is used to determine whether it has been accurately predicted. If so, the neuron at the corresponding position of the pixel is deactivated based on the feedback result. Then, based on the inactivation judgment value, a mask map of the corresponding network layer is constructed. By using the feature maps of each network layer With mask image Multiplication updates the feature maps of each network layer Select neurons with an inactivation probability of 0; Then the normalized update To ensure that after selective inactivation, the network layers have the same mean and variance; Repeat the operation until The value changes from 1 to 5 to complete the output of all network layers.
8. The method according to claim 7, characterized in that, The formula for the inactivation judgment value is as follows: in, This is the probability feedback threshold.
Citation Information
Patent Citations
Method and system for melanoma image tissue segmentation based on deep neural network
CN108510502A
Image segmentation method based on MaskFormer network
CN118470323A