Cell segmentation method for M-ROSE picture
By applying the 2D deep supervision decomposed segmentation network and pixel difficulty perception adaptive network layer deep feedback mechanism on M-ROSE pictures, the problem of insufficient cell segmentation speed and accuracy of M-ROSE pictures in the prior art is solved, and faster and more accurate cell segmentation is achieved, reducing the work burden of clinicians.
Patent Information
- Application Number
- CN202510247841.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The prior art is difficult to quickly and accurately perform cell segmentation of M-ROSE pictures, which leads to clinicians spending more time and energy on image interpretation, which increases work burden.
The 2D deep supervision can decompose the segmentation network and pixel difficulty perception adaptive network layer depth feedback mechanism is adopted, and the network layer depth is adaptively adjusted according to the difficulty of segmentation of pixels in the M-ROSE picture to realize intelligent cell segmentation of M-ROSE pictures.
It improves the speed and accuracy of cell segmentation of M-ROSE pictures, reduces the work burden of clinicians, and achieves faster and more accurate pathogen detection.
Smart Images

Figure CN120182291A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of cell segmentation, and particularly relates to a method for cell segmentation of M-ROSE images. Background Art
[0002] Pulmonary infection is a common disease of the respiratory system, mostly caused by different pathogens such as bacteria, fungi, viruses, etc. Patients in the intensive care unit are mostly in the stage of severe pneumonia, with characteristics such as high incidence, rapid disease progression, high fatality rate, and poor prognosis. An international study on the prevalence and outcomes of infections in intensive care units shows that infection is the main cause of morbidity and mortality in ICU wards, and 64% of the infections are pulmonary infections. Therefore, there is an urgent clinical need to timely judge the changes in the patient's condition and quickly identify the pathogen so as to formulate a precise treatment plan. However, the most common current clinical method for pathogen detection is to obtain a sputum specimen from the patient for smear and culture to identify the pathogenic bacteria. This method has the disadvantages of long time consumption and low positive culture rate. Rapid on-site microbiological evaluation (M-ROSE) obtains specimens of bronchoalveolar lavage fluid or tracheal aspirate from the patient through a bronchoscope for rapid on-site specimen preparation and staining, makes a preliminary judgment on the infectious pathogen based on the staining and microscopic examination results, and further evaluates the patient's condition and treatment effect through cell background such as cell classification count and ratio, so as to guide the patient to receive individualized treatment and improve the prognosis.
[0003] The basic steps of M-ROSE include that patients in the intensive care unit often obtain clinical specimens (bronchoalveolar lavage fluid, tracheal aspirate, etc.) through a bedside bronchoscope. For tracheal aspirate, take a typical specimen and smear it evenly on a glass slide with appropriate thickness. For bronchoalveolar lavage fluid, 3 - 5 ml can be taken and placed in a centrifuge. After centrifuging for 10 minutes, discard the supernatant. Mix the remaining specimen and then make a smear. After the specimen is air-dried, perform staining: (1) Diff staining: Place the glass slide successively in solution A (about 30 s), rinse with phosphate buffer solution (PBS), solution B (about 10 s), rinse with water, and air-dry before microscopic examination. (2) Gram staining: Add crystal violet solution to the glass slide and stain for 10 s, then rinse the stain off with water; add iodine stain for 10 s and rinse the stain off with water; add decolorizing solution (95% ethanol) for about 10 s until the purple color fades, then wash with water; add counterstain for 10 s, wash with water, and air-dry before microscopic examination. During the microscopic examination process, first observe the whole slide under the low-power microscope, and then switch to the high-power microscope to judge the cell morphology, quantity, pathogenic microorganisms, etc.
[0004] The advantages of M-ROSE include:
[0005] (1) Identification of infection and non-infection: First, based on the results of identifying the morphology, classification, and proportion of cells, the quality of the specimen and whether it meets the interpretation criteria are judged. In a qualified bronchoalveolar lavage fluid sample, the proportion of squamous cells among all cells (excluding red blood cells) under low-power microscopy is less than 1%, the proportion of columnar epithelial cells is less than 5%, and the proportion of red blood cells is less than 10% (excluding trauma or bleeding-related factors). Secondly, the slides are examined in qualified specimens to distinguish infection, colonization, and contamination. The proportion of neutrophils in normal bronchoalveolar lavage fluid is <3%. If the proportion of neutrophils >50%, it often indicates acute lung injury, aspiration pneumonia, or suppurative infection. If the neutrophil phagocytosis phenomenon >5% also exists, there is an infection, and the phagocytosed pathogen is the pathogenic bacterium, thus enabling targeted treatment.
[0006] (2) Preliminary determination of the infecting pathogen: Nosocomial infections often occur in patients in the intensive care unit. Common pathogenic bacteria include Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, Staphylococcus aureus, etc. Under Gram staining, Gram-positive bacteria appear purple, and Gram-negative bacteria appear red. Based on observing the color, morphology, and arrangement of pathogens under the microscope, M-ROSE can distinguish Gram-positive bacteria, Gram-negative bacteria, and fungi.
[0007] (3) Judging the treatment effect and prognosis
[0008] Based on the M-ROSE results, the pathogenic bacteria of the patient can be preliminarily determined, thus quickly formulating an individualized anti-infection treatment plan and reducing the occurrence of antibiotic abuse. At the same time, M-ROSE has high repeatability and can be carried out at different stages of the patient's disease. According to the M-ROSE results, it can assist doctors in further judging the changes in the patient's condition.
[0009] However, after obtaining the patient's M-ROSE pictures, it often relies on professional laboratory experts to read the slides. The timeliness, rapidity, and bedside characteristics of M-ROSE will inevitably lead to most of this work being completed by clinicians. Clinicians need to spend more time and energy to achieve the same accuracy as laboratory experts, undoubtedly increasing the workload of clinicians. Therefore, an artificial intelligence technology is needed to replace clinicians for picture interpretation and realize automatic analysis of specimens through deep learning technology. To achieve this intelligent effect, first, the M-ROSE pictures must be intelligently segmented. However, the quality of M-ROSE pictures is usually affected by many factors, such as the sampling process, specimen quality, patient individual differences, staining level, etc. Therefore, different segmentation requirements exist for each picture.
[0010] At the same time, the speed and accuracy of the balanced network model are a major problem in current segmentation methods. Shallow networks are fast, but it is difficult to segment difficult signals. Deep networks have improved accuracy, but have a huge amount of calculations and parameters, and are slow in operation. Existing deep learning methods often focus on the design of different models, rarely starting from the pixel perspective, ignoring the difficulty levels of different pixels, and using the same model to uniformly process all pixels. There is serious over - calculation on a large number of easily segmented pixels, which will reduce the operation speed. Summary of the Invention
[0011] To solve the above - mentioned technical problems, the present invention provides a method for cell segmentation of M - ROSE pictures, including:
[0012] Obtaining an adaptive feedback mechanism for the network layer depth required by pixels of different difficulties according to the segmentation difficulty of pixels in the M - ROSE picture;
[0013] Constructing a 2D deep - supervised decomposable segmentation network to output the prediction information of each layer of the network for each pixel in the M - ROSE picture;
[0014] Based on the adaptive feedback mechanism for the network layer depth required by pixels of different difficulties, constructing a pixel - difficulty - aware adaptive network layer depth feedback mechanism, perceiving the segmentation difficulty of pixels according to the output results of each layer of the network, and feeding it back to the 2D deep - supervised decomposable segmentation network, deactivating the neurons at the pixel positions that have been correctly segmented in the current layer, and realizing the cell segmentation of the image.
[0015] Preferably, the 2D deep - supervised decomposable segmentation network adopts an encoding and decoding segmentation method, uses long - short connection paths, combines information of different layer depths, and performs segmentation.
[0016] Preferably, the 2D deep - supervised decomposable segmentation network is an end - to - end network structure, the input is an image and a segmentation label, and the output is five network prediction results of the same size as the input image, and each result corresponds to the output of each layer of the network.
[0017] Preferably, the 2D deep - supervised decomposable segmentation network includes 16 convolutional network layers, a maximum pooling layer, and a transposed convolutional kernel;
[0018] Among them, 11 convolutional network layers in the convolutional network layer use 3×3 convolutional kernels for image feature extraction, and 5 1×1 convolutions are used to reduce the final output channels to 1; the convolutional network channels for feature extraction include 8, 16, and 32;
[0019] The maximum pooling layer is used to reduce the size of the input image by using the maximum pooling method, reduce the amount of calculation, and at the same time reduce the mean shift caused by parameter errors;
[0020] The deconvolution kernel is a 2×2 deconvolution kernel, which is used to convert the downsampled data into an image of the same size as the input image.
[0021] Preferably, the 2D deep supervision decomposable segmentation network adopts a deep supervision strategy, and respectively outputs and calculates the loss functions corresponding to the segmentation results of 5 layers of the network, and uses the sum of all loss functions as the final loss function.
[0022] Preferably, the process of using the sum of all loss functions as the final loss function includes:
[0023] Adopt a mixed function composed of Dice loss and weighted cross-entropy loss as the segmentation loss function L;
[0024] The corresponding formula expression includes:
[0025]
[0026] Among them, the parameter α1 is a variable constant, which is used to balance the loss function L dicek and L cenk ; k represents any network layer number, and the value ranges from 1 to 5; j is the index value of any pixel V j ; M is the total number of pixels of the 2D image; g j is the input label corresponding to the pixel V j , 1 represents the foreground, and 0 represents the background; p jk is the foreground prediction probability of the kth layer network for the pixel V j , and the value is in [0, 1]; ε is a constant.
[0027] Preferably, the process of perceiving the difficulty of pixel segmentation according to the output result of each layer of the network includes:
[0028] Construct a pixel difficulty perception feedback network layer depth mechanism, and fuse it with the 2D deep supervision decomposable segmentation network to obtain the pixel difficulty perception adaptive network layer depth feedback mechanism;
[0029] Based on the pixel difficulty perception adaptive network layer depth feedback mechanism, the network sequentially judges whether each layer of the network has correctly predicted each pixel according to the output result of each layer. If the shallow network has achieved accurate prediction of a certain pixel, the neurons at the corresponding position are inactivated using the artificial neuron inactivation mechanism, and the subsequent network layers no longer calculate this pixel.
[0030] Preferably, the process of the network sequentially judging whether each layer of the network has correctly predicted each pixel according to the output result of each layer includes:
[0031] Based on the pixel difficulty perception adaptive network layer depth feedback mechanism, respectively for the feature map F of each network layer k1 , k∈[1,5] and the output result Fk2 Act, first normalize the F value of the output result k2 to obtain the predicted probability map Prob of the signal k ∈[0,1]; according to the predicted probability of each pixel v, determine whether it has been accurately predicted. If so, inactivate the neuron at the corresponding position of the pixel according to the feedback result;
[0032] After that, construct the mask map mask of the corresponding network layer according to the inactivation judgment value k , by multiplying the feature map F of each network layer k1 with the mask map mask k to update the feature map F of each network layer k1 , and select the neurons at the positions where the inactivation probability is 0;
[0033] Then normalize the updated F k1 to ensure that after selective inactivation, the network layer has the same mean and variance;
[0034] Repeat the operation until the value of k changes from 1 to 5 to complete the output of all network layers.
[0035] Preferably, the formula expression of the inactivation judgment value is:
[0036]
[0037] where p th is the probability feedback threshold.
[0038] Compared with the prior art, the present invention has the following advantages and technical effects:
[0039] The present invention obtains an adaptive feedback mechanism for the network layer depth required by different difficulty pixels according to the segmentation difficulty of pixels in the M-ROSE picture; constructs a 2D deep supervision decomposable segmentation network to output the prediction information of each network layer for each pixel in the M-ROSE picture; based on the adaptive feedback mechanism for the network layer depth required by different difficulty pixels, constructs a pixel difficulty perception adaptive network layer depth feedback mechanism, perceives the segmentation difficulty of pixels according to the output result of each network layer, and feeds it back to the 2D deep supervision decomposable segmentation network to inactivate the neurons at the positions of the pixels that have been correctly segmented in the current layer, realizing the cell segmentation of the image. Based on the above ultra-fast segmentation network, the present invention quickly segments the cells in the M-ROSE picture, improving the segmentation speed while balancing the accuracy and accuracy of the segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0041] Figure 1 This is a schematic structural diagram of the 2D deep supervision decomposable segmentation network according to an embodiment of the present invention. Detailed implementation manners
[0042] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0043] It should be noted that the steps shown in the flowchart of the drawings may be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0044] As Figure 1 shown, in this embodiment, a method for cell segmentation of M-ROSE images is provided, including:
[0045] Obtaining an adaptive feedback mechanism for the network layer depth required by pixels of different difficulties according to the segmentation difficulty of pixels in the M-ROSE image;
[0046] Constructing a 2D deep supervision decomposable segmentation network to output the prediction information of each layer of the network for each pixel in the M-ROSE image;
[0047] Based on the adaptive feedback mechanism for the network layer depth required by pixels of different difficulties, constructing a pixel difficulty-aware adaptive network layer depth feedback mechanism, perceiving the segmentation difficulty of pixels according to the output results of each layer of the network, and feeding it back to the 2D deep supervision decomposable segmentation network, deactivating the neurons at the pixel positions that have been correctly segmented in the current layer, and realizing the cell segmentation of the image.
[0048] Furthermore, in this embodiment, a simple 2D deep supervision decomposable segmentation network (High Speed DeepSupervision Net, abbreviated as HSDS-Net) is constructed. HSDS-Net adopts an encoding and decoding segmentation method, uses long and short connection paths, combines information of different layer depths, and improves the segmentation accuracy. Its structure is as Figure 1 shown.
[0049] Furthermore, the 2D deep supervision decomposable segmentation network is an end-to-end network structure, the input is an image and a segmentation label, and the output is five network prediction results of the same size as the input image, and each result corresponds to the output of each layer of the network.
[0050] Furthermore, the 2D deep supervision decomposable segmentation network includes 16 convolutional network layers, a maximum pooling layer, and a deconvolution kernel;
[0051] Among them, in the convolutional network layer, 11 convolutional networks use 3×3 convolutional kernels for image feature extraction, and 5 1×1 convolutions are used to reduce the final output channels to 1; the convolutional network channels for feature extraction include 8, 16, and 32;
[0052] The max pooling layer is used to reduce the input image size by the max pooling method, reduce the computational amount, and at the same time reduce the mean shift caused by parameter errors;
[0053] The deconvolution kernel is a 2×2 deconvolution kernel, which is used to transform the downsampled data into an image of the same size as the input image.
[0054] Furthermore, the 2D deep supervision decomposable segmentation network adopts a deep supervision strategy, and respectively outputs and calculates the loss functions corresponding to the segmentation results of 5 layers of networks, and uses the sum of all loss functions as the final loss function.
[0055] Furthermore, the process of using the sum of all loss functions as the final loss function includes:
[0056] Adopt a mixed function composed of Dice loss and weighted cross-entropy loss as the segmentation loss function L;
[0057] The corresponding formula expressions include:
[0058]
[0059] Among them, the parameter α1 is a variable constant, which is used to balance the loss function L dicek and L cenk , the embodiment is set to 0.5; k represents any network layer number, and the value ranges from 1 to 5; j is the index value of any pixel V j ; M is the total number of pixels in the 2D image; g j is the input label corresponding to the pixel V j , 1 represents the foreground, and 0 represents the background; p jk is the foreground prediction probability of the kth layer network for the pixel V j , and the value is in [0, 1]; ε is a constant, and the value is 1.
[0060] Furthermore, the process of perceiving the difficulty of pixel segmentation according to the output results of each layer of network includes:
[0061] Construct a deep mechanism of the pixel difficulty perception feedback network layer and fuse it with the 2D deep supervision decomposable segmentation network to obtain a deep feedback mechanism of the pixel difficulty perception adaptive network layer;
[0062] Based on the pixel-difficulty-aware adaptive network layer depth feedback mechanism, the network successively determines whether each layer of the network has correctly predicted each pixel according to the output result of each layer. If a certain shallow network has achieved accurate prediction of a pixel, the neurons at specific positions are inactivated using the artificial neuron inactivation mechanism, and subsequent network layers no longer calculate that pixel.
[0063] Furthermore, the process of the network successively determining whether each layer of the network has correctly predicted each pixel includes:
[0064] Based on the pixel-difficulty-aware adaptive network layer depth feedback mechanism, the feature map F of each network layer is respectively k1 , k ∈ [1, 5] and the output result F k2 acted on. First, the value of F of the output result is normalized k2 to obtain the predicted probability map Prob k ∈ [0, 1]; according to the predicted probability of each pixel v, it is judged whether it has been accurately predicted. If so, the neurons at the corresponding positions of the pixel are inactivated according to the feedback result;
[0065] After that, according to the inactivation judgment value, the mask map mask of the corresponding network layer is constructed k , and the feature map F of each network layer is updated by multiplying k1 the feature map F with the mask map mask k , and the neurons with an inactivation probability of 0 are selected; k1 F
[0066] F k1 = F k1 · mask k
[0067] Then, the updated F k1 is normalized to ensure that after selective inactivation, the network layer has the same mean and variance;
[0068] F k1 = F k1 · count(mask k ) / count_ones(mask k )
[0069] The operation is repeated until the value of k changes from 1 to 5 to complete the output of all network layers.
[0070] Furthermore, the formula expression of the inactivation judgment value is:
[0071]
[0072] where p this the probability feedback threshold. To ensure the accuracy of pixel estimation, in this embodiment, p th is greater than 0.95.
[0073] Based on the above ultra-fast segmentation network, this embodiment quickly segments cells in the M-ROSE image, improving the segmentation speed while balancing the accuracy and precision of segmentation.
[0074] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for cell segmentation of M-ROSE images, characterized in that: include: According to the segmentation difficulty of pixels in the M-ROSE image, an adaptive feedback mechanism is used to obtain the network layer depth required for pixels of different difficulty levels; Construct a 2D deep supervised decomposable segmentation network and output the prediction information of each layer of the network for each pixel in the M-ROSE image; Based on the adaptive feedback mechanism of the network layer depth required for pixels of different difficulty, a pixel difficulty-aware adaptive network layer depth feedback mechanism is constructed. The difficulty of pixel segmentation is perceived according to the output results of each layer of the network, and the feedback is given to the 2D deep supervised decomposable segmentation network to deactivate the neurons at the pixel positions that have been correctly segmented in the current layer, thereby realizing cell segmentation of the image.
2. The method according to claim 1, characterized in that The 2D deep supervised decomposable segmentation network adopts an encoding and decoding segmentation method, using long and short connection paths and combining information of different layer depths to perform segmentation.
3. The method according to claim 1, characterized in that The 2D deep supervised decomposable segmentation network is an end-to-end network structure, the input is an image and a segmentation label, and the output is five network prediction results of the same size as the input image, each result corresponds to the output of each layer of the network.
4. The method according to claim 1, characterized in that: The 2D deep supervised decomposable segmentation network includes 16 convolutional network layers, a maximum pooling layer, and a deconvolution kernel; Among them, 11 layers of convolutional networks in the convolutional network layer use 3×3 convolution kernels to extract image features, and 5 layers of 1×1 convolution are used to reduce the number of final output channels to 1; the number of convolutional network channels for feature extraction includes 8, 16 and 32; The maximum pooling layer is used to reduce the size of the input image by using the maximum pooling method, reduce the amount of calculation, and reduce the mean shift caused by parameter errors; The deconvolution kernel is a 2×2 deconvolution kernel, which is used to convert the downsampled data into an image of the same size as the input image.
5. The method according to claim 1, characterized in that The 2D deep supervised decomposable segmentation network adopts a deep supervision strategy to output and calculate the loss functions corresponding to the segmentation results of the 5-layer network respectively, and the sum of all loss functions is used as the final loss function.
6. The method according to claim 5, characterized in that The process of taking the sum of all loss functions as the final loss function includes: A hybrid function consisting of Dice loss and weighted cross entropy loss is used as the segmentation loss function L; The corresponding formula expressions include: Among them, the parameter α1 is a variable constant used to balance the loss function L dicek and L cenk ; k represents any number of network layers, ranging from 1 to 5; j is any pixel V j The index value of the 2D image; M is the total number of pixels of the 2D image; g j is the pixel V j The corresponding input label is 1 for foreground and 0 for background; p jk For the k-th layer network, pixel V j The prospect prediction probability of , the value is [0, 1]; ε is a constant.
7. The method according to claim 1, characterized in that The process of perceiving the difficulty of pixel segmentation based on the output results of each layer of the network includes: Constructing a pixel difficulty perception feedback network layer deep mechanism, and fusing it with the 2D deep supervised decomposable segmentation network to obtain the pixel difficulty perception adaptive network layer deep feedback mechanism; Based on the pixel difficulty-aware adaptive network layer deep feedback mechanism, the network determines whether each layer of the network has correctly predicted each pixel according to the output results of each layer in turn. If the shallow network has achieved accurate prediction of a certain pixel, the artificial neuron inactivation mechanism is used to inactivate the neurons at the corresponding position, and the subsequent network layers will no longer calculate the pixel.
8. The method according to claim 7, characterized in that The network determines whether each layer has correctly predicted each pixel based on the output results of each layer in turn. The process includes: Based on the pixel difficulty-aware adaptive network layer deep feedback mechanism, the feature map F of each network layer is respectively k1 ,k∈[1,5] and the output result F k2 To work, first normalize the output result F k2 Value, get the predicted probability map Prob of the signal k ∈[0,1]; According to the prediction probability of each pixel v, determine whether it has been accurately predicted. If so, deactivate the neuron at the corresponding position of the pixel according to the feedback result; Then, according to the deactivation judgment value, the mask of the corresponding network layer is constructed. k , by taking the feature map F of each network layer k1 With mask k Multiply and update the feature map F of each network layer k1 , select the neuron with inactivation probability of 0; Then normalize the updated F k1 , ensuring that after selective inactivation, the network layers have the same mean and variance; Repeat the operation until the value of k changes from 1 to 5, completing the output of all network layers.
9. The method according to claim 8, characterized in that The formula expression of the deactivation judgment value is: Among them, p th is the probability feedback threshold.
Citation Information
Patent Citations
Method and system for melanoma image tissue segmentation based on deep neural network
CN108510502A
Unsupervised domain adaptive remote sensing road semantic segmentation method based on GAN network
CN113888547A
Multi-organ segmentation method based on parallel depth U-shaped network and probability density map
CN115222748A
Lightweight retinal vessel segmentation method based on attention mechanism
CN115760872A
Semantic segmentation method and system for learning image structure difficulty information, and storage medium
CN116721252A