AI-based M-ROSE medical image pathogen detection method and device
By combining fluorescence microscopy and deep learning models, using ResNet-50 and EfficientNet for fungal classification and YOLOX for Mycobacterium tuberculosis detection, the problems of long time consumption and high misdiagnosis rate in traditional M-ROSE technology are solved, and efficient and accurate automated pathogen detection is achieved.
Patent Information
- Application Number
- CN202510548038.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-09-16
AI Technical Summary
Traditional M-ROSE technology relies on manual microscope observation, which is time-consuming and has a high misdiagnosis rate. It is difficult to accurately identify micron-level pathogens, especially fungi and Mycobacterium tuberculosis, which have problems of missed detection and false positives. There is a lack of deep learning automatic recognition systems adapted to the characteristics of fungal/tuberculosis staining images.
Combining fluorescence microscopy and deep learning models, ResNet-50 and EfficientNet were used for fungal classification, and YOLOX was used for Mycobacterium tuberculosis detection. Automated analysis was achieved through dynamic patch selection and improved Focal Loss optimization.
It significantly improves the detection efficiency and accuracy of fungi and Mycobacterium tuberculosis, reduces manual intervention, shortens the detection time by more than 10 times, and meets the needs of rapid bedside diagnosis in the ICU.
Smart Images

Figure CN120655569A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical imaging artificial intelligence and rapid pathogen diagnosis technology, specifically an M-ROSE (Microbiology Rapid On-Site Evaluation) pathogen detection method based on multimodal microscopic imaging (fluorescence staining / acid-fast staining) and artificial intelligence analysis. By integrating deep learning models with computer vision technology, this method automates fungal type identification and Mycobacterium tuberculosis detection, addressing the manual and inefficient nature of traditional microscopic examinations. The method is suitable for rapid bedside pathogen assessment in ICUs. Background Art
[0002] Traditional M-ROSE technology is a rapid pathogen assessment method for lung infection and ICU critical care. Its core relies on manual microscopic observation, but it exposes multiple bottlenecks in clinical practice. On the one hand, the operation process is highly dependent on experienced doctors. It takes several hours from specimen preparation to film reading and analysis, and the quality of the specimen is directly affected by the standardization of the operation. For example, the misdiagnosis rate of bronchoalveolar lavage fluid (BALF) due to improper collection or processing remains high. On the other hand, the limitations of manual interpretation are significant: micron-sized pathogens (such as bacteria) are difficult to accurately identify due to their small morphology and low staining contrast. Although fungal detection is more accurate, Cryptococcus and Pneumocystis still require the assistance of immunofluorescence technology; acid-fast staining of Mycobacterium tuberculosis is easy to miss due to its aggregation or fracture characteristics, and false positives caused by uneven staining or background noise further affect the reliability of diagnosis.
[0003] Existing AI technologies have shown initial success in pathology image diagnosis, such as using convolutional neural networks (CNNs) to classify and segment tumor sections. However, there is still a lack of a deep learning automatic recognition system specifically designed for M-ROSE scenarios and adapted to the characteristics of fungal and tuberculosis staining images. Therefore, there is an urgent need to build a comprehensive solution that combines staining microscopic imaging, AI recognition models, target detection algorithms, and structured output capabilities to achieve efficient, standardized, and intelligent analysis of microbial pathogens. Summary of the Invention
[0004] The present invention aims to overcome the above-mentioned shortcomings of the prior art and provide an AI-based M-ROSE medical imaging pathogen detection method and device.
[0005] The first aspect of the present invention relates to an AI-based M-ROSE medical imaging pathogen detection method, which combines fluorescence microscopy, AI models, and computer vision technology to achieve automated analysis of fungal and tuberculosis detection, including the following steps:
[0006] Step 1: Build a deep learning-based medical imaging fungal fluorescence image classification method, automatically scan and identify the fungal samples after fluorescent staining, and automatically identify and classify them through the AI model.
[0007] Step 2: Build a deep learning-based medical imaging tuberculosis detection method. After the tuberculosis bacteria sample is stained with acid-fast staining or auramine O staining, it enters the automatic scanning and recognition system, and is combined with the target detection model to automatically determine whether the tuberculosis bacteria are present.
[0008] Step 3: Obtain the automatic identification result to determine the type of fungus present in the sample or whether tuberculosis bacteria are present.
[0009] Deep learning-based classification of medical imaging fungal fluorescence images, wherein the basic method described in step 1 specifically includes:
[0010] Step 11: Collect body fluid samples from the patient as medical imaging samples, prepare slides, and perform fluorescent staining to make specific types of fungi show high-contrast fluorescent signals under a fluorescence microscope, and collect medical imaging data;
[0011] Step 12: Use the residual network model ResNet-50 to perform the first stage classification of the microscope slide image to determine whether it contains fungi. Negative samples are directly archived and do not enter the subsequent analysis process. Positive samples will be processed for subsequent fungal classification.
[0012] Step 13: For positive samples, the patch extraction algorithm is used to accurately screen out the 60 key patch areas where fungi are most likely to exist, and the patches are preprocessed using color normalization, noise suppression, and contrast enhancement.
[0013] Step 14: Use the convolutional neural network EfficientNet as the core classification model to classify the screened patches to identify specific fungal types; EfficientNet adopts a composite scaling strategy: combining depthwise separable convolution and SE mechanism.
[0014] Step 15: Output the classification results and generate a medical imaging fungus classification report based on the classification results.
[0015] Preferably, the key patch area screening method described in step 13: adopts Adaptive Patch Selection Based on Local Variance Maximization, and preferentially selects fungus aggregation areas as the patch areas to be used.
[0016] Preferably, the EfficientNet described in step 14 includes:
[0017] The convolutional neural network model consists of a front 32-channel convolutional layer and multiple MBConv components, with a total of 9 stages, including 8 groups of MBConv structures and 1 final fully connected layer (FC); each MBConv component consists of a front convolutional layer and a group of cyclic network structures. The cyclic structure includes a convolutional layer, a batch normalization layer, a Swish activation function and a residual connection shortcut. Each component is stacked according to different numbers of repetitions (1,1,2,2,3,3,4,1) to ensure layer-by-layer extraction of deep features; in the MBConv structure, the number of channels is first expanded by 1x1 convolution, and then k×k depthwise convolution is used for spatial feature extraction, and the channel weights are adaptively adjusted through the SE module. The SE module uses global average pooling (GAP) to extract channel information, and combines the fully connected layer with dimensionality reduction and the Swish activation function to process the feature expression of the image. It then generates channel attention weights through dimensionality increase and the Sigmoid activation function in the fully connected layer, and multiplies them with the original features to enhance important features and suppress redundant information. Finally, MBConv uses 1x1 convolution to reduce the dimension and restore the number of channels. When the input and output shapes are consistent, a shortcut residual connection is introduced to optimize gradient propagation. DropConnect is used for regularization, which is only effective when shortcuts are present. During the feature downsampling process, the stride parameter of the pre-convolution of the first five groups of MBConv modules is set to 2, and the stride parameter of the remaining convolutions is set to 1. Therefore, the downsampling rate of the entire neural network is 32, achieving effective dimensionality reduction and expression of multi-level features. In addition, a shortcut residual connection is used between the two convolutional layers of each component to optimize computational efficiency and gradient propagation.
[0018] The network's feature aggregation and classification decision-making components include global pooling and fully connected layers, resulting in a final output channel count of 1280, ensuring sufficient extraction of high-level semantic information. During training, the batch size is set to 16, the initial learning rate is 0.01, and a step-wise learning rate reduction strategy is used for dynamic updates. Furthermore, to improve detection accuracy, the Focal Loss loss function is employed to balance the amount of data across different categories. Focal Loss introduces a category weighting factor α and a scaling factor γ based on the standard cross-entropy loss to enhance learning for difficult-to-classify samples.
[0019] Focal_Loss(p t )=-α(1-p t ) γlog(p t ) (1)
[0020] Among them, p t Represents the model's predicted probability of the true category, that is, when the sample belongs to the positive class p t =p, and when the sample belongs to the negative class p t =1-p, the category weight factor α is used to balance the ratio of positive and negative samples, ensuring that the minority class will not be ignored by the model due to too little data; the adjustment factor γ controls the loss contribution of easy-to-classify samples, reducing the impact of the loss function on high-confidence samples and giving higher weights to low-confidence samples, thereby guiding the model to pay more attention to samples that are difficult to classify.
[0021] The method for detecting tuberculosis in medical images based on deep learning described in step 2 specifically includes:
[0022] Step 21: Collect bronchoalveolar lavage fluid (BALF), bronchoscopic biopsy tissue, sputum, or other biological samples from the patient; use acid-fast staining or acridine orange staining, and acquire medical imaging data using a microscope or digital pathology scanner;
[0023] Step 22: Data preprocessing: Preprocess the medical images, including color standardization, background suppression, morphological filtering, and edge enhancement, to reduce noise interference during the staining process and improve the visualization characteristics of pathogens;
[0024] Step 23: Staining classification: Use the ResNet-50 deep learning model to classify the medical images to distinguish between acid-fast staining data and auramine O staining data, ensuring that the correct analysis strategy is used for subsequent target detection;
[0025] Step 24: Mycobacterium tuberculosis target detection: The classified image is input into the YOLOX target detection model for Mycobacterium tuberculosis identification. The model uses the SimOTA label assignment strategy for target optimization to improve the accuracy of Mycobacterium tuberculosis target detection. The adaptive multi-scale detection mechanism is used to enhance the ability to identify Mycobacterium tuberculosis of different sizes. Feature alignment technology is used to improve the stability and robustness of Mycobacterium tuberculosis detection.
[0026] Step 25: Test result output: Generate a medical imaging tuberculosis test report based on the test results.
[0027] Preferably, step 21 further includes color normalization, morphological filtering, adaptive contrast enhancement and background suppression.
[0028] Preferably, the staining mode classification described in step 23 includes: using the residual network ResNet-50 to classify acid-fast staining and auramine O staining; extracting the overall feature information of the image through the global feature extraction module, and using the local feature learning module to capture the detailed features in the image, so as to more comprehensively characterize the input data; finally using the SoftMax classification layer for final classification, and according to the classification results, assigning the input data to the acid-fast staining detection model or the auramine O staining detection model.
[0029] Preferably, the YOLOX model uses CSPDarknet as the backbone network. The entire deep learning architecture is composed of multiple network components, which includes a total of 9 stages. The backbone network is responsible for basic feature extraction, SPBottleneck is used as the feature enhancement module Neck, and YoloHead is used as the final detection head Head to jointly build a complete YOLOX target detection framework.
[0030] The backbone network CSPDarknet includes: first, the input feature is segmented through the Focus layer, the input image size is 224×224, and after Focus processing, it enters the 32-channel 3×3 standard convolution (Conv3×3), and then enters multiple CSPLayer structures for feature extraction; in Stage 2, CSPLayer (3×3 convolution) is used for convolution calculation, the input size is 112×112, the number of channels is 16, and the number of repetitions is 1; then, in Stage 3, CSPLayer (3×3 convolution) is continued to be used for feature extraction, the input size is 112×112, the number of channels is expanded to 24, and the number of repetitions of this stage is 2; in Stage 4, CSPLayer (5×5 convolution) is used, the input size is reduced to 56×56, the number of channels is increased to 40, and the structure is repeated twice to ensure the full extraction of deep features; in Stage 5. The model regression uses the CSPLayer (3×3 convolution) structure, the input size is reduced to 28×28, the number of channels is further expanded to 80, and this part is repeated 3 times; in Stage 6, the CSPLayer (5×5 convolution) structure is used, the input size is reduced to 14×14, the number of channels is increased to 112, and it is repeated 3 times; in Stage 7, the CSPLayer (5×5 convolution) structure is continued to be used, the input size remains 14×14, the number of channels is increased to 192, and it is repeated 4 times; in Stage 8, the CSPLayer (3×3 convolution) structure is used, the input size is reduced to 7×7, and the number of channels is increased to 320. This stage is only executed once;
[0031] The feature enhancement layer Neck includes: using SPPBottleneck to improve the multi-scale expression ability of features; SPPBottleneck consists of spatial pyramid pooling (SPP) and 1×1 convolution. The SPP structure extracts spatial information through pooling kernels of different scales (5×5, 9×9, 13×13), concatenates the multi-scale pooling results with the original features, and finally compresses the features through 1×1 convolution;
[0032] The detection head YoloHead includes: feature refinement through multiple 3×3 Conv2D_BN_SiLU convolutional layers, with a final output channel number of 1280, and outputs classification Cls, regression Reg, and target confidence Obj results respectively; the structure adopts an anchor-free design and combines a multi-scale feature fusion strategy with a downsampling strategy; the classification loss is used to calculate the error between the predicted category probability and the true category, and the final total loss function is the weighted sum of GIoU loss, confidence loss, and classification loss.
[0033] Furthermore, the GIoU loss not only considers the overlapping area of the predicted box and the true box, but also introduces the influence of the non-overlapping area, which can more accurately measure the similarity between the two boxes and has scale invariance. The calculation formula of GIoU is as follows:
[0034]
[0035] Among them, IoU is the intersection-over-union ratio of the predicted box and the real box, A c is the area of the minimum enclosed region containing the predicted box and the true box, and u is the union area of the predicted box and the true box. The final GIoU loss is:
[0036] L GIoU =1-GIoU (3)
[0037] Confidence loss is used to measure the existence of the target in the prediction box. For the prediction box containing the target (positive sample), the loss function is:
[0038]
[0039] Among them, C i is the confidence of the prediction, is the confidence of the true label (1 for positive samples), Is an indicator function, indicating whether the predicted box is a positive sample.
[0040] For the predicted box that does not contain the target (negative sample), the loss function is:
[0041]
[0042] Among them, λnoobj Is a weight coefficient used to balance the loss of positive and negative samples.
[0043] The classification loss is used to measure the difference between the predicted category probability and the true category. For each predicted box containing the target, the classification loss is:
[0044]
[0045] Among them, p i (c) is the predicted class probability, is the true class probability (1 for the target class and 0 for the rest), Is an indicator function, indicating whether the predicted box is a positive sample.
[0046] The second aspect of the present invention relates to an AI-based M-ROSE medical imaging pathogen detection device, which is characterized in that it includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement an AI-based M-ROSE medical imaging pathogen detection method of the present invention.
[0047] The working principle of the present invention is:
[0048] This invention is based on the AI-based M-ROSE medical imaging pathogen detection method, which combines fluorescence microscopy, acid-fast staining technology, deep learning models and computer vision algorithms to achieve automated detection of fungi and Mycobacterium tuberculosis. Its core process includes:
[0049] 1. Fungal detection:
[0050] (1) The fluorescently stained samples were scanned in their entirety, positive samples were preliminarily screened using ResNet-50, and then the fungal aggregation areas were focused using a dynamic patch selection algorithm.
[0051] (2) An improved EfficientNet model (including SE module, compound scaling strategy and Focal Loss optimization) was used for high-precision classification to solve the problem of complex fungal morphology and low contrast.
[0052] 2. Mycobacterium tuberculosis detection:
[0053] (1) ResNet-50 is used to distinguish staining types and achieve comprehensive recognition of data with different staining methods.
[0054] (2) High-sensitivity target recognition is achieved through the YOLOX model (including SimOTA label allocation and adaptive multi-scale detection), the GIoU loss function is introduced to optimize the detection box regression, and the Anchor-Free design is combined to improve the model's detection ability for small-sized pathogens.
[0055] The innovative features of the present invention are:
[0056] 1. Dynamic Patch Selection Algorithm:
[0057] (1) An Adaptive Patch Selection method based on local maximum variance is proposed to prioritize the screening of fungal clusters, reduce invalid calculations, and improve detection efficiency (which is superior to the time-consuming manual screening problem in traditional ROSE technology).
[0058] 2. Improved EfficientNet and YOLOX models:
[0059] (1) Integrating the SE module and Focal Loss in EfficientNet to enhance the focus on difficult-to-classify samples and solve the problem of imbalanced fungal categories;
[0060] (2) Feature space alignment technology was introduced into YOLOX to improve the robustness of Mycobacterium tuberculosis detection under different staining conditions.
[0061] 3. Multi-task joint optimization framework:
[0062] (1) Integrate fungal classification with Mycobacterium tuberculosis detection tasks, achieve parallel analysis of multiple pathogens through unified preprocessing and model architecture, and reduce repeated operations (traditional M-ROSE requires step-by-step processing).
[0063] 3. Intelligent combination of auramine O staining and acid-fast staining in the identification of Mycobacterium tuberculosis:
[0064] (1) ResNet-50 is used to automatically distinguish staining types and avoid manual misjudgment (traditional methods rely on experience).
[0065] The advantages of the present invention are:
[0066] 1. Dynamic feature extraction and efficient classification architecture
[0067] (1) Optimized combination of residual network and EfficientNet:
[0068] ResNet-50 was used for preliminary sample screening (negative / positive classification), combined with EfficientNet's composite scaling strategy (depthwise separable convolution + SE module) to achieve multi-level feature fusion for fungal classification. The SE module dynamically adjusts weights through a channel-wise attention mechanism, prioritizing high-discriminative features.
[0069] (2) Patch dynamic selection algorithm:
[0070] The Adaptive Patch Selection algorithm based on local maximum variance automatically screens fungal clusters (such as hyphae and spore-dense areas), reducing 80% of ineffective calculations compared to traditional full-film scanning while improving the classification accuracy of key areas.
[0071] 2. Structural innovation of target detection model
[0072] (1) Customized improvements of YOLOX:
[0073] In the detection of Mycobacterium tuberculosis, the YOLOX model introduces the SimOTA label allocation strategy (dynamic matching of positive and negative samples) and the adaptive multi-scale detection mechanism (optimizing the positioning of bacilli of different sizes), combined with the CSPDarknet backbone network to enhance the small target detection capability, effectively improving the sensitivity of Mycobacterium tuberculosis detection.
[0074] (2) Anchor-free box design and GIoU loss function:
[0075] The Anchor-Free structure is used to avoid the problem of preset box deviation. Combined with the GIoU loss function (comprehensively considering overlapping and non-overlapping areas), the regression error of the predicted box is effectively reduced, which is especially suitable for the precise positioning of micron-level pathogens.
[0076] 3. Efficient automation:
[0077] (1) Full process automation (from sample scanning to report generation) significantly reduces manual intervention and improves detection efficiency by more than 10 times compared to traditional M-ROSE.
[0078] 4. Real-time and clinical applicability:
[0079] (1) Combining fluorescent staining with AI models, preliminary fungal diagnosis can be completed within 10 minutes (traditional M-ROSE takes 2-3 hours), meeting the needs of rapid bedside diagnosis in the ICU. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] Figure 1 This is a flow chart of the deep learning-based medical imaging fungal fluorescence image classification method of the present invention.
[0081] Figure 2 Schematic diagram of the ResNet-50 deep residual network module of the present invention.
[0082] Figure 3 Schematic diagram of the deep learning-based medical imaging tuberculosis detection method of the present invention.
[0083] Figure 4 It is a structural diagram of the YOLOX model of the present invention. Figure 5 It is a schematic diagram of the device of the present invention. DETAILED DESCRIPTION
[0084] The technical solution of the present invention is further described below with reference to the accompanying drawings.
[0085] This paper proposes a full-process fungal fluorescence image classification solution that integrates deep learning and microscopic image processing (flow chart as shown in the figure) Figure 1 Its core innovation lies in the construction of a two-stage analysis framework: global initial screening and local precision classification. First, a residual network (ResNet-50) is used to rapidly screen whole slide images for negative or positive samples, eliminating over 90% of negative samples. Next, a local variance maximization algorithm is used to dynamically select key patches, combined with an efficient convolutional network (EfficientNet) to achieve fine-grained species-level classification. This process overcomes the technical limitations of traditional whole-slide scanning, effectively improving the classification accuracy of common clinical species such as Cryptococcus and Aspergillus through adaptive region selection and multi-scale feature fusion.
[0086] The AI-based M-ROSE medical imaging pathogen detection method of this embodiment combines fluorescence microscopy, AI models, and computer vision technology to achieve automated analysis of fungal and tuberculosis detection, including the following steps:
[0087] Step 1: Build a deep learning-based medical imaging fungal fluorescence image classification method, automatically scan and identify the fungal samples after fluorescent staining, and automatically identify and classify them through the AI model.
[0088] Step 2: Build a deep learning-based medical imaging tuberculosis detection method. After the tuberculosis bacteria sample is stained with acid-fast staining or auramine O staining, it enters the automatic scanning and recognition system, and is combined with the target detection model to automatically determine whether the tuberculosis bacteria are present.
[0089] Step 3: Obtain the automatic identification result to determine the type of fungus present in the sample or whether tuberculosis bacteria are present.
[0090] The step 1 of constructing a deep learning-based medical imaging fungal fluorescence image classification method includes:
[0091] 1. Imaging and Pretreatment of Fluorescently Stained Samples
[0092] Clinically collected samples, such as bronchoalveolar lavage fluid (BALF) and bronchial biopsy tissue, were stained with Calcofluor White, which reveals a specific blue-green fluorescence signal from the β-glucan component of the fungal cell wall at an excitation wavelength of 365 nm. Sample preparation followed a standardized centrifugation protocol (1600 × g for 10 minutes). Sediment preparation was performed using a high-resolution digital pathology scanner (such as the OLYMPUS VS200) for whole-slide imaging. The images were exported as SVS files, preserving micron-level details at 0.25 μm / pixel.
[0093] To adapt the input requirements of the deep learning model, the original image was cut into 1024×1024 pixel sub-images using a sliding window. Background threshold filtering (sub-image pixel variance <90) was used to exclude irrelevant regions, significantly reducing ineffective computational load. During this process, the dynamic range and contrast of the fluorescence signal were pre-calibrated to ensure consistency, preventing staining batch differences from interfering with model inference.
[0094] 2. Fungal Positive Pre-screening Module Based on Deep Residual Network
[0095] This system adopts an architecture to achieve high-precision pre-screening of fungal positives for whole-slide fluorescence microscopy images. Its core innovation is to optimize the stability of deep network training through residual blocks, effectively overcoming the gradient vanishing problem of traditional convolutional neural networks in processing large-scale medical images (typical size 4912×3684 pixels). The residual structure is as follows Figure 2 As shown in the figure, the ResNet-50 model has a relatively deep network structure, but due to its bottleneck structure, the computational complexity is relatively small. This allows gradients to be directly backpropagated from the back layers to the front layers, thereby improving training stability. Through the skip connection structure, the ResNet-50 model allows gradients to be directly transferred from the back layers to the front layers during backpropagation, improving training stability and accelerating convergence.
[0096] The residual block adds identity mapping to the original convolutional layer, that is, the input data can be directly used as part of the output, thereby avoiding the gradient vanishing problem and enabling the network to be trained deeper and more stably.
[0097]
[0098] As shown in formula (7), the residual block output y consists of two parts: the input data x is transformed by nonlinear transformation The result after a series of convolutional layers (i.e., batch normalization layers and activation functions (ReLU)) and the original input x. This structure ensures that even if the gradient is small, it can still be directly passed to the previous layer through x, thus alleviating the gradient vanishing problem.
[0099]
[0100] When the dimensions of input and output do not match, the structure of formula (8) can be used, and W needs to be introduced s Perform linear transformation to ensure that skip connections can be performed correctly. This is especially important in the deeper ResNet-50 to ensure that feature maps with different channel numbers can be added correctly.
[0101] During the processing of input data by the ResNet-50 model, the model's computational process can be divided into four stages, gradually extracting semantic features from low-level to high-level to achieve accurate fungal classification. First, in the low-level feature extraction stage, the model uses the initial convolutional layers and pooling layers to extract basic visual features, such as edge information, color distribution, and local texture patterns, providing a foundation for subsequent deep feature learning. Subsequently, in the intermediate feature extraction stage, deeper convolution operations can learn more recognizable structured information, including morphological features, tissue arrangement patterns, and potential pathological features, further enhancing the model's understanding of complex images. Finally, in the high-level feature extraction stage, the deep features of the network are classified through a fully connected layer to determine whether the input image contains fungi.
[0102] In this module, the ResNet-50 input is a normalized 448×448 pixel image patch. It extracts features step by step through a four-stage residual module: the first stage uses 7×7 convolutions and max pooling to capture the hyphal edges and fluorescence intensity distribution. The second through fourth stages utilize a bottleneck structure (a sequence of 1×1-3×3-1×1 convolutions) to extract high-order spatial correlations within spore clusters. The output layer uses a sigmoid activation function to generate positive probability values, and a focal loss function (α=0.25, γ=2) is applied to mitigate the class imbalance caused by the lack of positive samples in clinical data. A threshold is set based on the probability values output by the ResNet-50: if the classification probability is greater than the threshold, the image is classified as fungus-positive and enters the subsequent processing module; if the classification probability is less than the threshold, the image is classified as fungus-negative and is archived without further analysis.
[0103] 3. Key Patch Area Screening Module
[0104] like Figure 1 As shown in the figure, for samples that were pre-screened as positive, the Adaptive Patch Selection Based on Local Variance Maximization algorithm was used to accurately locate the active fungal area from the whole slide. The algorithm first divided the slide image into 1024×1024 pixel sub-regions and calculated the grayscale variance of each region (formula: σ 2 =∑(x i -μ) 2 / N, where μ is the mean value of regional pixels). The variance value reflects the degree of aggregation of fungal spore distribution and the heterogeneity of fluorescence signals. The top 60 regions with the highest variance values were selected by descending order, effectively avoiding interference areas such as the alveolar macrophage phagocytosis area (the mean variance was 38.6% lower than that of the fungal area). The selected regions were further subjected to standardization preprocessing: (1) Color normalization: The Macenko method was used to eliminate the differences in staining batches and unify the hue of the fluorescence signal to the standard spectral range; (2) Noise suppression: A 3×3 median filter was used to eliminate discrete noise points (such as glass scratches and bubble artifacts); (3) Contrast enhancement: The CLAHE algorithm (grid 8×8, clipping coefficient 2.0) was used to enhance the hyphae-background contrast, effectively improving the clarity of the spore edge.
[0105] 4. Fungal classification module
[0106] The key patch area data obtained through the above steps, such as Figure 1 As shown, the image is input to the fungus classification model for further classification judgment. The present invention uses EfficientNet as the core network structure for fungus classification. EfficientNet adopts a composite scaling strategy that simultaneously adjusts the depth, width and input image resolution of the network to achieve the best balance of network performance.
[0107] One of EfficientNet's core innovations is its basic building block, MBConv, which combines depthwise separable convolution with the Squeeze-and-Excitation mechanism to significantly reduce computational costs while enhancing feature expression capabilities. Furthermore, EfficientNet introduces the Swish activation function, further enhancing the network's nonlinear expression capabilities.
[0108] Depthwise convolution decomposes the standard convolution into two independent steps: depthwise convolution and pointwise convolution. Depthwise convolution applies a convolution kernel to each input channel independently without mixing the information between channels. Pointwise convolution uses a 1×1 convolution kernel to convolve the output of depthwise convolution in order to mix the information between channels. The main advantage of depthwise separable convolution is that by decomposing the operation, the amount of computation is significantly reduced, making the model more suitable for running on mobile devices or embedded systems. Despite the reduced computation, depthwise separable convolution is still able to capture rich feature information and maintain the expressive power of the model.
[0109] The Squeeze-and-Excitation mechanism is a channel attention mechanism used to enhance the performance of convolutional neural networks (CNNs). This mechanism enables the network to adaptively recalibrate channel feature responses by explicitly modeling the interdependencies between channels, thereby improving the network's representation capabilities. The core idea of this mechanism is to achieve feature recalibration through two main steps: compression (Squeeze) and excitation (Excitation). The compression operation compresses the spatial dimension (H×W) of the feature map into a channel descriptor (channel descriptor) through global average pooling (GAP), thereby capturing global information. For a given feature map Channel descriptor generated by compression operation Calculated by formula (3):
[0110]
[0111] Where c represents the channel index.
[0112] The excitation operation learns the weights of each channel through a small fully connected network (usually a two-layer multi-layer perceptron, MLP), thereby achieving adaptive recalibration of channel features. First, the channel descriptor z is mapped to a lower dimensional space through a dimensionality reduction fully connected layer, and then nonlinearity is introduced through the ReLU activation function:
[0113] y=ReLU(FC(z)) (10)
[0114] Next, y is mapped back to the original channel dimension through a fully connected layer with increased dimension, and the weight s of each channel is generated through the Sigmoid activation function:
[0115] s=σ(FC(y)) (11)
[0116] Among them, σ represents the Sigmoid activation function. Finally, the generated channel weight s is multiplied by the original feature map U to achieve adaptive recalibration of the features:
[0117]
[0118] Here, ⊙ represents the multiplication of the principal elements. The Squeeze-and-Excitation mechanism can adaptively emphasize important features and suppress unimportant features by explicitly modeling the interdependencies between channels, thereby improving the network's representation capability.
[0119] Specifically, the EfficientNet model architecture consists of the following components: The network first performs preliminary feature extraction on the input image through a pre-processed standard convolutional layer with 32 channels, a 3×3 kernel size, and a stride of 2. The MBConv module, the core structure of EfficientNet, includes 1×1 convolutions for channel dimensionality expansion; 3×3 depthwise convolutions for spatial feature extraction; and finally, 1×1 convolutions for dimensionality reduction back to the original number of channels. Each MBConv module also includes batch normalization, Swish activation, SE attention, and residual connections. The model consists of eight MBConv modules, forming nine stages. The number of MBConv module repetitions in each stage is (1, 1, 2, 2, 3, 3, 4, 1). At the end of the network, a global average pooling layer is used to convert feature maps into feature vectors. Finally, a 1280-channel fully connected layer is used to achieve multi-category recognition of fungal subtypes. During the model training phase, the batch size was set to 16, the initial learning rate was 0.01, and a step-wise learning rate decay strategy was used for dynamic adjustment during the training process. This embodiment introduces the Focal Loss function during the model training process.
[0120] During training, the model leverages a large amount of labeled data, encompassing a variety of fungal slide images, to ensure high-precision classification capabilities. The model is capable of identifying a variety of clinically relevant fungi, including: Aspergillus, commonly seen in respiratory infections, with typical branching hyphae observed in pathological slides; Candida, a widespread yeast that can cause oral candidiasis, vaginitis, and other infections; Penicillium marneffei, primarily found in tropical regions and often infecting immunocompromised individuals; Cryptococcus, which can cause severe meningitis, primarily in immunosuppressed patients; and Mucor, a filamentous fungus that can cause mucormycosis, a highly invasive infection. Based on the automatic feature extraction capabilities of deep learning, the EfficientNet model effectively classifies the morphological features of different fungi, providing a precise, intelligent auxiliary tool for clinical pathology analysis and fungal infection diagnosis.
[0121] The EfficientNet classifier is used to achieve accurate classification of common clinical fungi (including Aspergillus, Candida, Cryptococcus, Mucor, and Penicillium marneffei) and obtain the probability confidence of the corresponding category. Classification result fusion analysis and structured report generation stage: The EfficientNet network can output the corresponding fungal category probability for each key Patch area. The present invention adopts the following result fusion strategy: Single Patch prediction result: Each Patch independently predicts the fungal category, and the network outputs the corresponding category probability. Slide-level fusion judgment strategy: Summarize the prediction results of all patches in the slide, count the number and proportion of patches where each type of fungus appears, and set the slide-level positive judgment threshold. Situation where the judgment conditions are not met: If no fungal category meets the above-mentioned proportion and confidence requirements, the slide image result is considered to be negative.
[0122] In the classification task of fungal fluorescence images, the deep learning model constructed in this study achieved an overall accuracy of 83.64% on the test set. The classification confusion matrix is detailed in Table 1. Table 2 presents the precision and recall of the six fungal classifications. Classification performance analysis shows that Candida albicans, Mucor, and Cryptococcus all achieved precision and recall rates exceeding 80%, demonstrating excellent performance. The recall rates for C. marneffei and Cryptococcus were relatively low, likely due to the small number of samples of these fungi in the training set.
[0123] Table 1. Confusion matrix of six fungal classifications
[0124]
[0125] Table 2. Precision and recall of six fungi
[0126] category Precision Recall Aspergillus 0.8571 0.9231 Candida 0.6667 0.8000 Mucor 0.8333 1.0000 Cryptococcus 1.0000 0.6667 Malnife 1.0000 0.7778 Negative 1.0000 0.7500
[0127] 5. Structured report generation
[0128] After the above determination process, the system automatically generates a structured fungal classification report. This report includes the following specific content: basic sample information (such as patient ID, sampling time, specimen type, etc.); clear fungal classification results and corresponding probability confidence levels; image labeling and visualization (annotated boxes) of fungal-positive areas; prompt information and recommended review strategies for difficult or uncertain samples; and the system's automatic diagnosis conclusion and recommended clinical treatment suggestions. This report can be directly imported into the hospital's pathology information management system (such as PACS / HIS / LIS) for immediate review by clinicians, assisting in rapid diagnostic decision-making, thereby significantly improving clinical work efficiency and diagnostic accuracy.
[0129] The step 2 of constructing a deep learning-based medical imaging tuberculosis detection method includes:
[0130] This paper proposes a full-process tuberculosis detection solution that integrates deep learning and microscopic image processing (flow chart as shown in the figure) Figure 3 Its core innovation lies in building a multi-stain compatible intelligent analysis framework: First, a residual network (ResNet-50) is used to classify the stain type of whole-slide images, quickly distinguishing between acid-fast and auramine O staining data. A local variance maximization algorithm is then used to dynamically filter high-information regions, using patch data as a standardized input model. Finally, a target detection model (YOLOX) is used to accurately locate Mycobacterium tuberculosis. This process overcomes the limitations of traditional microscopy methods that rely on manual interpretation. While maintaining high detection sensitivity, it effectively shortens the analysis time of a single sample and provides a standardized AI solution for rapid bedside diagnosis of tuberculosis.
[0131] 21. Imaging and Pretreatment of Acid-Fast and Auramine O Fluorescent Staining
[0132] Input Data Preparation and Staining Processing Methods In this embodiment, the process for acquiring medical imaging data is as follows: First, a clinician collects biomedical samples from the patient, including but not limited to bronchoalveolar lavage fluid, sputum, and tissue biopsy samples obtained via bronchoscopy. Subsequently, based on the detection requirements of Mycobacterium tuberculosis, pathological staining processing is performed using acid-fast staining or auramine O staining to enhance the identification contrast of Mycobacterium tuberculosis in microscopic images. In pathological image analysis, different staining methods can highlight the different structural features of cells and microorganisms. Therefore, the input data is classified according to the staining method, and detection models are constructed for acid-fast staining and auramine O staining, respectively.
[0133] Acid-fast staining is a staining method commonly used to detect acid-fast bacteria (such as Mycobacterium tuberculosis) and certain fungi. This method causes acid-fast microorganisms to appear red or purple under a microscope, while the background typically appears blue or other contrasting colors, thereby improving the visualization of microbial detection. In this study, an acid-fast staining detection model was used to analyze acid-fast staining images to identify the presence of pulmonary tuberculosis with acid-fast characteristics. The main challenge with acid-fast staining is staining consistency, which can lead to uneven color distribution due to differences in laboratory conditions, staining reagent concentration, and microscope parameters. To address this issue, color normalization techniques were used during model training to reduce the impact of different experimental conditions on model performance. Acridine Orange (Acridine Orange) is a fluorescent staining technique widely used to detect fungi and certain pathogenic microorganisms. This method binds to polysaccharides in the fungal cell wall, resulting in blue or green fluorescence when observed under a fluorescence microscope. In this study, an Acridine Orange staining detection model was used to analyze Acridine Orange input images to identify and classify the presence of pulmonary tuberculosis. Auramine O staining has the advantage of high sensitivity, enabling detection of low concentrations of fungi. However, the fluorescence signal can be affected by factors such as background noise and autofluorescence interference. Therefore, during data preprocessing, background suppression and morphological filtering are performed on the fluorescence signal to improve the detection accuracy of the model.
[0134] After staining, the sample sections were scanned at high resolution using a digital pathology scanner to obtain high-quality digital medical image data. Image data can be obtained as panoramic pathology scan files (glass slide format, such as SVS) or directly exported as PNG images. To facilitate subsequent staining analysis, the slide data was cropped, downsampled, or converted to a 448×448 pixel thumbnail image suitable for model input.
[0135] 22. Intelligent identification of dyeing types
[0136] Since there are significant differences in the image features of medical images using different staining methods (acid-fast staining and auramine O staining), the second step of the present invention automatically classifies the staining methods of the input image to ensure that the subsequent Mycobacterium tuberculosis target detection step adopts the most matching analysis strategy.
[0137] like Figure 3As shown, the specific implementation method of the staining method classification step is as follows: First, the input medical image data is uniformly adjusted to a standard size of 448×448 pixels; the adjusted medical image is input into a pre-trained ResNet-50 convolutional neural network. The network structure is the same as the ResNet-50 structure in the aforementioned embodiment 1, including multiple convolution stages and a fully connected layer (output layer); the ResNet-50 network automatically extracts and learns the features of the input image, and after multiple layers of convolution and pooling, it finally achieves accurate classification of acid-fast staining and auramine O staining through the fully connected layer; the network output is a binary classification probability value of the staining type, and the Softmax activation function is used to calculate the probability of the staining category. If the model output probability value is greater than 0.5, it is determined to be an acid-fast staining image; otherwise, it is determined to be an auramine O staining image; according to the staining classification results, they are respectively input into the corresponding subsequent Mycobacterium tuberculosis target detection model for analysis.
[0138] 23. Key Patch Area Screening and Preprocessing
[0139] like Figure 3 As shown in the figure, PNG data can be directly scaled to 1024×1024 for the next stage of analysis. For slide data, it is necessary to use the key patch area selection algorithm based on local variance maximization (Adaptive Patch Selection Based on Local Variance Maximization) to accurately locate the high-density area of Mycobacterium tuberculosis from the whole slide. The algorithm first divides the slide image into multiple 1024×1024 pixel sub-regions and calculates the grayscale variance of each patch area (formula: σ 2 =Σ(x_i-μ) 2 / N, where μ is the mean pixel value of the region). The variance value reflects the degree of tuberculosis aggregation to a certain extent. By setting a threshold, we exclude some patch areas with variance values less than the threshold, effectively avoiding some interference areas containing impurities.
[0140] Perform necessary data preprocessing on the input image (acid-fast or auramine O-stained image), including the following: color normalization; background suppression (using image thresholding and morphological filtering); adaptive histogram equalization; and image sharpening using the Sobel operator or Laplacian filter.
[0141] 24. Tuberculosis detection method
[0142] Different tuberculosis detection models were developed for different staining methods. After the staining method classification model, the acid-fast staining data was fed into the model trained on acid-fast staining for tuberculosis detection, while the auramine O data was processed by the model trained on auramine O.
[0143] The tuberculosis detection model is trained based on the YOLOX model. YOLOX is a high-performance target detection model. Figure 4 The YOLOX model structure is presented. This model inherits the core concept of the YOLO series single-stage detection, which is to directly predict the bounding box and category probability of the target, while significantly improving detection accuracy and efficiency through a series of innovative designs.
[0144] The YOLOX model mainly includes three core modules: (1) Backbone network: CSPDarknet network structure. This embodiment selects CSPDarknet as the backbone network, and its network structure is as follows: The model input is an RGB medical image of size 640×640×3. The model first uses the Focus layer to process the input features, improves channel utilization by cutting and rearranging, and outputs a feature size of 320×320×12. It is then compressed into a feature map of 320×320×64 through a 3×3 standard convolutional layer (Conv2D_BN_SiLU). The feature map then passes through multiple convolution and CSP modules in sequence: the first stage: CSPLayer structure, output size 160×160, channel 128; the second stage (Stage3): CSPLayer structure, output size 80×80, channel 256; the third stage (Stage4): CSPLayer structure, output size 40×40, channel 512; the fourth stage (Stage5): contains an SPPBottleneck module for multi-scale feature fusion, output size 20×20, channel 1024. (2) Feature enhancement module: The specific design details of the SPPBottleneck multi-scale feature fusion module are as follows: first, the feature channels are unified through a convolution layer; then, the features are spatially pooled at scales of 5×5, 9×9, and 13×13 to extract spatial context information at different scales; then, the original features are spliced and fused with the above multi-scale pooled features in the channel dimension; finally, the fused features are channel compressed through a convolution layer to output a high-quality feature map for the input of the subsequent target detection head. (3) Object Detection Head: In this embodiment, the YOLOX object detection head predicts three tasks separately: object classification, object position regression, and object confidence. The specific implementation details are as follows: The object detection head first refines the feature map through multiple 3×3 convolutional layers; then, through independent branch networks, it implements object classification, position regression, and object confidence prediction to predict whether the current location is a Mycobacterium tuberculosis category; and directly predicts the coordinates of the target center point and the width and height of the bounding box to predict whether the location contains an object.
[0145] YOLOX's main innovations include: adopting an anchor-free design, abandoning the complexity of predefined anchor boxes in traditional YOLO and directly predicting bounding boxes; introducing a decoupled head to separate classification and regression tasks into different branch networks, thereby improving the model's specialization and convergence speed; adopting the SimOTA label allocation strategy to dynamically optimize the distribution of positive and negative samples to further improve training efficiency; and combining advanced data augmentation technologies such as Mosaic and MixUp to enhance the model's generalization capabilities.
[0146] The loss function of the YOLOX model consists of three main parts: bounding box regression loss (Box Loss), confidence loss (Confidence Loss) and classification loss (Classification Loss). The following is a detailed analysis and formula description of its loss function:
[0147] (1) Bounding Box Regression Loss (Box Loss). YOLOX uses an improved IoU (Intersection over Union) loss function, namely GIoU (Generalized IoU) loss. GIoU loss not only considers the overlapping area of the predicted box and the true box, but also introduces the influence of the non-overlapping area. It can more accurately measure the similarity between the two boxes and has scale invariance. The calculation formula of GIoU is as follows:
[0148]
[0149] Among them, IoU is the intersection-over-union ratio of the predicted box and the real box, A c is the area of the minimum closure region containing the predicted box and the true box, and u is the union area of the predicted box and the true box.
[0150] The final GIoU loss is:
[0151] L GIoU =1-GIoU (3)
[0152] (2) Confidence Loss. Confidence loss is used to measure the existence of the target in the prediction box. For the prediction box containing the target (positive sample), the loss function is:
[0153]
[0154] Among them, C i is the confidence of the prediction, is the confidence of the true label (1 for positive samples), Is an indicator function that indicates whether the predicted box is a positive sample. For the predicted box that does not contain the target (negative sample), the loss function is:
[0155]
[0156] Among them, λ noobj Is a weight coefficient used to balance the loss of positive and negative samples.
[0157] (3) Classification Loss. Classification loss is used to measure the difference between the predicted category probability and the true category. For each predicted box containing the target, the classification loss is:
[0158]
[0159] Among them, p i (c) is the predicted class probability, is the true class probability (1 for the target class and 0 for the rest), Is an indicator function, indicating whether the predicted box is a positive sample. (4) Total loss function. The total loss function of YOLOX is the weighted sum of the above three parts. The specific formula is as follows:
[0160]
[0161] In YOLOX's loss function design, GIoU loss can better handle scale differences between bounding boxes and converge faster during training than traditional IoU loss. YOLOX uses a decoupled head design to separate classification and confidence prediction into different branches, allowing the model to focus more on each task. The SimOTA label allocation strategy dynamically optimizes the distribution of positive and negative samples, further improving training efficiency. Through these improvements, YOLOX significantly improves the accuracy and robustness of object detection while maintaining fast inference speed.
[0162] After obtaining the tuberculosis detection model's test results, we set two hyperparameters related to determining the positive or negative nature of the input data: the test result confidence threshold η and the threshold for the number of tuberculosis cases identified, φ. First, we use the confidence threshold η to filter out test results for which the model cannot confidently predict. Data containing fewer tuberculosis cases than the threshold φ is then labeled as negative. Conversely, data containing more tuberculosis cases than the threshold φ is labeled as positive.
[0163] In the acid-fast staining analysis task for Mycobacterium tuberculosis, Tables 3 and 4 demonstrate the performance of the present invention in identifying acid-fast stained Mycobacterium tuberculosis. The model constructed in this study achieved a classification accuracy of 93.44% at the PNG image level, with a detection sensitivity of 94.50% for positive samples and a specificity of 93.23% for negative samples. In the slide-level binary classification task, the model achieved an accuracy of 94.87%, with a detection sensitivity of 50% for positive samples and a specificity of 97.3% for negative samples. These results demonstrate that the proposed method is highly effective in identifying Mycobacterium tuberculosis.
[0164] Tables 5 and 6 show the performance of the proposed model in the task of analyzing auramine O-stained Mycobacterium tuberculosis. At the PNG image level, the proposed model achieved a classification accuracy of 98.90%, a sensitivity of 92.70%, and a specificity of 98.95%. These metrics demonstrate that the proposed model performs excellently at the image level of auramine O-stained glass slides, effectively identifying Mycobacterium tuberculosis-positive samples.
[0165] Table 3. Classification results of acid-fast stained Mycobacterium tuberculosis PNG data level
[0166]
[0167] Table 4. Slide-level classification results of acid-fast stained Mycobacterium tuberculosis
[0168]
[0169] Table 5. Classification results of auramine O-stained Mycobacterium tuberculosis PNG data
[0170] Count item: yin and yang Prediction results Yin and Yang Positive Negative total Positive 228 17 245 Negative 5 1754 1759 total 233 1771 2004
[0171] Table 6. Slide-level classification results of Mycobacterium tuberculosis stained with auramine O
[0172]
[0173] 25.Structured report generation
[0174] To meet the needs of clinical applications, this embodiment specifically designs a structured report generation mechanism: A threshold is set for the test results (e.g., target confidence ≥ 0.7, number of targets ≥ 5), and a positive result is automatically determined; if the threshold is not reached, the result is marked as negative or suspicious. The report content is formatted: the report automatically includes basic information about the pathology slide (patient name, sample number, collection date), visual annotation of the test image (detection box and heat map display), the number of detected targets, confidence score, diagnostic recommendations, and clinical explanations.
[0175] Automatic report export: Automatically output reports in structured file formats (such as PDF or XML), which can be directly imported into clinical hospital information systems (such as LIS / PACS / HIS systems) for immediate reference by clinicians.
[0176] Example 2
[0177] like Figure 5 This embodiment relates to an AI-based M-ROSE medical image pathogen detection device, including a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they are used to implement an AI-based M-ROSE medical image pathogen detection method of Example 1.
[0178] The contents described in the embodiments of this specification are merely an enumeration of the implementation forms of the inventive concept. The scope of protection of the present invention should not be regarded as limited to the specific forms described in the embodiments. The scope of protection of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.
Claims
1. An AI-based M-ROSE medical imaging pathogen detection method, characterized in that: Combining fluorescence microscopy, AI models, and computer vision technology to achieve automated analysis of fungal and tuberculosis detection includes the following steps: Step 1: The fungal sample is fluorescently stained, automatically scanned and identified, and automatically identified and classified using an AI model; Step 2: After the Mycobacterium tuberculosis sample is stained with acid-fast staining or auramine O staining, it enters the automatic scanning and recognition system, and automatically determines whether Mycobacterium tuberculosis is present in combination with the target detection model; Step 3: Obtain the automatic identification result to determine the type of fungus present in the sample or whether tuberculosis bacteria are present.
2. The AI-based M-ROSE medical imaging pathogen detection method according to claim 1, characterized in that: Step 1 involves a deep learning-based method for classifying fungal fluorescence images in medical imaging, specifically including: Step 11: Collect body fluid samples from the patient as medical imaging samples, prepare slides, and perform fluorescent staining to make specific types of fungi show high-contrast fluorescent signals under a fluorescence microscope, and collect medical imaging data; Step 12: Use the residual network model ResNet-50 to perform the first stage classification of the microscope slide image to determine whether it contains fungi. Negative samples are directly archived and do not enter the subsequent analysis process. Positive samples will be processed for subsequent fungal classification. Step 13: For positive samples, the patch extraction algorithm is used to accurately screen out the 60 key patch areas most likely to contain fungi, and the patches are preprocessed using color normalization, noise suppression, and contrast enhancement. Step 14: Use the convolutional neural network EfficientNet as the core classification model to classify the selected patches to identify the specific fungal type. EfficientNet adopts a composite scaling strategy: combining depthwise separable convolution and SE mechanism. Step 15: Output the classification results and generate a medical imaging fungus classification report based on the classification results.
3. The AI-based M-ROSE medical imaging pathogen detection method according to claim 2, characterized in that: The key patch area screening method described in step 13: Adaptive Patch Selection Based on Local Variance Maximization is adopted to preferentially select fungus aggregation areas as the patch areas to be used.
4. The AI-based M-ROSE medical imaging pathogen detection method according to claim 2, characterized in that: The EfficientNet described in step 14 includes: The convolutional neural network model consists of a front 32-channel convolution layer and multiple MBConv components, with a total of 9 stages, including 8 groups of MBConv structures and 1 final fully connected layer FC; each MBConv component consists of a front convolution layer and a set of cyclic network structures. The cyclic structure includes a convolution layer, a batch normalization layer, a Swish activation function, and a residual connection shortcut. Each component is stacked according to a different number of repetitions (1,1,2,2,3,3,4,1) to ensure layer-by-layer extraction of deep features; in the MBConv structure, the number of channels is first expanded by 1x1 convolution, and then k×k depthwise convolution is used. Convolution performs spatial feature extraction and adaptively adjusts channel weights through the SE module; the SE module uses global average pooling (GAP) to extract channel information, and combines the fully connected layer with dimensionality reduction and Swish activation function to process the feature expression of the image. Subsequently, the fully connected layer is used for dimensionality increase and the Sigmoid activation function to generate channel attention weights, which are multiplied with the original features to enhance important features and suppress redundant information; finally, MBConv uses 1x1 convolution to reduce the dimension and restore the number of channels, and introduces shortcut residual connections to optimize gradient propagation when the input and output shapes are consistent. DropConnect is used for regularization, which is only effective when shortcuts are present; during the feature downsampling process, the stride parameters of the pre-convolution of the first five groups of MBConv modules are set to 2, and the stride parameters of the remaining convolutions are set to 1, so the downsampling rate of the entire neural network is 32, achieving effective dimensionality reduction and expression of multi-level features; in addition, shortcut residual connections are used between the two convolutional layers of each component to optimize computational efficiency and gradient propagation; The network's feature aggregation and classification decision-making components include global pooling and fully connected layers, resulting in a final output channel count of 1280, ensuring sufficient extraction of high-level semantic information. During training, the batch size is set to 16, the initial learning rate is 0.01, and a step-wise learning rate reduction strategy is used for dynamic updates. Furthermore, to improve detection accuracy, a Focal Loss loss function is used to balance the amount of data across different categories. Focal Loss introduces a category weighting factor α and a tuning factor γ based on the standard cross-entropy loss to enhance learning for difficult-to-classify samples. Focal_Loss(p t )=-α(1-p t ) γ log(p t ) (1) Among them, p t Represents the model's predicted probability of the true category, that is, when the sample belongs to the positive class p t =p, and when the sample belongs to the negative class p t =1-p, the category weight factor α is used to balance the ratio of positive and negative samples, ensuring that the minority class will not be ignored by the model due to too little data; the adjustment factor γ controls the loss contribution of easy-to-classify samples, reducing the impact of the loss function on high-confidence samples and giving higher weights to low-confidence samples, thereby guiding the model to pay more attention to samples that are difficult to classify.
5. The AI-based M-ROSE medical imaging pathogen detection method according to claim 1, characterized in that: Step 3 involves a deep learning-based method for detecting tuberculosis in medical images, specifically including: Step 31: Collect bronchoalveolar lavage fluid (BALF), bronchoscopic biopsy tissue, sputum, or other biological samples from the patient; use acid-fast staining or acridine orange staining, and acquire medical imaging data using a microscope or digital pathology scanner; Step 32: Data preprocessing: Preprocess the medical images, including color standardization, background suppression, morphological filtering, and edge enhancement, to reduce noise interference during the staining process and improve the visualization characteristics of pathogens; Step 33: Staining classification: Use the ResNet-50 deep learning model to classify the medical images to distinguish between acid-fast staining data and auramine O staining data, ensuring that the correct analysis strategy is used for subsequent target detection; Step 34: Mycobacterium tuberculosis target detection: The classified image is input into the YOLOX target detection model for Mycobacterium tuberculosis identification. The model uses the SimOTA label assignment strategy for target optimization to improve the accuracy of Mycobacterium tuberculosis target detection. The adaptive multi-scale detection mechanism is used to enhance the ability to identify Mycobacterium tuberculosis of different sizes. Feature alignment technology is used to improve the stability and robustness of Mycobacterium tuberculosis detection. Step 35: Output of test results: Generate a medical imaging tuberculosis test report based on the test results.
6. The method according to claim 5, characterized in that Step 31 also includes color normalization, morphological filtering, adaptive contrast enhancement and background suppression.
7. The method according to claim 5, characterized in that The staining method classification described in step 33 includes: using the residual network ResNet-50 to classify acid-fast staining and auramine O staining; extracting the overall feature information of the image through the global feature extraction module, and using the local feature learning module to capture the detailed features in the image, so as to more comprehensively characterize the input data; finally using the SoftMax classification layer for final classification, and according to the classification results, assigning the input data to the acid-fast staining detection model or the auramine O staining detection model.
8. The method according to claim 5, characterized in that The YOLOX model uses CSPDarknet as the backbone network. The entire deep learning architecture is composed of multiple network components, which includes a total of 9 stages. The backbone network is responsible for basic feature extraction, SPBottleneck is used as the feature enhancement module, and YoloHead is used as the final detection head to jointly build the complete YOLOX target detection framework. The backbone network CSPDarknet includes: first, the input feature is segmented through the Focus layer, the input image size is 224×224, and after Focus processing, it enters the 32-channel 3×3 standard convolution Conv3×3, and then enters multiple CSPLayer structures for feature extraction; in Stage 2, CSPLayer3×3 convolution is used for convolution calculation, the input size is 112×112, the number of channels is 16, and the number of repetitions is 1; then, in Stage 3, CSPLayer3×3 convolution is continued for feature extraction, the input size is 112×112, the number of channels is expanded to 24, and the number of repetitions in this stage is 2; in Stage 4. Using CSPLayer5×5 convolution, the input size is reduced to 56×56, and the number of channels is increased to 40. This structure is repeated twice to ensure sufficient extraction of deep features. In Stage 5, the model regression uses the CSPLayer3×3 convolution structure, the input size is reduced to 28×28, and the number of channels is further expanded to 80. This part is repeated 3 times. In Stage 6, the CSPLayer5×5 convolution structure is used, the input size is reduced to 14×14, the number of channels is increased to 112, and it is repeated 3 times. In Stage 7, the CSPLayer5×5 convolution structure is continued to be used, the input size remains at 14×14, the number of channels is increased to 192, and it is repeated 4 times. In Stage 8, the CSPLayer3×3 convolution structure is used, the input size is reduced to 7×7, and the number of channels is increased to 320. This stage is only executed once. The feature enhancement layer Neck includes: using SPPBottleneck to improve the multi-scale expression ability of features; SPPBottleneck consists of spatial pyramid pooling (SPP) and 1×1 convolution. The SPP structure extracts spatial information through pooling kernels of different scales (5×5, 9×9, 13×13), concatenates the multi-scale pooling results with the original features, and finally compresses the features through 1×1 convolution; The detection head YoloHead includes: feature refinement through multiple 3×3 Conv2D_BN_SiLU convolutional layers, with a final output channel number of 1280, and outputs classification Cls, regression Reg, and target confidence Obj results respectively; the structure adopts an anchor-free design and combines a multi-scale feature fusion strategy with a downsampling strategy; the classification loss is used to calculate the error between the predicted category probability and the true category, and the final total loss function is the weighted sum of GIoU loss, confidence loss, and classification loss.
9. The method according to claim 8, wherein The GIoU loss not only considers the overlapping area of the predicted box and the true box, but also introduces the influence of the non-overlapping area, which can more accurately measure the similarity between the two boxes and has scale invariance. The calculation formula of GIoU is as follows: Among them, IoU is the intersection-over-union ratio of the predicted box and the real box, A c is the area of the minimum closure region containing the predicted box and the true box, and u is the union area of the predicted box and the true box; the final GIoU loss is: L GIoU =1-GIoU (3) Confidence loss is used to measure the existence of the target in the prediction box; for the prediction box containing the target (positive sample), the loss function is: Among them, C i is the confidence of the prediction, is the confidence of the true label, which is 1 for positive samples, Is an indicator function, indicating whether the predicted box is a positive sample; For the predicted box that does not contain the target, that is, the negative sample, the loss function is: Among them, λ noobj Is a weight coefficient used to balance the loss of positive and negative samples; The classification loss is used to measure the difference between the predicted category probability and the true category; for each predicted box containing the target, the classification loss is: Among them, p i (c) is the predicted class probability, is the true class probability, which is 1 for the target class and 0 for the rest, Is an indicator function, indicating whether the predicted box is a positive sample.
10. An AI-based M-ROSE medical imaging pathogen detection device, characterized in that: The invention comprises a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement an AI-based M-ROSE medical imaging pathogen detection method according to any one of claims 1 to 9.
Citation Information
Cited By
Deep learning-driven sputum smear pathogen image recognition method and system
CN122116355A