Nasal secretion smear diagnosis method and system based on deep learning
Through deep learning-based nasal secretion smear diagnostic methods, sliding window blocking and cell detection/classification models are used to solve the problems of low efficiency and strong subjectivity of traditional diagnostic methods, and more efficient and accurate cell counting is achieved.
Patent Information
- Application Number
- PCT/CN2024/115074
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-09
- Filing Date
- 2024-08-28
- Publication Date
- 2025-06-19
AI Technical Summary
Traditional nasal secretion smear diagnosis methods have problems such as strong subjectivity, low diagnostic efficiency, and prone to cell leakage or repeated counting.
The deep learning-based nasal secretion smear diagnosis method is used to read the smear images in pieces through sliding windows, and combine cell detection and classification models to automatically detect and classify cells in nasal secretions.
It improves diagnostic efficiency and accuracy, reduces the doctor's dependence on subjective judgment, and ensures the accuracy of cell counts.
Smart Images

Figure CN2024115074_19062025_PF_FP_ABST
Abstract
Description
A nasal secretion smear diagnosis method and system based on deep learning Technical Field
[0001] The present invention relates to the technical field of nasal secretion smear diagnosis, and in particular to a nasal secretion smear diagnosis method and system based on deep learning. Background Art
[0002] The traditional nasal secretion smear diagnosis method mainly relies on doctors to examine the types and numbers of various cells in nasal secretions under a microscope. Due to the huge number of cells, which can be as many as tens of thousands of cells, doctors are unable to accurately count the number of each type of cells. They rely solely on experience and intuition, which has technical problems such as strong subjectivity and dependence on the doctor's experience. In addition, because cells are scattered on the smear and the single field of view of the microscope is limited, doctors need to constantly move the slide to view all the cells. This method is also prone to cell omissions or repeated counting, and the diagnostic efficiency is low.
[0003] Summary of the Invention
[0004] The present invention provides a deep learning-based nasal secretion smear diagnosis method and system to solve the problems existing in the above-mentioned prior art. The technical solution is as follows:
[0005] In one aspect, a deep learning-based nasal smear diagnosis method is provided, comprising:
[0006] S1. Read the nasal secretion smear to be diagnosed in blocks in a sliding window manner. Each time, the image of the window range is read, which is called the window image;
[0007] S2, performing preprocessing operations on the read window image;
[0008] S3. Input the preprocessed window image into the trained cell detection model. The cell detection model performs a series of feature extraction on the input image, and then enhances the pyramid feature map by bottom-up and top-down feature fusion. Finally, the coordinates and scores of the rectangular boxes of the detected targets are obtained according to the feature maps of each scale. According to the set score threshold, the rectangular boxes below the threshold are filtered out to obtain the rectangular box set R.
[0009] S4. According to the coordinates of the rectangular box in R, a corresponding image is cropped from the original window image and preprocessed, and the preprocessed corresponding image is input into a trained cell classification model, and the cell classification model outputs the cell category of the corresponding image;
[0010] S5. Post-process the cells in the corresponding image to obtain the cell types of the original nasal secretion smear to be diagnosed.
[0011] In another aspect, a deep learning-based nasal smear diagnostic system is provided, the system comprising:
[0012] A reading module is used to read the nasal secretion smear to be diagnosed in blocks in a sliding window manner, and each time the image of the window range is read, it is called a window image;
[0013] A preprocessing module is used to perform preprocessing operations on the read window image;
[0014] The cell detection module is used to input the preprocessed window image into the trained cell detection model. The cell detection model performs a series of feature extraction on the input image, enhances the pyramid feature map through bottom-up and top-down feature fusion, and finally obtains the coordinates and scores of the rectangular boxes of the detected targets based on the feature maps of each scale. According to the set score threshold, the rectangular boxes below the threshold are filtered out to obtain the rectangular box set R;
[0015] A cell classification module is used to crop the corresponding image from the original window image according to the coordinates of the rectangular box in R and perform preprocessing operations, and input the preprocessed corresponding image into a trained cell classification model, and the cell classification model outputs the cell category of the corresponding image;
[0016] The post-processing module is used to perform post-processing on the cells of the corresponding image to obtain the cell categories of the original nasal secretion smear to be diagnosed.
[0017] On the other hand, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned deep learning-based nasal secretion smear diagnosis method.
[0018] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned deep learning-based nasal secretion smear diagnosis method.
[0019] The beneficial effects brought about by the technical solution provided by the present invention include at least:
[0020] The present invention reads the nasal secretion smear to be diagnosed in blocks in a sliding window manner, divides an image into multiple blocks and inputs them into a deep learning model, thereby improving the efficiency of diagnosis. Moreover, through the deep learning model of the present invention (in order to overcome the problems of large computational complexity and poor real-time performance of the deep learning model, the present invention chooses to implement the efficient improved RTMDet as the detection network and the improved MobileNetV3 as the classification network, instead of using the same network for detection and classification as in the prior art), not only the doctor's diagnostic efficiency is greatly improved, but the accuracy is also higher than that of traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0022] FIG1 is a flow chart of a nasal secretion smear diagnosis method based on deep learning provided by an embodiment of the present invention;
[0023] FIG2 is a schematic diagram of an improved RTMDet network structure provided by an embodiment of the present invention;
[0024] FIG3 is a schematic diagram of the backbone network structure of the improved RTMDet provided in an embodiment of the present invention;
[0025] FIG4 is a schematic diagram of the neck network structure of the improved RTMDet provided in an embodiment of the present invention;
[0026] FIG5 is a schematic diagram of the structure of SPPELAN-E provided in an embodiment of the present invention;
[0027] FIG6 is a schematic diagram of the existing SPP structure;
[0028] FIG7 is a schematic diagram of the structure of the ECA attention mechanism provided by an embodiment of the present invention;
[0029] FIG8 is a schematic diagram of the improved MobileNetV3 network structure provided by an embodiment of the present invention;
[0030] FIG9 is a schematic diagram of the MobileNetV3 bneck-E structure provided by an embodiment of the present invention;
[0031] Figure 10 is a schematic diagram of the existing MobileNetV3 bneck structure;
[0032] FIG11 is a schematic diagram of window cell deduplication provided by an embodiment of the present invention;
[0033] FIG12 is a block diagram of a nasal secretion smear diagnosis system based on deep learning provided by an embodiment of the present invention;
[0034] FIG13 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0035] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0036] An embodiment of the present invention provides a deep learning-based nasal secretion smear diagnosis method, which can be implemented by an electronic device, such as a terminal or a server. As shown in Figure 1, a flowchart of a deep learning-based nasal secretion smear diagnosis method is provided. The processing flow of the method may include the following steps:
[0037] S1. Read the nasal secretion smear to be diagnosed in blocks in a sliding window manner. Each time, the image of the window range is read, which is called the window image;
[0038] The image of a nasal secretion scan is huge, generally 200,000x200,000 pixels in size (the effective physical area of a slide is approximately 50mmX24mm, and a secretion smear generally occupies half of that, i.e., 24mmX24mm. The mpp value of the shooting light path is approximately 0.1um / pix=1e-4mm / pix, 24mm / (1e-4mm / pix)=240,000pix). It cannot be read into the memory in its entirety, so it needs to be read in blocks. The embodiment of the present invention reads data in blocks in a windowed manner. The window size is 2048*2048, and the upper and lower and left and right overlaps between windows are 128. The image of the window range read each time is called a window image.
[0039] S2, performing preprocessing operations on the read window image;
[0040] Optionally, the preprocessing includes:
[0041] Scaling the window image to the size required by the cell detection model for the input image (the improved RTMDet in the embodiment of the present invention requires the input image size to be 640x640, so the window image is scaled to this size);
[0042] A predefined mean vector is subtracted from each pixel value in the image, and then divided by a predefined standard deviation vector (the embodiment of the present invention uses the mean and standard deviation calculated on the large public dataset ImageNet) to ensure that the pixel values in the image are evenly distributed, which is beneficial for subsequent model detection and classification.
[0043] Assume that the window image is a bgr image I of uint8 type with a size of 2048x2048x3. Scale the image I first, and then standardize the scaled image I to obtain image C = (I-(123.675,116.28,103.53)) / (58.395,57.12,57.375), where (123.675,116.28,103.53) is the predefined mean vector and (58.395,57.12,57.375) is the predefined standard deviation vector.
[0044] S3. Input the preprocessed window image into a trained cell detection model (the nasal secretion smear diagnosis method based on deep learning in an embodiment of the present invention is a technology that uses a deep learning algorithm to analyze and process nasal secretion smear scan images. By utilizing a large amount of nasal secretion scan image data, a large amount of nasal secretion cell annotation data, and the powerful pattern recognition ability of the deep learning algorithm, a cell detection model and a cell classification model can be trained to automatically detect and identify neutrophils, eosinophils, and other cells in nasal secretions, assisting doctors in disease diagnosis). After the cell detection model performs a series of feature extraction on the input image, it enhances the pyramid feature map through bottom-up and top-down feature fusion. Finally, the coordinates and scores of the rectangular boxes of the detected targets are obtained according to the feature maps at each scale. According to a set score threshold (e.g., 0.3), the rectangular boxes below the threshold are filtered out to obtain a rectangular box set R;
[0045] Optionally, as shown in FIG2 , the cell detection model is an improved RTMDet network structure, wherein the improved RTMDet includes: an input, a backbone network, a neck network and a head network, a loss function, and an output;
[0046] The input image enters the backbone network. As shown in Figure 3, the image first passes through four Conv Module structures to reduce the size, reducing the dimension while retaining important features. Then, it passes through three structures composed of CSPLayer and Conv Module to continue extracting important features of the context and generate two output feature maps of different sizes, recorded as output 2 and output 3, for subsequent feature fusion. Then, it passes through the improved spatial pyramid pooling structure SPPELAN-E to generate multi-scale features for detecting targets of different sizes. Finally, it passes through a CSPLayer structure to further extract features, enhance the performance of the model, and generate another output feature map, recorded as output 1.
[0047] The three outputs 1, 2, and 3 enter the neck network as the three inputs 1, 2, and 3 respectively. As shown in Figure 4, input 1 is first passed through a Conv Module structure and an upsampling operation to adjust the size and number of channels of the feature map, and then spliced and fused with input 2 to enrich the multi-scale information, and a CSPLayer structure is used to enhance the feature expression ability. Since the nasal secretion target is small, dense, and adhered, a lot of noise will be generated during feature fusion, which will affect the quality of the final fused feature map. Therefore, after each splicing and fusion and a CSPLayer structure operation, an ECA attention mechanism structure is added to suppress the fused noise and enhance the attention to nasal secretions. Then, a Conv Module structure and an upsampling operation are performed to adjust the size and number of channels of the feature map, and then spliced and fused with input 3 to further enrich the multi-scale information, and a CSPLayer structure is used to enhance the feature expression ability; then, an ECA attention mechanism structure is passed through, and two branches are generated at this time, one of which is used as the output 3 of the neck network, and the other branch passes through a Conv After the Module structure performs a downsampling operation to adjust the size, it is further fused with the previously generated feature map to enrich more target information; then it passes through a CSPLayer structure to enhance the feature expression ability, and then passes through the Conv Module structure to perform a downsampling operation to adjust the size, and then it is further fused with the previously generated feature map to enrich more target information. At this time, two branches are also generated, one of which is used as the output 2 of the neck network, and the other branch passes through a CSPLayer structure to enhance the feature expression ability, and then passes through the ECA attention mechanism structure, and finally outputs 1;
[0048] In the head network part, the three outputs 1, 2, and 3 of the neck network are decoded through some convolution modules, and finally the coordinates and scores of the rectangular box of the target (detected cells) are generated.
[0049] Optionally, the loss function includes classification loss and bounding box regression loss;
[0050] The classification loss is formulated as follows:
[0051] losscls=Quality Focal Loss
[0052] where p t Represents the predicted probability of the true category of the sample, α t and γ are adjustable weight parameters, y represents the actual label of the sample, and the scaling factor of Folcal Loss (1-p t ) γDuring training, it can reduce the weight of the simple class on the loss and quickly focus the model on the difficult class. t Used to adjust the ratio between positive and negative sample losses;
[0053] In order to solve the inconsistency problem between the training and testing phases, Quality Folcal Loss combines the positioning quality IOU value and the classification score on the basis of Folcal Loss. Since its label is a continuous value between 0 and 1, the two parts of Folcal Loss are improved. The formula is as follows:
[0054] Quality FocalLoss(σ)=-|y-σ| β ((1-y)log(1-σ)+ylog(σ))
[0055] Where y represents the actual label of the sample, and σ represents the label value obtained after combining the IOU value;
[0056] The bounding box regression loss uses GIoU Loss, which is used to calculate the relationship between the overlapping areas of two boxes. The larger the overlapping area, the smaller the loss, and vice versa. GIoU Loss is between [0, 2], and its value is limited to a small range, so the network will not fluctuate violently, and thus has better stability. The formula is as follows:
[0057] GIOU Loss=1-GIOU
[0058] Where A and B represent two bounding boxes, C is the smallest bounding box that can enclose them, and IOU is the intersection-union ratio of A and B.
[0059] Optionally, the improved spatial pyramid pooling structure SPPELAN-E is used to enable the network to process input images of different sizes without forcibly cropping or scaling the input images, thereby reducing information loss and improving model performance;
[0060] As shown in FIG5 , the SPPELAN-E retains one of the four branches of the existing SPP structure (as shown in FIG6 , the existing SPP structure performs four different operations on the four input branches, retains one of the branches without any change, and performs maximum pooling operations of different sizes on the remaining three branches to generate feature maps of different sizes, and finally splices the output results of the four different branches into the final output). It uses a dilated convolution to expand the receptive field and adds a residual structure to the output so that the inputs of the remaining branches come from the output of the previous branch. This expands the receptive field while enhancing the feature expression ability of the model, and finally splices the output results of the four different branches into the final output.
[0061] Compared to the existing SPP structure, the SPPELAN-E of the present invention retains more semantic information and performs better when detecting small targets such as nasal secretions. This is because the SPP structure repeatedly uses the max pooling operation, which only takes the maximum value of each element and ignores other elements, resulting in the loss of useful information in some feature maps. However, the dilated convolution expands the receptive field by introducing a dilation factor, without losing useful information. In addition, the continuous residual structure can also retain a certain amount of effective information.
[0062] Optionally, as shown in FIG7 (in the figure, X represents the input feature map, W represents the width of the feature map, H represents the height of the feature map, C represents the number of channels of the feature map, K represents the size of the convolution kernel used to calculate the attention weight of each channel, and σ represents the sigmoid activation function), the ECA attention mechanism structure first compresses the spatial dimension of the input feature map to 1 while keeping the number of channel dimensions unchanged;
[0063] Secondly, a one-dimensional convolution of size K is used to capture local cross-channel interaction information. This local cross-channel interaction strategy without dimensionality reduction helps to more effectively interact between channels while maintaining the correlation between channels, thereby improving the network's expressiveness and performance.
[0064] The output is then passed through a sigmoid activation function to ensure that the output is between 0 and 1, generating a channel attention weight.
[0065] Finally, the channel attention weight is multiplied by the original input feature map to obtain the final feature map.
[0066] Since ECA can capture the relationship between different channels, thereby improving the ability of feature representation, by utilizing this mechanism, the embodiments of the present invention can effectively enhance the representation ability of the network without increasing excessive parameters and computational costs, thereby increasing the accuracy of nasal secretion diagnosis.
[0067] S4. According to the coordinates of the rectangular box in R, a corresponding image is cropped from the original window image and preprocessed, and the preprocessed corresponding image is input into a trained cell classification model, and the cell classification model outputs the cell category of the corresponding image;
[0068] The cropped image of the embodiment of the present invention does not have a rectangular frame (that is, the image is cropped from the original image according to the coordinates of the rectangular frame). An image will only have one cell (the number of cells detected is the number of images, and each image has one cell). Then, the cell classification model classifies each cell. Specifically, a rectangular frame r = (x, y, w, h) is sequentially taken from the set R, and the rectangular frame r is cropped from the original window image I to obtain image A. The image A is scaled to a size of 224x224 (Mobi improved cell classification model) Image B is obtained, and image B is normalized to obtain image D = (B-(123.675, 116.28, 103.53)) / (58.395, 57.12, 57.375). Image D is input into the improved MobileNetV3 cell classification model to obtain the scores of the three categories of eosinophils, neutrophils, and other cells. The category corresponding to the highest score is selected as the category of the cell from the scores output by the model.
[0069] Optionally, as shown in FIG8 , the cell classification model is an improved MobileNetV3 network structure, including input, backbone network, head network, loss function and output;
[0070] The input image enters the backbone network for feature extraction. As shown in Figure 9, the backbone network of the improved MobileNetV3 includes multiple bneck-Es. For the multiple bneck structures of the existing MobileNetV3 backbone network (as shown in Figure 10, the backbone network of the existing MobileNetV3 has multiple bneck structures. When the step size is 1, the feature map first undergoes a 1×1 convolution to expand the number of channels, followed by a depth-separable convolution to reduce the model parameters while extracting features, and finally a 1×1 convolution is used to restore the original number of channels; when the step size is 2, the feature map will pass through two branches, where the main branch is similar to the step size of 1, except that An SE attention mechanism is added after the depthwise separable convolution to enhance the representation of important features, and the result is added to another branch to avoid gradient explosion and improve model performance. However, for nasal secretion images, due to differences in staining and scraping techniques, the existing network cannot effectively extract key features. When the step size is 1, no modification is made. In the bneck structure with a step size of 2, the SE attention mechanism (SE attention mechanism is similar to the ECA attention mechanism, but uses a fully connected layer to obtain channel information, which increases computational efficiency) is modified to the ECA attention mechanism to speed up computational efficiency and reduce the number of parameters. In addition, a branch is added after the depthwise separable convolution to enhance network performance.
[0071] After feature extraction, the network passes through the head, including some convolutional layers and fully connected layers, to calculate the classification loss value of the extracted feature map, and finally output the cell classification results (the output cell categories include eosinophils, neutrophils, and other cells).
[0072] Optionally, the loss function uses cross entropy loss, and the formula is as follows:
[0073] Where p represents the predicted probability of the sample in this category, and y represents the sample label.
[0074] S5. Post-process the cells in the corresponding image to obtain the cell types of the original nasal secretion smear to be diagnosed.
[0075] The post-processing includes:
[0076] Restore the coordinates of the cell detected in the window image with the coordinates of the upper left corner of the window as the origin to the original image. Assuming that the coordinates of the cell are r = (x, y, w, h) and the coordinates of the upper left corner of the window are (l, t), then the coordinates of the cell on the original image are (x+l, y+t, w, h);
[0077] The cells in the overlapping area of the window are deduplicated. For each cell, the IOU value between it and all cells detected in the adjacent window is calculated. If the IOU value is greater than 0.5, it is considered to be the same cell.
[0078] Figure 11 is a schematic diagram of window cell deduplication. As shown in the figure, the central area is the cells of the current window, and the triangular area is the overlapping cells with the eight adjacent windows. Therefore, in order to remove duplicate cells, it is necessary to calculate the IOU value with each cell in the surrounding eight windows to achieve the purpose of deduplication.
[0079] As shown in FIG12 , an embodiment of the present invention further provides a nasal secretion smear diagnosis system based on deep learning, the system comprising:
[0080] The reading module 1210 is used to read the nasal secretion smear to be diagnosed in blocks in a sliding window manner, and each time the image of the window range is read, it is called the window image;
[0081] A preprocessing module 1220 is used to perform preprocessing operations on the read window image;
[0082] The cell detection module 1230 is configured to input the preprocessed window image into a trained cell detection model. The cell detection model performs a series of feature extractions on the input image, then enhances the pyramid feature map through bottom-up and top-down feature fusion. Finally, the coordinates and scores of the rectangular boxes of the detected targets are obtained based on the feature maps at each scale. Based on a set score threshold, rectangular boxes with scores below the threshold are filtered out to obtain a set of rectangular boxes R.
[0083] A cell classification module 1240 is configured to crop a corresponding image from the original window image according to the coordinates of the rectangular frame in R and perform preprocessing operations, and input the preprocessed corresponding image into a trained cell classification model, and the cell classification model outputs the cell category of the corresponding image;
[0084] The post-processing module 1250 is configured to perform post-processing on the cells in the corresponding image to obtain the cell types of the original nasal secretion smear to be diagnosed.
[0085] An embodiment of the present invention provides a deep learning-based nasal secretion smear diagnosis system, whose functional structure corresponds to a deep learning-based nasal secretion smear diagnosis method provided by an embodiment of the present invention, and will not be repeated here.
[0086] Figure 13 is a structural schematic diagram of an electronic device 1300 provided in an embodiment of the present invention. The electronic device 1300 may have relatively large differences due to different configurations or performance, and may include one or more processors (central processing units, CPU) 1301 and one or more memories 1302, wherein the memory 1302 stores at least one instruction, and the at least one instruction is loaded and executed by the processor 1301 to implement the steps of the above-mentioned deep learning-based nasal secretion smear diagnosis method.
[0087] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions, which can be executed by a processor in a terminal to implement the above-mentioned deep learning-based nasal secretion smear diagnosis method. For example, the computer-readable storage medium can be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0088] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0089] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A nasal secretion smear diagnosis method based on deep learning, characterized in that: The method comprises: S1, reading the nasal secretion smear to be diagnosed in blocks in a sliding window manner, and reading the image of the window range each time is called the window image; S2, performing preprocessing operation on the read window image; S3, input the preprocessed window image into the trained cell detection model, the cell detection model performs a series of feature extraction on the input image, enhances the pyramid feature map by bottom-up and top-down feature fusion, and finally obtains the coordinates and score of the rectangular box of the detected target according to the feature map of each scale, and filters out the rectangular boxes below the threshold according to the set score threshold to obtain the rectangular box set R; S4, according to the coordinates of the rectangular box in R, cutting the corresponding image from the original window image and performing a preprocessing operation, inputting the preprocessed corresponding image into the trained cell classification model, and the cell classification model outputs the cell category of the corresponding image; S5. Post-process the cells of the corresponding image to obtain the cell categories of the original nasal secretion smear to be diagnosed.
2. The method according to claim 1, characterized in that The preprocessing comprises: Scaling the window image to the size required by the cell detection model for the input image; The scaled image is normalized by subtracting a predefined mean vector from each pixel value in the image and dividing it by a predefined standard deviation vector to make the pixel values in the image evenly distributed.
3. The method according to claim 1, characterized in that: The cell detection model is an improved RTMDet network structure, and the improved RTMDet includes: input, backbone network, neck network and head network, loss function and output; the input image enters the backbone network, and the image first passes through four Conv Module structures to reduce the size, reducing the dimension while retaining important features, and then passes through three structures composed of CSPLayer and Conv Module to continue to extract important features of the context, and generate two output feature maps of different sizes, recorded as output 2 and output 3, for subsequent feature fusion, and then passes through the improved spatial pyramid pool structure SPPELAN-E to generate multi-scale features for detecting targets of different sizes, and finally passes through a CSPLayer structure to further extract features, enhance the performance of the model, and generate another output feature map, recorded as output 1; The three outputs 1, 2, and 3 are respectively used as the three inputs 1, 2, and 3 to enter the neck network. First, input 1 passes through a Conv Module structure and an upsampling operation to adjust the size and number of channels of the feature map, and then is spliced and fused with input 2 to enrich the multi-scale information, and a CSPLayer structure is used to enhance the feature expression ability. Since the nasal secretion target is small, dense and adherent, a lot of noise will be generated during feature fusion, which will affect the quality of the final fused feature map. Therefore, after each splicing and fusion and a CSPLayer structure operation, an ECA attention mechanism structure is added to suppress the fused noise and enhance the attention to the nasal secretions. Then, a Conv Module structure and Upsampling operation, adjusts the size and number of channels of the feature map, and then concatenates and fuses it with the input 3 to further enrich the multi-scale information, and enhances the feature expression ability through a CSPLayer structure; then passes through an ECA attention mechanism structure, and two branches are generated at this time, one of which is used as the output 3 of the neck network, and the other branch is further fused with the previously generated feature map after a downsampling operation through a Conv Module structure to enrich more target information; then a CSPLayer structure is used to enhance the feature expression ability, and then the Conv Module structure is used to downsample and adjust the size, and then it is further fused with the previously generated feature map to enrich more target information. At this time, two branches are also generated, one of which is used as the output 2 of the neck network, and the other branch is enhanced by a CSPLayer structure to enhance the feature expression ability, and then passes through the ECA attention mechanism structure, and finally outputs 1; In the head network part, the three outputs 1, 2, and 3 of the neck network are respectively decoded through some convolution modules to finally generate the coordinates and score of the target rectangular box.
4. The method according to claim 3, characterized in that The loss function includes classification loss and bounding box regression loss; The classification loss is formulated as follows: loss cls =Quality Focal Loss; where p t Represents the predicted probability of the true category of the sample, α t and γ are adjustable weight parameters, y represents the actual label of the sample, and the scaling factor of Folcal Loss (1-p t ) γ During training, it can reduce the weight of the simple class on the loss and quickly focus the model on the difficult class. t Used to adjust the ratio between positive and negative sample losses; Quality Folcal Loss combines the positioning quality IOU value and the classification score on the basis of Folcal Loss. Since its label is a continuous value between 0 and 1, the two parts of Folcal Loss are improved. The formula is as follows: Quality Focal Loss(σ)=-|y-σ| β ((1-y)log(1-σ)+y log(σ)); Where y represents the actual label of the sample, and σ represents the label value obtained after combining the IOU value; The bounding box regression loss uses GIoU Loss, which is used to calculate the relationship between the overlapping areas of two boxes. The larger the overlapping area, the smaller the loss, and vice versa. In addition, GIoU Loss is between [0,2], and its value is limited to a smaller range, so the network will not fluctuate violently, and thus has better stability. The formula is as follows: GIOU Loss = 1-GIOU; Where A and B represent two bounding boxes, C is the smallest bounding box that can enclose them, and IOU is the intersection-union ratio of A and B.
5. The method according to claim 3, characterized in that: The improved spatial pyramid pooling structure SPPELAN-E enables the network to process input images of different sizes without forcibly cropping or scaling the input images, thereby reducing information loss and improving model performance; The SPPELAN-E retains one of the four branches of the existing SPP structure, uses a dilated convolution to expand the receptive field, and adds a residual structure to the output so that the inputs of the remaining branches come from the output of the previous branch. This not only expands the receptive field but also enhances the feature expression ability of the model. Finally, the output results of the four different branches are spliced into the final output.
6. The method according to claim 3, characterized in that The ECA attention mechanism structure first compresses the spatial dimension of the input feature map to 1 while keeping the number of channel dimensions unchanged; Secondly, a one-dimensional convolution of size K is used to capture local cross-channel interaction information. This local cross-channel interaction strategy without dimensionality reduction helps to more effectively interact between channels while maintaining the correlation between channels, thereby improving the network's expressiveness and performance. Then the output is passed through the sigmoid activation function to ensure that the output is between 0 and 1, and the channel attention weight is generated; Finally, the channel attention weight is multiplied by the input original feature map to obtain the final feature map.
7. The method according to claim 1, characterized in that The cell classification model is an improved MobileNetV3 network structure, including input, backbone network, head network, loss function and output; The input image enters the backbone network for feature extraction. The improved MobileNetV3 backbone network includes multiple bneck-Es. For the multiple bneck structures of the existing MobileNetV3 backbone network, no modification is made when the step size is 1. In the bneck structure with a step size of 2, the SE attention mechanism is modified to the ECA attention mechanism to speed up the calculation efficiency and reduce the number of parameters. In addition, a branch is added after the depthwise separable convolution to enhance the performance of the network. After feature extraction, the head network, including some convolutional layers and fully connected layers, calculates the classification loss value of the extracted feature map, and finally outputs the cell classification result.
8. The method according to claim 7, characterized in that The loss function uses cross entropy loss, and the formula is as follows: Where p represents the predicted probability of the sample in this category, and y represents the sample label.
9. The method according to claim 1, characterized in that: The post-processing comprises: Restore the coordinates of the cell detected in the window image with the coordinates of the upper left corner of the window as the origin to the original image. Assuming that the coordinates of the cell are r = (x, y, w, h), and the coordinates of the upper left corner of the window are (l, t), then the coordinates of the cell on the original image are (x+l, y+t, w, h); The cells in the overlapping area of the window are removed. For each cell, the IOU value between it and all cells detected in the adjacent window is calculated. If the IOU value is greater than 0.5, it is considered to be the same cell.
10. A nasal secretion smear diagnosis system based on deep learning, characterized in that: The system comprises: A reading module is used to read the nasal secretion smear to be diagnosed in blocks in a sliding window manner, and each time the image of the window range is read, it is called a window image; A preprocessing module, used for performing preprocessing operations on the read window image; The cell detection module is used to input the preprocessed window image into the trained cell detection model. The cell detection model performs a series of feature extraction on the input image, and then enhances the pyramid feature map by bottom-up and top-down feature fusion. Finally, the coordinates and scores of the rectangular boxes of the detected targets are obtained according to the feature maps of each scale. According to the set score threshold, the rectangular boxes below the threshold are filtered out to obtain the rectangular box set R. A cell classification module, used for cutting a corresponding image from the original window image and performing a preprocessing operation according to the coordinates of the rectangular frame in R, and inputting the preprocessed corresponding image into a trained cell classification model, and the cell classification model outputs the cell category of the corresponding image; The post-processing module is used to perform post-processing on the cells of the corresponding image to obtain the cell categories of the original nasal secretion smear to be diagnosed.
Citation Information
Patent Citations
Thyroid cytology multi-type cell detection method based on deep learning
CN114187277A
Cellular pathology image anomaly detection method based on improved YOLOv5 and OfficientNet
CN115937188A
Pasteur smear cervical cell image classification method based on CNN-SPPF and ViT
CN118097662A
Deep learning-based nasal secretion smear diagnosis method and system
CN118888124A
Mounting bracket for a suspension strut of a vehicle, suspension strut comprising the mounting bracket, damper assembly comprising the suspension strut and comprising a wheel carrier for a vehicle and vehicle comprising the damper assembly
KR1020230084047A
Cited By
Biological cell image classification method and system and storage medium
CN120564186A