A pollen detection method and system based on pollen scanning slides
Through the combination of the rtmdet network and Vit model, the automation and efficient family classification of pollen detection are achieved, and the error detection and classification accuracy of pollen detection in traditional methods are solved, meeting the needs of different fields.
Patent Information
- Application Number
- CN202510341880.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-21
AI Technical Summary
Traditional pollen detection methods rely on manual operations, making it difficult to accurately distinguish fine pollen samples, and are prone to missed detection and low classification efficiency and accuracy of family genera, especially when there are many types of pollen and high similarity.
The pollen detection model based on the rtmdet network is used for object detection, and the pollen filtering model and family classification model are established in combination with Vit. Through image processing and deep learning technology, pollen is automatically identified, filtered and classified, and the number of various family categories is counted.
It improves the accuracy and efficiency of pollen detection, reduces missed detection, ensures the accuracy of family classification, and is suitable for the needs of allergy research, ecological monitoring and agricultural production.
Smart Images

Figure CN119851050B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a pollen detection method and system based on pollen scanning slides. Background Art
[0002] The pollen detection method based on pollen scanning slides is a technical method that uses digital scanning technology and computer vision algorithms to automatically identify and analyze pollen images. This method first obtains a scanned image of a pollen sample through a microscope or other imaging device to generate a pollen scanning slide (Whole Slide Image, WSI). Then, with the help of deep learning, image processing, and pattern recognition technologies, the pollen in the scanning slide is automatically detected, filtered, classified, and statistically analyzed.
[0003] With the intensification of global climate change and air pollution, the rising pollen concentration has had an important impact on allergic diseases, agricultural production, and the ecological environment. An efficient and accurate pollen detection method can automatically and precisely perform full quantitative detection of pollen and genus classification counting to make up for the deficiencies of manual detection and counting, significantly improve the speed and accuracy of pollen monitoring, provide early warning information for allergy patients in a timely manner, and help optimize management in the agricultural and environmental fields.
[0004] However, most traditional pollen detection methods rely on manual operation. When dealing with complex pollen scanning slides, it is often impossible to accurately distinguish subtle pollen samples, and the phenomena of false detection and missed detection are likely to occur. Due to the large variety of pollen in the classification task and the high similarity between genera, the traditional methods have weak discrimination ability for these similar pollens, resulting in too low efficiency and accuracy of genus classification. Summary of the Invention
[0005] In order to solve the technical problems that most traditional pollen detection methods rely on manual operation, it is often impossible to accurately distinguish subtle pollen samples when dealing with complex pollen scanning slides, the phenomena of false detection and missed detection are likely to occur, due to the large variety of pollen in the classification task and the high similarity between genera, the traditional methods have weak discrimination ability for these similar pollens, resulting in too low efficiency and accuracy of genus classification, the present invention provides a pollen detection method and system based on pollen scanning slides.
[0006] The technical solutions provided by the embodiments of the present invention are as follows:
[0007] First aspect:
[0008] A pollen detection method based on pollen scanning slides provided by an embodiment of the present invention includes:
[0009] S1: Obtain the pollen scanning slide;
[0010] S2: Based on the rtmdet network, establish a pollen detection model;
[0011] S3: Input the pollen scan sheet into the pollen detection model, and output the coordinates, width, and height of the target pollen;
[0012] S4: Determine the cropped image of the target pollen according to the coordinates, width, and height of the target pollen;
[0013] S5: Based on Vit, establish a pollen filtering model;
[0014] S6: Input the cropped image into the pollen filtering model, and output the pollen category, where the pollen category includes clear pollen, blurred pollen, and non-pollen;
[0015] S7: Determine the coordinates, width, and height of the clear pollen according to the pollen category;
[0016] S8: Determine the cropped image of the clear pollen according to the coordinates, width, and height of the clear pollen;
[0017] S9: Based on Vit, establish a family and genus classification model;
[0018] S10: Input the cropped image of the clear pollen into the family and genus classification model, and determine the family and genus category of the clear pollen;
[0019] S11: Count the number of pollen in each family and genus category and the total number of pollen in the pollen scan sheet to complete pollen detection.
[0020] Second aspect:
[0021] A pollen detection system based on a pollen scan sheet provided by an embodiment of the present invention includes:
[0022] A processor;
[0023] A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the pollen detection method based on the pollen scan sheet as described in the first aspect is implemented.
[0024] Third aspect:
[0025] A computer-readable storage medium provided by an embodiment of the present invention, on which a computer program is stored, and when the program is executed by a processor, the pollen detection method based on the pollen scan sheet as described in the first aspect is implemented.
[0026] The beneficial effects brought by the technical solution provided by the embodiment of the present invention at least include:
[0027] In the embodiments of the present invention, by establishing a pollen detection model based on the rtmdet network, the targets in pollen images can be accurately detected, the identification ability of the model for pollen samples can be enhanced, the situations of false detection and missed detection can be reduced, thereby improving the accuracy of pollen detection. By using the pollen filtering model, invalid or blurred pollen can be further filtered according to the clarity and classification of the target pollen, thereby improving the accuracy of the final detection result and ensuring that valid data is used for subsequent analysis. By establishing a family and genus classification model, clear pollen can be accurately classified to clarify its family and genus categories, meeting the needs of different fields such as allergy research, ecological monitoring, and agricultural production. The present invention can adapt to the pollen detection requirements in various actual usage scenarios during the pollen detection process, improving the efficiency and accuracy of pollen detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0029] Figure 1 It is a schematic flow chart of a pollen detection method based on a pollen scanning slide provided by an embodiment of the present invention;
[0030] Figure 2 It is a schematic structural diagram of a pollen detection system based on a pollen scanning slide provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] The following will describe the technical solutions in the present invention with reference to the drawings.
[0032] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0033] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when their differences are not emphasized, their intended meanings are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when their differences are not emphasized, their intended meanings are the same.
[0034] In the embodiments of the present invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, their intended meanings are the same.
[0035] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0036] Refer to the attached Figure 1 figures, which show a schematic flowchart of a pollen detection method based on pollen scanning slides provided by an embodiment of the present invention.
[0037] The embodiments of the present invention provide a pollen detection method based on pollen scanning slides. The method includes:
[0038] S1: Obtain a pollen scanning slide.
[0039] Among them, a pollen scanning slide (Whole Slide Image, WSI) refers to a high-resolution digital image obtained by scanning and saving the entire microscopic section image of a pollen sample through digital microscopy technology. Such an image contains all the detailed information of the sample, can be magnified for viewing, and can be further analyzed by a computer.
[0040] It should be noted that the step of obtaining a pollen scanning slide is the basis of the pollen detection process, which can obtain high-resolution full-image data, provide more accurate information, avoid the limitations of manual observation under a traditional microscope, support efficient automated detection, and ensure higher detection accuracy and faster processing speed.
[0041] S2: Based on the rtmdet network, establish a pollen detection model.
[0042] Among them, the rtmdet (Real-Time Multi-Detection) network is a deep learning-based object detection framework, usually adopting an efficient convolutional neural network (CNN) structure to extract image features and perform classification and regression tasks, optimizing the use of computing resources, making the object detection speed faster, and being suitable for real-time processing tasks.
[0043] It should be noted that establishing a pollen detection model based on the RTMDet network can improve the efficiency and accuracy of object detection. RTMDet combines multi-layer convolutional feature extraction with efficient algorithms, enabling fast response in real-time detection, accurately identifying the coordinates and size of pollen at the same time, and being suitable for large-scale data processing and on-site applications.
[0044] Specifically, pollen scanning images usually have a high resolution. When pollen images are used as the input of the target detection network, they must be cropped or resized according to the memory of the hardware and the input size of the model. In the present invention, the rtmdet network is improved. For example, large kernel depth convolutions with a size of 5 × 5 are adopted in the Backbone and Neck layers, a dynamic soft label assignment strategy is designed, and a caching mechanism is used to increase the data volume. A shallower feature map with a higher resolution is introduced at the C2 stage of the Backbone, and a DSPPF module is designed at the C5 stage to make it more suitable for capturing the receptive field of pollen. At the Neck, the feature map is sent into the csp module through a connection operation and then convolved. When the feature map is downsampled by 32 times, the HDA module is added to strengthen the features again.
[0045] In a possible implementation manner, the pollen detection model includes a backbone network module and a feature network module.
[0046] Among them, the backbone network module (Backbone Network Module) is the core part of the deep learning model responsible for extracting features from the input data. In the pollen detection model, the backbone network module includes specific structural units for efficiently extracting important feature information from pollen scanning images. The feature network module (Feature Network Module) is used to further process and fuse the features extracted by the backbone network to refine pollen detection, and adjusts the scale, resolution, and content of the feature map to provide higher-quality detection results.
[0047] The backbone network module includes CSPNext Block units and DSPPF dilated spatial pyramid pooling fast units.
[0048] Among them, the CSPNext Block unit is a structural unit based on a convolutional neural network that reduces the computational burden through cross-stage partial connections, while maintaining smooth information flow and improving the efficiency and accuracy of the model. The DSPPF dilated spatial pyramid pooling fast unit is a dilated spatial pyramid pooling technology that can effectively capture large-scale context information in images through multi-scale pooling operations, which is crucial for pollen sample detection.
[0049] The feature network module includes CSPNextPAFPN units, HAD units, CSP units, and CARAFE upsampling operator units.
[0050] Among them, the CSPNextPAFP unit is an improved feature fusion network that combines a partial cross-stage convolutional network (CSP) and a feature pyramid network (FPN), effectively fusing features at different levels and enhancing the expression ability of multi-scale features. The HAD (Hybrid Attention Domain) unit combines multiple attention mechanisms, enhancing the model's ability to focus on features in different regions and at different scales, and improving the distinctiveness of image features. The CSP (Cross-Stage Partial) unit enhances the feature learning efficiency by splitting the network structure and using partial cross-stage connections, reducing the computational amount while improving the model's performance. The CARAFE (Content-Aware ReAssembly of FEatures) upsampling operator can reconstruct a low-resolution feature map into a high-resolution image in a content-aware manner, preserving the detailed information of the image and avoiding the blurring phenomenon introduced by traditional upsampling methods.
[0051] In a possible implementation, the mathematical calculation formula of the DSPPF dilated spatial pyramid pooling fast unit is specifically as follows:
[0052]
[0053] Among them, M i represents the maximum distance between two non-zero values in the i-th feature map, max represents maximization, M i+1 represents the maximum distance between two non-zero values in the (i + 1)-th feature map, r i represents the dilation rate of the i-th layer of convolution, f out1 represents the feature map after passing through the convolutional layer, f out2 represents the feature map after passing through the convolutional layer, f out3 represents the feature map after passing through the convolutional layer, f out4 represents the feature map after passing through the convolutional layer, f out represents the final output feature map, represents a convolutional layer with a size of 1×1, represents a convolutional layer with a size of 1×3 and a dilation rate of 1, represents a convolutional layer with a size of 3×3 and a dilation rate of 2, represents a convolutional layer with a size of 3×3 and a dilation rate of 3, f in represents the input feature map, SEModule represents the SE attention module, and Concat represents the concatenation operation.
[0054] In the present invention, a 1×1 ordinary convolution is adopted in the first layer, and then three 3×3 convolutions with dilation rates of 1, 2, and 3 respectively are applied to extract feature maps with different receptive fields. Finally, the feature maps of four different receptive fields are concatenated and dimensionally reduced, and the output is provided by the SE attention module.
[0055] The mathematical calculation formula of the CARAFE upsampling operator unit is specifically as follows:
[0056]
[0057]
[0058] Among them, represents the convolution kernel of the l-th feature point after upsampling, φ represents the kernel prediction module, N represents kernel normalization, that is, the Softmax activation function in the spatial domain, X represents the feature map, represents the feature map of the l-th feature point after upsampling, k encoder represents the convolution layer with a kernel size of k encoder The convolution layer, θ represents the content-aware recombination module, k up represents the size of the recombination kernel, m represents the column direction in which the window advances, n represents the row direction in which the window advances, , r represents half of k up and rounded down, i and j represent the coordinate positions of the l-th feature point after transformation, σ represents the sampling rate, and represent the original coordinate positions of the l-th feature point.
[0059] The mathematical calculation formula of the HAD unit is specifically as follows:
[0060]
[0061] Among them, f out represents the output feature map, ECAMoudle represents the ECA attention mechanism, f in represents the input feature map, represents the convolution layer with a size of 1×1, represents the convolution layer with a size of 3×3 and a dilation rate of 1, represents the convolution layer with a size of 3×3 and a dilation rate of 2, represents the convolution layer with a size of 3×3 and a dilation rate of 4.
[0062] In the present invention, after passing through the feature network module, the output feature maps are respectively {4, 8, 16, 32} times the size of the input downsampling. The HDA module is added after the feature map downsampled by 32 times to enhance the model's ability to interpret context information. The model reduces the channel dimension and module parameters through a 1×1 convolution. The output after dimensionality reduction is then fed in parallel to three convolutional branches and a residual attention branch. The dilation rates of the three parallel convolutions are 1, 2, and 4 respectively, and the residual attention consists of spatial attention and channel attention.
[0063] Specifically, the backbone network module provides strong feature extraction capabilities through the CSPNext Block and the DSPPF dilated spatial pyramid pooling fast unit, and can extract multi-scale information in complex backgrounds. The feature network module combines the CSPNextPAFPN, HAD unit, CSP unit, and CARAFE upsampling operator, effectively optimizing the feature fusion and upsampling processes, and improving the classification accuracy and detection speed of the model.
[0064] In a possible implementation, after S2, it further includes:
[0065] Annotate the pollen scanning film to determine the first annotation data.
[0066] Input the first annotation data into the pollen detection model for training until the loss function value is less than the preset loss function value.
[0067] It should be noted that by annotating the pollen scanning film and generating the first annotation data, a high-quality training data set can be provided for the pollen detection model. Inputting these annotation data into the pollen detection model for training enables the model to gradually optimize its recognition ability and improve the detection accuracy by learning the annotation information. During the training process, optimize the loss function value until the model performance reaches the preset target, which can ensure that the trained model has good generalization ability and accuracy, thus realizing the efficient detection and classification of pollen.
[0068] The calculation formula of the loss function value is specifically:
[0069]
[0070] Among them, C represents the total loss function, C cls represents the classification loss, C reg represents the regression loss, λ1, λ2, and λ3 all represent weight coefficients, C center represents the region loss, CE represents the cross-entropy loss, P represents the probability predicted by the model, Y soft represents the soft label, IoU represents an index to measure the overlap degree between the predicted box and the true box, α and β both represent, x pred represents the center coordinate of the predicted box, xgt Represents the center coordinates of the ground truth box.
[0071] In the present invention, α and β are defaulted to 10 and 3.
[0072] In the present invention, during the process of training the model, the initial learning rate is set to 0.003, the learning rate update strategy is cosine, the batch size is 32, the optimizer is Adamw, the number of epochs is set to 300, and the training is completed on two 4080s.
[0073] S3: Input the pollen scanning slice into the pollen detection model, and output the coordinates, width, and height of the target pollen.
[0074] In the present invention, the scanning slice is scaled to mpp = 1, and then windows of size 2024×2024 with an overlap of 128×128 are used to divide the scanning slice. These sliding windows are respectively input into the detection model. For each divided window, the model will output the position and score of the pollen on the window, retain the pollen with a score greater than 0.3, and return the coordinates of the pollen to the WSI, removing the pollen in the overlapping part between the windows.
[0075] S4: Determine the cropped image of the target pollen according to the coordinates, width, and height of the target pollen.
[0076] It should be noted that determining the cropped image according to the coordinates, width, and height of the target pollen can accurately extract the pollen area and remove irrelevant background information, thereby improving the efficiency of subsequent processing and analysis.
[0077] S5: Based on Vit, establish a pollen filtering model.
[0078] It should be noted that establishing a pollen filtering model can effectively screen out clear pollen samples, filter out blurred or non-pollen targets, improve the accuracy of subsequent classification and analysis, make the pollen detection process more automated, avoid the limitations of manual intervention, save time and cost, and enhance the generalization ability of the model.
[0079] In the present invention, the main framework of the pollen filtering model is vision transformer (Vit). Since the model based on Vit is usually several times slower than the convolutional network, and at the same time, to solve the problem of limited device resources such as edge devices, the present invention changes the inefficient design in Vit and proposes a dimension-consistent network with 4D feature implementation and 3D multi-head self-attention to offset the frequent reshape operations.
[0080] In a possible implementation manner, the pollen filtering model includes: MetaFormer module, Patch Embeding module, dimension-consistent processing module, multi-head attention module, and latency-driven lightweight module.
[0081] Among them, the MetaFormer module is a module based on the Transformer architecture, which combines a convolutional network and an attention mechanism, and can efficiently process multi-scale features in images. The MetaFormer module effectively improves the feature extraction ability through parallel processing and self-attention mechanism, and performs feature fusion at different scales to enhance important information in the image. The Patch Embedding module is a commonly used technique in image processing and Transformer models, which mainly cuts the input image into several small patches, and then maps each small patch into a feature vector through a linear transformation with consistent dimensions. The dimension consistency processing module is used to ensure the dimension alignment of feature maps between different layers or modules in the network. The multi-head attention module is a core component in the Transformer architecture. By calculating the attention mechanism in parallel on multiple "heads", the model can focus on different feature parts from multiple angles, enhancing the model's ability to capture global information. The latency-driven lightweight module reduces latency and improves processing speed by dynamically selecting appropriate network paths.
[0082] In the present invention, the essence of the dimension consistency processing module is to divide the MB module into two types of networks, 4D and 3D parts. Among them, linear projection and attention are performed on 3D, so that the global modeling ability can be enjoyed without sacrificing the efficiency of multi-head attention. The 4D is mainly applied to the beginning part of the network output and is mainly implemented in the form of a Conv-net. First, the input image will be subjected to patch embbeding by two 3×3 Convs with a stride of 2, and then through the MB 4D extract low-level semantics, and after processing all the MB 4D modules, perform reshape into the MB at one time 3D .
[0083] In a possible implementation manner, the formula for the pollen filtration model to perform pollen filtration is specifically:
[0084]
[0085] Among them, y represents the output of the pollen filtration model, MB i represents the i-th MB module, PatchEmbed represents image patch embedding, B represents the batch size, H represents the height of the input image, W represents the width of the input image, x0 represents the input image, represents the cumulative operation from i to m.
[0086] The mathematical calculation formula of the MetaFormer module is specifically:
[0087]
[0088] Among them, x i+1 represents the intermediate feature transferred to the (i + 1)-th MB module, and x i represents the intermediate feature transferred to the i-th MB module, where MB i represents the i-th MB module, MLP represents a multi-layer perceptron, TokenMixer represents a 4D multi-head self-attention module, BN represents batch normalization, Conv2D represents a two-dimensional convolution, RELU represents an activation function, Upsample represents upsampling, V represents the Value in the input, and V local represents local feature enhancement operation, which consists of a depthwise separable convolution and a BN layer. Both TH1 and TH2 represent convolutions of size 1×1, Softmax represents an activation function, Attention represents an attention operation, and DepthwiseConv2D represents a depthwise separable convolution.
[0089] The mathematical calculation formula of the Patch Embeding module is specifically:
[0090]
[0091] Among them, represents the feature map obtained after the PatchEmbed operation, and C j represents the number of channels in the j-th Stage.
[0092] The mathematical calculation formula of the dimension consistency processing module is specifically:
[0093]
[0094] Among them, I i represents the output of the i-th Stage, Pool represents a pooling operation, and Conv B represents connecting a BN layer after convolution, and Conv B,G represents Conv B followed by connecting a GeLU activation layer.
[0095] The mathematical calculation formula of the multi-head attention module is specifically:
[0096]
[0097] Among them, Linear represents a linear layer, MHSA represents multi-head self-attention, LN represents a LayerNorm layer, and Linear G represents connecting a GeLU activation layer after the linear layer, Q represents a query, K represents a key, V represents a value, b represents a bias, T represents a transpose operation.
[0098] The mathematical calculation formula of the delay-driven lightweight module is specifically as follows:
[0099]
[0100] Among them, MP r represents the meta-path in the r-th stage, represents the i-th 4D MetaFormer module in the meta-path, represents the i-th 3D MetaFormer module in the meta-path.
[0101] In a possible implementation manner, after S5, it further includes:
[0102] Label the target pollen cropped image to determine the second labeled data, where the second labeled data specifically includes: clear pollen, blurred pollen, and non-pollen.
[0103] Input the second labeled data into the pollen filtering model for training.
[0104] It should be noted that by labeling the target pollen and determining the second labeled data, it can help the pollen filtering model more accurately distinguish clear pollen, blurred pollen, and non-pollen, enabling the model to make more accurate judgments in complex backgrounds, improving the accuracy and robustness of pollen classification. In addition, the introduction of the second labeled data can strengthen the learning effect of the model, reduce misidentifications, improve the filtering effect, and ultimately optimize the overall performance of the pollen detection process.
[0105] In the present invention, the pollen filtering model adopts CrossEntropyLoss cross-entropy loss. During the training process, the initial learning rate is set to 0.003, the learning rate update strategy is cosine, the batchsize is 32, the optimizer is Adamw, and the number of iteration cycles is set to 100, and the training is completed on a single 4080.
[0106] S6: Input the cropped image into the pollen filtering model to output the pollen category, where the pollen category includes clear pollen, blurred pollen, and non-pollen.
[0107] It should be noted that inputting the cropped image into the pollen filtering model can ensure that the model only processes the image area containing clear pollen, removing irrelevant backgrounds and noises, improving the accuracy and efficiency of the model, reducing the waste of computing resources, and focusing on the processing of high-quality data.
[0108] S7: Determine the coordinates, width, and height of the clear pollen according to the pollen category.
[0109] In the present invention, after obtaining the coordinates, width, and height of the pollen sample from the detection model, the original image is cropped to obtain the corresponding target, and then the target is scaled to a size of 384×384 and used as the model input. The model outputs the probabilities of three categories: clear pollen, blurred pollen, and impurities, and the category with the highest probability is taken as the category of the pollen.
[0110] S8: Determine the cropped image of the clear pollen according to the coordinates, width, and height of the clear pollen.
[0111] It should be noted that by determining the cropped image according to the coordinates, width, and height of the clear pollen, the pollen area can be accurately extracted, background noise and irrelevant information can be removed. The cropped image reduces unnecessary computational complexity, improves the model processing speed and efficiency, and at the same time avoids interference from blurred or incorrect data, further improving the accuracy of pollen detection.
[0112] S9: Based on Vit, establish a family and genus classification model.
[0113] It should be noted that establishing a family and genus classification model can effectively and precisely classify clear pollen, clarify its family and genus categories. By deeply learning the subtle features of pollen, the accuracy of pollen classification is improved. The family and genus classification model can handle complex pollen samples, identify pollen of similar families and genera, and avoid errors caused by manual intervention.
[0114] In the present invention, the family and genus classification model uses vision transformer as the basic framework of the model, and the improvement of the model is the same as that of the above pollen filtering model, which will not be elaborated here.
[0115] In a possible implementation manner, after S8, it further includes:
[0116] Annotate the cropped image of the clear pollen to determine the third annotation data.
[0117] Input the third annotation data into the family and genus classification model for training.
[0118] In the present invention, the clear pollen obtained from the filtering model is divided into these categories: Quercus of Fagaceae, Artemisia of Asteraceae, Kochia of Chenopodiaceae, Amaranthus of Amaranthaceae, Chenopodium of Chenopodiaceae, Cyperus, Taraxacum of Asteraceae, Ambrosia of Asteraceae, Picea of Pinaceae, Pinus of Pinaceae, Corylus of Betulaceae, Humulus of Moraceae, Taxus of Taxaceae, Salix of Salicaceae, Salix of Salicaceae, Ulmus of Ulmaceae. During the training process of the model, the initial learning rate is set to 0.003, the learning rate update strategy is cosine, the batchsize is 64, the optimizer is Adamw, and the number of iteration cycles is set to 200, and the training is completed on a single 4080.
[0119] S10: Input the cropped image of the clear pollen into the family and genus classification model to determine the family and genus category of the clear pollen.
[0120] It should be noted that after inputting the clear pollen cropped image into the family and genus classification model, it can ensure that the model focuses on classifying high-quality pollen samples, avoiding interference factors. By accurately identifying the family and genus categories of pollen, the model can significantly improve the accuracy and reliability of classification, automate the family and genus analysis of pollen, reduce human errors, improve the processing speed, and at the same time provide more effective data support for pollen research and allergy prevention and control.
[0121] In the present invention, the coordinates, width, and height of the clear pollen are cropped from the original image and scaled to a size of 384×384 as the input of the classification model. The family and genus classification model outputs the probabilities of the pollen belonging to each family and genus, and the category with the highest probability is taken as the family and genus to which the pollen belongs.
[0122] S11: Count the number of pollen in each family and genus category and the total number of pollen on the pollen scanning slide to complete pollen detection.
[0123] It should be noted that by counting the number of pollen in each family and genus category and the total number of pollen scanning slides, comprehensive pollen distribution data can be provided, which helps to accurately evaluate the concentration of different pollen species.
[0124] In the present invention, clear pollen is obtained through the filtering model, the family and genus category of the clear pollen is obtained through the family and genus classification model, and the number of pollen in each family and genus category is added up to obtain the total number of pollen on the scanning slide.
[0125] The beneficial effects brought by the technical solution provided by the embodiment of the present invention at least include:
[0126] In the embodiment of the present invention, by establishing a pollen detection model based on the rtmdet network, the targets in the pollen image can be accurately detected, the recognition ability of the model for pollen samples can be enhanced, the situations of false detection and missed detection can be reduced, thereby improving the accuracy of pollen detection. Using the pollen filtering model, invalid or blurred pollen can be further filtered according to the clarity and classification of the target pollen, thereby improving the accuracy of the final detection result and ensuring that effective data is used for subsequent analysis. By establishing a family and genus classification model, the clear pollen can be accurately classified to clarify its family and genus category, meeting the needs of different fields such as allergy research, ecological monitoring, and agricultural production. The present invention can adapt to the pollen detection requirements in various actual usage scenarios during the pollen detection process, improving the efficiency and accuracy of pollen detection.
[0127] Refer to the attached Figure 2 illustrates a schematic structural diagram of a pollen detection system based on a pollen scanning slide provided by the present invention.
[0128] The present invention also provides a pollen detection system 20 based on a pollen scanning slide, which is applied to the above-mentioned pollen detection method based on a pollen scanning slide, and includes:
[0129] A processor 201.
[0130] A memory 202, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor 201, the pollen detection method based on a pollen scanning slide as in the method embodiment is implemented.
[0131] The pollen detection system 20 based on a pollen scanning slide provided by the present invention can execute the above-mentioned pollen detection method based on a pollen scanning slide and achieve the same or similar technical effects. To avoid repetition, the present invention will not elaborate further.
[0132] The beneficial effects brought by the technical solution provided in the embodiment of the present invention at least include:
[0133] In the embodiment of the present invention, by establishing a pollen detection model based on the rtmdet network, the target in the pollen image can be accurately detected, the identification ability of the model for pollen samples can be enhanced, the situations of false detection and missed detection can be reduced, thereby improving the accuracy of pollen detection. Using the pollen filtering model, invalid or blurred pollen can be further filtered according to the clarity and classification of the target pollen, thereby improving the accuracy of the final detection result and ensuring that valid data is used for subsequent analysis. By establishing a family and genus classification model, clear pollen can be accurately classified to clarify its family and genus categories, meeting the needs of different fields such as allergy research, ecological monitoring, and agricultural production. The present invention can adapt to the pollen detection requirements in various actual usage scenarios during the pollen detection process, improving the efficiency and accuracy of pollen detection.
[0134] It should be understood that the processor in the embodiment of the present invention may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0135] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0136] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. A computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0137] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.
[0138] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0139] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0140] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0141] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0142] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms.
[0143] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0144] In addition, the functional units in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0145] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0146] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the pollen detection method based on pollen scanning slides as in the method embodiment.
[0147] The computer-readable storage medium provided by the present invention can implement the steps and effects of the pollen detection method based on pollen scanning slides in the above method embodiment. To avoid repetition, the present invention will not elaborate further.
[0148] The beneficial effects brought by the technical solution provided by the embodiments of the present invention at least include:
[0149] In the embodiments of the present invention, by establishing a pollen detection model based on the rtmdet network, the targets in pollen images can be accurately detected, the identification ability of the model for pollen samples can be enhanced, the situations of false detection and missed detection can be reduced, thereby improving the accuracy of pollen detection. Using the pollen filtering model, invalid or blurred pollen can be further filtered according to the clarity and classification of the target pollen, thereby improving the accuracy of the final detection result and ensuring that valid data is used for subsequent analysis. By establishing a genus and species classification model, clear pollen can be accurately classified to clarify its genus and species categories, meeting the needs of different fields such as allergy research, ecological monitoring, and agricultural production. The present invention can adapt to the pollen detection requirements in various actual usage scenarios during the pollen detection process, improving the efficiency and accuracy of pollen detection.
[0150] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
[0151] The following points need to be explained:
[0152] (1)The accompanying drawings of the embodiments of the present invention only relate to the structures involved in the embodiments of the present invention, and other structures can refer to the general design.
[0153] (2)For clarity, in the accompanying drawings used to describe the embodiments of the present invention, the thickness of layers or regions is enlarged or reduced, that is, these drawings are not drawn to actual scale. It can be understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, the element can be "directly" on or under the other element or there can be intermediate elements.
[0154] (3)Without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.
[0155] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A pollen detection method based on pollen scanning slides, characterized in that, Including: S1: Obtain the pollen scanning film; S2: Based on the rtmdet network, establish a pollen detection model; S3: Input the pollen scanning film into the pollen detection model, and output the coordinates, width, and height of the target pollen; S4: Determine the cropped image of the target pollen according to the coordinates, width, and height of the target pollen; S5: Based on Vit, establish a pollen filtering model; S6: Input the cropped image into the pollen filtering model, and output the pollen category, where the pollen category includes clear pollen, blurred pollen, and non-pollen; S7: Determine the coordinates, width, and height of the clear pollen according to the pollen category; S8: Determine the cropped image of the clear pollen according to the coordinates, width, and height of the clear pollen; S9: Based on Vit, establish a family and genus classification model; S10: Input the cropped image of the clear pollen into the family and genus classification model to determine the family and genus category of the clear pollen; S11: Count the number of pollen in each family and genus category and the total number of pollen in the pollen scanning film to complete pollen detection; Among them, between the Backbone layer and the Neck layer of the rtmdet network, there is a large kernel depth convolution with a size of 5×5; The rtmdet network includes a dynamic soft label assignment strategy and a caching mechanism; The C2 stage of the Backbone layer includes a shallow feature map, and the C5 stage of the Backbone layer includes a DSPPF module; The Vit is specifically a dimensionally consistent network with 4D feature implementation and 3D multi-head self-attention.
2. The pollen detection method based on pollen scanning slides according to claim 1, wherein, The pollen detection model includes: a backbone network module and a feature network module; The backbone network module includes a CSPNext Block unit and a DSPPF dilated spatial pyramid pooling fast unit; The feature network module includes a CSPNextPAFPN unit, a HAD unit, a CSP unit, and a CARAFE upsampling operator unit.
3. The pollen detection method based on pollen scanning slides according to claim 2, characterized in that, The mathematical calculation formula of the DSPPF dilated spatial pyramid pooling fast unit is specifically: ; Among them, M i represents the maximum distance between two non-zero values in the i-th feature map, max represents maximization, and M i+1 represents the maximum distance between two non-zero values in the (i + 1)-th feature map, and r i represents the dilation rate of the i-th layer of convolution, and f out1 represents after the feature map of the convolutional layer, and f out2 represents after the feature map of the convolutional layer, and f out3 represents after the feature map of the convolutional layer, and f out4 represents after the feature map of the convolutional layer, and f out represents the final output feature map, represents a convolutional layer of size 1×1, represents a convolutional layer of size 1×3 with a dilation rate of 1, represents a convolutional layer of size 3×3 with a dilation rate of 2, represents a convolutional layer of size 3×3 with a dilation rate of 3, and f in represents the input feature map, SEModule represents the SE attention module, and Concat represents the concatenation operation; The mathematical calculation formula of the CARAFE upsampling operator unit is specifically: ; ; Among them, represents the convolution kernel of the l-th feature point after upsampling, φ represents the kernel prediction module, N represents kernel normalization, that is, the Softmax activation function in the spatial domain, X represents the feature map, represents the feature map of the l-th feature point after upsampling, k encoder represents the convolutional layer with a kernel size of k encoder The convolution layer, θ represents the content-aware recombination module, k up represents the size of the recombination kernel, m represents the column direction in which the window advances, n represents the row direction in which the window advances, , r represents half of k up rounded down, i and j represent the coordinate positions of the l-th feature point after transformation, σ represents the sampling rate, and represent the original coordinate positions of the l-th feature point; The mathematical calculation formula of the HAD unit is specifically: ; Among them, f out represents the output feature map, ECAMoudle represents the ECA attention mechanism, and f in represents the input feature map, represents a convolutional layer with a size of 1×1, represents a convolutional layer with a size of 3×3 and a dilation rate of 1, represents a convolutional layer with a size of 3×3 and a dilation rate of 2, represents a convolutional layer with a size of 3×3 and a dilation rate of 4.
4. The pollen detection method based on pollen scanning slides according to claim 1, wherein After S2, it also includes: Annotate the pollen scanning film to determine the first annotation data; Input the first annotation data into the pollen detection model for training until the loss function value is less than the preset loss function value.
5. The pollen detection method based on pollen scanning slides according to claim 1, wherein The calculation formula of the loss function value is specifically: ; Among them, C represents the total loss function, C cls represents the classification loss, C reg represents the regression loss, λ1, λ2, and λ3 all represent weight coefficients, C center represents the region loss, CE represents the cross-entropy loss, P represents the probability predicted by the model, Y soft represents the soft label, IoU represents an index measuring the overlapping degree of the predicted bounding box and the ground truth bounding box, α and β both represent hyperparameters, x pred represents the center coordinates of the predicted bounding box, x gt represents the center coordinates of the ground truth bounding box.
6. The pollen detection method based on pollen scanning slides according to claim 1, wherein The pollen filtering model includes: a Meta Former module, a Patch Embeding module, a dimensionally consistent processing module, a multi-head attention module, and a latency-driven lightweight module.
7. The pollen detection method based on pollen scanning slides according to claim 6, wherein The formula for the pollen filtering model to filter pollen is: ; where y represents the output of the pollen filtration model, MB i represents the i-th MB module, PatchEmbed represents image patch embedding, B represents the batch size, H represents the height of the input image, W represents the width of the input image, and x0 represents the input image, represents the cumulative operation from i to m; The mathematical calculation formula of the MetaFormer module is specifically: ; Among them, x i+1 represents the intermediate feature flowing into the (i + 1)-th MB module, and x i represents the intermediate feature flowing into the i-th MB module. MB i represents the i-th MB module, MLP represents a multi-layer perceptron, TokenMixer represents a 4D multi-head self-attention module, BN represents batch normalization, Conv2D represents a two-dimensional convolution, RELU represents an activation function, Upsample represents upsampling, V represents the Value in the input, and V local represents local feature enhancement operation, which consists of a depthwise separable convolution and a BN layer. Both TH1 and TH2 represent convolutions of size 1×1, Softmax represents an activation function, Attention represents an attention operation, and DepthwiseConv2D represents a depthwise separable convolution; The mathematical calculation formula of the Patch Embeding module is specifically: ; Among them, represents the feature map obtained after the PatchEmbed operation, and C j represents the number of channels in the j-th stage; The mathematical calculation formula of the dimensionally consistent processing module is specifically: ; Among them, I i represents the output of the i-th Stage, Pool represents the pooling operation, Conv B represents connecting the BN layer after convolution, Conv B,G represents Conv B followed by connecting the GeLU activation layer; The mathematical calculation formula of the multi-head attention module is specifically: ; Among them, Linear represents the linear layer, MHSA represents the multi-head self-attention, LN represents the LayerNorm layer, and Linear G represents that a GeLU activation layer is connected after the linear layer, Q represents the query, K represents the key, V represents the value, b represents the bias, T represents the transpose operation; The mathematical calculation formula of the latency-driven lightweight module is specifically: ; Among them, MP r represents the meta-path of the r-th stage, represents the i-th 4D MetaFormer module in the meta-path, represents the i-th 3D MetaFormer module in the meta-path.
8. The pollen detection method based on pollen scanning slides according to claim 1, characterized in that After S5, it also includes: Annotate the cropped image of the target pollen to determine the second annotation data, where the second annotation data specifically includes: clear pollen, blurred pollen, and non-pollen; Input the second annotation data into the pollen filtering model for training.
9. The pollen detection method based on pollen scanning slides according to claim 1, wherein After S9, it further includes: Annotate the cropped image of the clear pollen to determine the third annotation data; Input the third annotation data into the genus and species classification model for training.
10. A pollen detection system based on pollen scanning slides, characterized in that, It includes: A processor; A memory, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, the pollen detection method based on pollen scanning slides as described in any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Pollen image classification method and device
CN113723453A
YOLO network-based airborne pollen sensitization plant intelligent identification method
CN116958643A