Pollen target detection method and device and electronic equipment
By adopting a decoder structure of double branch attention in the object detection model, the adverse interaction problem between self-attention and cross-attention is solved, and the detection performance of the model is improved.
Patent Information
- Application Number
- CN202411929158.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-13
AI Technical Summary
The Transformer-based object detection model has an adverse interaction between self-attention and cross-attention during training, resulting in a degradation in model detection performance.
A decoder structure with double branch attention is adopted. Each decoder layer contains two parallel branches. The first branch performs self-attention and cross-attention calculations, and the second branch performs only cross-attention calculations. By fusing the outputs of the two branches, the impact of self-attention and cross-attention in the training process is balanced.
Through the decoder structure of double-branch attention, the adverse interaction between self-attention and cross-attention is reduced, and the detection performance of the model is improved.
Smart Images

Figure CN119992542A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and in particular to a pollen target detection method, device and electronic equipment. Background Art
[0002] Pollen allergy is a global health problem that affects nearly 20% of adults and 40% of children, leading to reduced quality of life and mental health issues. Currently, more than 150 pollen proteins have been identified as allergens. With urbanization and climate change, pollen allergy problems are expected to intensify. Therefore, the development of accurate pollen detection systems is crucial to help people prevent and treat allergy symptoms.
[0003] Traditional pollen identification relies on time-consuming manual processes with limited accuracy. In order to improve efficiency and accuracy, researchers have begun to explore new methods based on machine learning, among which the Transformer architecture has become a highly sought-after option.
[0004] However, there are some training challenges in Transformer-based object detection models such as DETR. Specifically, in the actual model training process, the cross-attention layer in the DETR decoder tends to focus more on multiple queries around a single object, while the self-attention makes these queries move away from each other. This interaction between self-attention and cross-attention in the actual model training process will have an adverse effect on model training and reduce the detection performance of the model. Summary of the invention
[0005] Based on the above technical problems, the present invention provides a pollen target detection method, device and electronic device.
[0006] According to one aspect of the present invention, a pollen target detection method is provided, comprising: Construct an object detection model, which includes a backbone network, a Transformer encoder, a dual-branch attention decoder, and a prediction head; the decoder includes multiple decoder layers, each decoder layer includes two parallel branches, the first branch performs self-attention and cross-attention calculations, and the second branch only performs cross-attention calculations; Based on the backbone network, obtain the multi-scale feature map corresponding to the image to be detected; Input the multi-scale feature map into the Transformer encoder to obtain the enhanced feature map; Input the enhanced feature map and query into the decoder of the dual-branch attention, and obtain the target query output by the decoder; Based on the target query, the prediction head outputs the target detection result.
[0007] According to an aspect of the present invention, a pollen target detection method is provided, wherein the enhanced feature map and the query are input to a decoder of a dual-branch attention system, and a target query output by the decoder is obtained, including: The enhanced feature map and the initialized query Input to the decoder of the dual-branch attention l Decoder layer; In the l In the first branch of the decoder layer, the query Perform self-attention calculation to get the updated query , and the updated query Perform cross attention calculation with the feature map to get the first output ; In the l In the second branch of the decoder layer, the query Perform cross attention calculation with the feature map to get the second output ; Fusion first output and the second output , and get the third output ; The third output Input to l The fully connected layer of the decoder layer obtains the l Output of the decoder layer ; The first l Output of the decoder layer As the input query of the next decoder layer, and repeat the above process; After being processed by stacking multiple decoder layers, the target query is output.
[0008] According to the pollen target detection method of one aspect of the present invention, the updated query , first output , Second Output and the third output and l Output of the decoder layer Calculate according to the following formulas: =Norm(FFN( )) Among them, Self represents the calculation of self-attention, Que ( )、 Key ( )、 Value ( ) represent the extraction of query vector, key vector and value vector from the query, respectively. They represent the key vector and value vector extracted from the enhanced feature map respectively, Z is the enhanced feature map output by the encoder, Norm means normalizing the result, and Cross means calculating cross attention.
[0009] According to a pollen target detection method according to one aspect of the present invention, the target detection result includes: a predicted value of the intersection over union (IoU) between the predicted bounding box and the real box, a predicted classification, and a predicted bounding box; Based on the target query, the prediction head outputs the target detection results, including: Construct the intersection-over-union loss function, which is used to calculate the deviation between the predicted value of IoU and the true value of IoU; Based on the predicted value of the intersection-over-union ratio and the predicted classification, a classification loss function is constructed. The classification loss function is used to calculate the deviation between the probability of the predicted classification and the true classification label; Construct a loss function for the predicted bounding box. The loss function for the predicted bounding box is used to calculate the deviation between the predicted bounding box and the true bounding box. Optimize the parameters of the object detection model based on the loss function of the predicted bounding box, the intersection-over-union loss function, and the classification loss function; Based on the target detection model and target query after optimizing parameters, the prediction head outputs the target detection results.
[0010] According to the pollen target detection method of one aspect of the present invention, the intersection-over-union loss function is: in, To query the predicted value of the intersection-and-union ratio of q, is the true value of the intersection and union ratio of query q, where query q is an element in the target query.
[0011] According to one aspect of the present invention, a pollen target detection method is provided, which constructs a classification loss function based on the predicted value and predicted classification of the intersection-over-union ratio, including: Based on the predicted value of the intersection-over-union ratio and the probability of the predicted classification, a weighted geometric mean function is constructed; Divide the image to be detected into a foreground sample set and a background sample set; A classification loss function is constructed based on the foreground sample set, the background sample set, the predicted classification and the weighted geometric mean function.
[0012] According to the pollen target detection method of one aspect of the present invention, the weighted geometric mean function is: Among them, s is the probability of predicted classification, E(IoU) is the predicted value of intersection over union, and α is the weight parameter, which is used to adjust the influence of the probability of predicted classification and the predicted value of intersection over union on the weighted geometric mean function.
[0013] According to the pollen target detection method of one aspect of the present invention, the classification loss function is: Among them, Npos represents the foreground sample set, Nneg represents the background sample set, i represents the i-th element of the foreground sample set, j represents the j-th element of the background sample set, and BCE represents the binary cross entropy loss function.
[0014] According to another aspect of the present invention, there is provided a pollen target detection device, comprising: A model building module, used to build a target detection model, wherein the target detection model includes a backbone network, a Transformer encoder, a dual-branch attention decoder and a prediction head, wherein the decoder includes multiple decoder layers, each decoder layer includes two parallel branches, the first branch performs self-attention and cross-attention calculations, and the second branch only performs cross-attention calculations; A first acquisition module, used to acquire a multi-scale feature map corresponding to the image to be detected based on the backbone network; The second acquisition module is used to input the multi-scale feature map into the Transormer encoder to obtain the enhanced feature map; A third acquisition module is used to input the enhanced feature map and the query into the decoder of the dual-branch attention, and obtain the target query output by the decoder; The detection result output module is used to output the target detection result based on the target query by the prediction head.
[0015] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the above-mentioned pollen target detection method when executing the computer program.
[0016] The pollen target detection method, device and electronic device provided by the present invention obtain a multi-scale feature map corresponding to the image to be detected, input the multi-scale feature map into a Transformer encoder, obtain an enhanced feature map, and input the enhanced feature map and the query into a dual-branch attention decoder, through self-attention and cross-attention calculations and two parallel branches that only perform cross-attention calculations, and finally output the target detection result according to the target query output by the dual-branch attention decoder, which helps to balance the influence of self-attention and cross-attention in the training process, reduce the adverse interaction between them, and thereby improve the detection performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 This is one of the flow charts of the pollen target detection method provided by the present invention.
[0019] Figure 2 This is the second flow chart of the pollen target detection method provided by the present invention.
[0020] Figure 3 This is the third flow chart of the pollen target detection method provided by the present invention.
[0021] Figure 4 It is a functional block diagram of the pollen target detection device provided by the present invention.
[0022] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0024] Pollen allergy is a global health problem that affects a large number of adults and children, leading to reduced quality of life and mental health issues. The problem is expected to increase further with urbanization and climate change. Therefore, the development of accurate pollen detection systems is crucial to help people prevent and treat allergy symptoms.
[0025] Traditional pollen identification methods mainly rely on manual processes, which are not only time-consuming and labor-intensive, but also have limited accuracy. In order to improve efficiency and accuracy, researchers have begun to explore machine learning-based methods, which can be divided into two categories: non-learning and learning. Non-learning methods require complex preprocessing and feature construction, while learning methods can integrate positioning and classification into the model, especially convolutional neural network (CNN)-based methods, which can effectively capture the spatial structure information of images and perform well in pollen detection.
[0026] In recent years, the Transformer architecture has achieved remarkable results in the field of natural language processing and has gradually shown its potential in the field of computer vision. However, in the task of automatic pollen detection, there are still few Transformer-based methods.
[0027] In the actual machine model training process, the self-attention and cross-attention in the DETR model will have opposite effects on object queries, damaging the training effect. At the same time, the complete independence of the classification and positioning prediction heads also leads to mismatches in the prediction results (i.e., target detection results).
[0028] Specifically, self-attention and cross-attention are important components of DETR-like models. Both are crucial to DETR, but with the continuous study of the model, researchers have found that these two types of attention will have some opposite effects on object queries in the actual training process, which will damage the training effect. In the actual training process, the cross-attention layer in the DETR decoder tends to pay more attention to multiple queries around a single object, while self-attention will make these queries move away from each other. At the same time, in today's object inspection models, the positioning and classification of the detection target are performed separately by two parallel prediction heads, but the complete independence also makes the final prediction results of the two mismatched, that is, the query has a high classification confidence but poor prediction box quality; or, it has a low classification confidence but a high-quality prediction box. Under the expected circumstances, queries with high classification confidence should also have high-quality prediction boxes. Therefore, the problem of misalignment will reduce the detection performance of the model.
[0029] In order to solve the above problems, two methods are proposed: a dual-branch attention decoder and an IoU-guided alignment scheme. The dual-branch attention decoder is a new decoder with a new structure. Each decoder layer of the decoder has two parallel branches, one for self-attention and cross-attention, and the other for cross-attention only. This design reduces the opposite effects of self-attention and cross-attention on training. In the IoU-guided alignment scheme, an additional IoU prediction head is added to predict the IoU of the predicted bounding box of the query and the grand-truth (referring to the accurate labeling of the target object position), and the predicted IoU is used to calibrate the classification training. In the final detection result calculation part, the predicted IoU and classification confidence are used to score the query, thereby reducing the query misalignment phenomenon.
[0030] Figure 1 is one of the flow charts of the pollen target detection method provided by the present invention, such as Figure 1 As shown, the method includes the following: Step 101, construct a target detection model, which includes a backbone network, a Transformer encoder, a dual-branch attention decoder and a prediction head, wherein the decoder includes multiple decoder layers, each decoder layer includes two parallel branches, the first branch performs self-attention and cross-attention calculations, and the second branch only performs cross-attention calculations.
[0031] In one embodiment of the present invention, the backbone network is generally a deep convolutional neural network, which can be a classic CNN structure such as ResNet, VGG, etc., or a network structure specially designed for target detection tasks, which is used for feature extraction in the target detection model. The Transformer encoder is the part of the target detection model used to further extract features. It receives the feature map output by the backbone network and uses the self-attention mechanism to capture global context information. The encoder can process serialized feature maps and establish connections between features to enhance the model's understanding of the global information of the image. The dual-branch attention decoder is a deep learning model based on the attention mechanism, which can capture the dependencies between different scales and different positions when processing image features. The dual-branch attention decoder includes multiple decoder layers, each of which includes two parallel branches. The first branch performs self-attention and cross-attention calculations to capture the dependencies within the feature map and between the feature map and the query. The second branch only performs cross-attention calculations to further refine and extract key feature information. The prediction head is the last part of the target detection model, which is responsible for generating the final target detection result based on the output of the decoder.
[0032] Step 102: Based on the backbone network, obtain a multi-scale feature map corresponding to the image to be detected.
[0033] In one embodiment of the present invention, before performing pollen target detection, it is first necessary to collect a series of images containing pollen as images to be detected. These images may come from observations under a microscope, image libraries taken with a microscope, or images containing pollen obtained by other means. The image to be detected is input into the backbone network, and the backbone network is used to extract features of the image to be detected to obtain a feature map corresponding to the image to be detected. Then, the feature map obtained previously is scaled using a feature pyramid. This usually involves upsampling (enlarging) or downsampling (reducing) the feature map to generate feature maps of multiple different scales. The feature map of each scale contains image feature information at the corresponding scale. The feature map is a two-dimensional or three-dimensional array that contains key feature information in the image. Feature Pyramid is a multi-scale feature representation method that can extract image features at different scales.
[0034] After the above steps, a set of multi-scale feature maps are finally obtained. These feature maps correspond to the target information of different scales in the image. When performing subsequent target detection tasks, these multi-scale feature maps can be used to improve the accuracy and robustness of detection.
[0035] Step 103: Input the multi-scale feature map into the Transformer encoder to obtain an enhanced feature map.
[0036] In one embodiment of the present disclosure, each feature map in the multi-scale feature map obtained in the above step 101 is flattened into a one-dimensional feature vector, which contains the values of all pixels. At the same time, in order to retain the position information of the pixels in the feature map, the position code is added to the flattened feature vector. The flattened feature vector with the position code added is input into the Transformer encoder. The encoder further processes and enhances these feature maps through the self-attention mechanism to extract richer feature information. After being processed by the Transformer encoder, a set of enhanced feature maps are obtained. The enhanced feature map not only retains the original image information, but also incorporates the additional feature information extracted by the encoder, providing a more accurate and robust feature representation for subsequent target detection tasks.
[0037] Step 104: Input the enhanced feature map as a query to a decoder of a dual-branch attention system to obtain a target query output by the decoder.
[0038] In one embodiment of the present invention, a query is a set of learnable parameters that are learned and adjusted during training to better predict targets in an image. The enhanced feature map and the query are processed by a dual-branch attention decoder to generate a target query. The target query contains predictions about the location and category of potential targets in the image and other relevant information.
[0039] Step 105: Based on the target query, the prediction head outputs the target detection result.
[0040] In one embodiment of the present invention, the target query output by the decoder is then fed into a prediction head, which is responsible for converting the output of the decoder into a final prediction result, i.e., a target detection result.
[0041] In summary, according to the technical solution provided by the embodiment of the present invention, by obtaining a multi-scale feature map corresponding to the image to be detected, the multi-scale feature map is input into the Transformer encoder to obtain an enhanced feature map, and the enhanced feature map and the initialized query are input into the dual-branch attention decoder, through self-attention and cross-attention calculations and two parallel branches that only perform cross-attention calculations, and finally output the target detection result according to the target query output by the dual-branch attention decoder, which helps to balance the influence of self-attention and cross-attention in the training process, reduce the adverse interactions between them, and thereby improve the detection performance of the model.
[0042] Figure 2 FIG. 2 is a flow chart of the pollen target detection method provided by the present invention. Figure 2 As shown, the enhanced feature map and the query are input into the decoder of the dual-branch attention to obtain the target query output by the decoder, which specifically includes the following: Step 201: The enhanced feature map and the query Input to the decoder of the dual-branch attention l Decoder layer.
[0043] In one embodiment of the present invention, the query Including location query and content query, initialization query through hybrid query selection strategy , this hybrid query selection strategy only initializes the location query, but not the content query, so that the content query remains in a learnable state. The purpose of this is to enable the target detection model to adaptively learn the content query, so as to better capture the target information in the image. Among them, the location query is part of the query, which is usually initialized to contain location information to help the model locate potential targets in the image. The content query is another part of the query, which is usually kept in a learnable state and not initialized. The purpose of this is to enable the model to adaptively learn the content query, so as to better capture the target information in the image. Enhanced feature map and query As the decoder input data of the two-branch attention. l A decoder layer refers to a layer in a decoder.
[0044] Step 202, in the l In the first branch of the decoder layer, the query Perform self-attention calculation to get the updated query , and the updated query Perform cross attention calculation with the feature map to get the first output .
[0045] In one embodiment of the present invention, in the first branch, a self-attention calculation is first performed to update the query. The self-attention mechanism allows the query to focus on different parts of itself, thereby extracting more useful information. Next, a cross-attention calculation is performed to combine the information of the query and the feature map. The cross-attention mechanism allows the query to focus on different parts of the feature map, thereby extracting information related to the query, and finally the first branch outputs the first output.
[0046] For example, the updated query , first output Calculated using the following formula: Among them, Self represents the calculation of self-attention, Que ( )、 Key ( )、 Value ( ) represent the extraction of query vector, key vector and value vector from the query, Norm represents the normalization of the result, and Cross represents the calculation of cross attention.
[0047] Step 203, in the l In the second branch of the decoder layer, the query Perform cross attention calculation with the feature map to get the second output .
[0048] In one embodiment of the present invention, in the second branch, query Directly perform cross attention calculation with the feature map. Cross attention allows the query to focus on the part of the feature map that is relevant to the target and extract information that is useful for target detection. is updated to form the output of the second branch , which is the second output. The second output contains the information extracted from the feature map. The second output is part of the decoder layer and will be used for subsequent processing.
[0049] Exemplarily, the second output may be calculated according to the following formula: .
[0050] Step 204: Fusion of the first output and the second output , and get the third output .
[0051] In one embodiment of the present invention, the first output and the second output are fused to form a third output. Fusion generally involves adding the two outputs or merging them by other means (such as weighted averaging) to integrate information from the two branches. This fusion helps the model learn target features from different perspectives and improves the accuracy and robustness of detection.
[0052] Exemplarily, the third output may be calculated according to the following formula: .
[0053] Step 205: Output the third Input to l The fully connected layer of the decoder layer obtains the l Output of the decoder layer .
[0054] In one embodiment of the present invention, the third output contains the integrated information of the two branches of the first decoder layer. The fully connected layer (Feed-Forward Network, FFN) receives the third output and transforms the input through two linear transformations and a nonlinear activation function (such as ReLU). Expand to a higher dimension and then map back to the original or required dimension, thereby enhancing the model's expressiveness and learning complex features, and finally outputting the l Output of the decoder layer .
[0055] Step 206:l Output of the decoder layer It is used as the input query for the next decoder layer and the above process is repeated.
[0056] In one embodiment of the present invention, l Output of the decoder layer is used as the input query for the next decoder layer. This establishes a chain of information passing between decoder layers. This process is repeated in each decoder layer, with each layer further processing and refining the output of the previous layer, gradually improving the accuracy of the prediction.
[0057] Step 207: After being stacked through multiple decoder layers, the target query is output.
[0058] In one embodiment of the present invention, multiple decoder layers are stacked, and each decoder layer updates and optimizes the input query. This stacking can process information at different abstract levels and gradually improve the accuracy of target detection. After being processed by all decoder layers, the final query (i.e., target query) output contains detailed information of the target in the image, such as location, category, and other related information. The target query is the basis for target detection.
[0059] In summary, according to the technical solution provided by the embodiment of the present invention, the multi-scale feature map processed by the Transformer encoder can capture rich feature information of targets of different scales in the image; further, the dual-branch structure in the decoder allows the model to learn target features from two perspectives at the same time. The first branch can better understand the relationship within the target and the relationship between the target and the feature map through the combination of self-attention and cross-attention. The second branch only uses cross-attention to directly associate the query and the feature map, which helps the model focus on the extraction of target features. By fusing the outputs of the first branch and the second branch, the model can combine the advantages of the two attention mechanisms to obtain a more comprehensive target feature representation; further, the third output is used as the input query of the next decoder layer, allowing the model to transfer and refine information between multiple decoder layers. This layer-by-layer processing method helps to gradually improve the accuracy of target detection; further, after processing through multiple decoder layers, the output target query contains detailed information about the target in the image, such as location, category, and other related information. The target query is the basis for target detection. They are refined and optimized at multiple layers to improve the accuracy of detection. The above design helps to balance the influence of self-attention and cross-attention in the training process and reduce the adverse interactions between them, thereby improving the detection performance of the model.
[0060] Figure 3 FIG. 3 is a flow chart of the pollen target detection method provided by the present invention. Figure 3 As shown, the target detection results include: the predicted value of the intersection over union (IoU) between the predicted bounding box and the true box, the predicted classification, and the predicted bounding box.
[0061] Specifically, the predicted value of the intersection over union (IoU) between the predicted bounding box and the true bounding box is an indicator that measures the degree of overlap between the predicted bounding box and the true bounding box. Prediction classification is to classify the detected object and predict the category to which it belongs. The predicted bounding box is the location of the predicted object in the image, usually represented by the coordinates of the bounding box.
[0062] Based on the target query, the prediction head outputs the target detection result, which includes the following steps: Step 301: construct an intersection-over-union loss function, where the intersection-over-union loss function is used to calculate the deviation between the predicted value of the IoU and the true value of the IoU.
[0063] In one embodiment of the present invention, in addition to predicting the classification and predicting the bounding box, an IoU prediction value can also be introduced in the prediction head of the model. Specifically, in the target detection task, the model will output one or more predicted bounding boxes, and each predicted box will have a corresponding IoU prediction value, which is based on the model's estimate of the degree of overlap between the predicted box and the real box. At the same time, for each predicted box, an IoU true value can be calculated, which is calculated based on the actual degree of overlap between the predicted box and the real box. The core function of the intersection loss function is to compare the difference between the two values, that is, the deviation between the predicted value and the true value. This deviation can be positive or negative, depending on whether the predicted value overestimates or underestimates the degree of overlap of the real box. The purpose of the loss function is to minimize this deviation, that is, to adjust the model parameters through the training process so that the predicted IoU value is as close to the real IoU value as possible. During the model training process, the weights of the model are updated according to the gradient calculated by the IoU loss function through the back propagation algorithm, so that the model learns how to adjust its prediction so that the IoU prediction value is closer to the true value.
[0064] Exemplarily, the intersection-over-union loss function is: in, To query the predicted value of the intersection-and-union ratio of q, is the true value of the intersection and union ratio of query q, where query q is an element in the target query.
[0065] Step 302: Based on the predicted value of the intersection-over-union ratio and the predicted classification, a classification loss function is constructed. The classification loss function is used to calculate the deviation between the probability of the predicted classification and the true classification label.
[0066] In one embodiment of the present invention, the constructed classification loss function is intended to measure the difference between the classification probability predicted by the model and the true classification label. This loss function is particularly suitable for target detection tasks, in which the model not only predicts the category of the target, but also predicts the intersection-over-union (IoU) value of the target bounding box. By combining the IoU prediction value and the predicted classification, the classification loss function can more comprehensively evaluate the performance of the model on the classification task. Specifically, the classification loss function not only focuses on the accuracy of the classification, but also considers the degree of spatial overlap between the predicted bounding box and the true bounding box. This combination enables the classification loss function to more accurately reflect the comprehensive performance of the model in target detection, including the accuracy of category prediction and the accuracy of bounding box positioning. During the training process, the classification loss function guides the adjustment of model parameters by calculating the deviation between the predicted classification probability and the true classification label. The goal is to minimize this deviation so that the predicted classification probability is as close as possible to the probability of the true classification label, while improving the IoU value of the bounding box prediction, thereby improving the overall performance of the model.
[0067] Furthermore, based on the predicted value and predicted classification of the intersection-over-union ratio, a classification loss function is constructed, including: Based on the predicted value of the intersection-over-union ratio and the probability of the predicted classification, a weighted geometric mean function is constructed; Divide the image to be detected into a foreground sample set and a background sample set; A classification loss function is constructed based on the foreground sample set, the background sample set, the predicted classification and the weighted geometric mean function.
[0068] Specifically, based on the predicted value of the intersection over union ratio and the probability of the predicted classification, a weighted geometric mean function is constructed. The weighted geometric mean function combines the predicted value of IoU and the predicted classification, aiming to balance the classification accuracy and positioning accuracy. The image to be detected is divided into a foreground sample set and a background sample set. The foreground sample set is a set of regions that contain all areas marked as targets. The background sample set is a set of regions that contain all areas marked as background. Based on the foreground sample set, the background sample set, the predicted classification and the weighted geometric mean function, a classification loss function is constructed.
[0069] Exemplarily, the weighted geometric mean function is: Among them, s is the probability of predicted classification, E(IoU) is the predicted value of intersection over union, and α is the weight parameter, which is used to adjust the influence of the probability of predicted classification and the predicted value of intersection over union on the weighted geometric mean function.
[0070] The classification loss function is: Among them, Npos represents the foreground sample set, Nneg represents the background sample set, i represents the i-th element of the foreground sample set, j represents the j-th element of the background sample set, and BCE represents the binary cross entropy loss function.
[0071] The classification loss function is designed to consider both classification accuracy and target localization accuracy. By combining IoU and classification probability and calculating the loss separately on foreground and background samples, the model can optimize both aspects during training. The advantage of this method is that it can more comprehensively evaluate the performance of the model, especially in tasks that require precise localization and classification, such as target detection and image segmentation.
[0072] Step 303: construct a loss function for the predicted bounding box, where the loss function for the predicted bounding box is used to calculate the deviation between the predicted bounding box and the true bounding box.
[0073] In one embodiment of the present invention, the loss function for predicting bounding boxes is intended to quantify the difference between the bounding boxes predicted by the model and the actual annotated bounding boxes. This difference or deviation is a key indicator of the accuracy of the model's predictions. The loss function guides the model to adjust parameters during training to minimize the deviation between the predicted bounding boxes and the true bounding boxes. In this way, the model learns how to more accurately predict the location and size of the bounding boxes.
[0074] Step 304: Optimize the parameters of the object detection model based on the loss function of the predicted bounding box, the intersection-over-union loss function, and the classification loss function.
[0075] In one embodiment of the present invention, the loss of the predicted bounding box is the difference between the bounding box predicted by the model and the true bounding box. The bounding box regression loss can usually be used to measure this difference and guide the model to learn more accurate bounding box coordinates. The intersection loss function is used to optimize the model's prediction of the overlap between the predicted bounding box and the true bounding box. By minimizing the difference between the predicted IoU value and the true IoU value, the model can more accurately locate the target. The classification loss function is used to measure the difference between the model's predicted classification and the true category, and the classification loss ensures that the model can correctly identify the category of the target. During the training process of the target monitoring model, the three losses of the predicted bounding box loss, the intersection loss function and the classification loss function are combined to form a total loss function. The optimization goal of the model is to minimize this total loss function. In this way, the model is improved in predicting bounding boxes, estimating IoU values and classification tasks. The target query is a learnable parameter in the model for predicting classification categories, predicting bounding boxes, and the predicted intersection of union (IoU) between the predicted bounding box and the true box. By minimizing the total loss function, the target query is optimized to more accurately reflect the target information in the image. During the training process, the gradient of the total loss function with respect to the model weights is calculated through the back-propagation algorithm. These gradients guide the update of the weights, making the model's predictions more accurate. The above process is repeated in each iteration of model training. Over time, the model gradually learns how to better predict bounding boxes, estimate IoU values, and classify, thereby improving the performance of object detection.
[0076] Step 305: Based on the target detection model with optimized parameters and the target query, the prediction head outputs the target detection result.
[0077] In one embodiment of the present invention, a target detection model with optimized parameters is used to analyze the target query, and then the target detection result is output according to the analysis result of the model.
[0078] In one embodiment of the present invention, the parameter optimization of the target detection model is achieved by comprehensively considering the loss function of the predicted bounding box, the intersection-over-union loss function, and the classification loss function. The loss function of the predicted bounding box, the intersection-over-union loss function, and the classification loss function comprehensively evaluate the performance of the model in terms of bounding box positioning, IoU prediction, and classification accuracy. By minimizing the comprehensive loss function (i.e., the loss function of the predicted bounding box, the intersection-over-union loss function, and the classification loss function), the parameters of the model are optimized. In this way, the model can not only improve its classification accuracy, but also improve the accuracy of bounding box positioning, thereby achieving better overall performance in the target detection task. Through the above steps, the model can simultaneously optimize the classification and positioning tasks during the training process, and ultimately achieve the purpose of improving target detection performance.
[0079] For example, in the final target detection result calculation part, the target query is scored using the predicted value of the predicted IoU, the predicted classification, and the predicted bounding box. This scoring mechanism can reduce the phenomenon of target query misalignment, that is, reduce the mismatch between the bounding box predicted by the model and the real bounding box. By combining the predicted value of the predicted IoU, the predicted classification, and the predicted bounding box, the model can more accurately identify and locate the target object.
[0080] In summary, according to the technical solution provided by the embodiment of the present invention, by comprehensively considering the IoU prediction value, the prediction classification and the prediction bounding box, the model can more accurately identify and locate the target object, thereby improving the accuracy of target detection. Specifically, the parameters of the model are optimized by using the intersection-over-union loss function, the classification loss function and the prediction bounding box loss function, so that the model can learn more effective feature representation and bounding box prediction during the training process. In addition, the combination of multiple loss functions not only focuses on the accuracy of classification, but also considers the degree of spatial overlap between the predicted bounding box and the true bounding box, so that the evaluation of model performance is more comprehensive. By scoring the target query using the predicted value of the predicted IoU, the predicted classification and the predicted bounding box, the situation where the bounding box predicted by the model does not match the true bounding box is reduced. The model optimizes the classification and positioning tasks simultaneously during the training process, so that the model can maintain high detection performance when facing different targets and backgrounds. By minimizing the comprehensive loss function, the model learns how to better generalize to new and unseen data, thereby improving the applicability of the model in practical applications. The model can directly perform end-to-end training from the target detection results to the model weights through the back propagation algorithm, which simplifies the training process and improves the training efficiency. In summary, the above technical solution optimizes the target detection model by comprehensively utilizing multiple loss functions, thereby improving detection accuracy, optimizing model parameters, comprehensively evaluating model performance, and enhancing model robustness, ultimately achieving the goal of improving target detection performance.
[0081] Figure 4 It is a functional block diagram of the pollen target detection device provided by the present invention.
[0082] like Figure 4 As shown, the pollen target detection device 400 includes a model building module 401, a first acquisition module 402, a second acquisition module 403, a third acquisition module 404 and a detection result output module 405.
[0083] The model construction module 401 is used to construct a target detection model, which includes a backbone network, a Transformer encoder, a dual-branch attention decoder and a prediction head, wherein the decoder includes multiple decoder layers, each decoder layer includes two parallel branches, the first branch performs self-attention and cross-attention calculations, and the second branch only performs cross-attention calculations.
[0084] Furthermore, the first acquisition module 402 is used to acquire a multi-scale feature map corresponding to the image to be detected based on the backbone network; Furthermore, the second acquisition module 403 is used to input the multi-scale feature map into the Transormer encoder to obtain the enhanced feature map; Further, a third acquisition module 404 is used to input the enhanced feature map and the initialized query into a dual-branch attention decoder to obtain a target query output by the decoder; wherein the decoder includes a plurality of decoder layers, each decoder layer includes two parallel branches, the first branch performs self-attention and cross-attention calculations, and the second branch only performs cross-attention calculations; Furthermore, the detection result output module 405 is used to output the target detection result by the prediction head based on the target query.
[0085] Figure 5 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the pollen target detection method.
[0086] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0087] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the pollen target detection methods provided by the above methods.
[0088] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the pollen target detection method provided by the above-mentioned methods.
[0089] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0090] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A pollen target detection method, characterized in that: include: Constructing a target detection model, the target detection model includes a backbone network, a Transformer encoder, a dual-branch attention decoder and a prediction head, wherein the decoder includes multiple decoder layers, each of the decoder layers includes two parallel branches, the first branch performs self-attention and cross-attention calculations, and the second branch only performs cross-attention calculations; Based on the backbone network, a multi-scale feature map corresponding to the image to be detected is obtained; Inputting the multi-scale feature map into a Transformer encoder to obtain an enhanced feature map; Inputting the enhanced feature map and the query into a dual-branch attention decoder to obtain a target query output by the decoder; Based on the target query, the prediction head outputs a target detection result.
2. The pollen target detection method according to claim 1, characterized in that: The step of inputting the enhanced feature map and the query into a dual-branch attention decoder to obtain a target query output by the decoder includes: The enhanced feature map and query Input to the decoder of the dual-branch attention l Decoder layer; In the said l In the first branch of the decoder layer, the query Perform self-attention calculation to get the updated query , and query after update Perform cross attention calculation with the feature map to obtain the first output ; In the said l In the second branch of the decoder layer, the query Perform cross attention calculation with the feature map to obtain the second output ; Fusion first output and the second output , and get the third output ; The third output Enter into the l The fully connected layer of the decoder layer obtains the l Output of the decoder layer ; The said l Output of the decoder layer As the input query of the next decoder layer, and repeat the above process; After being stacked and processed by the multiple decoder layers, the target query is output.
3. The pollen target detection method according to claim 2, characterized in that: The updated query The first output The second output and the third output And the l Output of the decoder layer Calculate according to the following formulas: =Norm ( FFN ( )) Among them, Self represents the calculation of self-attention, Que ( )、 Key ( )、 Value ( ) represent the extraction of query vector, key vector and value vector from the query, respectively. They represent the key vector and value vector extracted from the enhanced feature map respectively, Z is the enhanced feature map output by the encoder, Norm means normalizing the result, and Cross means calculating cross attention.
4. The pollen target detection method according to claim 1, characterized in that: The target detection result includes: a predicted value of the intersection over union (IoU) between the predicted bounding box and the real box, a predicted classification, and a predicted bounding box; The outputting of the target detection result by the prediction head based on the target query includes: Constructing an intersection-over-union loss function, wherein the intersection-over-union loss function is used to calculate the deviation between the predicted value of the IoU and the true value of the IoU; Based on the predicted value of the intersection-over-union ratio and the predicted classification, constructing a classification loss function, wherein the classification loss function is used to calculate the deviation between the probability of the predicted classification and the true classification label; Constructing a loss function of the predicted bounding box, wherein the loss function of the predicted bounding box is used to calculate the deviation between the predicted bounding box and the true bounding box; optimizing the parameters of the target detection model based on the loss function of the predicted bounding box, the intersection-over-union loss function, and the classification loss function; Based on the target detection model with optimized parameters and the target query, the prediction head outputs the target detection result.
5. The pollen target detection method according to claim 4, characterized in that: The intersection-over-union loss function is: in, To query the predicted value of the intersection-and-union ratio of q, is the true value of the intersection and union ratio of query q, where query q is an element in the target query.
6. The pollen target detection method according to claim 4, characterized in that: The step of constructing a classification loss function based on the predicted value of the intersection-over-union ratio and the predicted classification includes: Constructing a weighted geometric mean function based on the predicted value of the intersection-over-union ratio and the probability of the predicted classification; Dividing the image to be detected into a foreground sample set and a background sample set; A classification loss function is constructed based on the foreground sample set, the background sample set, the predicted classification and the weighted geometric mean function.
7. The pollen target detection method according to claim 6, characterized in that: The weighted geometric mean function is: Wherein, s is the probability of predicted classification, E(IoU) is the predicted value of intersection over union, and α is a weight parameter used to adjust the influence of the probability of predicted classification and the predicted value of intersection over union on the weighted geometric mean function.
8. The pollen target detection method according to claim 7, characterized in that: The classification loss function is: Among them, Npos represents the foreground sample set, Nneg represents the background sample set, i represents the i-th element of the foreground sample set, j represents the j-th element of the background sample set, and BCE represents the binary cross entropy loss function.
9. A pollen target detection device, characterized in that: include: A model building module, used to build a target detection model, wherein the target detection model includes a backbone network, a Transformer encoder, a dual-branch attention decoder and a prediction head, wherein the decoder includes multiple decoder layers, each of the decoder layers includes two parallel branches, the first branch performs self-attention and cross-attention calculations, and the second branch only performs cross-attention calculations; A first acquisition module, used to acquire a multi-scale feature map corresponding to the image to be detected based on the backbone network; A second acquisition module is used to input the multi-scale feature map into a Transormer encoder to obtain an enhanced feature map; A third acquisition module is used to input the enhanced feature map and the query into a dual-branch attention decoder to obtain a target query output by the decoder; wherein the decoder includes a plurality of decoder layers, each of the decoder layers includes two parallel branches, the first branch performs self-attention and cross-attention calculations, and the second branch only performs cross-attention calculations; A detection result output module is used to output the target detection result by the prediction head based on the target query.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the pollen target detection method according to any one of claims 1 to 8 is implemented.