Field watermelon flower body detection method based on color weighting
By using color weighting and a dual-branch network structure, the problems of missed detection and false detection in watermelon flower detection are solved, achieving efficient detection of small targets in complex backgrounds and improving detection accuracy and real-time performance.
Patent Information
- Application Number
- CN202510964051.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-11-21
AI Technical Summary
Existing deep learning-based flower detection methods are prone to missed detections and false detections when the background is complex, the target objects are small and vary greatly, and lack effective evaluation and detection of flowers at different growth stages.
A color-weighted field watermelon flower detection method is adopted. RGB images are acquired and yellow-weighted to generate ExY images. Features of RGB and ExY images are extracted separately using a dual-branch backbone network. The DyCDSA and RepWTC3 modules are combined for feature fusion and interaction to improve feature extraction capability. The RT-DETR decoder is used to generate categories and bounding boxes.
It improves the detection accuracy and sensitivity of small targets, enhances the detection performance of watermelon flowers under complex backgrounds, and meets the detection needs of actual agricultural pollination scenarios.
Smart Images

Figure CN120997661A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a color-weighted method for detecting watermelon flowers in the field, belonging to the field of computer vision target detection. Background Technology
[0003] In recent years, deep learning-based machine vision has become an important technology for flower head recognition and detection. It has been widely used in tasks involving the recognition of single or adjacent flowers in simple backgrounds, demonstrating high recognition speed and accuracy. However, with technological advancements, the field of view has expanded, and the area of the target object in the captured image has decreased. The focus has shifted from medium to large targets to small targets in complex backgrounds. At this point, many studies have utilized RGB+ strategies based on deep learning-based machine vision, achieving flower head recognition even in complex scenes. Although current deep learning-based flower head detection methods have improved in both speed and accuracy, they are still prone to false negatives and missed detections when dealing with naturally grown watermelon flowers that are small in scale, highly variable, and have low contrast with the background, resulting in less than ideal detection performance. Summary of the Invention
[0004] To address the issues of missed detections and false detections caused by complex backgrounds and small, highly variable target sizes in existing deep learning-based object detection methods, as well as the lack of evaluation and detection capabilities for flowers at different growth stages, this invention provides a color-weighted field watermelon flower detection method. The technical solution is as follows:
[0005] Step 1: Obtain the RGB image of the watermelon flower to be detected;
[0006] Step 2: Apply yellow weighting to the RGB image of the watermelon flower to be detected to obtain the super yellow operator ExY image;
[0007] Step 3: Perform target detection on the RGB and ExY images using a watermelon flower body detection network;
[0008] The watermelon flower detection network includes: a backbone network, an encoder, and a decoder;
[0009] The backbone network includes two branches, which extract features from the RGB image and the ExY image respectively. The features are then fused and input into the encoder and decoder to achieve target detection.
[0010] The method for constructing the ExY image of the super yellow operator is as follows:
[0011]
[0012] In the formula, G represents the green weight, R represents the red weight, and B represents the blue weight.
[0013] Optionally, each branch in the backbone network includes a convolutional layer with secondary connections, a max pooling layer, and multiple BasicBlock modules. The features of each branch are synchronously input into the DyCDSA module after passing through the BasicBlock module. The output of the DyCDSA module is added to the output features of the previous BasicBlock module of each branch and then input into the next BasicBlock module, thereby realizing bimodal information interaction.
[0014] Optionally, the encoder includes a RepWTC3 module, which uses wavelet convolution WTConv to extract multi-frequency domain features.
[0015] Optionally, the decoder adopts the RT-DETR decoder structure, which consists of multiple decoder layers. Each layer includes a self-attention mechanism and a cross-attention mechanism. It also iteratively optimizes the object query through an auxiliary prediction head to generate categories and bounding boxes.
[0016] Optionally, the detection results of watermelon flowers in the field include: open male flowers, obscured unknown flowers, closed male flowers, open female flowers, and closed female flowers.
[0017] A second objective of this invention is to provide a field watermelon flower detection system, the system being used to implement the field watermelon flower detection method as described in any of the preceding claims, the system comprising:
[0018] The image acquisition module is configured to acquire the RGB image of the watermelon flower to be detected;
[0019] The weighted fusion module is configured to perform yellow weighting on the RGB image of the watermelon flower to be detected, so as to obtain the super yellow operator ExY image;
[0020] The detection module is configured to perform target detection on the RGB image and ExY image using a watermelon flower body detection network;
[0021] The watermelon flower detection network includes: a backbone network, an encoder, and a decoder;
[0022] The backbone network includes two branches, which extract features from the RGB image and the ExY image respectively. The features are then fused and input into the encoder and decoder to achieve target detection.
[0023] Optionally, the image acquisition module uses a CMOS camera to acquire images.
[0024] A third objective of this invention is to provide a field watermelon flower detection device, comprising a memory and a processor;
[0025] The memory is used to store computer programs;
[0026] The processor is configured to, when executing the computer program, implement the field watermelon flower detection method as described in any of the preceding claims.
[0027] A fourth objective of the present invention is to provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the field watermelon flower detection method as described in any of the preceding claims.
[0028] The beneficial effects of this invention are:
[0029] The watermelon flower detection method of this invention extracts the yellow features of the watermelon flower through a color-weighted fusion strategy and outputs ExY, which is a three-channel grayscale image. In the network input part, ExY is concatenated with an RGB image to input a six-channel image for feature extraction, thus supplementing the color feature information of the RGB image. A dual-branch backbone network (DB module) is designed in the backbone of the detection network, allowing the backbone network to extract features of different modalities separately. This maximizes the acquisition and supplementation of important feature information such as the color, texture, and shape of the watermelon flower, and avoids the feature redundancy and confusion problems that occur when a single backbone processes dual-modal features.
[0030] In one embodiment of the present invention, a DyCDSA module is further embedded in the DB module to combine the feature channels of RGB and ExY and adapt different attention mechanisms to feature maps at different network depths, thereby effectively enhancing bimodal information interaction and improving sensitivity to small targets.
[0031] In one embodiment of the present invention, wavelet convolution WTConv is further introduced into the RepC3 structure to form the RepWTC3 module, which endows the convolution with the ability to extract features in multiple frequency domains and significantly reduces the computational cost while effectively expanding the receptive field.
[0032] Experimental results demonstrate that the field watermelon flower detection model constructed in this invention can improve the detection accuracy of targets that are small in scale, highly varied, and have low contrast with the background, thereby enhancing the overall performance of the field watermelon flower detection model. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0034] Figure 1 This is a flowchart of a color-weighted field watermelon flower detection method according to the present invention.
[0035] Figure 2 This is a structural diagram of a color-weighted field watermelon flower detection model according to the present invention.
[0036] Figure 3 This is a schematic diagram of the color weighted fusion strategy in this invention.
[0037] Figure 4 This is a schematic diagram of the dual-branch backbone network in Embodiment 2 of the present invention.
[0038] Figure 5 This is a schematic diagram of the DyCDSA module in Embodiment 2 of the present invention.
[0039] Figure 6 This is a schematic diagram of the RepWTC3 module in Embodiment 2 of the present invention.
[0040] Figure 7 This is a schematic diagram of the watermelon flower types in this invention.
[0041] Figure 8 The image shows the detection results of watermelon flower bodies by various network models. Detailed Implementation
[0042] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0043] Example 1:
[0044] This embodiment provides a method for detecting watermelon flowers in the field, characterized in that the method includes:
[0045] Step 1: Obtain the RGB image of the watermelon flower to be detected;
[0046] Step 2: Apply yellow weighting to the RGB image of the watermelon flower to be detected to obtain the super yellow operator ExY image;
[0047] Step 3: Use the watermelon flower body detection network to perform target detection on the RGB and ExY images;
[0048] The watermelon flower detection network consists of a backbone network, an encoder, and a decoder.
[0049] The backbone network consists of two branches, which extract features from RGB and ExY images respectively. The features are then fused and input into the encoder and decoder to achieve object detection.
[0050] Example 2:
[0051] This embodiment provides a color-weighted method for detecting watermelon flowers in the field. (See [link to relevant documentation]). Figure 1 ,include:
[0052] Step 1: Divide watermelon flowers into open male flowers, occluded unknown flowers, closed male flowers, open female flowers, and closed female flowers. Collect watermelon flower images from the field according to the defect classification, and complete the annotation, format conversion, and dataset division.
[0053] Step 2: Construct a field watermelon flower detection model based on the color-weighted fusion strategy, DB module, DyCDSA module, and RepWTC3 module.
[0054] Step 3: Train the dataset using the field watermelon flower detection model to obtain the trained field watermelon flower detection model.
[0055] Step 4: Input the image to be detected into the trained field watermelon flower detection model to detect the target flower and obtain the final detection result.
[0056] The color-weighted fusion strategy addresses the poor detection results caused by low contrast between the subject and background. When the target and background colors are similar, and there is a lot of clutter and mutual occlusion, the network convolutional layers often fail to extract richer target feature information due to low distinction between subject and background features, severely impacting network recognition accuracy. Watermelon flower detection requires distinguishing the flower from the surrounding complex vine and leaf environment, and the aforementioned problem easily occurs during the detection process. Shape, color, and texture are the most commonly used feature information in detection. Among them, color features are highly robust to changes in size, motion blur, and image quality, and are the most intuitive visual feature possessed by every object. Images acquired using ordinary cameras typically have RGB three channels. Therefore, to address the accuracy degradation caused by factors such as similar colors, a color-weighted fusion strategy is designed to extract the yellow features of the watermelon flower and output an ExY image.
[0057]
[0058] In the formula, G represents the green weight; R represents the red weight; and B represents the blue weight.
[0059] The ExY image is a 3-channel grayscale image. Fusing the RGB and ExY images to obtain a 6-channel image highlights the yellow features of the flower, effectively improving the feature extraction efficiency of the network. Figure 3 As shown.
[0060] The structure of the field watermelon flower detection model is as follows: Figure 2 As shown, it includes a dual-branch backbone (DB), an encoder, and a decoder.
[0061] The structure of the DB module is as follows: Figure 2 and Figure 4 As shown, it includes two branches. While the color-weighted fusion strategy adds yellow feature information to the watermelon flower, it suppresses features such as texture and shape to some extent compared to the RGB image. To preserve color and texture features, the dual-branch structure can further split the fused 6-channel image into RGB and ExY images. The main branches of the two branches are used to extract features from the two modalities separately, thereby better preserving the richness of feature map information and making full use of the complementary information of the two modalities for target detection.
[0062] The workflow of the DB module is as follows:
[0063] (1) After the network inputs a six-channel image, the image is split into RGB and ExY images by the Multi-in module;
[0064] (2) After the dual-modal images enter their respective branches, they pass through the convolutional layer (Conv), the max pooling layer (MaxPooling), and the basic block (BasicBlock) to generate multi-scale feature maps in order to obtain and supplement important feature information such as color, texture and shape of watermelon flower to the greatest extent. In addition, the number of feature map channels of each branch is halved to reduce the number of parameters.
[0065] (3) The last three layers of bimodal feature maps are spliced together in the same layer and the result is input into the encoder.
[0066] The DB module performs bilinear feature extraction by splitting the 6-channel bimodal image. While preserving the original image features to a great extent, it also adds color feature information and achieves pixel-level fusion of bimodal feature information. In addition, the feature extraction strategy of halving the number of channels avoids the problems of excessive computation and parameters, achieving high performance with low computational cost.
[0067] To address the lack of information interaction between bimodal information, a Dynamic Channel-aware Depth-wise Spatial Attention (DyCDSA) module is embedded between the DB modules. The DyCDSA module combines the RGB and ExY feature channels and adapts different attention mechanisms to feature maps at different network depths, thereby effectively enhancing bimodal information interaction and improving sensitivity to small objects.
[0068] like Figure 2As shown, the dual-modal features of BasicBlock are synchronously input into the DyCDSA module, and the results are then added to the output features of the previous layer of each branch to enhance the information interaction between the two modalities.
[0069] As network depth increases, the spatial resolution of feature maps decreases layer by layer, while the receptive field gradually expands, causing the network to shift from capturing local details to modeling high-level global semantic information. Therefore, shallow features, with their high spatial resolution and rich local structural information, are more suitable for introducing spatial attention mechanisms to enhance the response of key regions; while deep features have stronger semantic expressive power and richer channel dimensions, making channel attention mechanisms more helpful in mining key semantic channels to improve feature discriminativeness.
[0070] The structure of the DyCDSA module is as follows: Figure 5 As shown, the bimodal features from the previous layer are cascaded through channels and then input into the H-dimensional discriminator for subsequent operations. The specific steps are as follows:
[0071] (1) Feature map information F for H>80 A First, spatial features are obtained through an Efficient Multi-scale Attention (EMA) module. Then, a Channel-wise Fully Connected Layer (Channel-FC) is used to achieve a non-linear mapping of global channel features. Finally, the result is weighted and multiplied by the elements of the Spatial Feature. For F... A The mathematical expression for the processing procedure is as follows:
[0072] F spa =EMA(F A )
[0073] W ch =σ(f2(ReLU(f1(GAP(F)) spa )))))
[0074] In the formula, H > 80; GAP is Global Average Pooling, ReLU is the ReLU activation function, σ is the Sigmoid function, and W... ch W represents the channel attention weights. ch ∈[0,1] C×1×1 .
[0075] In the shallow stage with a relatively high feature resolution, to avoid the over-suppression of channel attention on detailed information, a channel attention reconciliation mechanism is introduced. By linearly scaling and translating the original channel weights:
[0076] W = 0.5·W ch + 0.5
[0077]
[0078] where W is the linear reconciliation weight and W ∈ [0.5, 1] C×1×1 , F mix is the mixed feature map.
[0079] (2) For the feature map information with 40 < H ≤ 80 adopt a processing method similar to that of F A , but do not introduce the channel attention reconciliation mechanism. The obtained channel attention weights are directly multiplied by the elements of the spatial feature map:
[0080]
[0081] (3) For the feature map information with H < 40 first perform average pooling and max pooling respectively, then pass through a Multilayer Perceptron (MLP) respectively and add the result elements to generate a Channel Attention Map (CAMap). Finally, multiply CAMap by the input feature map information F C element-wise. Its mathematical expression is as follows:
[0082] CA(F C ) = σ(MLP(AvgPool(F C )) + MLP(MaxPool(F C )))
[0083]
[0084] where MaxPool is max pooling and AvgPool is average pooling.
[0085] (4) After passing through the dynamic channel perception module, a mixed-feature map (MFMap) is obtained and input into the depth-wise convolution module for subsequent operations. In this module, multi-scale depth convolutions are used to capture the spatial relationships between features, which enhances the ability to focus on the target region and reduces computational complexity. In addition, depth convolutions with long kernels of 1×5, 1×7, and 1×11 are used to finely adjust the receptive field, expanding the range of information aggregation while maintaining resolution. The tail of the depth convolution module uses 1×1 Conv pairs for convolution operations to generate the final refined spatial attention map (SAMap), and the SAMap and MFMap elements are multiplied together to obtain the multi-refined feature. The mathematical expression of its calculation process is as follows:
[0086]
[0087] In the formula, F mix For Mixed Feature Map (MFMap), DwConv is for depthwise convolution, and Branch is for... i ,i∈{0,1,2,3} is the i-th branch, where Branch0 is the identity connection, F mul It features multiple refining characteristics.
[0088] Finally, to address the issues of high-frequency noise suppression and low sensitivity to multi-scale targets, a RepWTC3 module is introduced into the encoder. The encoder consists of a hybrid encoder and a transform encoder with an auxiliary prediction head. The hybrid encoder transforms multi-scale features into a series of image features through intra-scale interaction and cross-scale fusion. Simultaneously, the encoder employs IoU-aware query selection, choosing a fixed number of features from the feature sequence output by the encoder as the initial object query for the decoder.
[0089] The structure of the RepWTC3 module is as follows: Figure 6 As shown, wavelet convolutions (WTConv) are used to improve the network's ability to extract features from multi-scale targets and to improve the problem of parameter dilation in the receptive field expansion of convolutional layers.
[0090] like Figure 6As shown, wavelet convolution first maps the input features to the frequency domain using wavelet transform (WT), then performs convolution operations on different frequency sub-bands, and finally returns to the spatial domain using inverse wavelet transform (IWT). This allows for targeted convolution operations based on the different high and low frequency components, capturing global contextual information using low-frequency components while preserving detailed features through high-frequency components, thus significantly improving the model's feature representation capability. Wavelet convolution is implemented using the DwConv operation, employing a custom orthogonal wavelet kernel for the input features. A two-dimensional transformation is performed to obtain four different frequency components [X]. LL ,X LH ,X HL ,X HH ] The mathematical expression process for enhancing the expressive power of features is as follows:
[0091]
[0092] [X LL ,X LH ,X HL ,X HH ] = DwConv([f LL ,f LH ,f HL ,f HH ],X)
[0093] In the formula, f LL For a low-pass filter, f LH For a horizontal high-pass filter, f HL For a vertical high-pass filter, f HH For diagonal high-pass filters; X LL For low-frequency components (which can provide structural information, such as smooth regions and contours), X LH For the horizontal high-frequency component (which can capture significant details of lateral variations, such as vertical edges), X HL For vertical high-frequency components (capturing vertical abrupt changes, such as horizontal edges), X HH These are diagonal high-frequency components (which can characterize weak textures and small targets, such as texture variations).
[0094] The decoder adopts the RT-DETR decoder structure, which consists of multiple decoder layers. Each layer contains two parts: self-attention and cross-attention. It also uses an auxiliary prediction head to iteratively optimize object queries and generate categories and bounding boxes.
[0095] like Figure 7As shown, the detection results in this embodiment include fully open male flowers, obscured flowers, closed male flowers, fully open female flowers, and closed female flowers. These flowers are the five flower categories that appeared most frequently and were most difficult to detect in actual planting, selected through actual surveys in watermelon fields, and all of them were actually open. The definition of "fully open" and "closed" is based on whether the pollination criteria are met: a flower is defined as "fully open" if the angle between the outer side of the bud and the vertical line is greater than 45° and the entire flower center can be seen when viewed vertically. Flowers that do not meet the above conditions are defined as "closed". The definition of "obscured" is based on whether the male and female flowers can be distinguished: a flower whose sex characteristics are completely obscured by vines, leaves, or other flowers (such as the stamen or a small watermelon below a female flower) is considered "obscured". A flower that is partially obscured but whose sex can still be identified is not defined as "obscured". In addition, in the determination process, regardless of distance, the principle is that the flower must be visible to a person in a natural environment at a fixed height and fixed downward angle.
[0096] To verify the effectiveness and technical effect of the detection method of the present invention, an experiment was conducted under the following conditions:
[0097] Using hardware devices such as CMOS cameras, lenses, and tripods, a watermelon flower defect dataset was constructed. The image resolution is 960×640, and there are a total of 3098 images. After acquisition, defects were labeled using Label Img and corresponding XML file information was generated.
[0098] The training platform for the experiment was Ubuntu 20.04.4LTS, utilizing the GPU acceleration of an NVIDIA GeForce RTX3090, based on PyTorch 1.12.1, and using CUDA 11.6 for computational acceleration. The batch size for training was set to 16, the epochs to 400, the initial learning rate to 0.0002, and the input image size to 960×640×6.
[0099] To verify the effectiveness of the improved model, this embodiment selects five common evaluation metrics for object detection, specifically:
[0100] (1) Average Precision (AP): The average recognition accuracy for each category;
[0101] (2) Mean Average Precision (mAP): The average recognition accuracy of all categories;
[0102] (3) FPS (Frames per second), the number of frames transmitted per second by the model, that is, the number of images that can be detected per second;
[0103] (4) Precision P, calculated using the following formula:
[0104]
[0105] (5) Recall R (Recall), the formula for which is as follows:
[0106]
[0107] In the formula, TP is the number of positive samples that are correctly detected, FP is the number of negative samples that are incorrectly detected as positive samples, and FN is the number of positive samples that are incorrectly detected as negative samples.
[0108] To verify the effectiveness of the color-weighted fusion strategy in this embodiment and explore its impact on the detection performance of watermelon flower targets, the influence of changing the color-weighted fusion strategy at different positions in the backbone network was analyzed, and the optimal fusion method was obtained. Different positions refer to fusion being completed either in the early stage of image data reading (Early Fusion) or after data reading during the feature extraction stage (Middle Fusion). The experimental results at different positions are shown in Table 1.
[0109] Table 1: Impact of different locations of the fusion strategy on network performance
[0110]
[0111] As shown in Table 1, after fusing the feature information of RGB and ExY, Early Fusion achieved a slight increase in mAP@0.5 while maintaining almost the same FPS. Except for a slight decrease in the accuracy of closed female flowers (Cf), the accuracy of all categories improved. Although Middle Fusion reduced the FPS due to the increase in the number of network layers, the strategy of halving the number of channels in each branch significantly reduced the computational cost of the network. Furthermore, because its feature information is fused during the extraction process, it fully utilizes the individual features of each modality and retains the advantages of each feature expression. Compared with Early Fusion, it has higher detection performance for the more complex open female flowers (Bf) and closed female flowers (Cf), and its mAP@0.5 is also more advantageous.
[0112] To verify the effectiveness of the watermelon flower detection network in this embodiment, and to analyze the impact of the DB module, DyCDSA module, and RepWTC3 module on the accuracy and speed of the network model, an ablation experiment was conducted under the same parameters. The experimental results are shown in Table 2.
[0113] Table 2: Ablation Experiment Results
[0114]
[0115] Table 2 shows that Experiment 2 verifies that the RepWTC3 module improves the detection accuracy of small targets while reducing the number of parameters, with significant improvements in the detection performance of three types of small targets: occluded unknowns, closed male flowers, and open female flowers. Experiment 3 demonstrates that the dual-backbone approach can significantly reduce model complexity while maintaining the model's real-time detection performance, and can extract bimodal information to improve detection performance, significantly improving the recognition accuracy of open male and open female flowers. Experiments 3 and 4 show that although the DyCDSA module sacrifices the model's detection rate, it strengthens the information interaction between different modalities and improves the network's feature extraction capabilities, resulting in improvements for all types of flowers except closed female flowers, especially open female flowers of varying scales and occluded unknowns. While DyCDSA and RepWTC3 reduce the model's detection speed, they still meet the overall requirements for real-time detection and improve the detection accuracy by 3.5% to 86.4%. The accuracy of detecting the main target objects, blooming male flowers and blooming female flowers, is improved by 3.4% and 4.2% respectively. Furthermore, the ability to capture small and weak occlusions of unknown objects is improved, resulting in an accuracy increase of 5.6%, effectively solving the problem of detecting small objects at multiple scales.
[0116] To further verify the superiority and effectiveness of the watermelon flower head detection network in this embodiment, under the same parameters, Cascade-RCNN, Cascade-RCNN with a replaced Mamba backbone, YOLOv8 to v12 series networks, Gold-YOLO, D-FINE-n, EfficientDet, DINO, RT-DETR series networks, and the algorithm of this embodiment were compared on a self-made field watermelon flower head dataset. The comparison was made using four metrics: accuracy (P), recall (R), mAP@0.5, and FPS. The comparison results are shown in Table 3.
[0117] Table 3: Comparison Results of Different Algorithms
[0118]
[0119] As shown in Table 3, the two-stage Cascade-RCNN has a large file size and low frame rate. After replacing the Mamba backbone, the mAP@0.5 decreased by 3.2 percentage points instead of increasing. The single-stage YOLO series networks and EfficientDet have lower computational cost but lower detection accuracy. Gold-YOLO and DINO have similar mAP@0.5, but their file size and frame rate are both lower than the average of the compared networks. In particular, DINO's large computational cost results in its frame rate being the lowest among the proposed algorithms. The D-FINE-n model has a similar file size and detection rate to this embodiment, but its accuracy is 1.2 percentage points lower. The detection network of this invention has the highest mAP@0.5, which is 3.5 percentage points higher than the baseline. Due to the increase in the number of network layers, the detection speed is lower than before the improvement, but it is still 31.3 frames / sec, which meets the requirements of real-time detection. The results show that the detection network of this invention enhances the network model's detection accuracy for small and multi-scale watermelon flowers, and can meet the detection task of watermelon flowers in actual agricultural pollination scenarios.
[0120] To visually illustrate the detection performance of the detection network in this embodiment, a direct comparison of the detection results from different networks is provided, such as... Figure 8 As shown in the figure. By comparing the results, it was found that (a4), (b4), (c4), (d4), and (e4) showed a significant improvement in detection accuracy compared to the original RT-DETR. The yellow weighting strategy can effectively reduce the false positive and false negative rates of the model. When the number of categories is small but there is interference from withered yellow leaves, (a2) falsely detected open male flowers, (a3) falsely detected occluded unknowns, and (a4) accurately identified the flowers after eliminating interference. (b2) missed occluded unknowns and closed male flowers, and (b3) missed occluded unknowns. When there are complex multi-layered scenes with many categories and different scales, (c2) and (c3) both missed occluded unknowns, and (c4) showed a significant improvement in detection accuracy. (d2) falsely detected closed male flowers and missed occluded unknowns, (d3) falsely detected closed female flowers, and (d4) accurately detected all flower categories. The results show that the network in this embodiment has better detection performance, which enhances the adaptive ability and feature extraction ability of low-contrast multi-scale targets in complex environments, and improves the ability to distinguish various flowers.
[0121] Some steps in the embodiments of the present invention can be implemented using software, and the corresponding software program can be stored in a readable storage medium, such as an optical disc or a hard disk.
[0122] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for detecting watermelon flower bodies in the field, characterized in that, The method includes: Step 1: Obtain the RGB image of the watermelon flower to be detected; Step 2: Apply yellow weighting to the RGB image of the watermelon flower to be detected to obtain the super yellow operator ExY image; Step 3: Perform target detection on the RGB and ExY images using a watermelon flower body detection network; The watermelon flower detection network includes: a backbone network, an encoder, and a decoder; The backbone network includes two branches, which extract features from the RGB image and the ExY image respectively. The features are then fused and input into the encoder and decoder to achieve target detection.
2. The method for detecting watermelon flowers in the field according to claim 1, characterized in that, The method for constructing the ExY image of the super yellow operator is as follows: In the formula, G represents the green weight, R represents the red weight, and B represents the blue weight.
3. The method for detecting watermelon flowers in the field according to claim 1, characterized in that, Each branch in the backbone network includes a convolutional layer with secondary connections, a max pooling layer, and multiple BasicBlock modules. The features of each branch are synchronously input into the DyCDSA module after passing through the BasicBlock module. The output of the DyCDSA module is added to the output features of the previous BasicBlock module of each branch and then input into the next BasicBlock module, thereby realizing bimodal information interaction.
4. The method for detecting watermelon flowers in the field according to claim 1, characterized in that, The encoder includes a RepWTC3 module, which uses wavelet convolution WTConv to extract multi-frequency domain features.
5. The method for detecting watermelon flowers in the field according to claim 1, characterized in that, The decoder adopts the RT-DETR decoder structure, which consists of multiple decoder layers. Each layer contains two parts: a self-attention mechanism and a cross-attention mechanism. It also uses an auxiliary prediction head to iteratively optimize object queries and generate categories and bounding boxes.
6. The method for detecting watermelon flowers in the field according to claim 1, characterized in that, The detection results of watermelon flowers in the field include: open male flowers, obscured (unknown), closed male flowers, open female flowers, and closed female flowers.
7. A field watermelon flower detection system, characterized in that, The system is used to implement the field watermelon flower detection method as described in any one of claims 1-6, the system comprising: The image acquisition module is configured to acquire the RGB image of the watermelon flower to be detected; The weighted fusion module is configured to perform yellow weighting on the RGB image of the watermelon flower to be detected, so as to obtain the super yellow operator ExY image; The detection module is configured to perform target detection on the RGB image and ExY image using a watermelon flower body detection network; The watermelon flower detection network includes: a backbone network, an encoder, and a decoder; The backbone network includes two branches, which extract features from the RGB image and the ExY image respectively. The features are then fused and input into the encoder and decoder to achieve target detection.
8. The field watermelon flower detection system according to claim 1, characterized in that, The image acquisition module uses a CMOS camera to acquire images.
9. A field watermelon flower detection device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is configured to implement the field watermelon flower detection method as described in any one of claims 1-6 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the field watermelon flower detection method as described in any one of claims 1-6.