Weed detection method based on semi-supervised diffusion generation network
By adopting a semi-supervised diffusion generation network in weed detection and using pseudo-image and dynamic attention mechanism, the problem of insufficient adaptability to labeled data and complex scenes in the prior art is solved, and a more efficient and robust weed detection effect is achieved.
Patent Information
- Application Number
- CN202510166386.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art relies heavily on labeled data in weed detection, making it difficult to adapt to complex scenarios, and the pseudo-label brings noise, and the real-time and universality are insufficient.
Weed detection method based on semi-supervised diffusion generation network is adopted, and high-quality pseudo-images are generated through the semi-supervised diffusion data generation module. Combining generation and detection, the dynamic generative attention mechanism and semi-diffusion loss function are used to optimize pseudo-label generation and feature extraction.
It significantly improves data utilization efficiency and model robustness, improves detection accuracy and generalization performance, solves the problem of over-dependence of traditional methods on labeled data, and enhances the real-time and universality of the model.
Smart Images

Figure CN120088455A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of generative learning, supervised learning, and object detection, and in particular, to a weed detection method based on a semi-supervised diffusion generative network. Background Art
[0002] Traditional deep learning methods based on supervised learning. Such technologies are represented by single-stage detection methods such as Faster R-CNN, SSD, YOLOv4, and YOLOv7. They strongly rely on labeled data, require a large amount of high-quality labeled data, have a high labeling cost, and it is difficult to obtain data in agricultural scenarios. In addition, methods such as Faster R-CNN have limited generalization ability. The dynamic agricultural environment is complex and affected by multiple factors such as weather, light, and weed growth stages, so the performance of the model in diverse scenarios is unstable. Another serious defect of such methods is that the model lacks attention to key details. Especially in complex backgrounds, the similarity between weeds and crops is high, and the model pays insufficient attention to weeds, resulting in easy misdetection or missed detection.
[0003] Semi-supervised methods based on pseudo-labels: Existing ones include adaptive pseudo-label strategies and multi-scale feature representation techniques; pseudo-labels may introduce incorrect labeling, that is, noise, which will seriously affect the model performance in complex scenarios; in terms of feature extraction of data, these two methods also have certain limitations. They are prone to feature loss or incompleteness problems in the feature extraction stage, resulting in insufficient ability to capture fine-grained features.
[0004] Optimized classical detection frameworks: Such as optimized models based on Faster R-CNN, using a feature extraction backbone network such as VGG19-CBAM. The defects are mainly reflected in that the detection speed of these models is slow. The average detection speed of the two models in the example is only 336 ms per image, lacking real-time performance and being difficult to meet the requirements of efficient operations; the adaptability is not strong enough. The above two models are only optimized for specific crops or scenarios and have poor generality, and their performance may decline in other weed or crop scenarios.
[0005] YOLO series methods including YOLOv7: The main problem of such methods is that the model needs to be continuously updated. Although the accuracy of the model is high, reaching an accuracy of 99.8%, it still needs to be frequently updated in actual deployment to adapt to the changing on-site environment; at the same time, it overly relies on hardware resources, and there are performance bottlenecks when the model runs on embedded or in-vehicle computers, but embedded or in-vehicle computers are quite common data acquisition tools in agricultural scenarios.
[0006] Currently, in terms of weed detection, the cost of obtaining large-scale labeled data in agricultural scenarios is high, so that the dependence of existing deep learning models on data has become a major factor hindering their development. At the same time, the morphological diversity and complexity of weeds pose challenges to traditional supervised learning models. In particular, the appearance of weeds and crops is often very similar, and the detection performance of traditional models is limited. Summary of the Invention
[0007] The purpose of the present invention is to provide a weed detection method based on a semi-supervised diffusion generation network, aiming at the problems commonly existing in the prior art, such as high dependence on labeled data, difficulty in adapting to complex scenarios, pseudo-labels bringing noise, and insufficient real-time performance and generality. Creatively, a semi-supervised framework is constructed based on a semi-supervised diffusion data generation module, and generation and detection are innovatively combined, significantly improving the data utilization efficiency and the robustness of the model.
[0008] To achieve the above purpose, the present invention provides a weed detection method based on a semi-supervised diffusion generation network, including the following steps:
[0009] Step 1: Data augmentation;
[0010] Step 2: Data input and multi-scale feature extraction;
[0011] Step 3: Pseudo-label generation and optimization;
[0012] Step 4: Feature fusion and target detection;
[0013] Step 5: Joint training and optimization.
[0014] Preferably, in Step 1, the image data is preprocessed through three data augmentation methods: cutmix, mixup, and random occlusion.
[0015] Cutmix is to paste a local rectangular area in one image into another image;
[0016] Mixup mixes two images and their corresponding labels according to a specific ratio;
[0017] Random occlusion is to simulate an occlusion scenario by randomly occluding a rectangular area in the image.
[0018] Preferably, the specific process of Step 2 is as follows:
[0019] The SAGE-Weed model first receives the input image data. After the input image is processed by the backbone network, it is subjected to feature extraction through the multi-scale feature extraction layer of the efficient hybrid encoder. Among them, the three feature extraction layers S3, S4, and S5 in the efficient hybrid encoder are used to extract low-level, intermediate-level, and high-level features respectively;
[0020] The efficient hybrid encoder uses an Adaptive Integrated Feature Interaction (AIFI) module internally. The AIFI module enhances information interaction between different scales through a dynamic generative attention mechanism. The calculation method of the AIFI module is as follows:
[0021] F AIFI = softmax(Q·K T )·V;
[0022] where Q, K, and V represent the query matrix, key matrix, and value matrix respectively, and F AIFI represents the output feature. The softmax function is used to normalize the attention weights;
[0023] Apply the SiLU activation function and batch normalization (BN) processing to each extracted feature layer. The output of the feature extraction layer is fused with the features generated by the Cross-Scale Compact Feature Fusion (CCFF) module and the semi-supervised diffusion data generation module, while gradually reducing the resolution of each layer of features.
[0024] Preferably, the specific process of step three is as follows:
[0025] Use the semi-supervised diffusion data generation module to generate pseudo-images from noise through the inverse diffusion process. On the generated pseudo-images, the SAGE-Weed model predicts the unlabeled data, calculates the confidence of each target region, and filters according to a preset confidence threshold, only retaining the regions with high confidence as pseudo-labels. By calculating the prediction uncertainty of each target region and selecting the target with the lowest uncertainty as the pseudo-label, the uncertainty is calculated according to the standard deviation σ:
[0026]
[0027] where, represents the predicted class probability for each iteration, represents the predicted average value;
[0028] Train the pseudo-labels together with the original labeled data to form a joint optimization process, and simultaneously minimize two losses through the semi-diffusion loss function, including the supervised loss of the labeled data and the pseudo-supervised loss of the unlabeled data.
[0029] Preferably, the specific process of step four is as follows:
[0030] The cross-scale compact feature fusion CCFF module first encodes the features generated by the semi-supervised diffusion data generation module through position embedding. The generated features are weighted and fused with the multi-scale features of the feature extraction network layer by layer. The weighted fusion process uses residual connections to add the generated features and the original image features to retain the original feature information and introduce supplementary information of the generated features. In addition, a dynamic generative attention mechanism is used to allocate weights to different features during the weighted fusion process. Finally, the fused feature results are sent to the decoder for further processing to restore the resolution, and category classification and bounding box regression are performed through the object detection head.
[0031] Among them, the calculation method of the weighted fusion weights is as follows:
[0032]
[0033] Among them, F gen and respectively represent the generated features and the image features, and W fusion is the fusion weight.
[0034] Preferably, the specific process of step five is as follows:
[0035] The pseudo-labels generated by the semi-supervised diffusion data generation module and the labeled data are jointly input into the detection network for feature extraction and object detection.
[0036] The detection network optimizes the network parameters by calculating the supervised loss of the labeled data and the pseudo-supervised loss of the pseudo-label data.
[0037] The pseudo-supervised loss of the pseudo-label data will be fed back to the semi-supervised diffusion data generation module to adjust and optimize the pseudo-label generation strategy.
[0038] Therefore, the present invention adopts the above-mentioned weed detection method based on a semi-supervised diffusion generation network, and the beneficial effects are as follows:
[0039] (1) The SAGE-Weed model provided by the present invention applies a semi-supervised diffusion data generation module, introduces a generative attention mechanism and a semi-diffusion loss function, and combines the characteristics of supervised learning and generative learning; effectively utilizes limited labeled data in weed detection to obtain robustness and detection accuracy beyond existing models.
[0040] (2) The present invention solves the problem of over-reliance on labeled data in traditional methods, and provides a new direction and basis for the research and application of complex object detection tasks in intelligent agriculture.
[0041] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. Description of the Drawings
[0042] Figure 1 It is the overall flowchart of an embodiment of the weed detection method based on a semi-supervised diffusion generation network of the present invention;
[0043] Figure 2 It is an example of the data augmentation method of an embodiment of the weed detection method based on a semi-supervised diffusion generation network of the present invention, where (a) is random occlusion, (b) is cutmix, and (c) is mixup;
[0044] Figure 3 It is a schematic diagram of the dataset samples of an embodiment of the weed detection method based on a semi-supervised diffusion generation network of the present invention, where A is Xanthium spinosum, B is Xanthium italicum, C is Amaranthus rudis, D is Setaria viridis, and E is Iva xanthifolia;
[0045] Figure 4 It is the overall structure diagram of SSDDGL of an embodiment of the weed detection method based on a semi-supervised diffusion generation network of the present invention;
[0046] Figure 5 It is a schematic diagram of the architecture of the dynamic generative attention mechanism of an embodiment of the weed detection method based on a semi-supervised diffusion generation network of the present invention;
[0047] Figure 6 It is the visual analysis of the weed detection results under different methods in an embodiment of the weed detection method based on a semi-supervised diffusion generation network of the present invention. Detailed implementation manners
[0048] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0049] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the art to which the present invention belongs. The "first", "second" and similar terms used in the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. The terms such as "comprising" or "including" mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0050] In actual production, to address the problems of complex dynamic environments and strong dependence on labeled data in weed detection, the SAGE-Weed model of the present invention (Semi-supervised Attention-driven Generative modEl for Weed Detection, which combines semi-supervised characteristics (Semi-supervised), attention mechanism (Attention-driven), generative ability (Generative), and application scenario (Weed Detection)) needs to address the following main technical challenges:
[0051] Weak dependence on labeled data: In the agricultural scenario, the cost of large-scale data annotation is relatively high. The model needs to make the most of the limited labeled data as much as possible to improve data utilization efficiency.
[0052] Overcoming environmental interference: The environment in agricultural applications is complex. In particular, weeds and crops usually have a high degree of similarity, resulting in a decrease in the accuracy of weed detection. The model should try to shield the interference.
[0053] Real-time performance: The model needs to have good real-time performance to help farmers make decisions promptly and quickly, and minimize losses in a timely manner.
[0054] Minimizing the introduction of pseudo-label noise: To weakly depend on labeled data, supervised learning must be utilized. However, during the application process, the noise problem caused by the use of pseudo-labels should be overcome as much as possible.
[0055] The present invention proposes a model called SAGE-Weed. By introducing few-shot learning and prototype attention mechanism, it meets the urgent needs of modern agriculture for efficient, accurate, and intelligent disease detection methods. The main steps include: data augmentation, multi-scale feature extraction, pseudo-label generation and optimization, feature fusion, and joint training. After the above steps, SAGE-Weed improves data utilization efficiency. In the case of limited labeled data, it makes full use of unlabeled data to improve detection accuracy. Based on the semi-supervised diffusion data generation module, high-quality data is generated to optimize the quality of pseudo-images, significantly enhancing the generalization performance of the model. It enhances the robustness of the model, and multi-scale feature fusion and dynamic loss balancing improve the detection performance in complex scenarios.
[0056] As Figure 1 shown, a weed detection method based on a semi-supervised diffusion generative network includes the following steps:
[0057] Step 1: Data augmentation;
[0058] Adopt as Figure 3For the shown dataset sample, when the dataset sample is small, the model is prone to overfitting, which affects the generalization ability of the model. In this embodiment, three data augmentation methods, namely cutmix, mixup, and random occlusion, are used to preprocess the image data, improve the generalization ability of the model, reduce overfitting, and at the same time play a role in increasing data samples to solve the problem of data scarcity.
[0059] As Figure 2 shown, examples of the three data augmentation methods are as follows:
[0060] Cutmix is to paste a local rectangular area in one image into another image; Mixup mixes two images and their corresponding labels according to a specific ratio to improve the robustness of the model to blurred boundaries and mislabeled samples; Random occlusion is to randomly occlude a rectangular area in the image to simulate an occlusion scenario.
[0061] Step 2: Multi-scale feature extraction. The specific process of Step 2 is as follows:
[0062] The SAGE-Weed model first receives the input image data. After the input image is processed by the backbone network, it is subjected to feature extraction through the multi-scale feature extraction layer of the efficient hybrid encoder. As Figure 1 shown, among them, the three key feature extraction layers S3, S4, and S5 included in the efficient hybrid encoder are used to extract low-level, intermediate-level, and high-level features respectively.
[0063] The adaptive integrated feature interaction AIFI module is used inside the efficient hybrid encoder to achieve cross-scale dynamic information interaction. The AIFI module enhances the information interaction between different scales through a dynamic generative attention mechanism. The calculation method of the AIFI module is expressed as follows:
[0064] F AIFI = softmax(Q·K T )·V;
[0065] Among them, Q, K, and V represent the query matrix, key matrix, and value matrix respectively, and F AIFI represents the output feature. The softmax function is used to normalize the attention weights. As Figure 5 shown, the dynamic generative attention mechanism is implemented through a multi-layer feature fusion network and consists of three modules: a generative feature embedding module, a feature fusion module, and a weighted output module.
[0066] Apply the SiLU activation function and batch normalization BN processing to each extracted feature layer to maintain the stability of gradient flow. The outputs of these feature extraction layers are fused with the features generated by the cross-scale compact feature fusion CCFF module and the semi-supervised diffusion data generation module, while gradually reducing the resolution of the features in each layer to adapt to the multi-scale object detection task.
[0067] Step 3. The specific process of pseudo-label generation and optimization is as follows:
[0068] In this embodiment, the semi-supervised diffusion data generation module generates high-quality pseudo-images from pure noise through the inverse diffusion process. During the generation process, the teacher-student architecture SSDDGL as shown in Figure 4 is used. This architecture consists of a teacher encoder, a student encoder, a generative diffusion loss module, and a discriminative loss module, and guides the student network to optimize the generated features through the stable supervision signal of the teacher network.
[0069] The generated pseudo-images are used for the training of the detection network. The generated pseudo-labels are screened for high-confidence regions through the uncertainty minimization query module to ensure the reliability of the pseudo-images. The specific process is as follows:
[0070] On the generated pseudo-images, the SAGE-Weed model predicts the unlabeled data, calculates the confidence of each target region, and filters according to a preset confidence threshold, only retaining the high-confidence regions as pseudo-labels. This process is achieved by calculating the prediction uncertainty of each target region and selecting the target with the lowest uncertainty as the pseudo-label to improve the reliability of the pseudo-labels.
[0071] The uncertainty is calculated according to the standard deviation σ:
[0072]
[0073] where represents the predicted class probability for each iteration, represents the predicted average value;
[0074] Next, these pseudo-labels are trained together with the original labeled data to form a joint optimization process. By minimizing two losses simultaneously through the semi-diffusion loss function designed in the present invention, including the supervised loss of the labeled data and the pseudo-supervised loss of the unlabeled data, the collaborative optimization of the labeled data and the pseudo-label data is achieved. The labeled data is used to calculate the supervised loss, while the pseudo-label data is used to calculate the pseudo-supervised loss. The feature learning of the detection network and the pseudo-image generation of the semi-supervised diffusion data generation module complement each other, and the semi-supervised diffusion data generation module optimizes the pseudo-label generation strategy to improve the detection performance.
[0075] In each iteration, the pseudo-labels are continuously updated and optimized to improve their accuracy and reliability, thereby gradually enhancing the model's performance on unlabeled data while ensuring the stability and robustness of training.
[0076] Step Four: Feature Fusion and Object Detection. The specific process is as follows:
[0077] The CCFF module in the detection network first encodes the features generated by the semi-supervised diffusion data generation module through position embedding to ensure that the generated features contain spatial information. Subsequently, these generated features are weighted and fused with the multi-scale features of the feature extraction network (efficient hybrid encoder) at each layer (on multiple scales). The weighted fusion process uses residual connections to add the generated features and the original image features to retain the original feature information while introducing supplementary information of the generated features. In addition, a dynamic generative attention mechanism is used in the weighted fusion process to allocate weights to different features, thereby optimizing the information interaction between the generated features and the real features. Finally, the fused feature results are sent into the decoder for further processing to restore the resolution, and class classification and bounding box regression are performed through the object detection head.
[0078] Among them, the calculation method of the weighted fusion weights is as follows:
[0079]
[0080] Among them, F gen and respectively represent the generated features and the image features, and W fusion is the fusion weight.
[0081] Step Five: Joint Training and Optimization. The specific process is as follows:
[0082] The pseudo-labels generated by the semi-supervised diffusion data generation module in the joint training are input together with the labeled data into the detection network for feature extraction and object detection to achieve the collaborative optimization of the labeled data and the pseudo-label data. Among them, the labeled data is used to calculate the supervised loss, while the pseudo-label data is used to calculate the pseudo-supervised loss; the feature learning of the detection network complements the image generation of the semi-supervised diffusion data generation module. The detection network optimizes the network parameters by calculating the supervised loss of the labeled data and the pseudo-supervised loss of the pseudo-label data; at the same time, the pseudo-supervised loss of the pseudo-label data is fed back to the semi-supervised diffusion data generation module to adjust and optimize the pseudo-label generation strategy, thereby improving the quality and credibility of the generated pseudo-images to improve the detection performance.
[0083] Final Output: The model finally outputs the detection results, including the classification information of the target class and the accurately located bounding box, and further improves the localization accuracy through the IoU optimization loss function.
[0084] Such asFigure 6 As shown, when applying the method provided by the present invention and other methods for weed detection, throughout the process, the labeled data can provide a stable supervision signal, and the pseudo-labeled data expands the distribution range of the training samples. The two cooperate and optimize to ensure that the model can still achieve a high-precision detection effect under the condition of limited labeled data.
[0085] Therefore, the present invention adopts the above-mentioned weed detection method based on a semi-supervised diffusion generation network, applies a semi-supervised diffusion data generation module, introduces a generation attention mechanism and a semi-diffusion loss function, combines the characteristics of supervised learning and generative learning, effectively utilizes limited labeled data in weed detection, obtains robustness and detection accuracy beyond existing models, solves the problem of over-reliance on labeled data in traditional methods, and provides a new direction and basis for the research and application of complex target detection tasks in intelligent agriculture.
[0086] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A weed detection method based on a semi-supervised diffusion generation network, characterized in that: The following steps are involved: Step 1: Data enhancement; Step 2: Data input and multi-scale feature extraction; Step 3: Pseudo label generation and optimization; Step 4: Feature fusion and target detection; Step 5: Joint training and optimization.
2. The weed detection method based on a semi-supervised diffusion generative network according to claim 1, characterized in that: In step 1, the image data is preprocessed by three data enhancement methods: shear blending, blending enhancement, and random occlusion; Cut blending is pasting a local rectangular area from one image into another image; Mixed enhancement mixes two images and their corresponding labels in a specific ratio; Random occlusion simulates occlusion scenes by randomly occluding a rectangular area in the image.
3. The weed detection method based on a semi-supervised diffusion generative network according to claim 2, characterized in that: The specific process of step 2 is: The SAGE-Weed model first accepts input image data. After the input image is processed by the backbone network, it is extracted through the multi-scale feature extraction layer of the efficient hybrid encoder. Among them, the three feature extraction layers S3, S4, and S5 in the efficient hybrid encoder are used to extract low-level, medium-level, and high-level features respectively; The efficient hybrid encoder uses the adaptive integrated feature interaction AIFI module. The AIFI module enhances the information interaction between different scales through a dynamic generative attention mechanism. The calculation method of the AIFI module is as follows: F AIFI =softmax(Q·K T )·V; Among them, Q, K, and V represent the query matrix, key matrix, and value matrix respectively, and F AIFI Represents the output features, and the softmax function is used to normalize the attention weights; The SiLU activation function and batch normalization (BN) processing are applied to each extracted feature layer. The output of the feature extraction layer is fused with the features generated by the semi-supervised diffusion data generation module through the cross-scale compact feature fusion (CCFF) module, while the resolution of each layer of features is gradually reduced.
4. The weed detection method based on a semi-supervised diffusion generative network according to claim 3, characterized in that: The specific process of step three is: The semi-supervised diffusion data generation module is used to generate pseudo images from noise through the inverse diffusion process. On the generated pseudo images, the SAGE-Weed model predicts the unlabeled data, calculates the confidence of each target area, and screens according to the preset confidence threshold. Only the areas with high confidence are retained as pseudo labels. The prediction uncertainty of each target area is calculated, and the target with the lowest uncertainty is selected as the pseudo label. The uncertainty is calculated according to the standard deviation σ: in, represents the predicted class probability at each iteration, represents the average value of the prediction; The pseudo labels are trained together with the original annotated data to form a joint optimization process, which simultaneously minimizes two losses, including the supervision loss of the annotated data and the pseudo supervision loss of the unlabeled data, through the semi-diffusion loss function.
5. The weed detection method based on a semi-supervised diffusion generative network according to claim 4, characterized in that: The specific process of step 4 is as follows: The cross-scale compact feature fusion CCFF module first encodes the features generated by the semi-supervised diffusion data generation module through position embedding, and performs weighted fusion of the generated features and the multi-scale features of the feature extraction network at each layer. The weighted fusion process uses residual connections to add the generated features to the original image features to retain the original feature information while introducing the supplementary information of the generated features. In addition, the weights of different features are assigned through a dynamic generative attention mechanism during the weighted fusion process, and the fused feature results are finally sent to the decoder for further processing to restore the resolution, and the category classification and bounding box regression are performed through the object detection head. Among them, the calculation method of weighted fusion weight is as follows: Among them, F gen and Represent the generated features and image features respectively, W fusion is the fusion weight.
6. The weed detection method based on a semi-supervised diffusion generative network according to claim 5, characterized in that: The specific process of step five is: Joint training: The pseudo labels generated by the semi-supervised diffusion data generation module are input into the detection network together with the annotated data for feature extraction and object detection; The detection network optimizes the network parameters by calculating the supervision loss of the labeled data and the pseudo-supervision loss of the pseudo-labeled data; The pseudo-supervised loss of the pseudo-label data will be fed back to the semi-supervised diffusion data generation module to adjust and optimize the pseudo-label generation strategy.