Field weed growing point detection model and method based on key points

Through the detection model based on key points, the detection problem of field weed growth points under natural environment occlusion and light changes is solved, and high-precision and real-time positioning of weed growth points is achieved.

CN120355983APending Publication Date: 2025-07-22SHANDONG AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510401078.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing field weed growth point detection method is not ideal in natural environments, especially in the case of shading and light changes, it is difficult to accurately locate weed growth points.

Method used

Key point-based detection model is adopted, including backbone network Backbone, neck network Neck and detection head head, to achieve accurate detection of weed categories, locations and confidence through multi-scale fusion and feature map transformation.

Benefits of technology

Maintain high detection accuracy in natural field environments, realize lightweight model, improve real-time detection capabilities, and accurately locate weed growth points.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355983A_ABST
    Figure CN120355983A_ABST
Patent Text Reader

Abstract

The invention discloses a field weed growing point detection model and method based on key points. Backbone is used for receiving a field weed image to realize learning of multi-scale weed image features and outputting feature maps of different stages; the Neck uses multi-scale fusion to fuse different stages of feature maps output by the Backbone, and the performance ability of weed features is enhanced; and the Head converts the feature map output by the Neck into specific information required by weed detection, wherein the specific information comprises category, position and confidence information of weeds. The constructed detection model can maintain high detection precision in a natural field environment, the model is light, the real-time detection capability is improved, and the growth points of weeds can be accurately identified in real time in the field environment. By adopting the key point detection method, the specific position coordinates of the growth points of the weeds can be directly output, and the problem that the traditional weed detection method can only detect the whole or part of the weeds and is difficult to determine the growth points is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of visual detection technology, and specifically relates to a detection model and method for the growth points of field weeds based on key points. Background Art

[0002] In agricultural production, weeds are one of the most serious threat factors, causing varying degrees of losses to crop production every year. Weed control has always been a long-standing problem in the agricultural field. In order to achieve the sustainable development of agriculture, the weed control methods are gradually developing towards precision, environmental protection and intelligence. More and more considerations are given to site-specific weed management in agriculture, and precision weeding robots have broad development prospects.

[0003] With the gradual application of deep learning in agriculture, object detection technology is of great significance for the end effector of the robot to achieve precision weeding. The key to the precision weed suppression technology is to accurately identify the growth points of weeds, that is, the apical meristems of weeds, so that the weeding robot can handle weeds more precisely. The cell division activity at the growth point is vigorous. Precise strikes on the growth point part of weeds can effectively inhibit the growth rate of weeds or directly prevent their growth, reduce their competition for key resources, and enable crop seedlings to absorb nutrients and water more fully. At present, most studies identify weeds by detection boxes or segmentation methods, which can only identify the whole or part of weeds and cannot determine the position of the weed growth points. On the other hand, there are problems such as uneven weed distribution, occlusion and light changes in the natural field environment, and the detection of weed growth points has certain challenges. Therefore, the research on the detection method of field weed growth points is of great significance for precision weeding.

[0004] Key point detection was initially applied to human pose estimation and has gradually been applied in the agricultural field in recent years. For example, key point detection can be used to achieve fruit branch pruning during fruit picking; the position of the tea picking point can be located according to the information of key points to achieve key point detection and picking positioning of tea buds in complex environments; the key point detection method can locate to a point, providing a feasible solution for the detection of weed growth points. As a model with better performance in the YOLO series, YOLOv8 has excellent performance in terms of accuracy and speed and has strong advantages in detection and deployment. However, in the natural environment, facing the problems of dense weeds, occlusion and light changes, the original YOLOv8-Pose has unsatisfactory detection effects for the growth points of field weeds. Summary of the Invention

[0005] In order to solve the above technical problems, this application proposes the following technical solutions:

[0006] In a first aspect, an embodiment of the present application provides a key-point-based field weed growth point detection model, which is characterized by including: a backbone network Backbone, a neck network Neck, and a detection head Head. The Backbone is used to receive field weed images to realize the learning of multi-scale weed image features and output feature maps at different stages; the Neck uses multi-scale fusion to fuse the feature maps at different stages output by the Backbone to enhance the performance ability of weed features; the Head converts the feature maps output by the Neck into specific information required for weed detection, including the category, location, and confidence information of the weeds.

[0007] In a possible implementation manner, the Backbone includes: a plurality of CBS modules, a C2f_RVB module, and a C2f_DWR module. The input of each C2f_RVB module and C2f_DWR module is cascaded with a CBS module; wherein the first CBS module receives the field weed image, and the output information is input into the first C2f_RVB module through the second CBS module. The output information of the first C2f_RVB module is input into the second C2f_RVB module through the third CBS module. The feature maps output by the second C2f_RVB module are respectively transmitted to the third CBS module and the Neck; the feature maps output by the third CBS module through the second C2f_RVB module are transmitted to the first C2f_DWR module. The output information of the first C2f_DWR module is respectively transmitted to the fourth CBS module and the Neck. The information output through the fourth CBS module is transmitted into the second C2f_DWR module. The information output by the second C2f_DWR module is transmitted to the Neck again through the spatial pyramid pooling layer SPPF.

[0008] In a possible implementation manner, the C2f_DWR module includes a CBS module and a DWR module. The C2f_DWR module uses residual features for learning to realize the extraction of weed and its growth point features; the C2f_RVB module includes a CBS module and an RVB module. The C2f_RVB module reduces structural parameters through structural reparameterization; the SPPF includes a CBS module and a max pooling layer Max Pool. The SPPF converts feature maps of any size into fixed-size feature vectors to fuse local features and global features.

[0009] In a possible implementation manner, the CBS module includes a convolution Conv, a batch normalization layer BN, and a SiLU activation function layer. The Conv extracts low-level spatial features of the image, accelerates the training convergence and stabilizes the gradient through BN processing, and then introduces non-linear expression through the SiLU activation function layer.

[0010] In a possible implementation, the Neck includes a top-down path FPN and a bottom-up path PAN. The feature maps output by the second C2f_RVB module, the first C2f_DWR module, and the SPPF are first processed by the FPN to transfer high-level semantics to low levels and enhance the detection of small target weeds. The output after the feature fusion by the FPN is downsampled by the PAN to transfer low-level details to high levels and improve the positioning accuracy of the weed growth points.

[0011] In a possible implementation, the FPN includes a first upsampling UpSample, a third C2f_RVB module, and a second UpSample. The first UpSample receives the feature map output by the SPPF, enlarges the low-resolution weed feature map to a high resolution, and matches the size of the output feature map of the Backbone for feature fusion. The feature map output by the first upsampling UpSample and the feature map output by the first C2f_RVB module are fused and then output to the third C2f_RVB module for optimization to reduce feature attenuation in the deep network and enhance the ability to retain details in small target detection. The optimized output information is respectively output to the PAN and the second UpSample. The second UpSample upsamples the output optimized by the third C2f_RVB module again and performs feature fusion with the output feature map of the second C2f_RVB module and then outputs to the PAN.

[0012] In a possible implementation, the PAN includes a fourth C2f_RVB module. The fourth C2f_RVB module receives the feature map from the FPN, and after passing through the fifth CBS module, it is fused with the output optimized by the third C2f_RVB module to achieve downsampling and reduce the high-resolution feature map to a low resolution. The feature map obtained by downsampling is processed by the fourth C2f_RVB module to enhance the multi-scale perception ability, and then after passing through the sixth CBS module, it is fused with the feature map output by the SPPF to achieve the second downsampling, and finally processed by the fifth C2f_RVB module. The fourth C2f_RVB module, the fifth C2f_RVB module, and the sixth C2f_RVB module in the PAN respectively output weed feature maps of different scales to the Head.

[0013] In a possible implementation, the Head includes three detection heads Detect with different scales, and different Detects receive weed feature maps with different scales; each Detect includes three branches, and each branch includes a CBS module, a CBS_SEAM module, and a Conv connected in sequence. The CBS_SEAM module uses the relationship between feature maps to supplement the occluded features of weeds, so as to emphasize the weed areas in the image and make up for the occluded features of weeds in the original features. The different branches of the Detect output differently, which are the class prediction cls loss, the bounding box regression Bbox loss, and the key point prediction poseloss respectively.

[0014] In a possible implementation, each key point contains the corresponding position and confidence: x, y, conf. There are a total of N×3 elements for N key points associated with an anchor point. For the pose detection branch with n key points, the prediction vector P at each anchor point v is defined as:

[0015]

[0016] In a second aspect, an embodiment of the present application provides a method for detecting the growth points of field weeds based on key points. Based on the model described in any possible implementation of the first aspect, it includes:

[0017] Collecting the image data of weeds in real time;

[0018] Inputting the weed image data into the model for detecting the growth points of field weeds based on key points to detect the growth points of weeds, and outputting the label information, confidence, and specific positions of the growth points of various weeds;

[0019] Performing confidence screening according to a preset threshold to obtain the final detection result of the growth points of weeds.

[0020] In the embodiment of the present application, the constructed detection model can maintain a high detection accuracy in a natural field environment with occlusion, density, and light changes. At the same time, it realizes the lightweight of the model, improves the real-time detection ability, and ensures that the growth points of weeds can be accurately identified in real time in the field environment. By using the method of key point detection, the specific position coordinates of the growth points of weeds can be directly output, solving the problem that traditional weed detection methods can only detect the whole or part of weeds and it is difficult to determine the growth points. Description of the Drawings

[0021] Figure 1 It is a schematic structural diagram of a model for detecting the growth points of field weeds based on key points provided by an embodiment of the present application;

[0022] Figure 2Schematic diagram of the C2f_DWR module provided by the embodiment of the present application;

[0023] Figure 3 Schematic diagram of the DWR module provided by the embodiment of the present application;

[0024] Figure 4 Schematic diagram of the C2f_RVB module provided by the embodiment of the present application;

[0025] Figure 5 Schematic diagram of the RVB module structure provided by the embodiment of the present application;

[0026] Figure 6 Schematic diagram of the SPPF structure provided by the embodiment of the present application;

[0027] Figure 7 Schematic diagram of the CBS module structure provided by the embodiment of the present application;

[0028] Figure 8 Schematic diagram of the Detect structure provided by the embodiment of the present application;

[0029] Figure 9 Schematic diagram of the CBS_SEAM structure provided by the embodiment of the present application;

[0030] Figure 10 Schematic diagram of the flow of a method for detecting the growth points of field weeds based on key points provided by the embodiment of the present application;

[0031] Figure 11 Original field image provided by the embodiment of the present application;

[0032] Figure 12 Comparison chart of detection results using different models provided by the embodiment of the present application. Detailed implementation manners

[0033] The following elaborates on this solution in combination with the accompanying drawings and specific implementation manners.

[0034] Refer to Figure 1 The key-point-based field weed growth point detection model provided in this embodiment includes: a backbone network Backbone, a neck network Neck, and a detection head Head. The Backbone is used to receive field weed images to implement the learning of multi-scale weed image features and output feature maps at different stages; the Neck uses multi-scale fusion to fuse the feature maps at different stages output by the Backbone to enhance the representation ability of weed features; the Head converts the feature maps output by the Neck into specific information required for weed detection, including the category, location, and confidence information of the weeds.

[0035] In this embodiment, the Backbone includes: a plurality of CBS modules, a C2f_RVB module, and a C2f_DWR module. The input of each C2f_RVB module and C2f_DWR module is cascaded with a CBS module. Among them, the first CBS module receives the field weed image, and the output information is input into the first C2f_RVB module through the second CBS module. The information output by the first C2f_RVB module is input into the second C2f_RVB module through the third CBS module. The feature maps output by the second C2f_RVB module are respectively transmitted to the third CBS module and the Neck. The feature map output by the third CBS module through the second C2f_RVB module is transmitted to the first C2f_DWR module. The information output by the first C2f_DWR module is respectively transmitted to the fourth CBS module and the Neck. The information output by the fourth CBS module is transmitted into the second C2f_DWR module. The information output by the second C2f_DWR module is transmitted to the Neck again through the Spatial Pyramid Pooling Layer (SPPF).

[0036] See Figure 2 , the C2f_DWR module includes a CBS module and a DWR module. The C2f_DWR module uses residual features for learning to extract the features of weeds and their growth points. When passing through the C2f_DWR module, the main branch extracts the deep features of weeds through multiple DWR modules. The structure of the DWR module is as Figure 3 shown. The feature extraction process of weeds and their growth points is divided into two steps. The first step is regional residualization (RR). The initial features of weeds are extracted through a 3×3 ordinary convolution, and the regional features are activated by combining the Batch Normalization (BN) layer and the activation function ReLU. The second step is semantic residualization (SR). First, the regional features are divided into several groups, and then different groups pass through extended convolutions at different rates. The first step can learn the required features according to the receptive field size of the second step to achieve reverse matching of the receptive field and effectively obtain the multi-scale context information of weeds to meet the detection requirements under dense growth and light changes. The sub-branch also directly transmits the original features and then fuses (Concat) the deep features of the main branch to enhance the multi-scale perception ability.

[0037] See Figure 4 , the C2f_RVB module includes a CBS module and an RVB module. The C2f_RVB module reduces the structural parameters through structural reparameterization. Specifically, when passing through the C2f_RVB module, the main branch extracts the deep features of weeds and their growth points through multiple RVB modules. The structure of the RVB module is as Figure 5As shown, a 3×3 depthwise convolution (DW) and a squeeze-and-excitation layer (SE) are placed in the first and second layers for the fusion of spatial information (i.e., token mixer). The 1×1 expansion layer in the third layer and the 1×1 projection layer in the fourth layer achieve the interaction between channels (i.e., channel mixer). Structural reparameterization is used to separate the token mixer and the channel mixer, thereby eliminating some related computational and memory costs during the inference process, realizing the saving of computing resources, and meeting the real-time requirement of field operations. The sub-branch directly passes the features, retaining their original information.

[0038] See Figure 6 , the SPPF includes a CBS module and a max pooling layer Max Pool. The SPPF converts a feature map of any size into a fixed-size feature vector, fusing local features and global features. The input to the SPPF structure captures the multi-scale context information of weeds and their growth points through cascaded pooling. First, it undergoes Conv(1×1) for channel compression to reduce the number of channels. Then, the max pooling operation is repeated three times to achieve an equivalent enlarged receptive field. The original features and the multi-scale pooling results are fused to complete the output of the backbone network.

[0039] See Figure 7 , the CBS module includes a convolution Conv, a batch normalization layer BN, and a SiLU activation function layer. The Conv extracts the low-level spatial features of the image, and after being processed by BN, it accelerates the training convergence and stabilizes the gradient. Then, through the SiLU activation function layer, non-linear expression is introduced.

[0040] Further see Figure 1 , in this embodiment, the Neck includes a top-down path FPN and a bottom-up path PAN. The feature maps output by the second C2f_RVB module, the first C2f_DWR module, and the SPPF are first processed by the FPN to transfer high-level semantics to the low level, enhancing the detection of small target weeds; the output after the FPN feature fusion undergoes downsampling by the PAN to transfer low-level details to the high level, improving the localization accuracy of weed growth points.

[0041] In this embodiment, the Neck uses multi-scale fusion to fuse the feature maps at different stages output by the backbone network, enhancing the representation ability of weed features. The structure follows the feature FPN and PAN, effectively integrating the top-down and bottom-up information flows in the network and enhancing the detection performance.

[0042] In this embodiment, the FPN includes a first upsampling (UpSample), a third C2f_RVB module, and a second UpSample. The first UpSample receives the feature map output by the SPPF, enlarges the low-resolution weed feature map to a high resolution, and matches the size of the output feature map of the Backbone for feature fusion. The feature map output by the first upsampling (UpSample) is fused with the feature map output by the first C2f_RVB module and then output to the third C2f_RVB module for optimization to reduce feature attenuation in the deep network and enhance the ability to retain details for small target detection. The optimized output information is respectively output to the PAN and the second UpSample. The second UpSample upsamples the output optimized by the third C2f_RVB module again, performs feature fusion with the output feature map of the second C2f_RVB module, and outputs it to the PAN.

[0043] The PAN includes a fourth C2f_RVB module. The fourth C2f_RVB module receives the feature map from the FPN, and after passing through the fifth CBS module, it is fused with the output optimized by the third C2f_RVB module to achieve downsampling, reducing the high-resolution feature map to a low resolution. The feature map obtained by downsampling is processed by the fourth C2f_RVB module to enhance the multi-scale perception ability, and then after passing through the sixth CBS module, it is fused with the feature map output by the SPPF to achieve the second downsampling, and finally processed by the fifth C2f_RVB module. The fourth C2f_RVB module, the fifth C2f_RVB module, and the sixth C2f_RVB module in the PAN respectively output weed feature maps of different scales to the Head.

[0044] Feature maps of different scales output by the backbone network are input into the Neck. The feature maps are first processed by the FPN to transfer high-level semantics to the low level and enhance the detection of small target weeds. The FPN upsamples (UpSample) the feature map (20×20) output by the SPPF in the backbone network, enlarges the low-resolution weed feature map to a high resolution, and makes its size match the output feature map (40×40) of the fourth layer (P4) in the backbone network for feature fusion. The fused feature is optimized by the third C2f_RVB module to reduce feature attenuation in the deep network and enhance the ability to retain details for small target detection. The optimized output is upsampled again and fused with the output feature map (80×80) of the third layer (P3) of the backbone network.

[0045] The output after feature fusion is downsampled by PAN, which transfers low-level details to the high level to improve the positioning accuracy of weed growth points. The PAN obtains a weed feature map (80×80) after passing the input through a convolutional operation (CBS), and fuses (Concat) it with the middle-level features (40×40) of the FPN to achieve downsampling and reduce the high-resolution feature map to a low resolution. The middle-level fused features are processed by the fourth C2f_RVB module to enhance the multi-scale perception ability, and then a feature map (40×40) is obtained through a convolutional operation (CBS) and fused (Concat) with the high-level features (20×20) of the FPN to achieve the second downsampling. The high-level fused features are processed by the fifth C2f_RVB again, and finally three enhanced weed feature maps are output.

[0046] In this embodiment, the Head includes three detection heads Detect with different scales, and different Detects receive weed feature maps with different scales. Refer to Figure 8 , each Detect includes three branches, and each branch includes a CBS module, a CBS_SEAM module, and a Conv connected in sequence. The CBS_SEAM module uses the relationship between feature maps to supplement the features of occluded weeds, achieving the purposes of emphasizing the weed area in the image and compensating for the occluded features of weeds in the original features. The outputs of different branches of the Detect are different, which are the class prediction cls loss, the bounding box regression Bbox loss, and the key point prediction poseloss respectively.

[0047] The detection head part introduces the CBS_SEAM to design the three branches in the Detect structure, replacing the second-layer convolution Conv in the CBS structure with the CBS_SEAM. The CBS_SEAM uses the relationship between feature maps to supplement the features of occluded weeds, and can achieve the two purposes of emphasizing the weed area in the image and compensating for the occluded features of weeds in the original features. The CBS_SEAM structure is as Figure 9 shown. The input is first processed by a channel and spatial mixing module (CSMM), and after the output, it is subjected to average pooling (Average Pooling) and then channel expansion (Channel exp). Finally, the output of the SEAM module is used as the attention to multiply with the original features to compensate for the occluded features of weeds in the original features, enabling the model to more effectively process the recognition of weed growth points in the case of occlusion.

[0048] In this embodiment, each key point contains the corresponding position and credibility: x, y, conf. There are a total of N×3 elements for N key points associated with an anchor point. For the pose detection branch with n key points, the prediction vector P v is defined as:

[0049]

[0050] Based on the keypoint-based detection model for the growth points of field weeds provided in the above embodiments, this embodiment also provides a method for detecting the growth points of field weeds based on keypoints.

[0051] Refer to Figure 10 , the method for detecting the growth points of field weeds based on keypoints in this embodiment includes:

[0052] S101, collect the image data of weeds in real time.

[0053] Refer to Figure 11 , which is the original image of the field collected. Before detecting the growth points of weeds using the keypoint-based detection model for the growth points of field weeds in the above embodiments, this embodiment needs to train the model.

[0054] Specifically, it includes: collecting weed images in the natural field environment. The image data contains weed images in different scenarios such as a wide variety of weed species, uneven weed distribution, dense weed growth, mutual occlusion, and light changes, as Figure 1 shown. The image is taken 60 cm perpendicular to the ground, and the shooting angle is vertically downward; the image is divided into a training set, a validation set, and a test set according to the ratio of 7:2:1.

[0055] Manually annotate the weed image dataset using the image annotation tool labelme. When annotating, the growth point is used as the keypoint, and the minimum bounding rectangle of the weed is used as the detection box. The annotated file is saved in JSON and then converted into a TXT file. Data augmentation is performed on the annotated images, and data is augmented through operations such as translation, horizontal flipping, fogging, and changing brightness, and the corresponding annotation files are generated synchronously. After data augmentation, the dataset has a total of 3380 images.

[0056] In this embodiment, the experimental platform: Lenovo legion desktop computer, the main hardware configuration is: Intel(R) Core(TM) i7 10700k CPU @ 3.8GHz, 16GB of memory, NVIDIA GeForce RTX 2080SUPER GPU. The software system environment is Windows 10 64-bit operating system, CUDA 10.0 version, CUDNN 7.1 version, Python 3.9.10 version, Pytorch 1.2.0 version. Hyperparameter settings: Batch is 16, Epoch is 500, the initial learning rate is 0.01, the momentum parameter is 0.937, and the stochastic gradient descent strategy SGD is used to optimize the network parameters during the training phase.

[0057] In this embodiment, in order to verify the superiority of the trained model, as shown in Table 1, the model uses the following control schemes: the original YOLOv8-Pose model A, model B with DWRM introduced into the backbone network, model C with RVB introduced into the backbone network and the neck network, model D with SEAM introduced into the detection head, models E, F, and G with two modules introduced respectively, and the SRD-YOLO model constructed completely proposed in this embodiment. Through comparative experiments, it is verified that the overall performance of the model in this embodiment is the best.

[0058] Table 1 Ablation Experiment Results

[0059]

[0060] As shown by the results of the ablation experiment in Table 1, each module is effective. Specifically, the mAP of the model kpt reaches 96.5%, an increase of 4.5%; the number of parameters is reduced by 8.7M compared with the original model, and the frames per second (FPS) is increased by 21 (frames / second). The comparison of the detection results is as Figure 12 shown, and the detection performance of the SRD-YOLO model is significantly better than that of the original model.

[0061] S102: Input the weed image data into the field weed growth point detection model based on key points to detect the weed growth points, and output the label information, confidence level, and specific positions of various weeds.

[0062] S103: Perform confidence level screening according to the preset threshold to obtain the final weed growth point detection result.

[0063] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent the situations of A existing alone, A and B existing simultaneously, and B existing alone. Where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the front and rear associated objects. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, and c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0064] The above is only the specific implementation manner of the present application. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. The protection scope of the present application shall be subject to the protection scope of the claimed rights.

Claims

1. A key-point-based detection model for the growth points of field weeds, characterized in that, It includes: Backbone network, Neck network, and Detection Head. The Backbone is used to receive field weed images to learn multi-scale weed image features and output feature maps at different stages; the Neck uses multi-scale fusion to fuse the feature maps at different stages output by the Backbone to enhance the representation ability of weed features; the Detection Head converts the feature maps output by the Neck into specific information required for weed detection, including weed category, location, and confidence information.

2. The key-point-based field weed growth point detection model according to claim 1, wherein The Backbone includes: a plurality of CBS modules, C2f_RVB modules, and C2f_DWR modules. Each C2f_RVB module and C2f_DWR module has a CBS module cascaded at its input; among them, the first CBS module receives field weed images, and the output information is input into the first C2f_RVB module through the second CBS module. The information output by the first C2f_RVB module is input into the second C2f_RVB module through the third CBS module. The feature maps output by the second C2f_RVB module are respectively transmitted to the third CBS module and the Neck; the feature maps output by the third CBS module through the second C2f_RVB module are transmitted to the first C2f_DWR module. The information output by the first C2f_DWR module is respectively transmitted to the fourth CBS module and the Neck. The information output through the fourth CBS module is transmitted into the second C2f_DWR module. The information output by the second C2f_DWR module is transmitted to the Neck again through the Spatial Pyramid Pooling Layer SPPF.

3. The key-point-based field weed growth point detection model according to claim 2, wherein, The C2f_DWR module includes a CBS module and a DWR module. The C2f_DWR module uses residual features for learning to extract weed and its growth point features; the C2f_RVB module includes a CBS module and an RVB module. The C2f_RVB module reduces structural parameters through structural reparameterization; the SPPF includes a CBS module and a Max Pooling layer. The SPPF converts feature maps of any size into fixed-size feature vectors to fuse local features and global features.

4. The key-point-based field weed growth point detection model according to claim 2 or 3, characterized in that, The CBS module includes a convolution Conv, a Batch Normalization layer BN, and a SiLU activation function layer. The Conv extracts low-level spatial features of the image, accelerates training convergence and stabilizes gradients through BN processing, and then introduces non-linear expressions through the SiLU activation function layer.

5. The key-point-based field weed growth point detection model according to claim 2, wherein The Neck includes a top-down path FPN and a bottom-up path PAN. The feature maps output by the second C2f_RVB module, the first C2f_DWR module, and the SPPF are first processed by the FPN to transfer high-level semantics to the low level to enhance the detection of small target weeds; the output after feature fusion by the FPN is downsampled by the PAN to transfer low-level details to the high level to improve the positioning accuracy of weed growth points.

6. The detection model of the growth point of field weeds based on key points according to claim 5, characterized in that, The FPN includes a first upsampling (UpSample), a third C2f_RVB module, and a second UpSample. The first UpSample receives the feature map output by the SPPF, enlarges the low-resolution weed feature map to a high resolution to match the size of the output feature map of the Backbone for feature fusion. The feature map output by the first UpSample is fused with the feature map output by the first C2f_RVB module and then output to the third C2f_RVB module for optimization, reducing feature attenuation in the deep network and enhancing the ability to retain details for small target detection. The optimized output information is respectively output to the PAN and the second UpSample. The second UpSample upsamples the output optimized by the third C2f_RVB module again, performs feature fusion with the output feature map of the second C2f_RVB module, and outputs to the PAN.

7. The key point-based field weed growth point detection model according to claim 5, wherein The PAN includes a fourth C2f_RVB module. The fourth C2f_RVB module receives the feature map from the FPN, and after passing through the fifth CBS module, it is fused with the output optimized by the third C2f_RVB module to achieve downsampling, reducing the high-resolution feature map to a low resolution. The feature map obtained by downsampling is processed by the fourth C2f_RVB module to enhance the multi-scale perception ability, and then after passing through the sixth CBS module, it is fused with the feature map output by the SPPF to achieve the second downsampling, and finally processed by the fifth C2f_RVB module. The fourth C2f_RVB module, the fifth C2f_RVB module, and the sixth C2f_RVB module in the PAN respectively output weed feature maps of different scales to the Head.

8. The key-point-based field weed growth point detection model according to claim 7, characterized in that The Head includes three detection heads (Detect) of different scales, and different Detects receive weed feature maps of different scales. Each Detect includes three branches, and each branch includes a CBS module, a CBS_SEAM module, and a Conv connected in sequence. The CBS_SEAM module uses the relationship between feature maps to supplement the occluded features of weeds, emphasizing the weed areas in the image and making up for the occluded features of weeds in the original features. The different branches of the Detect output differently, namely class prediction (cls loss), bounding box regression (Bbox loss), and key point prediction (pose loss).

9. The key-point-based detection model for the growth points of weeds in the field according to claim 8, characterized in that, Each key point contains the corresponding position and confidence: x, y, conf. There are a total of N×3 elements for N key points associated with an anchor point. For the pose detection branch with n key points, the prediction vector P at each anchor point v is defined as:

10. A method for detecting the growth points of field weeds based on key points, characterized in that, The model according to any one of claims 1-9 includes: Real-time collecting image data of weeds; Inputting the weed image data into the field weed growth point detection model based on key points for weed growth point detection, and outputting label information, confidence, and the specific position of the weed growth point of various weeds; Performing confidence screening according to a preset threshold to obtain the final weed growth point detection result.