Field cotton terminal bud high-speed dynamic identification and positioning method and system

By using a lightweight bud identification model and bud tracking algorithm, combined with binocular camera calibration for 3D coordinate transformation, the problems of high-precision, high-real-time identification and nonlinear motion adaptability of buds in complex cotton field backgrounds are solved, improving the efficiency and accuracy of topping operations.

CN122049692APending Publication Date: 2026-05-15JIANGSU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU UNIV
Filing Date
2026-03-24
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve high-precision, real-time identification and adaptability to nonlinear motion in cotton bud positioning within complex cotton field environments, resulting in low topping efficiency and a high risk of missed topping.

Method used

A lightweight bud recognition model is adopted, which combines a one-stage IOU (Intersection over Union) tracking algorithm with a two-stage 9-state vector acceleration Kalman filter. Combined with binocular camera calibration for 3D coordinate transformation, a lightweight bud recognition, tracking and localization module is constructed to achieve accurate and efficient perception of buds.

Benefits of technology

It significantly improves the accuracy of bud identification in complex backgrounds and the tracking robustness in dynamic scenarios, reduces the missed bud rate, and ensures the accuracy and reliability of bud removal operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122049692A_ABST
    Figure CN122049692A_ABST
Patent Text Reader

Abstract

The invention provides a field cotton terminal bud high-speed dynamic identification and positioning method and system, and the method comprises the following steps: obtaining a real-time image frame of a field cotton terminal bud, inputting the real-time image frame into a lightweight terminal bud identification model, and outputting the detection frame information of the terminal bud; detecting frame information is input to a target tracking module, the target tracking module maintains a 9-state vector containing the position, the speed and the acceleration for each terminal bud, cross-frame target tracking is achieved through a strategy combining first-stage IOU association and second-stage Kalman filtering prediction, and a unique identification ID is distributed to each tracked terminal bud; the terminal bud center point two-dimensional image coordinates with the ID output by the tracking module are input to a positioning module, and the positioning module combines internal and external parameters of the binocular camera to calculate and obtain three-dimensional space coordinates of the terminal bud in a camera coordinate system; and outputting data including the terminal bud ID and the corresponding three-dimensional space coordinates for guiding the topping execution mechanism to work, so that the topping operation is faster and more accurate, and the missing topping rate and damage are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent agricultural machinery technology, and in particular relates to a method and system for high-speed dynamic identification and positioning of cotton buds in the field. Background Technology

[0002] Cotton, as an important economic product, has a significant impact on the economy and people's livelihoods due to its cultivation. Topping, a key agronomic measure, plays a crucial role in improving yield and quality by suppressing apical dominance and promoting the rational distribution of nutrients. Precise topping is an important guarantee for achieving high yield and efficiency in cotton. Topping helps nutrient return, improves quality, increases seed cotton yield, reduces shedding and forms effective yield, and promotes early flowering and boll formation. However, several industry pain points exist: high cost and low efficiency of manual operation; chemical topping can cause reduced boll yield and regrowth; and mechanical topping is prone to missed topping. The key to mechanical topping lies in improving the accurate identification and positioning of cotton apical buds. Therefore, the core research topic is to develop a new visual topping technology that can balance topping effectiveness, operational efficiency, and cost control.

[0003] To address this problem, scholars have conducted extensive research on methods for identifying and locating cotton buds. In terms of identification, many scholars have designed target recognition models incorporating deep learning, but these suffer from high computational costs, a poor balance between accuracy and real-time performance, and a lack of consideration for dynamic situations. Regarding localization methods, early approaches using ultrasonic sensors, laser sensors, and contact-based contouring all suffer from accuracy issues and poor anti-interference capabilities. Subsequent research has employed methods combining target recognition and tracking-localization algorithms for bud localization, improving accuracy to some extent, but still exhibiting problems such as slow inference speed, simplistic mathematical models, and low robustness in tracking-localization.

[0004] The paper "A Top Bud Recognition and Positioning System and Method Applicable to Cotton Topping Machines in the Field" (Shandong Academy of Agricultural Machinery Sciences. A Top Bud Recognition and Positioning System and Method Applicable to Cotton Topping Machines in the Field: 202410113908.7 [P]. 2024-05-31.) proposes a top bud recognition model based on an improved YOLOv5 and the DeepSORT tracking algorithm, combined with binocular vision spatial positioning, to achieve real-time tracking and positioning of the top bud in field scenarios. However, this recognition model cannot effectively solve the interference in complex cotton field backgrounds, and the DeepSORT algorithm, based on a uniform linear motion mathematical model, cannot adapt to the nonlinear motion of the topping machine during field operations.

[0005] The paper "A Method and System for Extracting Three-Dimensional Coordinates of Cotton Apical Buds Based on Key Point Detection" (Institute of Aerospace Information Research, Chinese Academy of Sciences. A Method and System for Extracting Three-Dimensional Coordinates of Cotton Apical Buds Based on Key Point Detection: 202510016982.1 [P]. 2025-10-17.) proposes a method and system for extracting three-dimensional coordinates of cotton apical buds based on key point detection, and combines ROS for coordinate extraction. However, it observes the apical bud from a frontal viewpoint, determines the top of the cotton plant through the key point skeleton, and does not combine dynamic tracking algorithms, so it cannot solve the problem of real-time identification and positioning of apical buds in dynamic operation scenarios.

[0006] Therefore, there is a lack of a high-speed dynamic apical bud identification and positioning method in the current technology that can achieve high-precision, high-real-time identification in complex cotton field backgrounds and can adapt to nonlinear motion. Summary of the Invention

[0007] To address the aforementioned technical problems, this invention provides a method and system for high-speed dynamic identification and positioning of cotton apical buds in the field. By tracking the position of cotton apical buds in real time in the cotton field, topping operations can be made faster and more accurate, reducing the rate of missed topping and avoiding damage to plants and bolls.

[0008] This invention enables real-time identification and tracking of cotton bud targets in actual cotton field scenarios, making topping operations faster and more accurate, and effectively avoiding missed topping. The positioning system of this invention includes a bud identification model, a bud tracking algorithm, and a bud positioning module. The bud identification model includes a lightweight backbone network, a CA attention mechanism, an improved hybrid pooling SPPELAN, a PSA polarized self-attention mechanism, a C3_VIT context feature fusion module, and a BiFPN-like splicing architecture. The bud tracking algorithm includes a one-stage IOU (Intersection over Union) tracking and a two-stage 9-state vector acceleration Kalman filter tracking. The bud positioning module consists of a binocular camera calibration unit and a coordinate transformation unit. The binocular camera calibration unit acquires the intrinsic and extrinsic parameters of the binocular camera, and the coordinate transformation module combines the two-dimensional coordinates of the bud center obtained by the target tracking algorithm with the camera's intrinsic and extrinsic parameters to calculate the spatial position information of the bud. This invention can achieve accurate and efficient perception of bud targets.

[0009] The present invention achieves the above-mentioned technical objectives through the following technical means.

[0010] A method for high-speed dynamic identification and positioning of cotton terminal buds in the field includes the following steps:

[0011] Step S1: Obtain real-time image frames of cotton terminal buds in the field;

[0012] Step S2: Input the image frame into the lightweight apical bud recognition model and output the detection box information of the apical bud, the detection box information including the two-dimensional image coordinates of the center point of the apical bud;

[0013] Step S3: Input the detection box information into the target tracking module. The target tracking module maintains a 9-state vector containing position, velocity and acceleration for each apical bud, and realizes cross-frame target tracking through a strategy combining one-stage IOU association and two-stage Kalman filter prediction, and assigns a unique identifier ID to each tracked apical bud.

[0014] Step S4: Input the two-dimensional image coordinates of the apical bud center point with ID output by the tracking module to the positioning module. The positioning module combines the intrinsic and extrinsic parameters of the binocular camera and performs coordinate transformation through the principle of stereo vision to calculate the three-dimensional spatial coordinates of the apical bud in the camera coordinate system.

[0015] Step S5: Output data containing the apical bud ID and corresponding three-dimensional spatial coordinates to guide the apical bud actuator operation.

[0016] In the above scheme, the lightweight apical bud recognition model is built on the YOLO architecture, and its network structure includes a lightweight backbone network, an improved hybrid pooling module SPPELAN, a neck feature fusion network, a BiFPN-like multi-path feature fusion structure, and a detection head.

[0017] The lightweight backbone network adopts the MobileNetV3_large architecture and embeds a CA coordinate attention mechanism after its shallow and deep feature map outputs.

[0018] The improved hybrid pooling module SPPELAN is connected to the end of the lightweight backbone network to perform max pooling and variance pooling in parallel, and to concatenate the pooling results to extract the intensity features and texture complexity features of the apical bud target.

[0019] The neck feature fusion network includes a PSA polarization self-attention mechanism and a C3_VIT feature fusion extraction module. The PSA polarization self-attention mechanism uses a dual-branch structure of channel and space for polarization filtering. The C3_VIT module integrates a lightweight Vision Transformer on the basis of the C3 module to establish global context association.

[0020] The BiFPN-like multi-path feature fusion structure connects the detection head and the lightweight backbone network, forming a bidirectional cross-scale information flow channel with top-down, bottom-up, and same-level jump connections.

[0021] The detection head is used to output the bounding box, confidence level, and category of the apical bud.

[0022] Furthermore, the lightweight apical bud recognition model is trained using an adaptive weighted loss function, wherein the loss function... The weighted combination of ShapeIOU loss and NWD loss is expressed as:

[0023]

[0024] Where α is the balancing weight coefficient, which is used to dynamically adjust the weights of the two losses according to the target size; NWD

[0025] Loss function; This is the shapeiou loss function.

[0026] In the above scheme, the target tracking module adopts a two-stage tracking strategy, specifically including:

[0027] Phase 1 IOU tracking: Calculate the Intersection over Union (IOU) ratio between the tracker's predicted bounding box in the previous frame and the detection bounding box in the current frame, form a cost matrix, and use the Hungarian algorithm for preliminary association matching;

[0028] The second stage, 9-state Kalman filter tracking, involves using a 9-state vector acceleration model to predict motion and perform secondary association for unmatched targets after initial association.

[0029] Furthermore, the 9-state vector is represented as:

[0030]

[0031] Where P, V, and A represent the three-dimensional position, velocity, and acceleration components of the apical bud in the Cartesian coordinate system, respectively; Let be the coordinates of the apical bud in three-dimensional space. The speed at which the apical bud moves in three directions. For the top

[0032] The state transition matrix F of the Kalman filter for uniform acceleration in three directions is:

[0033]

[0034] Where Δt is the time increment.

[0035] In the above scheme, the target tracking module also includes a distance-compensation matching mechanism:

[0036] For trackers that still do not match after the first and second stages, calculate the Euclidean distance between their predicted position and the center of the unmatched detection box in the current frame. If the distance is less than a preset threshold, then perform forced association matching.

[0037] In the above scheme, the target tracking module also includes a smooth state update mechanism:

[0038] For a successfully matched tracker, alpha-beta filtering or low-pass filtering is used to smooth the velocity and acceleration observed in the current frame in order to update the 9-state vector and suppress detection box jitter.

[0039] In the above scheme, the coordinate transformation in step S4 specifically includes:

[0040] From the detection box information output by the recognition module, extract the two-dimensional pixel coordinates of the center point of the apical bud in the left image; search for matching points in the right image that correspond to the center point of the apical bud in the left image to form a matching two-dimensional point pair between the left and right images; using the pre-calibrated internal and external parameters of the binocular camera, based on the principle of stereo vision, map the matching two-dimensional point pair between the left and right images to three-dimensional points in the camera coordinate system; output the three-dimensional coordinates of the apical bud in the camera coordinate system, including the depth estimate z-coordinate.

[0041] The above solution also includes an occlusion handling mechanism:

[0042] When the target is occluded, causing the recognition module to fail to detect it briefly, the tracking module predicts the target position using Kalman filtering and re-associates the recognition box in subsequent frames using Intersection over Union (IOU) to deal with the situation where the target disappears briefly.

[0043] A high-speed dynamic identification and positioning system for cotton buds in the field to implement the method includes an image acquisition unit, a bud identification module, a target tracking module, a positioning module, and an output module;

[0044] The image acquisition unit is used to acquire real-time image frames of the top buds of cotton plants in the field;

[0045] The apical bud recognition module is used to input the image frame into the lightweight apical bud recognition model and output the detection box information of the apical bud, the detection box information including the two-dimensional image coordinates of the center point of the apical bud;

[0046] The target tracking module is connected to the bud identification module and is used to input the detection box information to the target tracking module. The target tracking module maintains a 9-state vector containing position, velocity and acceleration for each bud, and realizes cross-frame target tracking through a strategy combining one-stage IOU association and two-stage Kalman filter prediction, and assigns a unique identifier ID to each tracked bud.

[0047] The positioning module is connected to the target tracking module and is used to input the two-dimensional image coordinates of the apical bud center point with ID output by the tracking module to the positioning module. The positioning module combines the intrinsic and extrinsic parameters of the binocular camera and performs coordinate transformation through the principle of stereo vision to calculate the three-dimensional spatial coordinates of the apical bud in the camera coordinate system.

[0048] The output module is used to output data containing the apical bud ID and its corresponding three-dimensional spatial coordinates, which is used to guide the budding actuator in its operation.

[0049] Compared with the prior art, the beneficial effects of the present invention are:

[0050] 1. This invention systematically integrates the MobileNetV3 lightweight backbone, CA coordinate attention mechanism, SPPELAN hybrid pooling, PSA polarization self-attention, C3_VIT global context feature fusion, and BiFPN multi-scale feature fusion to construct a lightweight apical bud recognition model. This lightweight apical bud recognition model significantly improves the recognition accuracy of apical buds in complex cotton field backgrounds while maintaining real-time performance, effectively addressing interference from light variations, shading, and similar leaves. Experiments demonstrate that compared to mainstream YOLO series lightweight models, this invention significantly reduces the number of parameters and computational cost while maintaining high recognition accuracy.

[0051] 2. This invention employs a two-stage tracking strategy combining one-stage IOU association with two-stage 9-state vector acceleration Kalman filtering. By constructing a 9-state vector containing position, velocity, and acceleration, and a second-order acceleration motion model based on Taylor expansion, it can accurately describe and predict the nonlinear trajectory of the apical bud during field operations by the topping machine, effectively reducing ID switching and significantly improving tracking robustness in dynamic scenarios. Simultaneously, through a distance-compensation matching mechanism and a state smoothing update mechanism, the system's resistance to interference such as target occlusion and deformation blurring is further enhanced.

[0052] 3. This invention uses an identification module, a tracking module, and a positioning module to transmit the three-dimensional position information predicted by the tracking module to the positioning module in real time, thereby achieving accurate three-dimensional positioning in dynamic scenarios and ensuring that the jacking actuator obtains continuous and stable target position information, thus improving the accuracy and reliability of jacking operations. Attached Figure Description

[0053] Figure 1 This is a flowchart of an identification, tracking, and positioning system according to an embodiment of the present invention;

[0054] Figure 2 This is a framework diagram of an identification, tracking, and positioning system according to an embodiment of the present invention;

[0055] Figure 3 This is a network architecture diagram of a lightweight apical bud recognition model according to an embodiment of the present invention;

[0056] Figure 4 This is a diagram of an improved SPPELAN hybrid pooling network structure according to an embodiment of the present invention;

[0057] Figure 5This is a network structure diagram of the CA attention mechanism module according to an embodiment of the present invention;

[0058] Figure 6 This is a network structure diagram of a PSA polarization self-attention module according to an embodiment of the present invention;

[0059] Figure 7 This is a network structure diagram of the C3_VIT feature fusion and extraction module according to one embodiment of the present invention;

[0060] Figure 8 This is a flowchart illustrating the target tracking algorithm according to one embodiment of the present invention.

[0061] Figure 9 This is a target recognition effect diagram according to one embodiment of the present invention;

[0062] Figure 10 This is a heatmap of the identification model according to an embodiment of the present invention;

[0063] Figure 11 The image shows the target tracking effect of the tracking algorithm in one embodiment of the present invention in the previous and next frames. Detailed Implementation

[0064] The embodiments of the present invention are described in detail below. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0065] Figure 1 The above illustrates a preferred embodiment of the high-speed dynamic identification and positioning method for cotton apical buds in the field. The method includes the following steps:

[0066] Step S1: Obtain real-time image frames of cotton terminal buds in the field;

[0067] Step S2: Input the image frame into the lightweight apical bud recognition model and output the detection box information of the apical bud, the detection box information including the two-dimensional image coordinates of the center point of the apical bud;

[0068] Step S3: Input the detection box information into the target tracking module. The target tracking module maintains a 9-state vector containing position, velocity and acceleration for each apical bud, and realizes cross-frame target tracking through a strategy combining one-stage IOU association and two-stage Kalman filter prediction, and assigns a unique identifier ID to each tracked apical bud.

[0069] Step S4: Input the two-dimensional image coordinates of the apical bud center point with ID output by the tracking module to the positioning module. The positioning module combines the intrinsic and extrinsic parameters of the binocular camera and performs coordinate transformation through the principle of stereo vision to calculate the three-dimensional spatial coordinates of the apical bud in the camera coordinate system.

[0070] Step S5: Output data containing the apical bud ID and corresponding three-dimensional spatial coordinates to guide the apical bud actuator operation.

[0071] The lightweight apical bud recognition model is built on the YOLO architecture. Its network structure includes a lightweight backbone network, an improved hybrid pooling module SPPELAN, a neck feature fusion network, a BiFPN-like multi-path feature fusion structure, and a detection head. The lightweight backbone network adopts the MobileNetV3_large architecture and embeds a CA coordinate attention mechanism after its shallow and deep feature map outputs. The improved hybrid pooling module SPPELAN is connected to the end of the lightweight backbone network to perform max pooling and variance pooling in parallel, and concatenates the pooling results to extract the intensity and texture complexity features of the apical bud target. The neck feature fusion network includes a PSA polarization self-attention mechanism and a C3_VIT feature fusion extraction module. The PSA polarization self-attention mechanism uses a channel and spatial dual-branch structure for polarization filtering. The C3_VIT module integrates a lightweight Vision module based on the C3 module. The Transformer establishes global context association; the BiFPN-like multi-path feature fusion structure connects the detection head and the lightweight backbone network, forming a bidirectional cross-scale information flow channel with top-down, bottom-up, and same-level skip connections; the detection head is used to output the bounding box, confidence score, and category of the apical bud.

[0072] The lightweight apical bud recognition model is trained using an adaptive weighted loss function. The weighted combination of ShapeIOU loss and NWD loss is expressed as:

[0073]

[0074] Where α is the balancing weight coefficient, which is used to dynamically adjust the weights of the two losses according to the target size; NWD

[0075] Loss function; This is the shapeiou loss function.

[0076] The target tracking module employs a two-stage tracking strategy, specifically including:

[0077] Phase 1 IOU tracking: Calculate the Intersection over Union (IOU) ratio between the tracker's predicted bounding box in the previous frame and the detection bounding box in the current frame, form a cost matrix, and use the Hungarian algorithm for preliminary association matching;

[0078] The second stage, 9-state Kalman filter tracking, involves using a 9-state vector acceleration model to predict motion and perform secondary association for unmatched targets after initial association.

[0079] The 9-state vector is represented as follows:

[0080]

[0081] Where P, V, and A represent the three-dimensional position, velocity, and acceleration components of the apical bud in the Cartesian coordinate system, respectively; Let be the coordinates of the apical bud in three-dimensional space. The speed at which the apical bud moves in three directions. For the top

[0082] The state transition matrix F of the Kalman filter for uniform acceleration in three directions is:

[0083]

[0084] Where Δt is the time increment.

[0085] The target tracking module also includes a distance-compensation matching mechanism: for trackers that still fail to match after the first and second stages, the Euclidean distance between their predicted position and the center of the unmatched detection box in the current frame is calculated. If the distance is less than a preset threshold, forced association matching is performed.

[0086] The target tracking module also includes a state smoothing update mechanism: for a successfully matched tracker, Alpha-Beta filtering or low-pass filtering is used to smooth the velocity and acceleration observed in the current frame in order to update the 9-state vector and suppress detection box jitter.

[0087] The coordinate transformation in step S4 specifically includes: extracting the two-dimensional pixel coordinates of the center point of the apical bud in the left image from the detection box information output by the recognition module; searching for matching points in the right image that correspond to the center point of the apical bud in the left image to form a matching pair of two-dimensional points in the left and right images; using the pre-calibrated internal and external parameters of the binocular camera, based on the principle of stereo vision, mapping the matching pair of two-dimensional points in the left and right images to three-dimensional points in the camera coordinate system; and outputting the three-dimensional coordinates of the apical bud in the camera coordinate system, including the depth estimate z-coordinate.

[0088] It also includes an occlusion handling mechanism: when the target is occluded and the recognition module fails to detect it briefly, the tracking module predicts the target position through Kalman filtering and re-associates the recognition box in subsequent frames through IOU (Intersection over Union) to deal with the situation where the target disappears briefly.

[0089] A high-speed dynamic identification and positioning system for cotton buds in the field to implement the method includes an image acquisition unit, a bud identification module, a target tracking module, a positioning module, and an output module;

[0090] The image acquisition unit is used to acquire real-time image frames of the top buds of cotton plants in the field;

[0091] The apical bud recognition module is used to input the image frame into the lightweight apical bud recognition model and output the detection box information of the apical bud, the detection box information including the two-dimensional image coordinates of the center point of the apical bud;

[0092] The target tracking module is connected to the bud identification module and is used to input the detection box information to the target tracking module. The target tracking module maintains a 9-state vector containing position, velocity and acceleration for each bud, and realizes cross-frame target tracking through a strategy combining one-stage IOU association and two-stage Kalman filter prediction, and assigns a unique identifier ID to each tracked bud.

[0093] The positioning module is connected to the target tracking module and is used to input the two-dimensional image coordinates of the apical bud center point with ID output by the tracking module to the positioning module. The positioning module combines the intrinsic and extrinsic parameters of the binocular camera and performs coordinate transformation through the principle of stereo vision to calculate the three-dimensional spatial coordinates of the apical bud in the camera coordinate system.

[0094] The output module is used to output data containing the apical bud ID and its corresponding three-dimensional spatial coordinates, which is used to guide the budding actuator in its operation.

[0095] In one specific embodiment of the present invention, specifically:

[0096] like Figure 2 As shown, the system of this invention mainly comprises three core components: a terminal bud recognition model, a terminal bud tracking algorithm, and a terminal bud localization module. The system workflow is as follows: Figure 1 As shown, the specific steps are as follows:

[0097] Image Acquisition and Preprocessing: In a cotton field environment, the height of cotton plants was statistically analyzed to determine the ground clearance of the binocular camera. Real-time photos and video streams of the terminal buds were acquired from a bird's-eye view. The acquired images underwent data augmentation preprocessing, including horizontal and vertical flipping, random rotation, brightness and contrast adjustment, and random scaling, to simulate complex field conditions.

[0098] Dataset Preparation and Training: A large number of images of the tops of cotton plants were acquired using a stereo camera. The terminal buds in the images were manually labeled using tools such as LabelImg, with the label boxes tightly surrounding the buds. The labeled data was converted to YOLO format. The dataset was divided into training and validation sets in approximately an 8:2 ratio. A terminal bud recognition model, Easy_Cottonbud, was designed, imported into the dataset, and trained to obtain the terminal bud weight file.

[0099] Terminal bud recognition: Preprocessed image frames are input into the trained lightweight terminal bud recognition model Easy_Cottonbud. This model infers from the input image and outputs the bounding boxes of all cotton terminal buds in the image and their confidence scores. The recognition results include the coordinates of the center point of the terminal bud bounding box.

[0100] Bud tracking: The detection box sequence (detection box information) output by the recognition module is sent to the target tracking module. This module creates or associates a tracker for each detected bud target. Through a strategy combining one-stage IOU association and two-stage 9-state vector Kalman filter prediction, it achieves stable and continuous tracking of targets across frames and assigns a unique ID to each continuously tracked bud.

[0101] Bud localization: For bud targets with stable IDs output by the tracking module, extract the 2D image coordinates of the center point of their bounding box. Combining the pre-calibrated internal and external parameters of the binocular camera, and using the principle of stereo vision, map the 2D point pairs in the matched left and right images to 3D points in the camera coordinate system, thereby obtaining the precise position of each tracked bud in 3D space.

[0102] Information Output: The system ultimately outputs a string of structured data, typically including: bud ID and three-dimensional spatial coordinates. This information can be directly sent to the actuator of the bud-removing machine, guiding it to move precisely to the target bud position for bud removal.

[0103] like Figure 3 As shown, the apical bud recognition model of the present invention is based on the YOLO series target detection framework and is deeply improved. The overall structure is divided into three parts: backbone network, neck network and detection head, and is equipped with an improved adaptive loss function.

[0104] Model design and construction:

[0105] Backbone network replacement: MobileNetV3-Large is adopted as the lightweight backbone network to replace the heavier backbones such as YOLO's CSPDarknet. MobileNetV3 combines depthwise separable convolution, linear bottleneck and inverse residual structure, and integrates a lightweight SE attention module, which greatly reduces the amount of computation and parameters while maintaining high accuracy.

[0106] Compensating for the inability to extract core features:

[0107] CA (Coordinate Attention) mechanism integration: A CA module is embedded after the output of shallow and deep feature maps in MobileNetV3. For example... Figure 5As shown, the CA module captures precise location information and long-range dependencies across channels by decomposing channel attention into one-dimensional feature encoding along both height and width directions. This allows the network to focus more on the spatial location of the apical bud, which helps to accurately locate the apical bud in a background of dense foliage.

[0108] SPPELAN Multi-Scale Convergence Module: At the end of the backbone network, an improved SPPELAN module is accessed, with the following structure: Figure 4 This module performs max pooling and variance pooling in parallel and then concatenates the results. Variance pooling effectively characterizes the texture differences between the apical bud and the surrounding leaves, enhances the model's ability to capture subtle features of the apical bud, and improves its robustness against interference from similar backgrounds.

[0109] Enhanced neck and detection head:

[0110] PSA polarization self-attention mechanism: A PSA module is introduced in the neck feature pyramid network (FPN) path and before the detection head, such as... Figure 6 PSA performs soft filtering in the channel dimension and hard filtering in the spatial dimension through a "polarization" operation. This mechanism forces the network to focus on the most discriminative apical bud features in both the channel and spatial dimensions at a relatively low computational cost, effectively addressing complex field backgrounds and partial occlusion.

[0111] C3_VIT Feature Extraction Layer: Replace the standard C3 module in the neck network with the C3_VIT module, such as... Figure 7 This module retains the local feature extraction capabilities of C3 while incorporating a lightweight Vision Transformer block. The Transformer establishes global contextual relationships between all pixels within the feature map through a self-attention mechanism, helping the model to accurately distinguish similar pixels based on global semantics.

[0112] Feature fusion structure: A BiFPN-like structure is adopted for multi-scale feature fusion. This structure uses learnable weights to weight and fuse feature maps at different scales, enabling the network to adaptively prioritize the feature scales most useful for detecting apical buds.

[0113] Loss function design: Adaptive weighted ShapeIOU and NWD losses are used as the bounding box regression losses. Considering the small size and varied shapes of cotton bud targets, this combined loss function better balances shape matching and robustness to small center point shifts. The weight parameters can be dynamically adjusted according to the size of the target box; for smaller targets, NWD is assigned a higher weight.

[0114] Model Inference and Output: After training, the model is deployed to an embedded computing platform on a top-mounted device. During inference, the model processes each input frame of image in real time.

[0115] Apical bud tracking algorithm:

[0116] The target tracking algorithm receives the continuous frame recognition results output by the recognition model, is responsible for the allocation of base detection box association IDs, prevents repeated recognition between consecutive frames in dynamic situations, and forms a stable motion trajectory for moving targets. Target tracking can also continuously update the spatial position of the same bud during the top-piercing machine operation, thereby improving the accuracy and timeliness of bud positioning information.

[0117] like Figure 8 As shown, the working steps of the target tracking algorithm include:

[0118] (1) Target detection input: The list of apical bud detection boxes for each frame is used as the input to the tracking algorithm;

[0119] (2) One-stage IOU tracking: Calculate the IOU between all tracker predicted boxes and all detection boxes in the current frame to form a cost matrix. Use Hungarian algorithm or other methods for matching. Matching pairs with IOU higher than a set threshold are initially associated, which can quickly associate smooth and continuous targets.

[0120] (3) Two-stage 9-state Kalman filter tracking: Motion prediction is performed using a 9-state vector acceleration model Kalman filter. The prediction model takes acceleration into account, which can better describe the nonlinear motion that may occur in the cotton buds from the perspective of a high-speed traveling topping machine, and at the same time deal with the situation where the target disappears temporarily due to occlusion;

[0121] (4) Distance-compensated matching: For trackers that still fail to match after the above two stages (possibly due to momentary severe blurring or deformation), if the Euclidean distance between their predicted position and the center of a certain unmatched detection box is less than a relatively lenient threshold, forced association is performed to avoid losing the target due to brief interference.

[0122] (5) State observation and smooth update: For a tracker that has been successfully matched, the detection box information matched in the current frame is used to update its 9-state vector. During the update process, the observed velocity and acceleration are smoothed by Alpha-Beta filtering or low-pass filtering to suppress noise caused by detection box jitter and obtain a more stable and smooth motion trajectory.

[0123] Apical bud positioning module: The goal of the positioning model is to convert the two-dimensional coordinates of the tracked apical bud image into three-dimensional spatial coordinates that are instructive for the bud-growing robot arm.

[0124] The bud localization module includes a binocular camera calibration module and a coordinate transformation module. The binocular camera calibration module acquires the camera's intrinsic and extrinsic parameters; the coordinate transformation module, based on these parameters, converts the two-dimensional coordinates of the bud center obtained by the target tracking module into spatial position information. This process is a crucial step in achieving precise bud-picking robotic arm motion control.

[0125] like Figure 9 As shown, this model performs well in a complex cotton field scene. It's evident that the terminal buds are highly similar to the background green leaves, and their size is relatively small compared to the rest of the image, indicating they are small targets. Through the improvements of this invention, the designed recognition model accurately identifies all cotton terminal buds in the scene, with precise bounding boxes. The confidence level for each target is above 0.35, demonstrating the synergy between the designed modules, compensating for feature loss in small targets during training, and effectively addressing the suppression of green against a green background.

[0126] like Figure 10 As shown, this is a heat map of the sensory region in a complex cotton field scenario, from... Figure 10 It is evident that the model's main sensory region is concentrated on the apical bud, while remaining indifferent to cotton leaves. This demonstrates that the model effectively resists strong background interference and effectively focuses on the fine-grained features of the apical bud target.

[0127] like Figure 11 As shown, the target tracking effect of the tracking algorithm of this model is shown in the previous and next frames. In the previous frame, four apical targets were identified and assigned target IDs. The next frame is the identification under the camera's field of view after dynamic movement. It can be seen that the four apical targets identified in the previous frame have sent a certain displacement due to relative motion, but there is no inter-frame re-identification in the next frame, the identification ID has not changed, and new IDs are assigned to newly entered targets.

[0128] To verify the technical effect of the present invention, tests were conducted under the following experimental conditions:

[0129] Hardware platform: NVIDIA RTX 2060 GPU, ZED 2i depth camera, laptop (Windows 11 system).

[0130] Software environment: CUDA 11.1, Python 3.8, PyTorch 1.8.2 deep learning framework.

[0131] Images of cotton buds were collected in cotton fields, resulting in 3500 original images. The dataset was randomly divided into a training set (2800 images) and a validation set (700 images) at an 8:2 ratio, with all images having a uniform resolution of 1280*1280.

[0132] A comparative experiment was conducted using the same hyperparameter settings: batch size 4, learning rate 0.01, momentum 0.937, weight decay 0.0005, optimizer SGD, and 300 training epochs.

[0133] The comparison models include: YOLOv5n, YOLOv5s, YOLOv8s, YOLOv11s, and YOLOv12s.

[0134] After 300 rounds of training, the target recall R of the lightweight apical bud recognition model of this invention reached 0.828, the recognition accuracy mAP50 reached 0.845, the mAP50-90 reached 0.6, the number of parameters was 3.8M, and the computational cost was 6.9G FLOPs.

[0135] The number of parameters and computational cost of each comparative model are as follows:

[0136] YOLOv5s: 7.2M, 15.9G FLOPs

[0137] YOLOv5su: 9.12M, 24G FLOPs

[0138] YOLOv8s: 11.2M, 28.8G FLOPs

[0139] YOLOv11s: 9.5M, 21.7G FLOPs

[0140] YOLOv12s: 9.28M, 21.7G FLOPs

[0141] Experimental results show that the number of parameters and computational cost of the lightweight apical bud recognition model of this invention are significantly reduced compared with the comparative models, demonstrating a clear lightweight effect while maintaining high recognition accuracy.

[0142] With an NVIDIA RTX 2060 graphics card and a ZED 2i depth camera, the lightweight bud recognition model of this invention achieves a real-time recognition frame rate of 28 frames per second (FPS), meeting the real-time requirements of field operations.

[0143] This invention is a target tracking and positioning algorithm for cotton bud identification and adaptation to nonlinear motion in complex field environments.

[0144] This invention revolves around two core objectives: balancing accuracy and speed, and enhancing features. It systematically integrates multiple advanced modules to form a complementary and synergistic network architecture. First, it employs the MobileNetV3 lightweight backbone module, characterized by high performance, fast inference speed, low computational resource requirements, and balanced accuracy, to address the difficulties of deploying the original backbone network on resource-constrained devices. The original model also suffers from high computational load and low real-time frame rates during real-time detection. A CA coordinate attention mechanism is adopted, which can establish long-range dependencies in both channel and spatial dimensions with minimal computational overhead. Added to the lightweight backbone network, the CA coordinate attention mechanism can focus more on the subtle texture features of the apical bud and its spatial coordinates, compensating for the backbone network's feature extraction capabilities with less computation, thereby further enhancing the texture and positional information of the apical bud. An improved hybrid pooling module, SPPELAN, is added to the tail of the main structure. It utilizes max pooling to extract the intensity features of the apical bud target in the image and variance pooling to extract the texture complexity features of the apical bud target. Combining these two methods into hybrid pooling enhances the recognition model's resistance to interference and addresses misjudgments caused by light reflections and shadow occlusion. A PSA polarization self-attention mechanism is added to the feature extraction head. Distinguishing the apical bud from similar leaves in extremely complex backgrounds is a challenging aspect. The PSA polarization self-attention mechanism performs "polarization" filtering through both channel and spatial dimensions, strengthening the most relevant feature channels and spatial locations while completely suppressing irrelevant ones. Its advantage lies in enabling the network to focus on unoccluded local features and understand the spatial structural relationship between the apical bud and the overall cotton plant, further suppressing interference from complex environmental disturbances. For the feature extraction module, a C3_VIT feature fusion extraction module is designed in conjunction with a lightweight Vision Transformer. This enables the model to establish global contextual relationships, allowing the network to not only focus on local features but also model long-distance dependencies between regions in the image. This integrates information from a global perspective, enhancing the overall understanding of the apical bud and its surrounding environment. Finally, a BiFPN-like multi-path feature fusion structure is used for multi-scale feature stacking and weighted fusion to complete the overall network architecture. This structure achieves full integration of deep features through efficient cross-scale feature connections, giving the network high semantic feature information. Based on the above design, this invention improves the recognition accuracy under partial occlusion of the apical bud by improving the trunk, neck, and head on the basis of the YOLO architecture, while also achieving a good balance between recognition accuracy and real-time performance. In the target tracking part, a two-stage tracking strategy is used. The target features are input through a deep learning model. First, a preliminary tracking is performed using an IOU tracking algorithm with minimal computational cost. Meanwhile, for high-speed operation, the second-stage tracking algorithm uses a 9-state Kalman filter, a second acceleration model, and a Taylor expansion mathematical model to predict motion, handle nonlinear motion, and reduce ID switching. For cases of excessively fast motion, it is converted to Euclidean distance matching to retrieve the target.In the correlation processing between consecutive video frames, this algorithm can effectively improve the problem of target recognition difficulties caused by occlusion and partial occlusion, and enhance the system's stable tracking capability in complex scenes.

[0145] It should be understood that although this specification is described according to various embodiments, not every embodiment contains only one independent technical solution. This way of describing the specification is only for clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

[0146] The detailed descriptions listed above are merely specific illustrations of feasible embodiments of the present invention and are not intended to limit the scope of protection of the present invention. All equivalent embodiments or modifications made without departing from the spirit of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for high-speed dynamic identification and positioning of cotton terminal buds in the field, characterized in that, Includes the following steps: Step S1: Obtain real-time image frames of cotton terminal buds in the field; Step S2: Input the image frame into the lightweight apical bud recognition model and output the detection box information of the apical bud, the detection box information including the two-dimensional image coordinates of the center point of the apical bud; Step S3: Input the detection box information into the target tracking module. The target tracking module maintains a 9-state vector containing position, velocity and acceleration for each apical bud, and realizes cross-frame target tracking through a strategy combining one-stage IOU association and two-stage Kalman filter prediction, and assigns a unique identifier ID to each tracked apical bud. Step S4: Input the two-dimensional image coordinates of the apical bud center point with ID output by the tracking module to the positioning module. The positioning module combines the intrinsic and extrinsic parameters of the binocular camera and performs coordinate transformation through the principle of stereo vision to calculate the three-dimensional spatial coordinates of the apical bud in the camera coordinate system. Step S5: Output data containing the apical bud ID and corresponding three-dimensional spatial coordinates to guide the apical bud actuator operation.

2. The high-speed dynamic identification and positioning method for cotton apical buds in the field according to claim 1, characterized in that, The lightweight apical bud recognition model is built on the YOLO architecture, and its network structure includes a lightweight backbone network, an improved hybrid pooling module SPPELAN, a neck feature fusion network, a BiFPN-like multi-path feature fusion structure, and a detection head. The lightweight backbone network adopts the MobileNetV3_large architecture and embeds a CA coordinate attention mechanism after its shallow and deep feature map outputs. The improved hybrid pooling module SPPELAN is connected to the end of the lightweight backbone network to perform max pooling and variance pooling in parallel, and to concatenate the pooling results to extract the intensity features and texture complexity features of the apical bud target. The neck feature fusion network includes a PSA polarization self-attention mechanism and a C3_VIT feature fusion extraction module. The PSA polarization self-attention mechanism uses a dual-branch structure of channel and space for polarization filtering. The C3_VIT module integrates a lightweight Vision Transformer on the basis of the C3 module to establish global context association. The BiFPN-like multi-path feature fusion structure connects the detection head and the lightweight backbone network, forming a bidirectional cross-scale information flow channel with top-down, bottom-up, and same-level jump connections. The detection head is used to output the bounding box, confidence level, and category of the apical bud.

3. The high-speed dynamic identification and positioning method for cotton apical buds in the field according to claim 2, characterized in that, The lightweight apical bud recognition model is trained using an adaptive weighted loss function. The weighted combination of ShapeIOU loss and NWD loss is expressed as: Where α is the balancing weight coefficient, which is used to dynamically adjust the weights of the two losses according to the target size; NWD Loss function; This is the shapeiou loss function.

4. The high-speed dynamic identification and positioning method for cotton apical buds in the field according to claim 1, characterized in that, The target tracking module employs a two-stage tracking strategy, specifically including: Phase 1 IOU tracking: Calculate the Intersection over Union (IOU) ratio between the tracker's predicted bounding box in the previous frame and the detection bounding box in the current frame, form a cost matrix, and use the Hungarian algorithm for preliminary association matching; The second stage, 9-state Kalman filter tracking, involves using a 9-state vector acceleration model to predict motion and perform secondary association for unmatched targets after initial association.

5. The high-speed dynamic identification and positioning method for cotton apical buds in the field according to claim 4, characterized in that, The 9-state vector is represented as follows: Where P, V, and A represent the three-dimensional position, velocity, and acceleration components of the apical bud in the Cartesian coordinate system, respectively; Let be the coordinates of the apical bud in three-dimensional space. The speed at which the apical bud moves in three directions. For the top The state transition matrix F of the Kalman filter for uniform acceleration in three directions is: Where Δt is the time increment.

6. The high-speed dynamic identification and positioning method for cotton apical buds in the field according to claim 4, characterized in that, The target tracking module also includes a distance-compensation matching mechanism: For trackers that still do not match after the first and second stages, calculate the Euclidean distance between their predicted position and the center of the unmatched detection box in the current frame. If the distance is less than a preset threshold, then perform forced association matching.

7. The high-speed dynamic identification and positioning method for cotton apical buds in the field according to claim 4, characterized in that, The target tracking module also includes a smooth state update mechanism: For a successfully matched tracker, alpha-beta filtering or low-pass filtering is used to smooth the velocity and acceleration observed in the current frame in order to update the 9-state vector and suppress detection box jitter.

8. The high-speed dynamic identification and positioning method for cotton apical buds in the field according to claim 1, characterized in that, The coordinate transformation in step S4 specifically includes: From the detection box information output by the recognition module, extract the two-dimensional pixel coordinates of the center point of the apical bud in the left image; search for matching points in the right image that correspond to the center point of the apical bud in the left image to form a matching two-dimensional point pair between the left and right images; using the pre-calibrated internal and external parameters of the binocular camera, based on the principle of stereo vision, map the matching two-dimensional point pair between the left and right images to three-dimensional points in the camera coordinate system; output the three-dimensional coordinates of the apical bud in the camera coordinate system, including the depth estimate z-coordinate.

9. The high-speed dynamic identification and positioning method for cotton apical buds in the field according to claim 1, characterized in that, It also includes an occlusion handling mechanism: When the target is occluded, causing the recognition module to fail to detect it briefly, the tracking module predicts the target position using Kalman filtering and re-associates the recognition box in subsequent frames using Intersection over Union (IOU) to deal with the situation where the target disappears briefly.

10. A high-speed dynamic identification and positioning system for cotton apical buds in the field for implementing the method of any one of claims 1 to 9, characterized in that, It includes an image acquisition unit, a bud recognition module, a target tracking module, a positioning module, and an output module; The image acquisition unit is used to acquire real-time image frames of the top buds of cotton plants in the field; The apical bud recognition module is used to input the image frame into the lightweight apical bud recognition model and output the detection box information of the apical bud, the detection box information including the two-dimensional image coordinates of the center point of the apical bud; The target tracking module is connected to the bud identification module and is used to input the detection box information to the target tracking module. The target tracking module maintains a 9-state vector containing position, velocity and acceleration for each bud, and realizes cross-frame target tracking through a strategy combining one-stage IOU association and two-stage Kalman filter prediction, and assigns a unique identifier ID to each tracked bud. The positioning module is connected to the target tracking module and is used to input the two-dimensional image coordinates of the apical bud center point with ID output by the tracking module to the positioning module. The positioning module combines the intrinsic and extrinsic parameters of the binocular camera and performs coordinate transformation through the principle of stereo vision to calculate the three-dimensional spatial coordinates of the apical bud in the camera coordinate system. The output module is used to output data containing the apical bud ID and its corresponding three-dimensional spatial coordinates, which is used to guide the budding actuator in its operation.