Multi-target tracking method, dynamic target screening operation method, operation equipment and readable storage medium

Through the multi-target tracking method that combines the Hungarian algorithm and deep learning, combined with the odometry transformation matrix and sliding window screening, the stability and accuracy problems of the multi-target tracking method in complex environments are solved, and efficient and accurate target recognition and task execution are achieved.

CN120765690APending Publication Date: 2025-10-10新疆极目机器人科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510823776.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing multi-target tracking methods lack stability under factors such as lighting, target motion, and background changes, resulting in low target recognition and tracking accuracy. Traditional target screening methods are prone to introducing errors, making it difficult to achieve efficient and accurate task execution.

Method used

The Hungarian algorithm is used to match the target tracking frame, combined with deep learning target detection and odometry transformation matrix for coordinate conversion, target matching is determined by Euclidean distance, and a sliding window is used to dynamically eliminate targets that are out of view. The dynamic target screening method is combined to accurately determine the target to be operated.

Benefits of technology

It improves the stability and accuracy of multi-target tracking, optimizes the target screening process, and improves the efficiency and accuracy of operation execution. It is suitable for various application scenarios such as smart agriculture, industrial automation and unmanned driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765690A_ABST
    Figure CN120765690A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-target tracking method, a dynamic target screening operation method, operation equipment and a readable storage medium. The method comprises the steps of target initialization, target detection, coordinate conversion, target matching, pixel coordinate updating, sliding window screening and the like, and can realize efficient and stable target tracking in a complex environment. Target detection is carried out by adopting a YoV8 algorithm, and a target is converted from a camera coordinate system to a world coordinate system through 6D pose information and an odometer transformation matrix, so that accurate positioning of the target in a three-dimensional space is ensured. And matching the targets through the Euclidean distance, and screening the targets beyond the visual field range by using a sliding window algorithm. In combination with a dynamic target screening method, a laser is controlled based on a camera coordinate of a target to realize accurate topping operation. According to the method, the target tracking precision and the topping efficiency are remarkably improved, and the missing topping rate and the repeated topping rate are reduced by about 20% respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and target tracking, and in particular to a multi-target tracking method, a dynamic target screening operation method, an operation device based on multi-target tracking, and a readable storage medium. Background Art

[0002] In the fields of computer vision and automated control, multi-target tracking technology is widely used in tasks such as target recognition, intelligent monitoring, automated navigation, and precision strike. Existing multi-target tracking methods often rely on a single target detection model, which is susceptible to interference from lighting, target motion, and background changes, resulting in insufficient tracking stability. Furthermore, in specific application scenarios such as agriculture and industry, the accuracy of target selection and task execution is crucial to the ultimate performance.

[0003] However, traditional target screening methods typically use fixed thresholds or simple rule matching, which can easily introduce errors and make it difficult for target operation equipment to accurately perform tasks. Furthermore, because most current mainstream multi-target tracking algorithms rely solely on image features, they place high demands on image quality and require significant computing resources. They struggle to achieve high accuracy under varying viewing angles, lighting conditions, high target similarity, and high real-time requirements. Especially in complex outdoor lighting environments, existing multi-target tracking methods can easily lose targets or cause mismatches during related operations. Summary of the Invention

[0004] The present invention aims to provide a multi-target tracking method that can stably and accurately detect and track multiple targets, and combines it with a dynamic target screening method to improve the operating accuracy of operating equipment. The present invention also provides an operating device based on multi-target tracking to achieve automated and precise operations. The main technical solutions of the present invention are as follows:

[0005] The technical solutions of the present invention are as follows:

[0006] A multi-target tracking method, comprising:

[0007] (1) Target tracking:

[0008] Using preset rules, the tracking target is extracted from multiple consecutive frames of images, a target tracking frame is generated, and a tracking queue is constructed; wherein, each designated reference point in the target tracking frame corresponds to a world coordinate;

[0009] (2) Target detection:

[0010] Performing target detection on the multiple frames of image using a preset detection algorithm, generating target detection frames, and constructing a target detection frame queue arranged in time series, wherein the arrangement direction of the target detection frame queue matches the queue sorting direction of the tracking queue;

[0011] (3) Coordinate transformation:

[0012] Calculate the pixel coordinates of the preset detection position in each target detection frame, and obtain the corresponding camera coordinates by combining the camera intrinsic parameters, depth information and timestamp;

[0013] Convert the preset detection position to world coordinates based on the camera coordinates and the odometry transformation matrix;

[0014] (4) Target matching:

[0015] Calculate the Euclidean distance between the world coordinates of the preset detection position at the corresponding timestamp and the world coordinates of the specified reference point in the tracking queue;

[0016] If the Euclidean distance is less than or equal to the preset threshold, the match is considered successful;

[0017] (5) Pixel coordinate update:

[0018] Obtaining the tracking target corresponding to each detected target, and updating the pixel coordinates of each detected target as the pixel coordinates of the corresponding tracking target to obtain updated pixel coordinates of the tracking target; adding the detected targets that are determined to be unmatched as new tracking targets to the tracking queue, and updating the pixel coordinates of the new tracking targets based on the body coordinates of the corresponding detected targets;

[0019] (6) Sliding window screening:

[0020] The sliding window algorithm is used to dynamically eliminate tracking targets whose pixel coordinates are beyond the sliding window range.

[0021] Specifically,

[0022] The target tracking step includes:

[0023] The Hungarian algorithm is used to match the target in each frame image and generate the target tracking frame; the target tracking frame is sorted according to the time series to build the target tracking queue;

[0024] The target detection step includes:

[0025] The target detection model is used to perform target detection on multiple frames of images and generate target detection frames.

[0026] Specifically, the target detection step includes:

[0027] The coordinate conversion step comprises:

[0028] (1) Get the transformation matrix:

[0029] According to the image timestamp, extract the transformation matrix from the camera coordinate system to the world coordinate system that matches the corresponding timestamp from the odometry transformation matrix queue;

[0030] (2) Calculate the camera coordinates:

[0031] The camera coordinates of the target are calculated based on the target pixel coordinates, the image center pixel coordinates, the depth value and the focal length, satisfying:

[0032]

[0033] Where p_c is the camera coordinate of the preset detection position; u, v are the horizontal and vertical pixel coordinates of the target respectively, fx, fy are the focal lengths of the camera on the x and y axes respectively, cx, cy are the pixel coordinates of the image center on the x and y axes respectively, and d is the depth value;

[0034] (3) Calculate world coordinates:

[0035] According to the camera coordinates and transformation matrix, the target is transformed into the world coordinate system to meet the following requirements:

[0036] p_w=T wc *p_c,

[0037] Where p_w is the world coordinate of the preset detection position in the target detection frame; T wc is the transformation matrix from the camera coordinate system to the world coordinate system at the corresponding moment.

[0038] Specifically, the target matching step includes:

[0039] Calculate the Euclidean distance between the world coordinates of the preset detection position at the corresponding timestamp and the world coordinates of the specified reference point in the tracking queue, satisfying:

[0040]

[0041] Where p_w is the world coordinate of the preset detection position in the target detection frame, obj.p_w is the world coordinate of the preset tracking position in the tracking queue, and err_dist is the Euclidean distance between the two in the world coordinate system;

[0042] If err_dist ≤ the preset threshold, the target is determined to be matched, otherwise the corresponding detected target is used as a new tracking target in the tracking queue.

[0043] Specifically, updating the pixel coordinates of the new tracking target based on the body coordinates of the corresponding detection target includes:

[0044] (1) Calculate the body coordinates of the detection target corresponding to the new tracking target, satisfying:

[0045]

[0046] Among them, p b mT is the coordinate of the detection target in the body coordinate system; wc is the transformation matrix from the camera coordinate system to the world coordinate system; p_c is the camera coordinate of the detection target; T wb is the transformation matrix from the body coordinate system to the world coordinate system at the current moment;

[0047] (2) Use the body coordinates of the detected target to calculate the pixel coordinates of the new tracking target, satisfying:

[0048]

[0049] Among them, p uv are the pixel coordinates of the new tracking target; cx and cy represent the image center in the camera intrinsic parameters, and fx and fy represent the focal length in the camera intrinsic parameters.

[0050] The present invention also provides a method for dynamic target screening, which comprises:

[0051] (1) Target screening:

[0052] Based on the tracking target coordinate information obtained by the above method, and the operation flag of each tracking target, the operation flag is used to reflect whether the tracking target has been operated;

[0053] Traverse the tracking targets in the sliding window to filter out targets that do not have a preset operation flag;

[0054] (2) Target positioning:

[0055] Get the camera coordinates of the target to be operated;

[0056] (3) Equipment control:

[0057] Based on the camera coordinates of the target, the working direction of the working equipment is adjusted and the work is performed.

[0058] In the above-mentioned dynamic target screening operation method, the operation equipment includes a laser, and the method further includes:

[0059] (1) Calculate the laser control amount:

[0060] The XY galvanometer control quantity of the laser is calculated using the camera coordinates of the target to be operated and the odometer transformation matrix;

[0061] (2) Control laser output:

[0062] Adjust the emission direction of the laser according to the XY galvanometer control value to operate on the target;

[0063] (3) Target job status update:

[0064] After the job is completed, the job status of the target is marked to reflect that the target has been worked on.

[0065] The present invention also proposes an operating device based on multi-target tracking, comprising:

[0066] (1) a target tracking device for detecting and tracking targets in a continuous multi-frame image based on the multi-target tracking method;

[0067] (2) A target screening operation device is used to screen and operate the tracking target based on the dynamic target screening operation method described above.

[0068] Specifically, the operating equipment further includes:

[0069] (1) Visual inertial odometry, used to obtain the camera's 6D pose;

[0070] (2) a laser, used to control the laser output according to the camera coordinates of the target to be operated;

[0071] (3) A transport platform for carrying the target tracking device, the laser and / or the target screening device.

[0072] Finally, the present invention also proposes a readable storage medium, which stores computer-executable instructions, which are used to execute the above-mentioned multi-target tracking method or the above-mentioned dynamic target screening operation method, and can be run on a computer, embedded device, automation system or robot to achieve target detection, target screening and operation control.

[0073] Compared with the prior art, the present invention has the following advantages:

[0074] (1) Improve the stability and accuracy of target tracking:

[0075] The present invention adopts a target detection algorithm based on deep learning, which can efficiently and accurately identify targets from complex background images and extract the precise pixel position of the target; combining the depth information and timestamp of each target, using the odometry transformation matrix to perform coordinate transformation, accurately mapping the target from the image coordinate system to the world coordinate system, thereby effectively eliminating the target positioning error caused by camera motion and perspective change; further, the target matching step is optimized through the Hungarian algorithm, based on the Euclidean distance between targets, effectively reducing the situation of misidentification or repeated identification during the target recognition and tracking process, thereby improving the overall stability and accuracy of the multi-target tracking process.

[0076] (2) Optimize the target screening process and improve the efficiency and accuracy of operation execution:

[0077] The present invention adopts a dynamic sliding window screening method to determine in real time whether the target is within the effective field of view of the camera based on the updated pixel coordinate information. Through this dynamic target management strategy, targets that leave the field of view can be efficiently eliminated, avoiding the waste of computing resources caused by the continuous accumulation of targets, and improving the real-time performance of the system. In addition, the present invention also proposes a dynamic target screening operation method, which accurately determines the target to be operated based on the status of the target operation flag bit, realizes the efficient positioning of the target operation, effectively avoids the phenomenon of repeated operations or missed operations, and greatly improves the execution efficiency and accuracy of the operation equipment.

[0078] (3) It has wide applicability and can be flexibly adapted to various application scenarios:

[0079] Because the present invention achieves a highly stable and accurate target detection, tracking, and dynamic target management method, this universal technical framework can be effectively extended to multiple automation fields. For example, in the field of smart agriculture, such as cotton laser topping operations, the present invention can accurately locate and manage multiple targets in real time, improving the automation accuracy and efficiency of agricultural operations; in the field of industrial automation, such as automatic welding and laser cutting, it can quickly and accurately continuously track and locate multiple target positions with high precision; in addition, in scenarios such as unmanned driving and intelligent monitoring, the present invention significantly improves the system's intelligence and real-time response capabilities through accurate multi-target detection and tracking. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] Figure 1 This is a flowchart of a specific implementation of the multi-target tracking method of the present invention, which clearly shows the key steps of the method, including: (1) target tracking, (2) target detection, (3) coordinate conversion, (4) target matching, (5) pixel coordinate update, and (6) sliding window screening. The execution logic and sequential relationship between each step.

[0081] Figure 2 This is a schematic diagram of the arrangement and relative positions of the topping modules in the multi-target tracking operation equipment of the present invention, wherein each topping module includes a camera module, a laser and a fill light, showing the relative position relationship between the camera module and the laser.

[0082] Figure 3 FIG1 is a flow chart of a multi-target tracking method based on world coordinates and sliding windows according to the present invention. The arrows in the figure indicate the logical flow of each step.

[0083] Figure 4 The figure is a schematic diagram of updating, pushing and popping tracking queue targets based on the sliding window in the present invention. The figure shows the process of removing tracking targets in the sliding window.

[0084] Figure 5 This is a schematic diagram of laser focusing on the same cotton top at different times during the movement of the multi-target tracking operation equipment of the present invention.

[0085] Figure 6 This is a flowchart for the dynamic laser topping process using the multi-target tracking system of the present invention. The diagram illustrates the key steps in dynamic topping, including target tracking, laser control variable calculation, and topping status update. All operations are clearly presented using annotations and flow arrows.

[0086] FIG7( a ) shows the target detection frame and tracking queue status within the camera field of view at time t-1 when the present invention performs multi-target tracking;

[0087] FIG7( b ) shows the status update of the target tracking queue at time t when the present invention performs multi-target tracking as the camera moves, including newly added targets and updated targets;

[0088] FIG7( c ) shows that when the present invention performs multi-target tracking, at time t+1, targets beyond the sliding window range are eliminated, and new targets enter the tracking queue within the field of view. DETAILED DESCRIPTION

[0089] The present invention provides a multi-target tracking method based on computer vision and target tracking, a dynamic target screening method, and a multi-target tracking operation equipment based on the method. The present invention specifically provides a cotton laser topping system, which can achieve stable and efficient target tracking and topping operations in complex environments.

[0090] Reference Figure 1 The implementation process of the multi-target tracking method provided in this embodiment is as follows:

[0091] Step (1): Target tracking: using the Hungarian algorithm to match and extract the tracking target from multiple consecutive frames of images, generating a target tracking frame, sorting the target tracking frames in time series, and constructing a tracking queue; wherein, each designated reference point in the target tracking frame corresponds to a world coordinate of a preset tracking position, and the preset tracking position is the designated reference point in the target tracking frame.

[0092] Step (2): Target detection: Use a preset target detection algorithm (such as YoloV8) to perform target detection on the image sequence, generate target detection frames, and construct a target detection frame queue arranged in time series. The arrangement direction of the target detection frame queue matches the queue sorting direction of the tracking queue.

[0093] Step (3): Coordinate conversion, calculate the pixel coordinates of the preset detection position in each target detection frame, and obtain the corresponding camera coordinates by combining the camera intrinsic parameters, depth information and timestamp; according to the camera coordinates and the odometry transformation matrix, convert the preset detection position to the world coordinates.

[0094] Step (4): Target matching, calculate the Euclidean distance between the world coordinates of the preset detection position at the corresponding timestamp and the world coordinates of the preset tracking position in the specified reference point in the tracking queue; if the Euclidean distance is less than or equal to the preset threshold, the match is determined to be successful.

[0095] Step (5): pixel coordinate update, obtain the tracking target matched by each detection target, and update the pixel coordinates of each detection target as the pixel coordinates of the corresponding tracking target to obtain the updated pixel coordinates of the tracking target. Based on the camera coordinates in the coordinate information after matching and updating described in step (4), calculate the latest pixel coordinates of the target in the tracking queue to achieve stable tracking of the tracking target; for the detection target that is determined to be unmatched, it is treated as a new tracking target and added to the tracking queue, and the pixel coordinates of the new tracking target are updated based on the body coordinates corresponding to the detection target.

[0096] Step (6): Sliding window screening, using the sliding window algorithm to dynamically remove pixel coordinates to screen out the tracking targets that are beyond the scope of the sliding window.

[0097] Through the above steps, the present invention realizes accurate detection and real-time stable tracking of multiple dynamic targets, and effectively reduces misidentification and missed identification during the target tracking process, ensures the real-time and accuracy of target data, and significantly improves the work efficiency and accuracy of subsequent target operation equipment.

[0098] When implementing the above multi-target tracking method, the present invention also provides a cotton laser topping system based on multi-target tracking to track multiple cotton top targets and perform topping operations on the cotton tops. The cotton laser topping system may include:

[0099] Target tracking unit: by establishing a target tracking queue, using a target detection algorithm for target detection to obtain coordinate information, to update the tracking target coordinates in the tracking queue, and using a sliding window method to propose the relevant tracking targets in the tracking queue, to ensure accurate tracking of the target.

[0100] Target screening operation unit: based on the tracking target coordinate information and the obtained operation flag of each tracking target, the operation flag is used to reflect whether the tracking target has been operated, so as to ensure that the operation equipment only operates on unoperated targets, such as only on untopped targets.

[0101] Device control module: based on the spatial position of the unoperated tracking target, calculate the laser XY galvanometer control amount, dynamically adjust the laser power, and improve the topping accuracy.

[0102] Visual inertial navigation system: integrated IMU, GNSS, to improve the target coordinate conversion accuracy.

[0103] As shown in the following Figure 2 For the existing "one film six rows" cotton planting mode scene, a cotton laser topping system composed of three groups of topping modules is provided, including a first group of topping modules 1, a second group of topping modules 2 and a third group of topping modules 3. Each group of topping modules has two lasers 5 and a camera module 4, and the camera module 4 is located in the middle position of the front and rear (up and down in the figure) two lasers 5. Further, light supplementing lamps 6 can also be provided on both sides of the laser 5.

[0104] Each topping module does not interfere with each other and performs the topping task respectively. Among them, the camera module is responsible for target positioning and real-time 6D pose calculation. Through the calibration parameters between the camera module and the laser, the real-time 6D pose information of the target can be converted into angle information in the coordinate system of the laser. According to the calibration relationship between the laser angle and the voltage control amount, the control voltage value required for the actual laser to hit the target can be obtained. The 6D pose information represents the position (x, y, z) and attitude (roll, pitch, yaw) of the target in three-dimensional space, which is obtained in real time through combined navigation or visual inertial odometry.

[0105] In this embodiment, the camera frame rate is preferably matched with the moving frequency of the sliding window to ensure that the laser can accurately capture the real-time position of each tracking target and effectively top. Specifically, the frame rate of the camera can be set to 20Hz, that is, 20 frames of images are captured per second, to ensure continuous capture of target images in a moving state and avoid target loss or unstable tracking.

[0106] It can be understood that the world coordinates are used in the process of tracking the cotton top and dynamically topping the cotton top in this embodiment, and the origin of the world coordinate system is the location of the device point when the cotton laser topping system is powered on, and the coordinate axis directions are west, north, and sky respectively. Figure 4 、 Figure 5 , the X-axis points to the north, the Y-axis points to the west, and the Z-axis points to the sky.

[0107] 2. The multi-target tracking method of the present invention mainly includes target tracking, target detection, target matching, coordinate conversion, sliding window screening and other steps, such as Figure 3 As shown, the specific process is as follows:

[0108] 1) Target tracking:

[0109] The tracking queue is a double-ended queue, with targets joining from the end and leaving from the head. A global target ID counter, point_id, is also defined, with an initial value of 0. The tracking queue uses a double-ended queue structure, with targets joining from the end and leaving from the head, ensuring efficient management.

[0110] 2) Input video image frame: According to the timestamp t of the image frame, find the transformation matrix T whose timestamp is closest to the image frame t from the odometry transformation matrix queue odom_buf wc , T wc Represents the change from the camera coordinate system to the world coordinate system. Preferably, the Hungarian algorithm is used to match the target in each frame to generate a target tracking frame, and the target tracking frames are sorted in time sequence to construct a target tracking queue. Each designated reference point in the target tracking frame corresponds to a world coordinate.

[0111] More specifically, the position of the target in the world coordinate system is relatively fixed. Based on this, the pixel coordinates (2D coordinates) of the target can be obtained by image recognition, that is, the 2D coordinates of the specified reference point. Thereafter, the coordinates are converted based on the depth information and internal and external parameters of the image to obtain the camera coordinates (3D coordinates) of the specified reference point. It can be understood that multiple frames of images correspond to the camera coordinates of the corresponding specified reference points. Thereafter, each camera coordinate is combined with the real-time 6D pose obtained by the odometer to convert the corresponding camera coordinates to the world coordinate system. After matching with the Hungarian algorithm, the world coordinates of each tracking target in the tracking queue, that is, the world coordinates of each specified reference point, can be obtained.

[0112] 3) Object detector extracts the object detection frame:

[0113] The deep learning model YoloV8 is used to detect objects in the input image frames, generating 2D detection frames within the image. The dataset used in this implementation is derived from a series of cotton bud images collected in experimental fields. The dataset includes samples collected under both daytime natural light and nighttime supplementary lighting conditions. A total of approximately 8,000 frames were collected, covering a variety of scenarios, including different times of day, lighting conditions, and cotton plant density. Each frame includes depth information and a timestamp, making it suitable for multi-target 3D detection and tracking.

[0114] 4) Sort the target detection boxes:

[0115] When the topping machine moves forward, the target detection frames are arranged in descending order; when the topping machine moves backward, the target detection frames are arranged in ascending order.

[0116] 5) Calculate the camera coordinates and world coordinates of the center point of the target detection frame:

[0117] Calculate the target's coordinates in the camera coordinate system and convert them to the world coordinate system based on the 6D pose information and the odometry transformation matrix to ensure stable tracking of the target at different time points. Specifically:

[0118] Each detected target contains information such as the height and width of the target detection frame, the pixel coordinates u, v of the detection frame center, the camera coordinates p_c, the world coordinates p_w, the change matrix mTwc, the target ID point_id, etc. Therefore, the camera coordinates of the target can be calculated based on the target pixel coordinates, the image center pixel coordinates, the depth value and the focal length, satisfying:

[0119]

[0120] Where p_c is the camera coordinate of the preset detection position; u, v are the horizontal and vertical pixel coordinates of the target, fx, fy are the focal lengths of the camera on the x and y axes, cx, cy are the pixel coordinates of the image center on the x and y axes, and d is the depth value.

[0121] According to the camera coordinates and transformation matrix, the target is converted to the world coordinates to meet the following requirements:

[0122] p_w=T wc *p_c,

[0123] Where p_w is the world coordinate of the preset detection position; T wc T is the transformation matrix from the camera coordinate system to the world coordinate system at the corresponding moment. wc =[R,T]4x4, R represents the rotation matrix and T represents the translation vector.

[0124] 6) Target Matching

[0125] Calculate the Euclidean distance between the world coordinates of the preset detection position at the corresponding timestamp and the world coordinates of the specified reference point in the tracking queue. If the tracking queue is empty, assign the ID of the target in 5) to point_id and add them to the tracking queue in turn; if the tracking queue is not empty, perform Euclidean space similarity matching between the target in [5] and the target obj in the tracking queue.

[0126]

[0127] Among them, p_w is the world coordinate of the current detection target, obj.p_w is the world coordinate of the target obj in the tracking queue, and err_dist is the Euclidean distance between the two in the world coordinate system.

[0128] The current detection target and the tracking target in the tracking queue are traversed to calculate err_dist. When err_dist is less than the preset threshold, the two are considered to be the same target, that is, the target is determined to match. At this time, the pixel coordinates and camera coordinates of the tracking target obj in the tracking queue are updated to the pixel coordinates and camera coordinates corresponding to the current detection target, and the transformation matrix mT of the target obj is updated. wc Updated to T wc , the point_id of the target obj remains unchanged.

[0129] When err_dist is greater than the threshold, the current target cannot be matched with all targets in the tracking queue. Then, the current detection target is added to the tracking queue as a new tracking target, and a new point_id, pixel coordinates, camera coordinates, world coordinates, timestamp, etc. are assigned to the new tracking target.

[0130] Specifically, the pixel coordinates of the new tracking target are updated based on the body coordinates of the detection target, including: calculating the body coordinates of the detection target corresponding to the new tracking target, satisfying:

[0131]

[0132] Among them, p b mT is the coordinate of the detection target in the body coordinate system; wc is the transformation matrix from the camera coordinate system to the world coordinate system; p_c is the camera coordinate of the detection target; T wb It is the transformation matrix from the body coordinate system to the world coordinate system at the current moment.

[0133] Afterwards, the pixel coordinates of the new tracking target are calculated using the body coordinates of the detected target, satisfying:

[0134]

[0135] Among them, p uv are the pixel coordinates of the tracked target; cx and cy represent the image center in the camera intrinsics, and fx and fy represent the focal length in the camera intrinsics.

[0136] It is understandable that in order to ensure matching accuracy and stability, the matching judgment adopts a preset distance threshold. The setting of this threshold is combined with the characteristics of the cotton planting row structure in the agricultural scene: in the conventional cotton planting model, the minimum plant spacing of cotton in the same row is usually 7 cm. Therefore, in order to avoid the top buds between adjacent plants being mistaken for the same target, the matching distance threshold is actually set to 5 cm as a safety boundary to reduce the false matching rate.

[0137] 7) Eliminate targets that exceed the limit in the tracking queue through a sliding window:

[0138] Figure 4 It shows how the sliding window can be dynamically adjusted to ensure that objects within the camera's field of view are managed efficiently.

[0139] A sliding window method is used to filter the tracking queue and remove targets that are out of view. The tracking queue is traversed and the corresponding pixel coordinates are calculated based on the target's latest camera coordinates in the current image frame. If the pixel coordinates exceed the image boundary range set by the sliding window, the target is considered to have left the current field of view and removed from the tracking queue.

[0140] The sliding window is a two-dimensional window defined on the image pixel plane, and its range is determined by the camera's visible area and the set viewing distance threshold. Figure 4 As shown in the figure, during the topping process, the cotton top target is relatively stationary, while the camera is moving. The sliding window is used to dynamically judge and eliminate targets that are out of the field of view, ensuring that only targets within the sliding window are tracked, thereby reducing the probability of target tracking loss or mismatch, and supporting subsequent related devices to perform related operations such as topping on each tracked target. Specifically, the width of the window can be set to 600 pixels and the height can be set to 480 pixels. Accordingly, at the pixel coordinate p of the tracking target, the camera can be positioned at the top of the target. uv When the sliding window range [0, 0, 600, 480] is exceeded, the corresponding tracking target is removed from the tracking queue. The moving frequency of the sliding window can be set to 15 Hz, which is adapted to the aforementioned camera frame rate.

[0141] Figure 4In the figure, at the initial moment, i.e., t0, three cotton top targets with ID: 1, ID: 2, and ID: 3 are pushed into the tracking queue in the sliding window. At the next moment, i.e., t1, target ID: 1 leaves the sliding window range. At this time, target ID: 4 enters the sliding window range, so target ID: 4 is pushed into the tracking queue. At the same time, target ID: 1 is removed from the tracking queue. Targets ID: 2 and ID: 3 update their pixel coordinates, while the camera coordinates, transformation matrix, timestamp, and world coordinates remain unchanged.

[0142] As the camera continues to move forward, the sliding window field of view keeps changing, new targets enter the queue, and tracking targets that exceed the sliding window field of view are removed from the queue; this cycle repeats, and the field of view continues to slide in the forward direction. Through the sliding window mechanism, it is ensured that the tracking targets in the tracking queue can be continuously tracked within a certain period of time until the coordinates of the tracking targets exceed the sliding window range.

[0143] 3. Dynamic target screening operation method

[0144] (1) Initialize the target topping flag

[0145] Dynamic topping operation mainly involves the laser continuously ablating the cotton top during the movement of the topping vehicle. Figure 5 shown.

[0146] During the target tracking phase, the state of the newly added target topping flag in the tracking queue is false, that is, not topping the target.

[0147] During the dynamic topping stage, each time a cotton plant is toppled, the corresponding target topping flag is set to true.

[0148] (2) Specific process of dynamic topping

[0149] like Figure 6 As shown, the steps may include:

[0150] (1) Traversing the targets in the tracking queue, specifically traversing the tracking targets in the sliding window, so as to support the subsequent screening of targets with the operation flag being in the unoperated state as the targets to be operated, such as supporting the subsequent screening of targets that have not been topped as the targets to be topped;

[0151] (2) Determine whether the target has been topped. Specifically, the top flag of the target (obj) in the tracking queue can be used to determine whether the target has been topped. If the flag is true, it means that the target has been topped; if the flag is false, it means that the target has not been topped.

[0152] (3) Get the transformation matrix T of the current world in the odometry odom_buf wc;

[0153] (4) Calculate the coordinates p_c of the target at the current moment in the camera coordinate system;

[0154] (5) Calculate the XY galvanometer control quantities c_trx and c_try corresponding to the target coordinate p_c to control the laser output

[0155] (6) The laser output direction is controlled by the XY galvanometer control quantities c_trx and c_try, so that the laser emission direction is determined to support subsequent operations on the target to be operated. Here, the calibration accuracy of the camera and laser is that the X-galvanometer and Y-galvanometer are 1 meter in the depth direction (Z axis), and the recognition target and the laser focus target position are less than 5mm.

[0156] (7) Delay of 20ms, that is, the laser starts outputting laser after the direction is locked for 20ms to avoid target deviation caused by communication delay or insufficient stabilization time of the galvanometer. During ablation, the laser is controlled to continue ablation for 20ms.

[0157] (8) Repeat steps (1) to (7) for N times until the loop ends.

[0158] (9) After the cycle is completed, the laser is turned off and the target topping status flag is set to true. That is, after the operation is completed, the operation status of the operated target is marked to reflect that the target has been operated. Then, the next target is removed for topping.

[0159] In this implementation, the actions of tracking the cotton top (target) to generate a tracking queue and dynamically topping the cotton top are executed by two different threads. The tracking queue, as a global variable, can be read and written consistently across the two threads using a mutex lock. For example, in the corresponding threads, a mutex lock can be used to protect basic operations such as push_back (adding an element, corresponding to adding a new tracking target) and pop_front (removing an element, corresponding to removing a tracking target). This avoids complex data manipulation and IO operations, and does not introduce significant delays.

[0160] It is understood that when implementing the above-mentioned multi-target tracking-based operation method, it can be implemented using multi-target tracking-based operation equipment, which may include a target tracking device and a target screening operation device, wherein:

[0161] The target tracking device is used to detect and track targets in continuous multi-frame images based on the aforementioned multi-target tracking method;

[0162] The target screening operation device is used to screen the tracking target and perform operations based on the dynamic target screening operation method.

[0163] Here, the working device can further include: a visual inertial odometer, mainly used for obtaining the 6D pose of the camera; a laser, mainly used for controlling the laser output according to the camera coordinates of the target to be worked; and a carrying platform, mainly used for carrying the target tracking device, the laser, and / or the target screening working device. The carrying platform can be a multi-wheel self-propelled vehicle body, which can drive the components, devices, and / or equipment mounted thereon to move. Here, based on the cotton laser topping working scene of the present application, the corresponding working device can be a cotton laser topping machine.

[0164] It can be understood that the present application also provides a readable storage medium storing computer executable instructions for executing the multi-target tracking method as described above or the dynamic target screening working method as described above, and the readable storage medium can be run on a computer, an embedded device, an automated system, or a robot to realize target detection, target screening, and working control.

[0165] 4. Time series analysis of laser topping module and target tracking

[0166] A camera and two lasers of the cotton laser topping machine form a topping module, and the blue outer frame in the camera field of view range is the laser topping range limit frame. The limit frame is divided into left and right parts by the center red line, and the left and right parts correspond to the topping areas of the left and right lasers in a set of topping modules, as shown in FIG. 7. One laser corresponds to one tracking queue, the red frame is the target tracking frame, and the yellow frame is the target detection frame. Taking the target tracking situation of the left tracking queue in FIG. 7 as an example:

[0167] In FIG. 7(a), at t-1, there are 2 target tracking frames and 1 target detection frame on the left side; and there are 1 target tracking frame and 1 target detection frame on the right side; during this period, the topping machine continues to travel.

[0168] In FIG. 7(b), at t, there are 2 target tracking frames and 2 target detection frames on the left side; and there are also 2 target tracking frames and 2 target detection frames on the right side.

[0169] In FIG. 7(c), at t+1, there are 2 target tracking frames and 3 target detection frames on the left side, and the newly added target detection frame can be used as a new target tracking frame and added to the tracking queue; and there are 2 target tracking frames and 1 target detection frame on the right side.

[0170] It is understandable that during the movement of the topping machine, the cotton tops cannot be identified in every frame. In order to enable the laser to continuously ablate the tops of the same cotton plant to achieve topping operations, the present invention continuously tracks each cotton top in a time series, associates historical frame information, combines world coordinates and camera coordinates, and ensures accurate positioning and tracking of the target at different time points, thereby greatly improving the accuracy and efficiency of topping.

[0171] Through actual operation tests in the experimental field, the topping accuracy of the topping method using the multi-target tracking method of the present invention is 85.2%, while the topping accuracy of the dynamic topping using the traditional target tracking method is 78.6%. By comparison, it can be seen that the present invention improves the tracking accuracy and the topping accuracy.

Claims

1. A multi-target tracking method, characterized in that: The method comprises: (1) Target tracking: Using preset rules, the tracking target is extracted from multiple consecutive frames of images, a target tracking frame is generated, and a tracking queue is constructed; wherein, each designated reference point in the target tracking frame corresponds to a world coordinate; (2) Target detection: Performing target detection on the multiple frames of image using a preset detection algorithm, generating target detection frames, and constructing a target detection frame queue arranged in time series, wherein the arrangement direction of the target detection frame queue matches the queue sorting direction of the tracking queue; (3) Coordinate transformation: Calculate the pixel coordinates of the preset detection position in each target detection frame, and obtain the corresponding camera coordinates by combining the camera intrinsic parameters, depth information and timestamp; Convert the preset detection position to world coordinates based on the camera coordinates and the odometry transformation matrix; (4) Target matching: Calculate the Euclidean distance between the world coordinates of the preset detection position at the corresponding timestamp and the world coordinates of the specified reference point in the tracking queue; If the Euclidean distance is less than or equal to the preset threshold, the match is considered successful; (5) Pixel coordinate update: Obtaining the tracking target corresponding to each detected target, and updating the pixel coordinates of each detected target as the pixel coordinates of the corresponding tracking target to obtain updated pixel coordinates of the tracking target; adding the detected targets that are determined to be unmatched as new tracking targets to the tracking queue, and updating the pixel coordinates of the new tracking targets based on the body coordinates of the corresponding detected targets; (6) Sliding window screening: The sliding window algorithm is used to dynamically eliminate tracking targets whose pixel coordinates are beyond the sliding window range.

2. The multi-target tracking method according to claim 1, wherein: The target tracking step includes: The Hungarian algorithm is used to match the target in each frame image and generate the target tracking frame; the target tracking frame is sorted according to the time series to build the target tracking queue; The target detection step includes: The target detection model is used to perform target detection on multiple frames of images and generate target detection frames.

3. The multi-target tracking method according to claim 1, wherein: The coordinate conversion step comprises: (1) Get the transformation matrix: According to the image timestamp, extract the transformation matrix from the camera coordinate system to the world coordinate system that matches the corresponding timestamp from the odometry transformation matrix queue; (2) Calculate the camera coordinates: The camera coordinates of the target are calculated based on the target pixel coordinates, the image center pixel coordinates, the depth value and the focal length, satisfying: Where p_c is the camera coordinate of the preset detection position; u, v are the horizontal and vertical pixel coordinates of the target respectively, fx, fy are the focal lengths of the camera on the x and y axes respectively, cx, cy are the pixel coordinates of the image center on the x and y axes respectively, and d is the depth value; (3) Calculate world coordinates: According to the camera coordinates and transformation matrix, the target is transformed into the world coordinate system to meet the following requirements: p_w=T wc *p_c, Where p_w is the world coordinate of the preset detection position in the target detection frame; T wc is the transformation matrix from the camera coordinate system to the world coordinate system at the corresponding moment.

4. The multi-target tracking method according to any one of claims 1 to 3, wherein: The target matching step includes: Calculate the Euclidean distance between the world coordinates of the preset detection position at the corresponding timestamp and the world coordinates of the specified reference point in the tracking queue, satisfying: Where p_w is the world coordinate of the preset detection position in the target detection frame, obj.p_w is the world coordinate of the preset tracking position in the tracking queue, and err_dist is the Euclidean distance between the two in the world coordinate system; If err_dist ≤ the preset threshold, the target is determined to be matched, otherwise the corresponding detected target is used as a new tracking target in the tracking queue.

5. The multi-target tracking method according to any one of claims 1 to 3, wherein: Updating the pixel coordinates of the new tracking target based on the body coordinates of the detection target includes: (1) Calculate the body coordinates of the detection target corresponding to the new tracking target, satisfying: Among them, p b mT is the coordinate of the detection target in the body coordinate system; wc is the transformation matrix from the camera coordinate system to the world coordinate system; p_c is the camera coordinate of the detection target; T wb is the transformation matrix from the body coordinate system to the world coordinate system at the current moment; (2) Use the body coordinates of the detected target to calculate the pixel coordinates of the new tracking target, satisfying: Among them, p uv are the pixel coordinates of the new tracking target; cx and cy represent the image center in the camera intrinsic parameters, and fx and fy represent the focal length in the camera intrinsic parameters.

6. A method for dynamic target screening, characterized in that: The method comprises: (1) Target screening: The tracking target coordinate information obtained based on the method according to any one of claims 1 to 5, and the operation flag of each tracking target, wherein the operation flag is used to reflect whether the tracking target has been operated; Traverse the tracking targets in the sliding window to filter out the targets with the operation flag bit in the unoperated state as the targets to be operated; (2) Target positioning: Get the camera coordinates of the target to be operated; (3) Equipment control: Based on the camera coordinates of the target, the working direction of the working equipment is adjusted and the work is performed.

7. The dynamic target screening operation method according to claim 6, characterized in that: The operation equipment includes a laser, and the method further includes: (1) Calculate the laser control amount: The XY galvanometer control quantity of the laser is calculated using the camera coordinates of the target to be operated and the odometer transformation matrix; (2) Control laser output: Adjust the emission direction of the laser according to the XY galvanometer control value to operate on the target; (3) Target job status update: After the job is completed, the job status of the target is marked to reflect that the target has been worked on.

8. An operating device based on multi-target tracking, characterized in that: include: (1) A target tracking device for detecting and tracking targets in a plurality of consecutive frames of images based on the multi-target tracking method according to any one of claims 1 to 5; (2) A target screening operation device for screening and operating the tracking target based on the dynamic target screening operation method according to any one of claims 6 to 7.

9. The operating equipment according to claim 8, characterized in that: Also includes: (1) Visual inertial odometry, used to obtain the camera's 6D pose; (2) a laser, used to control the laser output according to the camera coordinates of the target to be operated; (3) A transport platform for carrying the target tracking device, the laser and / or the target screening device.

10. A readable storage medium storing computer-executable instructions, wherein the instructions are used to execute the multi-target tracking method described in any one of claims 1 to 5 or the dynamic target screening operation method described in any one of claims 6 to 7, and can be run on a computer, an embedded device, an automation system or a robot to achieve target detection, target screening and operation control.