Target tracking method, device and system based on vehicle end image, vehicle and medium

By extracting the preset frame number of the vehicle-side image as the detection frame, using the pre-trained model for object detection and motion trajectory calculation, the problem of large hardware computing power overhead is solved by frame-by-frame detection and tracking, and efficient and real-time target tracking is achieved.

CN120013997AActive Publication Date: 2025-05-16CHONGQING CHANGAN AUTOMOBILE CO LTD

Patent Information

Application Number
CN202510085122.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

Various models are used to detect and track the video images frame by frame, which is expensive to calculate the vehicle hardware, resulting in high production costs.

Method used

By obtaining the car end image of the preset frame number, extracting several frames as detection frames, using the pre-trained object detection model to detect the detection frame, calculating the target motion trajectory of the target, and calculating the coordinates and bounding box size of the target in the interval frame based on the motion trajectory, and finally projecting the bounding box into the interval frame to achieve target tracking.

Benefits of technology

It effectively reduces the amount of calculation required for target tracking, improves tracking accuracy and real-time, reduces hardware overhead, and reduces production costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013997A_ABST
    Figure CN120013997A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent driving, and discloses a vehicle end image-based target tracking method, device and system, a vehicle and a medium, and the method comprises the steps: extracting a plurality of detection frames from a preset frame number of vehicle end images, carrying out the target detection of each detection frame through a target detection model, and carrying out the target tracking based on the coordinates of a current target in each detection frame; calculating the movement track of the current target, calculating the coordinates of the current target in all interval frames based on the movement track of the current target, and calculating the coordinates of the current target in all the interval frames based on the collection moment corresponding to the current interval frame, the bounding box size of the current target in the detection frame adjacent to the current interval frame and the corresponding collection moments; and calculating the size of the bounding box of the current target in the current interval frame, and finally projecting the bounding box into the current interval frame according to the coordinate of the current target in the current interval frame and the size of the bounding box to obtain the coordinates and bounding boxes of all targets in the preset frame number of vehicle end images, so that the calculation amount required by target tracking can be effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent driving technology, and in particular to a target tracking method, device, system, vehicle and medium based on vehicle-side images. Background Art

[0002] Multi-Object Tracking (MOT) technology plays a vital role in the fields of autonomous driving, intelligent transportation and security monitoring. Multi-object tracking of vehicle-side images refers to capturing scene information such as roads, pedestrians, and vehicles through on-board cameras during vehicle driving, tracking multiple target objects (such as pedestrians, vehicles, bicycles, etc.) in real time, and predicting their movement trajectories. This technology provides the necessary perception capabilities for autonomous driving, helping vehicles to plan paths and avoid obstacles safely and intelligently.

[0003] In the autonomous driving system, visual perception is one of the key technical modules. The vehicle can "see" the surrounding environment through the camera, identify and track dynamic objects on the road, such as other vehicles, pedestrians, bicycles, obstacles, etc. Multi-target tracking technology analyzes the image sequence of continuous frames, identifies the target and tracks it consistently, ensuring that all important targets are monitored and analyzed in complex traffic environments.

[0004] With the rapid development of deep learning technology, object detection algorithms based on convolutional neural networks (CNNs) (such as YOLO, SSD, and Faster RCNN) have significantly improved detection accuracy and speed, which provides more accurate input for multi-target tracking. At the same time, tracking methods based on deep learning are also gradually emerging, using deep neural networks to learn the appearance features and motion patterns of targets, which can better cope with complex scenes and target interactions.

[0005] However, the on-board camera is installed on a moving vehicle, which causes significant movement and perspective changes between image frames. Secondly, the scene complexity of the vehicle-side image is high, including fast-moving targets, complex backgrounds, lighting changes, dynamic shadows, etc. Finally, the real-time requirement is high. When the vehicle is driving at high speed, it needs to process a large amount of image data in real time, perform target detection and tracking, and make decisions quickly. Therefore, if you rely too much on pre-trained models, such as using various models to detect and track video images frame by frame, the vehicle hardware computing power overhead accounts for a large proportion, which will result in higher production costs. Summary of the invention

[0006] In view of this, the present invention provides a target tracking method, device, system, vehicle and medium based on vehicle-side images to solve the problem that frame-by-frame detection and tracking of video images using various models accounts for a large proportion of vehicle hardware computing power overhead and will result in high production costs.

[0007] In a first aspect, the present invention provides a target tracking method based on vehicle-side images, the method comprising: acquiring a preset number of vehicle-side images, and extracting a plurality of frames of vehicle-side images from the preset number of vehicle-side images as detection frames, wherein the preset number of vehicle-side images are images continuously collected at a preset frame rate within a preset period; performing target detection on each detection frame using a pre-trained target detection model, obtaining labeling results corresponding to all targets in each detection frame, and calculating a motion trajectory of the current target based on the coordinates of the current target in each detection frame, wherein the labeling results at least include the coordinates corresponding to the target and a bounding box used to label the target; based on the motion trajectory of the current target , calculate the coordinates of the current target in all interval frames, where the interval frames are vehicle-side images of the preset number of frames, excluding the detection frames; calculate the size of the bounding box of the current target in the current interval frame based on the acquisition time corresponding to the current interval frame, the size of the bounding box of the current target in the detection frames adjacent to the current interval frame, and the corresponding acquisition times; project the bounding box into the current interval frame according to the coordinates of the current target in the current interval frame and the size of the bounding box, traverse all interval frames, detection frames and all targets, obtain the coordinates and bounding boxes of all targets in the vehicle-side images of the preset number of frames, and obtain the target tracking result based on the coordinates and bounding boxes of the target.

[0008] The target tracking method based on vehicle-side images provided in the present embodiment obtains vehicle-side images of a preset number of frames, and extracts several frame images from the vehicle-side images of the preset number of frames as detection frames, performs target detection on each detection frame using a target detection model, obtains the labeling results corresponding to all targets in each detection frame, and calculates the motion trajectory of the current target based on the coordinates of the current target in each detection frame, and then calculates the coordinates of the current target in all interval frames based on the motion trajectory of the current target, calculates the size of the bounding box of the current target in the current interval frame based on the acquisition time corresponding to the current interval frame, the size of the bounding box of the current target in the detection frames adjacent to the current interval frame, and the corresponding acquisition times, and finally, according to the coordinates of the current target in the current interval frame and the size of the bounding box, the bounding box is projected into the current interval frame to obtain the coordinates and bounding boxes of all targets in the vehicle-side images of the preset number of frames. By adopting a target detection method that mainly uses mathematical calculations and is supplemented by models, the amount of calculation required for target tracking can be effectively reduced, and higher tracking accuracy can be achieved, thereby improving the real-time performance of tracking.

[0009] In an optional embodiment, the calculation of the motion trajectory of the current target based on the coordinates of the current target in each detection frame includes: obtaining the first coordinates of the current target in the initial detection frame, the initial detection frame being the frame image in which the current target is first detected; obtaining the second coordinates of the current target in the current detection frame, the current detection frame being any frame in all detection frames except the initial detection frame; calculating the translation matrix and rotation matrix of the current target in the current detection frame compared with the initial detection frame based on the first coordinates and the second coordinates; calculating the absolute coordinates of the current target in the current detection frame with the vehicle coordinate system of the current target in the initial detection frame as the absolute coordinate system based on the translation vector and the rotation matrix of the current target in the current detection frame compared with the initial detection frame and the second coordinates; calculating the motion trajectory of the current target based on the absolute coordinates of the current target in each detection frame.

[0010] The present invention takes into account that the coordinate system corresponding to the target is moving. By calculating the absolute coordinates of the current target in the same coordinate system in each detection frame and using the absolute coordinates to calculate the motion trajectory of the current target, the practicality and accuracy of the motion trajectory calculation can be improved.

[0011] In an optional embodiment, the absolute coordinates are three-dimensional coordinates, and the motion trajectory of the current target is calculated based on the absolute coordinates of the current target in each detection frame, including: fitting a motion trajectory curve based on the absolute three-dimensional coordinates of the current target in all detection frames including the second detection frame, and determining the predicted three-dimensional coordinates of the current target in all detection frames based on the fitted motion trajectory curve; calculating the weighted loss function value corresponding to each three-dimensional coordinate component based on the predicted three-dimensional coordinates and the absolute three-dimensional coordinates of the current target in all detection frames; summing the weighted loss function values ​​corresponding to each three-dimensional coordinate component to obtain the sum of the loss functions; adjusting the fitted motion trajectory curve until the sum of the loss functions is less than a preset loss function threshold, to obtain the adjusted motion trajectory of the current target.

[0012] The present invention fits the motion trajectory by the least square method and adjusts the motion trajectory by using a weighted loss function to obtain the optimal motion trajectory of the current target, thereby improving the accuracy of subsequent coordinate prediction of the current target in the interval frame.

[0013] In an optional embodiment, the coordinates of the current target in the current interval frame are absolute coordinates of the vehicle coordinate system of the current target in the initial detection frame as the absolute coordinate system, and the projecting of the bounding box into the current interval frame according to the coordinates of the current target in the current interval frame and the size of the bounding box includes: calculating the absolute coordinates of the current target in the current interval frame based on the motion trajectory of the current target, and calculating the translation matrix and rotation matrix of the current target in the current interval frame compared with the initial detection frame based on the absolute coordinates of the current target in the current interval frame and the first coordinates of the current target in the initial detection frame; calculating the third coordinates of the current target in the current interval frame using the translation matrix and rotation matrix of the current target in the current interval frame compared with the initial detection frame and the first coordinates of the current target in the initial detection frame; calculating the boundary point coordinates of the bounding box based on the size of the bounding box and the third coordinates according to the rule that the coordinates of the current target in the current interval frame are the midpoint coordinates of the bounding box; and projecting the bounding box into the current interval frame based on the boundary point coordinates of the bounding box.

[0014] The present invention realizes target position prediction by calculating the size of the bounding box of the current target in the current interval frame based on the size of the bounding box of the current target in the detection frame adjacent to the current interval frame, and projects the bounding box into the current interval frame based on the coordinates of the current target in the current interval frame.

[0015] In an optional implementation, the annotation results corresponding to all targets in each detection frame also include boundary point coordinates of a bounding box, and the method further includes: calculating a fourth coordinate of the current target in the current detection frame and a size of the first bounding box based on the motion trajectory of the current target; calculating the boundary point coordinates of the first bounding box based on the size of the first bounding box according to the rule that the fourth coordinate of the current target in the current detection frame is the midpoint coordinate of the bounding box; comparing whether the difference between the boundary point coordinates of the first bounding box and the boundary point coordinates of the corresponding bounding box of the current target detected in the current detection frame using the pre-trained target detection model is greater than a preset difference threshold; if the difference between the boundary point coordinates of the first bounding box and the boundary point coordinates of the corresponding bounding box of the current target detected in the current detection frame using the pre-trained target detection model is greater than the preset difference threshold, adjusting the boundary point coordinates of the corresponding bounding box of the current target detected in the current detection frame using the pre-trained target detection model using the difference between the boundary point coordinates of the first bounding box and the boundary point coordinates of the corresponding bounding box of the current target detected in the current detection frame using the pre-trained target detection model to obtain the final boundary point coordinates of the first bounding box in the current detection frame.

[0016] If the bounding box position of the current target in the current detection frame predicted by the motion trajectory does not match the bounding box position detected by the target detection model, the bounding box position can be appropriately corrected to obtain the final bounding box position of the current target in the current detection frame, thereby avoiding sudden changes in the bounding box tracking trajectory.

[0017] In an optional embodiment, the method further includes: if the current target is not detected in the current detection frame, still based on the motion trajectory of the current target, calculating the coordinates of the current target in the current detection frame; calculating the size of the bounding box of the current target in the current detection frame based on the acquisition time corresponding to the current detection frame, the size of the bounding box of the current target in the detection frame adjacent to the current detection frame, and the corresponding acquisition times; projecting the bounding box into the current detection frame according to the coordinates of the current target in the current detection frame and the size of the bounding box, until the duration for which the current target is not detected is greater than a preset duration threshold.

[0018] Even if the current target is not detected in the vehicle-side image of the current detection frame, the present invention still predicts the coordinates of the current target in the current detection frame according to the motion trajectory of the current target, and then draws a boundary box to ensure real-time tracking of the current target.

[0019] In an optional embodiment, the annotation results corresponding to all targets in each detection frame also include the depth of the target and the category of the target, and the method also includes: if the pre-trained target detection model does not detect the depth of the current target, based on the detected current target category, obtaining the actual feature size and the size of the bounding box of the current target; based on the actual feature size of the current target, the size of the bounding box and the shooting parameters of the image acquisition device for acquiring the vehicle-side image, calculating the depth of the current target.

[0020] If the present invention determines that the target detection model has not detected the depth of the current target, it can obtain the actual size of the current target based on the subdivision category of the current target, and calculate the depth of the current target according to the bounding box size of the current target in the detection frame and the camera focal length, so as to ensure the integrity of the target information in the detection frame.

[0021] In an optional embodiment, the annotation results corresponding to all targets in each detection frame also include an identifier of the target, and the method also includes: if no identifier information is detected for the current target in the current detection frame, obtaining all targets and their corresponding bounding boxes in adjacent detection frames; normalizing the size of the bounding box corresponding to the current target in the current detection frame and the size of the bounding boxes corresponding to all targets in adjacent detection frames to the same size; calculating the similarity between the picture in the bounding box corresponding to the current target and the picture in the bounding box of the same size corresponding to all targets in adjacent detection frames; marking the identifier of the corresponding target in the adjacent detection frame with the highest picture similarity and greater than a preset similarity threshold as the identifier of the current target in the current detection frame.

[0022] After the identifier information of the current target is not detected, the present invention calculates the similarity between the pictures in each boundary box in the adjacent detection frame and the picture in the boundary box of the current target, and marks the identifier of the corresponding target in the adjacent detection frame with the highest similarity and greater than a preset similarity threshold as the identifier of the current target in the current detection frame, thereby improving the accuracy of subsequent calculation of the motion trajectory according to the coordinates of the same target in all detection frames.

[0023] In a second aspect, the present invention provides a target tracking device based on vehicle-side images, the device comprising: a detection frame extraction module, used to obtain a preset number of vehicle-side images, and extract a plurality of frames of vehicle-side images from the preset number of vehicle-side images as detection frames, wherein the preset number of vehicle-side images are images continuously collected at a preset frame rate within a preset period; a target detection module, used to perform target detection on each detection frame using a pre-trained target detection model, obtain the labeling results corresponding to all targets in each detection frame, and calculate the motion trajectory of the current target based on the coordinates of the current target in each detection frame, the labeling results at least including the coordinates corresponding to the target and a bounding box used to label the target; a target coordinate calculation module, used to calculate the motion trajectory of the current target based on the current target The motion trajectory of the target is used to calculate the coordinates of the current target in all interval frames, where the interval frame is the vehicle-side image of the preset number of frames, except the detection frame; a bounding box determination module is used to calculate the size of the bounding box of the current target in the current interval frame based on the acquisition time corresponding to the current interval frame, the size of the bounding box of the current target in the detection frame adjacent to the current interval frame, and the corresponding acquisition times; a target tracking module is used to project the bounding box into the current interval frame according to the coordinates of the current target in the current interval frame and the size of the bounding box, traverse all interval frames, detection frames and all targets, obtain the coordinates and bounding boxes of all targets in the vehicle-side image of the preset number of frames, and obtain the target tracking result based on the coordinates and bounding boxes of the target.

[0024] In a third aspect, the present invention provides a target tracking system based on vehicle-side images, the system includes a controller, the controller includes: a memory and a processor, the memory and the processor are communicatively connected to each other, computer instructions are stored in the memory, and the processor executes the target tracking method based on vehicle-side images of the above-mentioned first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0025] In a fourth aspect, the present invention provides a vehicle, the vehicle comprising the target tracking system based on vehicle-side image according to the third aspect

[0026] In a fifth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the vehicle-side image-based target tracking method of the above-mentioned first aspect or any corresponding embodiment thereof.

[0027] The target tracking method based on vehicle-side images provided by the present invention obtains vehicle-side images of a preset number of frames, and extracts several frame images from the vehicle-side images of the preset number of frames as detection frames, performs target detection on each detection frame using a target detection model, obtains the labeling results corresponding to all targets in each detection frame, and calculates the motion trajectory of the current target based on the coordinates of the current target in each detection frame, and then calculates the coordinates of the current target in all interval frames based on the motion trajectory of the current target, calculates the size of the bounding box of the current target in the current interval frame based on the acquisition time corresponding to the current interval frame, the size of the bounding box of the current target in the detection frames adjacent to the current interval frame, and the corresponding acquisition times, and finally projects the bounding box into the current interval frame according to the coordinates of the current target in the current interval frame and the size of the bounding box, and obtains the coordinates and bounding boxes of all targets in the vehicle-side images of the preset number of frames. By adopting a target detection method that mainly uses mathematical calculations and is supplemented by models, the amount of calculation required for target tracking can be effectively reduced, and higher tracking accuracy can be achieved, thereby improving the real-time performance of tracking. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0029] Figure 1 is a flow chart of a target tracking method based on vehicle-side images according to an embodiment of the present invention;

[0030] Figure 2a is a schematic diagram of a VCS coordinate system of a vehicle according to an embodiment of the present invention;

[0031] Figure 2b is a schematic diagram of a camera coordinate system of a vehicle according to an embodiment of the present invention;

[0032] Figure 3 is a schematic diagram of a bounding box according to an embodiment of the present invention;

[0033] Figure 4 is a schematic diagram of the tracking state of the current target in the video picture according to an embodiment of the present invention;

[0034] Figure 5 is a flow chart of another target tracking method based on vehicle-side images according to an embodiment of the present invention;

[0035] Figure 6 is a specific flow chart of a target tracking method based on vehicle-side images according to an embodiment of the present invention;

[0036] Figure 7 is a structural block diagram of a target tracking device based on vehicle-side images according to an embodiment of the present invention;

[0037] Figure 8 is a structural block diagram of a target tracking system based on vehicle-side images according to an embodiment of the present invention;

[0038] Fig. 9 is a structural block diagram of a vehicle according to an embodiment of the present invention;

[0039] Fig.10 Schematic diagram of the hardware structure of the controller according to the embodiment of the present invention. DETAILED DESCRIPTION

[0040] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0041] Early multi-target tracking methods usually rely on the Tracking-by-Detection framework. This method first identifies potential targets in the image through a target detector, and then uses filters (such as Kalman filters) or data association algorithms (such as the Hungarian algorithm) to associate these detected targets in consecutive frames to achieve tracking. However, these methods often perform poorly when dealing with complex scenes such as occlusion, target deformation, illumination changes, and multi-target interactions.

[0042] In recent years, with the rapid development of deep learning technology, object detection algorithms based on convolutional neural networks (CNNs) (such as YOLO, SSD, and Faster RCNN) have significantly improved detection accuracy and speed, which provides more accurate input for multi-target tracking. At the same time, tracking methods based on deep learning are also gradually emerging, using deep neural networks to learn the appearance features and motion patterns of targets, which can better cope with complex scenes and target interactions.

[0043] Compared with traditional video surveillance, multi-target tracking in vehicle-side images presents special challenges. First, the on-board camera is installed on a moving vehicle, resulting in significant motion and perspective changes between image frames. Second, the scene complexity of vehicle-side images is high, including fast-moving targets, complex backgrounds, lighting changes, dynamic shadows, etc. Finally, the real-time requirements are high. When the vehicle is driving at high speed, it needs to process a large amount of image data in real time, perform target detection and tracking, and make decisions quickly. Therefore, if you rely too much on pre-trained models, such as using various models to detect and track video images frame by frame, the vehicle hardware computing power overhead accounts for a large proportion, which will result in higher production costs.

[0044] According to an embodiment of the present invention, an embodiment of a target tracking method based on vehicle-side images is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0045] In this embodiment, a target tracking method based on vehicle-side images is provided, which can be used in the above-mentioned controller. Figure 1 is a flow chart of a target tracking method based on vehicle-side images according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0046] Step S101, obtaining a preset number of vehicle-side images, and extracting a plurality of vehicle-side image frames from the preset number of vehicle-side images as detection frames.

[0047] Among them, the vehicle-side images with a preset frame number are images continuously collected at a preset frame rate within a preset period.

[0048] In the embodiments of the present invention, the commonly used detection targets on the vehicle side are mainly divided into dynamic targets and static targets. Static targets include obstacles, traffic signs, traffic cones, etc., and their tracking is relatively simple, mainly related to the motion trajectory of the vehicle. Dynamic targets include dynamic targets such as vehicles and pedestrians. The coordinate information of various targets during the movement of the vehicle is detected and calculated, and reflected in the bounding box information in the picture. This is just an example. The vehicle-side image with a preset frame number can be a vehicle-side image with a preset frame number continuously collected at a preset frame rate within a preset period, or it can be a road video collected during the driving process of the vehicle. For the video, frame-by-frame detection is often required. For example, a frame rate of 40 frames per second means that 40 frames of vehicle-side images need to be detected per second. Among them, the method for obtaining the vehicle-side image with a preset frame number can be collected by a sensor installed on the vehicle side, such as a camera or a video acquisition device, or it can be collected by an image acquisition device such as a camera on the road. This is just an example.

[0049] After acquiring a preset number of vehicle-end images, an embodiment of the present invention can extract several frames of vehicle-end images from the preset number of vehicle-end images as detection frames, wherein the method of extracting the detection frames is not limited, and several frames of vehicle-end images can be randomly extracted from the preset number of vehicle-end images, or several frames of vehicle-end images can be extracted from the preset number of vehicle-end images based on a preset frame interval as detection frames, wherein, generally, in order to ensure the accuracy of the motion trajectory, at least three frames of vehicle-end images can be extracted as detection frames, and the setting of the frame interval is not limited, and the corresponding interval can be set based on factors such as the accuracy of subsequent calculation of the target motion trajectory or reducing the cost of model detection. For example, the frame interval can be set to 10 frames, and the first frame of the camera video can be set as the starting frame, and one frame is extracted from every 10 frames as the detection frame, so that 4 frames of vehicle-end images can be extracted as detection frames, which is just an example.

[0050] Step S102, using the pre-trained target detection model to perform target detection on each detection frame, obtain the labeling results corresponding to all targets in each detection frame, and calculate the motion trajectory of the current target based on the coordinates of the current target in each detection frame.

[0051] The labeling result at least includes the coordinates corresponding to the target and a bounding box used to label the target.

[0052] In the embodiment of the present invention, a pre-trained multi-task target detection model can be used to perform target detection on each detection frame to obtain the labeling results corresponding to all targets in each detection frame, wherein the target results may include the coordinates of the target in each detection frame and the bounding box BBOX used to label the target. The commonly used coordinates in the intelligent driving system are as follows: Figure 2a The vehicle coordinate system (VCS) shown in Figure 2bThe camera coordinate system shown in the figure, the vehicle coordinate system is used for subsequent target tracking in the subsequent embodiments, and the shape of the bounding box BBOX is not limited, such as Figure 3 As shown, for example, a rectangular frame may be used to surround the target in the image, and its width and height respectively represent the maximum width and maximum height of the target, which is only an example.

[0053] The embodiment of the present invention does not limit the method for calculating the motion trajectory of the current target based on the coordinates of the current target in each detection frame. A mathematical method can be used to estimate the intermediate position through known coordinate points to form a smooth curve, that is, the motion trajectory of the current target. The motion trajectory of the current target can also be drawn using the plot function in MATLAB, where the horizontal coordinate of the motion trajectory of the current target can be the acquisition time corresponding to the frame image, and the vertical coordinate is the coordinate of the current target, such as Figure 4 As shown, the dotted curve represents the motion trajectory of the current target, t0, t1, t2, and t3 represent the time corresponding to each detection frame, and the target detection model can be used to detect the bounding box of the current target in each detection frame, which is only used as an example.

[0054] Step S103, based on the motion trajectory of the current target, calculate the coordinates of the current target in all interval frames.

[0055] The interval frames are vehicle-end images of frames other than the detection frames among vehicle-end images of a preset number of frames.

[0056] An embodiment of the present invention can determine other frame images in the vehicle-side images of all frames except the detection frame as interval frames, and then calculate the coordinates of the current target in all interval frames from the motion trajectory of the current target based on the acquisition time corresponding to the interval frame, as an example only.

[0057] Step S104, calculating the size of the bounding box of the current target in the current interval frame based on the acquisition time corresponding to the current interval frame, the size of the bounding box of the current target in the detection frame adjacent to the current interval frame, and the corresponding acquisition times.

[0058] The embodiment of the present invention can calculate the size of the bounding box of the current target in the current interval frame by using an interpolation algorithm based on the acquisition time corresponding to the current interval frame, the size of the bounding box of the current target in the detection frame adjacent to the current interval frame, and the corresponding acquisition time. Generally, the size of the bounding box in the current interval frame is similar to that in the detection frame closer to the current interval frame, such as Figure 4 As shown, t2.1, t2.2, and t2.3 are the acquisition moments corresponding to each interval frame. The bounding box size of the current target in all interval frames between t2 and t3 is calculated by using the bounding box size in the detection frame corresponding to the moments t2 and t3, which is only used as an example.

[0059] Step S105, according to the coordinates of the current target in the current interval frame and the size of the bounding box, the bounding box is projected into the current interval frame, all interval frames, detection frames and all targets are traversed, the coordinates and bounding boxes of all targets in the vehicle-side image of a preset number of frames are obtained, and the target tracking result is obtained based on the coordinates and bounding box of the target.

[0060] The embodiment of the present invention can use the coordinates of the current target in the current interval frame and the size of the bounding box to project the bounding box into the current interval frame. For example, the coordinates of the current target in the current interval frame are used as the coordinates of the vertices of the bounding box, and then the bounding box is drawn according to the size of the bounding box. Finally, all interval frames and detection frames are traversed to obtain the following: Figure 4 The coordinates and bounding boxes of the current target in all vehicle-side images are shown, and then the target tracking results are obtained, so that the vehicle-side can perform autonomous driving based on the target tracking results. This is just an example.

[0061] The target tracking method based on vehicle-side images provided in the present embodiment obtains vehicle-side images of a preset number of frames, and extracts several frame images from the vehicle-side images of the preset number of frames as detection frames, performs target detection on each detection frame using a target detection model, obtains the labeling results corresponding to all targets in each detection frame, and calculates the motion trajectory of the current target based on the coordinates of the current target in each detection frame, and then calculates the coordinates of the current target in all interval frames based on the motion trajectory of the current target, calculates the size of the bounding box of the current target in the current interval frame based on the acquisition time corresponding to the current interval frame, the size of the bounding box of the current target in the detection frames adjacent to the current interval frame, and the corresponding acquisition times, and finally, according to the coordinates of the current target in the current interval frame and the size of the bounding box, the bounding box is projected into the current interval frame to obtain the coordinates and bounding boxes of all targets in the vehicle-side images of the preset number of frames. By adopting a target detection method that mainly uses mathematical calculations and is supplemented by models, the amount of calculation required for target tracking can be effectively reduced, and higher tracking accuracy can be achieved, thereby improving the real-time performance of tracking.

[0062] In this embodiment, a target tracking method based on vehicle-side images is provided, which can be used in the above-mentioned controller. Figure 5 is a flow chart of a target tracking method based on vehicle-side images according to an embodiment of the present invention. Figure 5 As shown, the process includes the following steps:

[0063] Step S501, obtain a preset number of vehicle-side images, and extract a number of vehicle-side images from the preset number of vehicle-side images as detection frames. The preset number of vehicle-side images are images continuously collected at a preset frame rate within a preset period. For details, please refer to Figure 1 Step S101 of the illustrated embodiment will not be described in detail here.

[0064] Step S502, use the pre-trained target detection model to perform target detection on each detection frame, obtain the labeling results corresponding to all targets in each detection frame, and calculate the motion trajectory of the current target based on the coordinates of the current target in each detection frame. The labeling results at least include the coordinates corresponding to the target and the bounding box used to label the target.

[0065] In an optional embodiment, the annotation results corresponding to all targets in each detection frame also include the depth of the target and the category of the target. If the pre-trained target detection model does not detect the depth of the current target, the actual feature size and the size of the bounding box of the current target are obtained based on the detected current target category; the depth of the current target is calculated based on the actual feature size of the current target, the size of the bounding box and the shooting parameters of the image acquisition device that obtains the vehicle-side image.

[0066] The target detection model used in the embodiment of the present invention can be used to identify the common major target categories in the field of intelligent driving, including vehicles, pedestrians, signs and traffic lights, etc., and supports the subdivision of major categories. For example, vehicles can be divided into SUVs, sedans, tricycles, etc. The target detection model performs target detection on each detection frame and records the detected target information, mainly including the major target category, subdivided category, identifier (track_id), distance (or depth) and direction, etc., as an example only.

[0067] After the target detection model completes the detection of all detection frames, if it is determined that the target detection model has not detected the depth of the current target, it can be determined first whether there is one or more cameras in the vehicle that collect the vehicle-side image. If the vehicle has multiple cameras, the depth of the current target can be estimated based on the position of the bounding box of the current target in different calibrated cameras using the binocular ranging principle, as an example only; if the vehicle only uses a single camera to collect the vehicle-side image, because the actual sizes of targets under the same subdivision category are relatively close, the subdivision category of the current target can be obtained, and the actual feature size w of the current subdivision category can be determined, including length, width, height, etc., and then the bounding box size p of the current target in the detection frame can be obtained, and the depth of the current target can be calculated using the following formula:

[0068]

[0069] Wherein, d represents the depth of the current target, and f represents the shooting parameters of the image acquisition device, such as the focal length of the camera.

[0070] In the embodiment of the present invention, the actual feature size and the bounding box size features can be selected according to actual needs. For example, if the error in the height direction of the vehicle in this subdivided category is small, the feature can select the height in the actual feature size and the height in the bounding box size, or multiple feature sizes can be used to calculate the depth separately, and finally the weighted summation method is used to calculate the depth of the current target.

[0071] If the present invention determines that the target detection model has not detected the depth of the current target, it can obtain the actual size of the current target based on the subdivision category of the current target, and calculate the depth of the current target according to the bounding box size of the current target in the detection frame and the camera focal length, so as to ensure the integrity of the target information in the detection frame.

[0072] In an optional embodiment, the annotation results corresponding to all targets in each detection frame also include an identifier of the target. If the identifier information is not detected for the current target in the current detection frame, all targets in adjacent detection frames and their corresponding bounding boxes are obtained; the size of the bounding box corresponding to the current target in the current detection frame and the size of the bounding boxes corresponding to all targets in adjacent detection frames are normalized to the same size; the similarity between the picture in the bounding box corresponding to the current target and the picture in the bounding box of the same size corresponding to all targets in the adjacent detection frames is calculated; and the identifier of the corresponding target in the adjacent detection frame with the highest picture similarity and greater than a preset similarity threshold is marked as the identifier of the current target in the current detection frame.

[0073] In the embodiment of the present invention, the identifier track_id is a unique identifier of the target, and the same target can be identified in multiple detection frames. If the identifier information of the current target in the current detection frame is not detected, all targets and corresponding bounding boxes in the adjacent detection frames can be obtained, and the bounding box size of the current target in the current detection frame and the bounding box size corresponding to all targets in the adjacent detection frames can be normalized to the same size to improve the accuracy of similarity calculation. Then, the similarity between the picture in the bounding box corresponding to the current target and the picture in the bounding box of the same size corresponding to all targets in the adjacent detection frames can be calculated, and the identifier of the corresponding target in the adjacent detection frame with the highest picture similarity and greater than the preset similarity threshold is marked as the identifier of the current target in the current detection frame.

[0074] After the identifier information of the current target is not detected, the present invention calculates the similarity between the pictures in each boundary box in the adjacent detection frame and the picture in the boundary box of the current target, and marks the identifier of the corresponding target in the adjacent detection frame with the highest similarity and greater than a preset similarity threshold as the identifier of the current target in the current detection frame, thereby improving the accuracy of subsequent calculation of the motion trajectory according to the coordinates of the same target in all detection frames.

[0075] Specifically, obtain the first coordinate of the current target in the initial detection frame, where the initial detection frame is a frame image in which the current target is detected for the first time; obtain the second coordinate of the current target in the current detection frame, where the current detection frame is any frame in all detection frames except the initial detection frame; based on the first coordinate and the second coordinate, calculate the translation matrix and rotation matrix of the current target in the current detection frame compared to the initial detection frame; based on the translation vector and the rotation matrix and the second coordinate of the current target in the current detection frame compared to the initial detection frame, calculate the absolute coordinates of the current target in the current detection frame with the vehicle coordinate system of the current target in the initial detection frame as the absolute coordinate system; calculate the motion trajectory of the current target based on the absolute coordinates of the current target in each detection frame.

[0076] In the embodiment of the present invention, since the vehicle is moving, the origin of the corresponding VCS coordinate system is also moving, so it is necessary to determine the absolute coordinates of the target in each frame for calculating the motion trajectory of the target. The detection frame in which the current target is first detected can be used as the initial detection frame (wherein, the initial detection frame and the above-mentioned starting frame may be the same frame or may be different frames), and the VCS coordinate system of the current target in the initial detection frame can be used as the absolute coordinate system until the current target disappears. The absolute coordinates of the current target in subsequent frames all use this coordinate system as a reference to represent its position.

[0077] For example, if the initial detection frame is the start frame, the coordinates of the current target in the VCS coordinate system (absolute coordinate system) at time t0 (initial detection frame) are P0 = (x0, y0, z0), and the coordinates of the current target at t i The coordinates of the current detection frame in the VCS coordinate system corresponding to the moment are P i =(x i ,y i ,z i ), the current target is between t0 and t i The motion trajectory S can be expressed as the translation vector T(t i ) and the rotation matrix R(t i ) indicates that the current target is at t i The absolute coordinate of the moment is P ωi =R(t i )·P i +T(t i ), the absolute coordinates of the current target in each detection frame in the same coordinate system can be obtained, and the motion trajectory of the current target can be calculated based on the absolute coordinates of the current target in each detection frame.

[0078] The present invention takes into account that the coordinate system corresponding to the target is moving. By calculating the absolute coordinates of the current target in the same coordinate system in each detection frame and using the absolute coordinates to calculate the motion trajectory of the current target, the practicality and accuracy of the motion trajectory calculation can be improved.

[0079] Specifically, the absolute coordinates are three-dimensional coordinates, and the above step S502 includes:

[0080] Step S5021, fitting a motion trajectory curve based on the absolute three-dimensional coordinates of the current target in all detection frames including the second detection frame, and determining the predicted three-dimensional coordinates of the current target in all detection frames based on the fitted motion trajectory curve.

[0081] Step S5022, based on the predicted three-dimensional coordinates and absolute three-dimensional coordinates of the current target in all detection frames, calculate the weighted loss function value corresponding to each three-dimensional coordinate component.

[0082] Step S5023, summing up the weighted loss function values ​​corresponding to each three-dimensional coordinate component to obtain the sum of the loss functions.

[0083] Step S5024, adjusting the fitted motion trajectory curve until the sum of the loss functions is less than a preset loss function threshold, thereby obtaining an adjusted motion trajectory of the current target.

[0084] The embodiment of the present invention does not limit the method of calculating the motion trajectory of the current target by using the coordinates of the current target in each detection frame, and can use a nonlinear regression function to fit the motion trajectory. Due to the uncertainty of the motion of the current target, in the process of fitting the motion trajectory, it should be considered that the trajectory of the current target is greatly affected by the most recent coordinate position. The earlier the position, the smaller the influence on its trajectory equation. Therefore, a weighted least squares polynomial fitting method can be used to reduce the influence of early positions on the fitting effect, and the coordinate points of the current target in the interval frame can be more accurately predicted. During vehicle operation, the movement of objects is relatively simple, and the accuracy requirements for fitting are not high. In addition, due to the use of weighted least squares fitting, the distant coordinates have little effect on the fitting effect. Considering that high-order polynomials have a high computing consumption, the degree of the fitting function can be selected as a 3-4th order polynomial to cover most fitting requirements.

[0085] The embodiment of the present invention tracks each target appearing on the frame screen separately, and starts fitting the motion trajectory from the second detection frame in which a certain detection target appears. First, the component matrix and the loss weight matrix corresponding to each three-dimensional coordinate component and the target vector are constructed, wherein the component matrix constructed for the x component is:

[0086]

[0087] Among them, t k represents the detection frame corresponding to the kth moment, m represents the number of detection frames, and n represents the degree of the polynomial.

[0088] The loss weight matrix is:

[0089]

[0090] Among them, ω k is the weight corresponding to the coordinate point of the current target in the kth detection frame; coefficient component a x =(X T WX) - 1 X T Wx, the above method is also used for the y component and the z component, which will not be repeated here.

[0091] For the three-dimensional coordinate components, define an n-order polynomial to fit, and assume that the polynomials are:

[0092]

[0093]

[0094] Among them, a ix 、a iy 、a iz are the polynomial coefficients corresponding to the coordinate components; i represents the i-th order polynomial; t i Represents a parametric variable in a polynomial.

[0095] The embodiment of the present invention is based on the above-mentioned polynomial fitting motion trajectory curve, and based on the fitted motion trajectory curve, the predicted three-dimensional coordinates of the current target in all detection frames are determined, and then based on the predicted three-dimensional coordinates and absolute three-dimensional coordinates of the current target in all detection frames, the weighted loss function values ​​corresponding to each three-dimensional coordinate component are calculated, wherein the weighted loss function of a single coordinate component is defined as:

[0096]

[0097] Among them, L x (a x ) represents the weighted loss function value of the x component; a x =[a x0 ,a x1 ,a x2 ,…,a xn ] represents the coefficient vector to be optimized, a y and a z Similar; f x (t k ) represents the predicted coordinates of the x component, x k represents the actual coordinate of the x component; ω k The weight corresponding to the coordinate point of the current target in the k-th detection frame reflects the influence of the coordinate point of the current target in the k-th detection frame on the fitting process and is defined as:

[0098] ω k =e -α(k-2)

[0099] Among them, α>0 is a decay factor that controls the decreasing speed of the weight, indicating that the loss weights corresponding to different detection frames are negatively correlated with the acquisition order corresponding to the detection frames, that is, the earlier the position of the target point of the detection frame is acquired, the smaller the corresponding weight.

[0100] To fit the three coordinate components, the total loss function L(a) can be defined as the sum of the loss functions of each component:

[0101] L(a)=L x (a x )+L y (a y )+L z (a z )

[0102] The embodiment of the present invention can adjust the fitted motion trajectory curve until the sum of the loss functions is less than a preset loss function threshold, thereby obtaining an adjusted motion trajectory of the current target.

[0103] The present invention fits the motion trajectory by the least square method and adjusts the motion trajectory by using a weighted loss function to obtain the optimal motion trajectory of the current target, thereby improving the accuracy of subsequent coordinate prediction of the current target in the interval frame.

[0104] Step S503, based on the motion trajectory of the current target, calculate the coordinates of the current target in all interval frames. The interval frame is the vehicle-side image of the preset number of frames, excluding the detection frame. For details, please refer to Figure 1 Step S103 of the illustrated embodiment will not be described in detail here.

[0105] Step S504, based on the acquisition time corresponding to the current interval frame, the frame size of the current target in the detection frame adjacent to the current interval frame, and the corresponding acquisition time, calculate the size of the boundary box of the current target in the current interval frame. Figure 1 Step S104 of the illustrated embodiment will not be described in detail here.

[0106] Step S505, according to the coordinates of the current target in the current interval frame and the size of the bounding box, the bounding box is projected into the current interval frame, all interval frames, detection frames and all targets are traversed, the coordinates and bounding boxes of all targets in the vehicle-side image of a preset number of frames are obtained, and the target tracking result is obtained based on the coordinates and bounding box of the target.

[0107] Specifically, the coordinates of the current target in the current interval frame are absolute coordinates with the vehicle coordinate system of the current target in the initial detection frame as the absolute coordinate system. Step S505 includes:

[0108] Step S5051, calculate the absolute coordinates of the current target in the current interval frame based on the motion trajectory of the current target, and calculate the translation matrix and rotation matrix of the current target in the current interval frame compared with the initial detection frame based on the absolute coordinates of the current target in the current interval frame and the first coordinates of the current target in the initial detection frame.

[0109] Step S5052, calculating the third coordinate of the current target in the current interval frame using the translation matrix and rotation matrix of the current target in the current interval frame compared to the initial detection frame and the first coordinate of the current target in the initial detection frame.

[0110] Step S5053 , according to the rule that the coordinates of the current target in the current interval frame are the midpoint coordinates of the bounding box, based on the size of the bounding box and the third coordinates, the coordinates of the boundary points of the bounding box are calculated.

[0111] Step S5054: projecting the bounding box into the current interval frame based on the coordinates of the boundary points of the bounding box.

[0112] The embodiment of the present invention uses the fitted motion trajectory curve to calculate the absolute coordinates of the current target in the current interval frame, and then can inversely solve its coordinates in the VCS coordinate system of the current interval frame according to the absolute coordinates of the current interval frame and the motion trajectory of the vehicle from the initial detection frame to the current interval frame. The motion trajectory of the current target is known, and the absolute coordinates of the current target in the current interval frame are calculated. Then, based on the absolute coordinates of the current target in the current interval frame and the first coordinates of the current target in the initial detection frame, the translation matrix R of the current target in the current interval frame compared to the initial detection frame is calculated. 0→j and the rotation matrix t 0→j , where j represents the interval frame corresponding to the jth moment:

[0113]

[0114] The position vector of the current target in the initial detection frame is P0 = [X0, Y0, Z0, 1] T , that is, the position vector P in the VCS coordinate system of the current interval frame in the current interval frame j for:

[0115]

[0116] In the embodiment of the present invention, based on the rule that the coordinates of the current target in the current interval frame are the midpoint coordinates of the bounding box, the size of the bounding box and P j, calculate the coordinates of the boundary points of the bounding box, where the shape of the bounding box can be a rectangular box, a circle or an irregular shape, etc. The corresponding boundary points can be four vertices, boundary points corresponding to the diameter of the circle or randomly selected boundary points. Generally, the bounding box is a rectangular box, and the coordinates of the four vertices of the bounding box are calculated accordingly. Finally, based on the coordinates of the four vertices of the bounding box, the camera external parameter and internal parameter matrix are used to project the bounding box onto the screen of the current interval frame.

[0117] The present invention realizes target position prediction by calculating the size of the bounding box of the current target in the current interval frame based on the size of the bounding box of the current target in the detection frame adjacent to the current interval frame, and projects the bounding box into the current interval frame based on the coordinates of the current target in the current interval frame.

[0118] In an optional implementation, the annotation results corresponding to all targets in each detection frame also include boundary point coordinates of the bounding box, and based on the motion trajectory of the current target, the fourth coordinate of the current target in the current detection frame and the size of the first bounding box are calculated; according to the rule that the fourth coordinate of the current target in the current detection frame is the midpoint coordinate of the bounding box, based on the size of the first bounding box, the boundary point coordinates of the first bounding box are calculated; compare whether the difference between the boundary point coordinates of the first bounding box and the boundary point coordinates of the corresponding bounding box of the current target in the current detection frame detected by the pre-trained target detection model is greater than a preset difference threshold; if the difference between the boundary point coordinates of the first bounding box and the boundary point coordinates of the corresponding bounding box of the current target in the current detection frame detected by the pre-trained target detection model is greater than the preset difference threshold, then the boundary point coordinates of the corresponding bounding box of the current target in the current detection frame detected by the pre-trained target detection model are adjusted using the difference between the boundary point coordinates of the first bounding box and the boundary point coordinates of the corresponding bounding box of the current target in the current detection frame detected by the pre-trained target detection model, so as to obtain the final boundary point coordinates of the first bounding box in the current detection frame.

[0119] In the embodiment of the present invention, in addition to calculating the coordinates and bounding box of the target in the interval frame using the above-mentioned mathematical method, when the next detection frame appears in the order of the acquisition time, the motion trajectory curve of the current target can be used to calculate the coordinates of the current target in the current detection frame, and the size of the predicted bounding box in the current detection frame can be calculated according to the size of the bounding box of the current target in the adjacent detection frame. Finally, according to the rule that the coordinates of the current target in the current detection frame are the midpoint coordinates of the bounding box, the coordinates of the boundary points of the predicted bounding box are calculated. Then, the difference between the boundary point coordinates of the predicted bounding box and the boundary point coordinates of the corresponding bounding box of the current target in the current detection frame detected by the target detection model is compared to see whether it is greater than a preset difference threshold. If the difference is greater than the preset difference threshold, the difference between the boundary point coordinates of the predicted bounding box and the corresponding bounding box of the current target in the current detection frame detected by the target detection model can be used to adjust the boundary point coordinates of the corresponding bounding box of the current target in the current detection frame to obtain the final bounding box position of the current target in the current detection frame. As an example only, the interpolation method or the Kalman filtering method can also be used to correct the bounding box to obtain the final bounding box position of the current target in the current detection frame.

[0120] If the bounding box position of the current target in the current detection frame predicted by the motion trajectory does not match the bounding box position detected by the target detection model, the bounding box position can be appropriately corrected to obtain the final bounding box position of the current target in the current detection frame, thereby avoiding sudden changes in the bounding box tracking trajectory.

[0121] In an optional implementation, if the current target is not detected in the current detection frame, the coordinates of the current target in the current detection frame are calculated based on the motion trajectory of the current target; the size of the bounding box of the current target in the current detection frame is calculated based on the acquisition time corresponding to the current detection frame, the size of the bounding box of the current target in the detection frame adjacent to the current detection frame, and the corresponding acquisition times; the bounding box is projected into the current detection frame according to the coordinates of the current target in the current detection frame and the size of the bounding box, until the duration for which the current target is not detected is greater than a preset duration threshold.

[0122] In an embodiment of the present invention, if the current target is not detected in the current detection frame, wherein the current target is not detected in the current frame, for example, the current target is lost, blocked, or exceeds the range of the vehicle-side image screen, the motion trajectory of the current target is still used to calculate the coordinates of the current target in the current detection frame, and based on the acquisition time corresponding to the current detection frame, the size of the bounding box of the current target in the detection frame adjacent to the current detection frame, and the corresponding acquisition times, the size of the bounding box of the current target in the current detection frame is calculated, and the bounding box is projected into the current detection frame according to the coordinates of the current target in the current detection frame and the size of the bounding box, until after the time delay set by the system, the drawing of the bounding box of the current target is stopped, wherein the delay set by the system, that is, the duration of the continuous failure to detect the current target, can be set according to factors such as the driving state of the vehicle and road conditions.

[0123] Even if the current target is not detected in the vehicle-side image of the current detection frame, the present invention still predicts the coordinates of the current target in the current detection frame according to the motion trajectory of the current target, and then draws a boundary box to ensure real-time tracking of the current target.

[0124] In a specific embodiment, Figure 6As shown, a vehicle-side image of a preset number of frames is obtained, a starting frame and an interval frame number are determined, and the next frame after each interval frame number is determined as a detection frame to obtain a number of detection frames, and then a multi-task target detection model is used to detect the starting frame and the detection frame to obtain the labeling results of all targets in all detection frames. If it is determined that the target detection model cannot identify the depth of the current target, the depth of the current target is calculated based on the actual feature size of the current target, the size of the bounding box and the shooting parameters of the image acquisition device that obtains the vehicle-side image, the coordinates of the current target are obtained from all detection frames, and the motion trajectory curve of the current target is calculated according to the coordinates of the current target in all detection frames, wherein the method of calculating the motion trajectory curve of the current target can use the least squares method and the loss function to fit the motion trajectory of the current target, and then based on the motion of the current target The motion trajectory predicts the coordinates of the current target in all interval frames, and then calculates the size of the boundary box of the current target in the current interval frame based on the acquisition time corresponding to the current interval frame, the border size of the current target in the detection frame adjacent to the current interval frame, and the corresponding acquisition time. According to the coordinates of the current target in the current interval frame and the size of the boundary box, the boundary box is projected into the current interval frame until the next detection frame. In the next detection frame, the motion trajectory of the current target can also be used to calculate the coordinates of the current target. If the tracking target can no longer be detected in the current detection frame, the boundary box of the current target is still drawn based on the coordinates of the current target and the size of the boundary box until the time threshold is reached and the tracking of the current target is stopped. This method can be applied to pure visual intelligent driving systems, has less reliance on other perceptions, and can be widely popularized. In addition, the mathematical scheme can effectively reduce the amount of calculation required for tracking, and can have a higher tracking accuracy, improve the real-time performance of tracking. For detailed descriptions, please refer to the above embodiments, which will not be repeated here.

[0125] In this embodiment, a target tracking device based on vehicle-side images is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and will not be repeated hereafter. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0126] This embodiment provides a target tracking device based on vehicle-side images, such as Figure 7As shown, it includes: a detection frame extraction module 701, which is used to obtain a preset number of vehicle-side images, and extract a plurality of frames of vehicle-side images from the preset number of vehicle-side images as detection frames, wherein the preset number of vehicle-side images are images continuously collected at a preset frame rate within a preset period; a target detection module 702, which is used to perform target detection on each detection frame using a pre-trained target detection model, obtain the annotation results corresponding to all targets in each detection frame, and calculate the motion trajectory of the current target based on the coordinates of the current target in each detection frame, wherein the annotation results at least include the coordinates corresponding to the target and the bounding box used to annotate the target; a target coordinate calculation module 703, which is used to calculate the current target based on the motion trajectory of the current target. The coordinates of all interval frames, the vehicle-side images of the vehicle-side images with a preset number of interval frames, and the vehicle-side images of other frames except the detection frames; a bounding box determination module 704, which is used to calculate the size of the bounding box of the current target in the current interval frame based on the acquisition time corresponding to the current interval frame, the border size of the current target in the detection frame adjacent to the current interval frame, and the corresponding acquisition times; a target tracking module 705, which is used to project the bounding box into the current interval frame according to the coordinates of the current target in the current interval frame and the size of the bounding box, traverse all interval frames, detection frames and all targets, obtain the coordinates and bounding boxes of all targets in the vehicle-side images with a preset number of frames, and obtain the target tracking result based on the coordinates and bounding boxes of the target.

[0127] In some optional embodiments, the target detection module 702 includes: a first coordinate acquisition unit, used to acquire the first coordinates of the current target in the initial detection frame, the initial detection frame being the frame image in which the current target is first detected; a second coordinate acquisition unit, used to acquire the second coordinates of the current target in the current detection frame, the current detection frame being any frame other than the initial detection frame in all detection frames; a matrix calculation unit, used to calculate the translation matrix and rotation matrix of the current target in the current detection frame compared to the initial detection frame based on the first coordinate and the second coordinate; an absolute coordinate calculation unit, used to calculate the absolute coordinates of the current target in the current detection frame with the vehicle coordinate system in the initial detection frame as the absolute coordinate system based on the translation vector and rotation matrix and the second coordinate of the current target in the current detection frame compared to the initial detection frame; a motion trajectory calculation unit, used to calculate the motion trajectory of the current target based on the absolute coordinates of the current target in each detection frame.

[0128] In some optional embodiments, the motion trajectory calculation unit includes: a trajectory curve fitting subunit, which is used to fit the motion trajectory curve based on the absolute three-dimensional coordinates of the current target in all detection frames including the second detection frame, and determine the predicted three-dimensional coordinates of the current target in all detection frames based on the fitted motion trajectory curve; a loss function calculation unit, which is used to calculate the weighted loss function values ​​corresponding to each three-dimensional coordinate component based on the predicted three-dimensional coordinates and absolute three-dimensional coordinates of the current target in all detection frames; summing the weighted loss function values ​​corresponding to each three-dimensional coordinate component to obtain the sum of the loss functions; and a trajectory adjustment unit, which is used to adjust the fitted motion trajectory curve until the sum of the loss functions is less than a preset loss function threshold, thereby obtaining the adjusted motion trajectory of the current target.

[0129] In some optional embodiments, the coordinates of the current target in the current interval frame are absolute coordinates with the vehicle coordinate system of the current target in the initial detection frame as the absolute coordinate system, and the target tracking module 705 includes: a matrix calculation unit, which is used to calculate the absolute coordinates of the current target in the current interval frame based on the motion trajectory of the current target, and calculate the translation matrix and rotation matrix of the current target in the current interval frame compared with the initial detection frame based on the absolute coordinates of the current target in the current interval frame and the first coordinates of the current target in the initial detection frame; a third coordinate calculation unit, which is used to calculate the third coordinates of the current target in the current interval frame using the translation matrix and rotation matrix of the current target in the current interval frame compared with the initial detection frame and the first coordinates of the current target in the initial detection frame; a bounding box coordinate calculation unit, which is used to calculate the boundary point coordinates of the bounding box based on the size of the bounding box and the third coordinates according to the rule that the coordinates of the current target in the current interval frame are the midpoint coordinates of the bounding box; a bounding box projection unit, which is used to project the bounding box into the current interval frame based on the boundary point coordinates of the bounding box.

[0130] In some optional embodiments, the annotation results corresponding to all targets in each detection frame also include the coordinates of the boundary points of the bounding box, and the device also includes: a bounding box size calculation module, which is used to calculate the fourth coordinates of the current target in the current detection frame and the size of the first bounding box based on the motion trajectory of the current target; a boundary point coordinate calculation module, which is used to calculate the boundary point coordinates of the first bounding box based on the size of the first bounding box according to the rule that the fourth coordinates of the current target in the current detection frame are the midpoint coordinates of the bounding box; a bounding box comparison module, which is used to compare the boundary point coordinates of the first bounding box with the coordinates of the boundary points of the current target in the current detection frame detected by the pre-trained target detection model. whether the difference between the boundary point coordinates of the first boundary box and the boundary point coordinates of the corresponding boundary box of the current target detected by the pre-trained target detection model in the current detection frame is greater than a preset difference threshold; a boundary box adjustment module is used to adjust the boundary point coordinates of the corresponding boundary box of the current target detected by the pre-trained target detection model in the current detection frame by using the difference between the boundary point coordinates of the first boundary box and the boundary point coordinates of the corresponding boundary box of the current target detected by the pre-trained target detection model in the current detection frame, so as to obtain the final boundary point coordinates of the first boundary box in the current detection frame.

[0131] In some optional embodiments, the device also includes: a coordinate calculation module, which is used to calculate the coordinates of the current target in the current detection frame based on the motion trajectory of the current target if the current target is not detected in the current detection frame; a bounding box size calculation module, which is used to calculate the size of the bounding box of the current target in the current detection frame based on the acquisition time corresponding to the current detection frame, the bounding box size of the current target in the detection frame adjacent to the current detection frame, and the corresponding acquisition times; a bounding box projection module, which is used to project the bounding box into the current detection frame according to the coordinates of the current target in the current detection frame and the size of the bounding box, until the duration for which the current target is not detected is greater than a preset duration threshold.

[0132] In some optional embodiments, the annotation results corresponding to all targets in each detection frame also include the depth of the target and the category of the target, and the device also includes: a target category determination module, which is used to obtain the actual feature size and bounding box size of the current target based on the detected current target category if the pre-trained target detection model does not detect the depth of the current target; a target depth calculation module, which is used to calculate the depth of the current target based on the actual feature size of the current target, the size of the bounding box and the shooting parameters of the image acquisition device for acquiring the vehicle-side image.

[0133] In some optional embodiments, the annotation results corresponding to all targets in each detection frame also include an identifier of the target, and the device also includes: a target identifier detection module, which is used to obtain all targets and their corresponding bounding boxes in adjacent detection frames if no identifier information is detected for the current target in the current detection frame; a bounding box size normalization module, which is used to normalize the bounding box size corresponding to the current target in the current detection frame and the bounding box size corresponding to all targets in adjacent detection frames to the same size; a similarity calculation module, which is used to calculate the similarity between the picture in the bounding box corresponding to the current target and the picture in the bounding box of the same size corresponding to all targets in adjacent detection frames; an identifier marking module, which is used to mark the identifier of the corresponding target in the adjacent detection frame with the highest picture similarity and greater than a preset similarity threshold as the identifier of the current target in the current detection frame.

[0134] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0135] The target tracking device based on vehicle-side images in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0136] The embodiment of the present invention also provides a target tracking system based on vehicle-side images, such as Figure 8 As shown, there is a controller for executing the above-mentioned target tracking method based on vehicle-side images. For detailed description, please refer to the above-mentioned embodiment, which will not be repeated here.

[0137] The embodiment of the present invention further provides a vehicle, such as Fig. 9 As shown, the vehicle includes the above-mentioned vehicle-side image-based target tracking system.

[0138] See also Fig.10 , Fig.10 is a schematic diagram of the structure of a controller provided by an optional embodiment of the present invention, such as Fig.10As shown, the controller includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the controller, including instructions stored in or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple controllers can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Fig.10 A processor 10 is taken as an example.

[0139] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0140] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0141] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required by at least one function; the data storage area may store data created according to the use of the controller, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the controller via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0142] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0143] The controller also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Fig.10The example of connecting through bus is taken in the following.

[0144] The input device 30 can receive input digital or character information and generate key signal input related to the user settings and function control of the controller, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.

[0145] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0146] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A target tracking method based on vehicle-side images, characterized in that: The method comprises: Acquire vehicle-end images of a preset number of frames, and extract a plurality of frames of vehicle-end images from the vehicle-end images of the preset number of frames as detection frames, wherein the vehicle-end images of the preset number of frames are images continuously acquired at a preset frame rate within a preset period; Perform target detection on each detection frame using a pre-trained target detection model to obtain labeling results corresponding to all targets in each detection frame, and calculate the motion trajectory of the current target based on the coordinates of the current target in each detection frame, wherein the labeling results at least include the coordinates corresponding to the target and a bounding box used to label the target; Based on the motion trajectory of the current target, the coordinates of the current target in all interval frames are calculated, and the interval frames are vehicle-side images of other frames except the detection frame in the vehicle-side images of the preset number of frames; Calculate the size of the bounding box of the current target in the current interval frame based on the acquisition time corresponding to the current interval frame, the size of the bounding box of the current target in the detection frame adjacent to the current interval frame, and the corresponding acquisition times; According to the coordinates of the current target in the current interval frame and the size of the bounding box, the bounding box is projected into the current interval frame, and all interval frames, detection frames and all targets are traversed to obtain the coordinates and bounding boxes of all targets in the vehicle-side image of a preset number of frames, and the target tracking result is obtained based on the coordinates and bounding boxes of the target.

2. The method according to claim 1, characterized in that The calculating the motion trajectory of the current target based on the coordinates of the current target in each detection frame includes: Obtaining the first coordinate of the current target in the initial detection frame, where the initial detection frame is a frame image in which the current target is detected for the first time; Obtaining the second coordinate of the current target in the current detection frame, where the current detection frame is any frame other than the initial detection frame among all the detection frames; Based on the first coordinates and the second coordinates, calculating a translation matrix and a rotation matrix of the current target in the current detection frame compared to the initial detection frame; Based on the translation vector and rotation matrix of the current target in the current detection frame compared with the initial detection frame and the second coordinate, calculate the absolute coordinates of the current target in the current detection frame with the vehicle coordinate system of the current target in the initial detection frame as the absolute coordinate system; Based on the absolute coordinates of the current target in each detection frame, the motion trajectory of the current target is calculated.

3. The method according to claim 2, characterized in that The absolute coordinates are three-dimensional coordinates. Based on the absolute coordinates of the current target in each detection frame, the motion trajectory of the current target is calculated, including: Fitting a motion trajectory curve based on the absolute three-dimensional coordinates of the current target in all detection frames including the second detection frame, and determining the predicted three-dimensional coordinates of the current target in all detection frames based on the fitted motion trajectory curve; Based on the predicted 3D coordinates and absolute 3D coordinates of the current target in all detection frames, the weighted loss function values ​​corresponding to each 3D coordinate component are calculated; The weighted loss function values ​​corresponding to each three-dimensional coordinate component are summed to obtain the sum of the loss functions; The fitted motion trajectory curve is adjusted until the sum of the loss functions is less than a preset loss function threshold, thereby obtaining an adjusted motion trajectory of the current target.

4. The method according to claim 2, characterized in that: The coordinates of the current target in the current interval frame are absolute coordinates with the vehicle coordinate system of the current target in the initial detection frame as the absolute coordinate system, and the projecting of the bounding box to the current interval frame according to the coordinates of the current target in the current interval frame and the size of the bounding box includes: Calculating the absolute coordinates of the current target in the current interval frame based on the motion trajectory of the current target, and calculating the translation matrix and rotation matrix of the current target in the current interval frame compared with the initial detection frame based on the absolute coordinates of the current target in the current interval frame and the first coordinates of the current target in the initial detection frame; Calculate the third coordinate of the current target in the current interval frame by using the translation matrix and rotation matrix of the current target in the current interval frame compared to the initial detection frame and the first coordinate of the current target in the initial detection frame; According to the rule that the coordinates of the current target in the current interval frame are the midpoint coordinates of the bounding box, based on the size of the bounding box and the third coordinates, the coordinates of the boundary points of the bounding box are calculated; Based on the coordinates of the boundary points of the bounding box, the bounding box is projected into the current interval frame.

5. The method according to claim 1, characterized in that The annotation results corresponding to all the targets in each detection frame also include the coordinates of the boundary points of the bounding box. The method also includes: Based on the motion trajectory of the current target, calculating a fourth coordinate of the current target in the current detection frame and a size of the first bounding box; According to the rule that the fourth coordinate of the current target in the current detection frame is the midpoint coordinate of the bounding box, based on the size of the first bounding box, the coordinates of the boundary points of the first bounding box are calculated; Compare whether the difference between the boundary point coordinates of the first boundary box and the boundary point coordinates of the corresponding boundary box of the current target detected by the pre-trained target detection model in the current detection frame is greater than a preset difference threshold; If the difference between the boundary point coordinates of the first bounding box and the boundary point coordinates of the corresponding bounding box of the current target detected by using the pre-trained target detection model in the current detection frame is greater than a preset difference threshold, the boundary point coordinates of the first bounding box and the boundary point coordinates of the corresponding bounding box of the current target detected by using the pre-trained target detection model in the current detection frame are adjusted using the difference between the boundary point coordinates of the first bounding box and the boundary point coordinates of the corresponding bounding box of the current target detected by using the pre-trained target detection model to obtain the final boundary point coordinates of the first bounding box in the current detection frame.

6. The method according to claim 1, characterized in that The method further comprises: If the current target is not detected in the current detection frame, the coordinates of the current target in the current detection frame are still calculated based on the motion trajectory of the current target; Calculate the size of the bounding box of the current target in the current detection frame based on the acquisition time corresponding to the current detection frame, the size of the bounding box of the current target in the detection frame adjacent to the current detection frame, and the corresponding acquisition times; According to the coordinates of the current target in the current detection frame and the size of the bounding box, the bounding box is projected into the current detection frame until the duration for which the current target is not detected is greater than a preset duration threshold.

7. The method according to claim 1, characterized in that The annotation results corresponding to all the targets in each detection frame also include the depth of the target and the category of the target. The method also includes: If the pre-trained target detection model does not detect the depth of the current target, obtain the actual feature size and bounding box size of the current target based on the detected current target category; The depth of the current target is calculated based on the actual characteristic size of the current target, the size of the bounding box, and the shooting parameters of the image acquisition device for acquiring the vehicle-side image.

8. The method according to claim 1, characterized in that The labeling results corresponding to all the targets in each detection frame also include identifiers of the targets, and the method further includes: If the identifier information is not detected for the current target in the current detection frame, all targets and their corresponding bounding boxes in adjacent detection frames are obtained; Normalizing the size of the bounding box corresponding to the current target in the current detection frame and the sizes of the bounding boxes corresponding to all targets in adjacent detection frames to the same size; Calculate the similarity between the picture in the bounding box corresponding to the current target and the pictures in the bounding boxes of the same size corresponding to all targets in adjacent detection frames; The identifier of the corresponding target in the adjacent detection frame with the highest picture similarity and greater than a preset similarity threshold is marked as the identifier of the current target in the current detection frame.

9. A target tracking device based on vehicle-side images, characterized in that: The device comprises: A detection frame extraction module is used to obtain a preset number of vehicle-end images and extract a number of vehicle-end images as detection frames from the preset number of vehicle-end images, wherein the preset number of vehicle-end images are images continuously collected at a preset frame rate within a preset period; The target detection module is used to perform target detection on each detection frame using a pre-trained target detection model, obtain the labeling results corresponding to all targets in each detection frame, and calculate the motion trajectory of the current target based on the coordinates of the current target in each detection frame. The labeling results at least include the coordinates corresponding to the target and the bounding box used to label the target; A target coordinate calculation module, used to calculate the coordinates of the current target in all interval frames based on the motion trajectory of the current target, wherein the interval frames are vehicle-side images of the preset number of frames, excluding the detection frame; A bounding box determination module, used to calculate the size of the bounding box of the current target in the current interval frame based on the acquisition time corresponding to the current interval frame, the size of the bounding box of the current target in the detection frame adjacent to the current interval frame, and the corresponding acquisition times; The target tracking module is used to project the bounding box into the current interval frame according to the coordinates of the current target in the current interval frame and the size of the bounding box, traverse all interval frames, detection frames and all targets, obtain the coordinates and bounding boxes of all targets in the vehicle-side image of a preset number of frames, and obtain the target tracking result based on the coordinates and bounding boxes of the target.

10. A target tracking system based on vehicle-side images, characterized in that: The system includes a controller, which includes a memory and a processor. The memory and the processor are communicatively connected to each other. Computer instructions are stored in the memory. The processor executes the target tracking method based on vehicle-side images as described in any one of claims 1 to 8 by executing the computer instructions.

11. A vehicle, characterized in that: The vehicle includes the vehicle-side image-based target tracking system as described in claim 10.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the target tracking method based on vehicle-side images according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Target tracking method and device

    CN109859239A

  • Face tracking method and system based on multi-feature fusion

    CN112215155A

  • Multi-target tracking method and device, computer equipment and storage medium

    CN112529942A

  • Video-based target association method and device and readable storage medium

    CN114219828A

  • Structured target detection method and device, equipment and storage medium

    CN114663648A

Cited By

  • Prediction bounding box dynamic correction method based on multi-dimensional kinematics characteristics

    CN120976501A