A method, device and system for target tracking

The target motion trajectory is obtained through sensors and the target switching is automatically identified using neural networks and Kalman filtering technology. Combined with computer vision verification, the problem of tracking discontinuity after target switching is solved, achieving efficient and accurate target tracking.

CN113326719BActive Publication Date: 2025-07-08HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010427448.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-28
Filing Date
2020-05-19
Publication Date
2025-07-08
Estimated Expiration
2040-05-19

AI Technical Summary

Technical Problem

Existing target tracking methods cannot achieve continuous and accurate tracking when target switching (such as after suspicious persons change vehicles), resulting in discontinuous tracking and requires manual intervention, which occupies human resources.

Method used

The moving trajectory of the target is obtained through sensors, and the neural network and Kalman filtering technology are used to fuse multi-sensor data, automatically identify target switching and update tracking targets, and perform secondary verification with computer vision to ensure the continuity and accuracy of tracking.

Benefits of technology

It realizes automatic update of tracking targets after target switching, improves tracking efficiency, reduces calculation amount, ensures real-time and accuracy, and avoids resource utilization of manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113326719B_ABST
    Figure CN113326719B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device and system for target tracking, which are mainly applied to fields such as personnel tracking. Among them, the method includes: obtaining the movement trajectory of an initial tracking target (such as a pedestrian) and the movement trajectories of other targets (such as vehicles) in the scene where the initial tracking target is located through sensors; determining the switched tracking target according to the movement trajectory of the initial tracking target and the movement trajectories of other targets in the scene. The above method can help monitoring personnel master the whereabouts of the tracking target. When the tracking target changes the means of transportation (that is, the tracking target is switched), the switching and continuous tracking of the tracking target can also be realized. Compared with the prior art where manpower is used for analysis and judgment, the efficiency of target tracking can be greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent security technology, and particularly to a method, device and system for target tracking. More specifically, it is applicable to realizing the switching of a tracking target (such as switching the tracking target to a vehicle) after the tracked target (such as a suspicious person) changes the means of transportation (such as a car). Background Art

[0002] With the development of video surveillance technology, many cities at home and abroad have deployed high-precision intelligent cameras on various main roads to facilitate the monitoring of road information. In addition to warning criminals, the above-mentioned surveillance cameras are also one of the main tools for tracking suspicious persons and assisting in case detection.

[0003] The current automatic target tracking methods mainly include: (1) Single-lens target tracking: That is, tracking pedestrians or vehicles within the same camera. When the target disappears (such as being blocked), the target is re-tracked within the lens through the Person Re-identification (Person Re-ID, simply referred to as Re-ID) algorithm. (2) Cross-lens target tracking: That is, when the target leaves the shooting range of the current lens, the target is also identified in another lens through the Re-ID algorithm and re-tracked. Therefore, the current automatic tracking methods are limited to tracking the same target in each lens, which will result in discontinuous tracking, that is, the target will leave the monitor's field of view within a certain period of time. In addition, the environment (light, occlusion, and target posture, etc.) of the cross-lens scenario is very complex, and the accuracy and computational efficiency of the Re-ID algorithm still need to be considered.

[0004] In addition to the above-mentioned tracking scenarios for the same target, there are also cases of cross-target tracking in the actual scenario. The so-called cross-target tracking refers to the situation where the tracked target has switched during the monitoring process (such as the suspicious person as the initial tracked target changes the means of transportation), and it is necessary to track the switched target. Currently, for this application scenario of cross-target tracking, it is mainly through the monitor manually switching the target for tracking. For example, after the suspicious person gets on the bus, the monitor switches the tracked target from the person to the vehicle. This method realized through manual participation cannot guarantee the real-time nature of target tracking and occupies a lot of human resources.

[0005] In summary, there is currently a need for a method that can continuously and accurately track targets, so as to improve the efficiency of target tracking. Summary of the Invention

[0006] The present application provides a method, device and system for target tracking. After the tracked target changes the means of transportation (i.e., target switching occurs), the switching and continuous tracking of the tracked target can also be achieved, which enables the monitoring personnel to fully master the whereabouts of the tracked target and greatly improves the efficiency of target tracking.

[0007] The first aspect of the present application provides a method applied to target tracking. The method includes: obtaining the movement trajectories of the targets included in the scene where the first target is located through sensors, where the targets included in the scene where the first target is located include the first target and at least one other target other than the first target, and the first target is the initial tracked target; then determining a second target according to the movement trajectories of the first target and at least one other target, and using the second target as the new tracked target. It should be noted that the movement trajectory of a target within a period of time is composed of the positions of the target at each moment that constitutes this period of time. The position of the target can be represented by coordinates, and the coordinates can be coordinates in the east-north-up coordinate system or coordinates in the north-east-earth coordinate system. All embodiments of the present invention do not make specific limitations on the specific type of coordinate system. Through the above method, when the original tracked target switches (for example, the tracked target gets in the car), the video monitoring system automatically determines the updated tracked target, ensuring the continuity of the movement trajectory of the tracked target and not missing the position information of the tracked target in any time period, thus improving the tracking efficiency. In addition, the above method judges whether the original tracked target has a switching behavior through trajectory data, which reduces the computational amount compared with the method of directly using video data for intelligent behavior analysis and reduces the requirement for computing power.

[0008] In the above method, the "scene where the first target is located" refers to the real scene where the first target is located. Exemplarily, the range of the scene can be an area centered on the first target with a radius of 100 meters. The embodiments of the present application do not limit the specific range of the scene and shall be determined according to specific circumstances. It should be noted that the scene where the first target is located is constantly changing as the first target moves. When starting to track the first target, the trajectories of the first target and other surrounding targets are started to be obtained.

[0009] In a possible implementation, determining a second target based on the motion trajectory of a first target and the motion trajectories of at least one other target includes: determining a set of candidate targets, where the candidate targets are: the at least one other target, or among the at least one other target, the other targets whose distance from the first target is less than a preset threshold; for each candidate target, inputting the motion trajectory of the candidate target and the motion trajectory of the first target into a pre-trained first neural network to obtain the probability that the candidate target is the second target; determining the second target according to the probabilities of at least one candidate target. The above "candidate targets" include two cases: (1) all other targets in the scenario except the original tracking target; (2) other targets in the scenario whose distance from the original tracking target is less than the preset threshold. The above method inputs the motion trajectory of the first target and the motion trajectory of each candidate target into the pre-trained neural network respectively, and outputs the probability that each candidate target is suspected to be the second target, which ensures both accuracy and real-time performance.

[0010] Exemplarily, the first neural network can be a long short-term memory network (LSTM). Before using the first neural network, it needs to be trained. For example, some historical videos of people getting into a car can be manually selected, and the video data of people and the car in a period of time before getting into the car can be obtained to generate their trajectory data. Convert these trajectory data into trajectory feature pairs, label them as the training set, and train the first neural network. When using the first neural network, the input is the trajectory feature pairs of the targets included in the scene where the first target is located, and the output is the probability that each candidate target is suspected to be the second target.

[0011] It should be noted that in addition to inputting the trajectory data into the neural network, other artificial intelligence algorithms can also be used to determine the second target. For example, traditional classification models such as Support Vector Machine (SVM) can be used. In addition to the above-listed artificial intelligence algorithms, rules selected manually can also be used for judgment. Exemplarily, the distance between the first target and other targets and their corresponding speeds, etc. are obtained through the trajectory data, and the changes in distance and speed, etc. are used as reference indicators to determine the second target. In short, the embodiments of the present application do not make specific limitations on how to use the trajectory data to determine the second target.

[0012] In another possible implementation, determining the second target according to the obtained probability includes: when a first probability among the obtained probabilities is higher than a preset threshold, determining the target corresponding to the first probability as the second target. During the tracking process, the probability that each candidate target among other targets is the second target will be continuously calculated. At a certain moment, when the probability of a certain candidate target exceeds the preset threshold, it can be determined that the candidate target corresponding to this probability is the second target.

[0013] In another implementation, for each candidate target, inputting the motion trajectory of the first target and the motion trajectory of the candidate target into a pre-trained first neural network to obtain the probability that the candidate target is suspected to be the second target includes: for each candidate target, establishing at least one set of trajectory feature pairs based on the motion trajectory of the candidate target and the motion trajectory of the first target. Each set of trajectory feature pairs includes at least two trajectory feature pairs at consecutive moments. The trajectory feature pair at each moment includes the position and speed of the first target, the position and speed of other targets, and the included angle between the motion directions of the first target and the other target at this moment. Each candidate target establishes trajectory feature pairs with the first target respectively according to the above method, so that at least one set of trajectory feature pairs can be obtained. Inputting the at least one set of trajectory feature pairs into the first neural network can output the probability that each candidate target is the second target. Exemplarily, if you want to obtain the probability that a certain candidate target is the second target at the current moment, a set of feature pairs needs to be input. This set of feature pairs can include 10 feature pairs, which are the positions, speeds, and included angles of the other target and the first target at 10 moments before the current moment respectively. Inputting this set of trajectory feature pairs into the neural network can obtain the probability that the candidate target is the second target at the current moment. The above method establishes trajectory feature pairs between the trajectory data of the first target and the trajectory data of each other target respectively as the input of the neural network, and determines the second target from at least one other target by analyzing the feature relationship between the trajectory of each other target and the trajectory of the first target.

[0014] In another possible implementation, setting the moment when the first probability is higher than the preset threshold as the first moment, the method further includes: acquiring video frames for a period of time before and after the first moment, and the video frames include the first target; then inputting the video frames into a pre-trained second neural network, and determining a third target as the tracking target after switching according to the output result. The above video frames "including the first target" refer to the video frames in which the original tracking target appears in the video picture. Exemplarily, it can be a video frame of a person getting into a car taken from the side, or a video frame of the car door taken from the front. This implementation is a further verification based on trajectory judgment. When the third target and the second target are the same target, directly use this target as the target after switching; when the third target and the second target are not the same target, use the third target as the target after switching. The above method mainly uses computer vision methods to perform intelligent behavior analysis on video data to judge whether the first target has switched to the second target. On the basis of using trajectory judgment, relevant video data is extracted for secondary judgment, improving the accuracy of the final judgment.

[0015] In another possible implementation, the second neural network includes a convolutional neural network and a graph convolutional neural network. Determining a third target by inputting video data into the second neural network includes: inputting the video data into a pre-trained convolutional neural network to output the features and bounding boxes of all targets in the video data; constructing a graph model based on the features and bounding boxes of the targets included in the video data; inputting the graph model into a pre-trained graph convolutional neural network, and determining the third target as the switched tracking target according to the output result. By using the second neural network to extract the features of the targets in the video frames and generate the bounding boxes corresponding to the targets, a unique graph model is established, and then the graph convolutional neural network is used to judge the behaviors of the targets to determine the new switched target. By using the graph convolutional neural network, more spatio-temporal information is incorporated to improve the accuracy of judgment.

[0016] The second neural network includes a convolutional neural network and a graph convolutional neural network. The convolutional neural network is mainly used to extract the features of the targets in the video and generate the bounding boxes of the targets for constructing the graph model; the graph convolutional neural network is mainly used to judge whether the original tracking target has a cross-target behavior according to the constructed graph model. Before using each neural network, they need to be pre-trained. Exemplarily, for the convolutional neural network, the training set can be images with bounding boxes that have been manually labeled, including people and vehicles in the images. Exemplarily, for the graph convolutional neural network, a video of getting in the car needs to be manually selected and these videos are input into the above-mentioned convolutional neural network to extract the features of the targets and generate the bounding boxes, and the graph model is generated according to the method described above, and the graph model is labeled as the behavior of getting in the car, and this is used as the training set to train the graph convolutional neural network. It should be noted that in addition to using neural networks to judge whether there is a cross-target behavior of the targets in the video, traditional machine learning models can also be used for judgment, such as support vector machine (SVM), etc.

[0017] In another possible implementation, the sensor includes at least two sets of sensors, and the orientations of different sets of sensors are different. For each set of sensors in the at least two sets of sensors, a motion trajectory of each target corresponding to the set of sensors is generated according to the sensed data collected by the set of sensors, so as to obtain at least two motion trajectories of the target; the at least two motion trajectories of the target are fused to form a fused motion trajectory of the target. Among them, "orientation" refers to the position or direction of an object in the actual space. "The orientations of different sets of sensors are different" means that the actual physical positions of each set of sensors are far apart. Exemplarily, the distance between groups is at least two meters or more. Before fusing the motion trajectories of the target obtained from different orientations, it is necessary to associate the same target under different sensing ranges. Among them, "sensing range" refers to the spatial range that the sensor can sense, and the scene ranges that sensors in different orientations can sense are also different. Exemplarily, a clustering algorithm can be used to associate the trajectories (sequences of positions) of the first target under different sensing ranges, indicating that these trajectories (sequences of positions) belong to the same target (the first target), and then fusion is performed. When fusing the motion trajectories of different sets of sensors, the Kalman fusion method can be used. Exemplarily, at each moment, for the same target under different sensing ranges, there are multiple positions. A relatively optimal position is selected as the measurement position, and an estimated position is obtained by methods such as fitting. The estimated position and the measurement position are subjected to Kalman fusion to obtain the final position of the target at this moment. The above method fuses the motion trajectories of the target under different sensing ranges to form the final motion trajectory of the target. The camera perspectives at different positions are different, and the sensing ranges of radars at different positions are also different. When the target is blocked by a foreign object (such as a billboard) in a certain sensing range (perspective), sensors in other orientations can continue to provide the position data of the target, ensuring the continuity of the target trajectory. At the same time, due to uncontrollable factors such as the environment and light, if only one set of sensors is used, the target often gets lost. Therefore, fusing the trajectory data of multiple sets of sensors in different orientations improves the accuracy of the final motion trajectory of the target and also improves the efficiency of target tracking.

[0018] In another possible implementation, each group of sensors includes at least two types of sensors, namely a camera and at least one of the following two types of sensors: millimeter-wave radar and lidar, and the two types of sensors are in the same orientation. "A group of sensors in the same orientation" means that multiple types of sensors constituting the group are arranged at adjacent physical positions. For example, this group of sensors includes a camera and a millimeter-wave radar, and these two types of sensors are installed on the same utility pole on the street. In the embodiments of the present application, the distance between at least two types of sensors in the same group is not specifically limited, as long as the sensing ranges of the two types of sensors are roughly the same. It should be noted that "orientation" refers to the physical position of the sensor in the actual space, and "sensing range" refers to the spatial range that the sensor can sense. Exemplarily, when the sensor is a camera, its sensing range refers to the spatial range of the scene it can capture; when the sensor is a radar, its sensing range refers to the actual scene spatial range it can detect. For each type of sensor in the same group of sensors, a monitoring trajectory of the target corresponding to this type of sensor is generated according to the sensing data collected by this type of sensor, so as to obtain at least two monitoring trajectories of the target; the at least two monitoring trajectories are fused to form the motion trajectory of the target. Exemplarily, for the same target, the trajectory data collected by different types of sensors can be fused using an innovative Kalman fusion method: using a unified prior estimate value (the optimal estimate value at the previous moment), successively performing Kalman fusion with the measurement values of different types of sensors at the current moment to form the optimal estimate value at the current moment, and the optimal estimate value at each moment forms the final motion trajectory of the target. The above method fuses the data of at least two types of sensors at adjacent positions, improving the trajectory accuracy of the target. Moreover, if only the camera is used for tracking, it is inevitable to be affected by the external weather. By fusing the trajectory data provided by the radar, it can more ensure that the tracking target will not be lost and improve the tracking efficiency.

[0019] Through the above description, the target tracking method provided by the present application can automatically update the tracking target to the new target after the original tracking target is switched, enabling the monitoring personnel to always keep track of the whereabouts of the tracked object. In addition, the present application uses the motion trajectory data of the target to determine the new target after the original tracking target is switched, greatly reducing the calculation amount on the premise of ensuring accuracy. Further, a segment of video frame data is selected to establish a unique graph model, and a graph convolutional neural network is used to perform a secondary judgment on the switching behavior of the original tracking target, incorporating more spatio-temporal information and improving the reliability of target tracking.

[0020] Second aspect, the present application provides another method for target tracking, including: obtaining the motion trajectory of a first target through a sensor, where the first target is an initial tracking target; determining the moment when the motion trajectory of the first target disappears as a first moment, and obtaining video frames for a period of time before and after the first moment, where the video frames include the pictures in which the first target may be updated to a second target; inputting the video frames into a pre-trained neural network to determine the second target, and using the second target as the updated tracking target. The above method analyzes the characteristics of the trajectory to find the time point at which the original tracking target may switch behaviors (such as the tracking target getting in the car), and obtains relevant video data containing the pictures of the original tracking target according to the time information for behavior analysis to determine the new tracking target, which greatly reduces the computational amount compared with analyzing the behaviors of the original tracking target throughout the process.

[0021] In another possible implementation, determining the initial moment when the first target trajectory disappears as the first moment according to the motion trajectory of the first target includes: determining that there is no motion trajectory of the first target after the initial moment, and determining the initial moment as the first moment.

[0022] In another possible implementation, determining the second target according to the video frames and using the second target as the updated tracking target includes: inputting the video frames into a pre-trained second neural network to determine the second target, and using the second target as the updated tracking target. Using a neural network to detect behaviors in relevant videos improves the judgment accuracy.

[0023] In another possible implementation, the second neural network includes a convolutional neural network and a graph convolutional neural network. Inputting the video data into the second neural network to determine the second target includes: inputting the video data into a pre-trained convolutional neural network to output the features and bounding boxes of all targets in the video data; constructing a graph model according to the features of all targets in the video data and the bounding boxes; inputting the graph model into a pre-trained graph convolutional neural network, and determining the second target as the switched tracking target according to the output result. By using the second neural network to extract the features of the targets in the video frames and generate the bounding boxes corresponding to the targets, a unique graph model is established, and then the graph convolutional neural network is used to judge the behaviors of the targets to determine the new switched target. Using the graph convolutional neural network incorporates more spatio-temporal information and improves the judgment accuracy.

[0024] In another possible implementation, the sensor includes at least two groups of sensors, and the orientations of different groups of sensors are different. For each group of sensors among the at least two groups of sensors, a motion trajectory of a first target corresponding to the group of sensors is generated according to the sensing data collected by the group of sensors, so as to obtain at least two motion trajectories of the first target; the at least two motion trajectories of the first target are fused to form a motion trajectory of the fused first target. Among them, "orientation" refers to the position or direction of an object in the actual space. "The orientations of different groups of sensors are different" means that the actual physical positions of each group of sensors are far apart. Exemplarily, the distance is at least two meters or more. "Sensing range" refers to the spatial range that the sensor can sense, and the scene ranges that sensors in different orientations can sense are also different. Exemplarily, for a camera, "different sensing ranges" means shooting the target in the scene from different perspectives, and for a radar, it means detecting the target in the scene from different orientations. When fusing the motion trajectories of different groups of sensors, the Kalman fusion method can be used. Exemplarily, at each moment, there are multiple position data of the first target under different sensing ranges. A relatively optimal position is selected from them as the measurement position, and an estimated position is obtained by using methods such as fitting. The estimated position and the measurement position are subjected to Kalman fusion to obtain the final position of the first target at that moment. The above method fuses the motion trajectories of the first target from multiple perspectives to form the final motion trajectory of the first target. The camera perspectives at different positions are different, and the sensing ranges of radars at different positions are also different. When the first target is blocked by a foreign object (such as a billboard) in a certain sensing range (perspective), sensors in other orientations can continue to provide the position data of the target, ensuring the continuity of the first target's trajectory.

[0025] In another possible implementation, each group of sensors includes at least two types of sensors, namely a camera and at least one of the following two types of sensors: millimeter-wave radar and lidar, and the two types of sensors are in the same orientation. "A group of sensors in the same orientation" means that multiple types of sensors constituting the group are arranged at adjacent physical positions. For example, this group of sensors includes a camera and a millimeter-wave radar, and these two types of sensors are installed on the same utility pole on the street. In the embodiments of the present application, the distance between at least two types of sensors in the same group is not specifically limited, as long as the sensing ranges of the two types of sensors are approximately the same. It should be noted that "orientation" refers to the physical position of the sensor in the actual space, and "sensing range" refers to the spatial range that the sensor can sense. Exemplarily, when the sensor is a camera, its sensing range refers to the spatial range of the scene it can capture; when the sensor is a radar, its sensing range refers to the actual scene spatial range it can detect. For each type of sensor in the same group of sensors, a monitoring trajectory of the first target corresponding to this type of sensor is generated according to the sensing data collected by this type of sensor, so as to obtain at least two monitoring trajectories of the first target; the at least two monitoring trajectories are fused to form a motion trajectory of the first target. Exemplarily, the fusion trajectory can adopt an innovative Kalman fusion method: using a unified prior estimate value (the optimal estimate value at the previous moment), and sequentially performing Kalman fusion with the measurement values of different types of sensors at the current moment to form the optimal estimate value at the current moment, and the optimal estimate value at each moment forms the motion trajectory of the first target. The above method fuses the data of at least two types of sensors at adjacent positions, improving the trajectory accuracy of the first target. Moreover, if only the camera is used for tracking, it is inevitable to be affected by the external weather. By fusing the trajectory data provided by the radar, it can more ensure that the tracking target will not be lost and improve the tracking efficiency.

[0026] In a third aspect, the present application provides a device for target tracking, including: an acquisition module and a processing module; the acquisition module is used to acquire the sensing data of the targets included in the scene where the first target is located, the targets included in the scene where the first target is located include the first target and at least one other target except the first target, and the first target is the initial tracking target; the processing module is used to generate the motion trajectories of the first target and at least one other target according to the sensing data; the processing module is further used to determine a second target according to the motion trajectories of the first target and at least one other target, and use the second target as the updated tracking target.

[0027] In another possible implementation, the processing module is further configured to determine a set of candidate targets, where the candidate targets are: the at least one other target, or among the at least one other target, the other targets whose distance from the first target is less than a preset threshold; for each candidate target, input the motion trajectory of the first target and the motion trajectory of the candidate target into a pre-trained first neural network to obtain the probability that the candidate target is the second target; determine the second target according to the probability that the at least one candidate target is the second target.

[0028] In another possible implementation, the processing module is further configured to detect that a first probability among the probabilities that the at least one candidate target is the second target is higher than a preset threshold, and determine the target corresponding to the first probability as the second target.

[0029] In another possible implementation, the processing module is further configured to: for each candidate target, establish at least one set of trajectory feature pairs according to the motion trajectories of the candidate target and the first target, and each set of trajectory feature pairs includes the position and speed of the first target, the position and speed of the candidate target, and the included angle between the motion directions of the first target and the other target at at least two consecutive moments; input the at least one set of trajectory feature pairs into the first neural network, and output the probability that the candidate target is the second target.

[0030] In another possible implementation, when the moment when the first probability is higher than the preset threshold is the first moment, the processing module is further configured to select video frames before and after the first moment, and the video frames include the picture where the first target may be updated to the third target; input the video frames into a pre-trained second neural network, and determine the third target as the updated tracking target according to the output result.

[0031] In another possible implementation, the second neural network includes a convolutional neural network and a graph convolutional neural network. Specifically, the processing module is configured to input the video frames into the pre-trained convolutional neural network, and output the features and bounding boxes of the targets included in the video frames; construct a graph model according to the features and bounding boxes of the targets; input the graph model into the pre-trained graph convolutional neural network, and determine the third target as the switched tracking target according to the output result.

[0032] In another possible implementation, the sensor includes at least two groups of sensors, and the orientations of different groups of sensors are different. For each of the first target and the other targets, the processing module is specifically configured to: respectively generate at least two motion trajectories of the target according to the sensing data collected by the at least two groups of sensor modules; fuse the at least two motion trajectories of the target to form the motion trajectory of the target.

[0033] In another possible implementation, each group of sensors includes at least two types of sensors. The at least two types of sensors are a camera and at least one of the following two types of sensors: millimeter-wave radar and lidar, and the at least two types of sensors are in the same orientation. For each target included in the scene where the first target is located, the processing module is specifically configured to: respectively generate at least two monitoring trajectories of the target according to the sensing data collected by at least two types of sensor modules; fuse the at least two monitoring trajectories of the target to form the movement trajectory of the target.

[0034] In a fourth aspect, the present application provides another device for target tracking, including an acquisition module and a processing module. The acquisition module is configured to acquire sensing data of a first target through sensors; the processing module is configured to generate a movement trajectory of the first target according to the sensing data; determine an initial moment when the trajectory of the first target disappears as a first moment; acquire video data for a period of time before and after the first moment, where the video data includes a picture in which the first target may be updated to a second target; determine the second target according to the video data, and use the second target as the updated tracking target.

[0035] In another possible implementation, the processing module is further configured to determine that there is no movement trajectory of the first target after the initial moment, and determine the initial moment as the first moment.

[0036] In another possible implementation, the processing module is further configured to input the video data into a pre-trained second neural network to determine the second target, and use the second target as the updated tracking target.

[0037] In another possible implementation, the second neural network includes a convolutional neural network and a graph convolutional neural network. The processing module is further configured to input the video frame into the pre-trained convolutional neural network to output the features and bounding boxes of the targets included in the video data; construct a graph model according to the features and bounding boxes of the targets included in the video frame; input the graph model into the pre-trained graph convolutional neural network, determine the second target according to the output result, and use the second target as the updated tracking target.

[0038] In another possible implementation, the sensors include at least two groups of sensors, and the orientations of different groups of sensors are different. The processing module is specifically configured to: respectively generate at least two movement trajectories of the first target according to the sensing data collected by the at least two groups of sensor modules; fuse the at least two movement trajectories of the first target to form the movement trajectory of the first target.

[0039] In another possible implementation, each group of sensors includes at least two types of sensors, and the at least two types of sensors are cameras and at least one of the following two types of sensors: millimeter-wave radar and lidar, and the at least two types of sensors are in the same orientation. The processing module is specifically configured to: generate at least two monitoring tracks of the first target respectively according to the sensing data collected by at least two types of sensor modules; fuse at least two monitoring tracks of the first target to form the motion track of the first target.

[0040] In a fifth aspect, the present application provides a device for target tracking. The device includes a processor and a memory, wherein: computer instructions are stored in the memory; the processor executes the computer instructions to implement the method according to any one of the first aspect and its possible implementations described above.

[0041] In a sixth aspect, the present application provides a device for target tracking. The device includes a processor and a memory, wherein: computer instructions are stored in the memory; the processor executes the computer instructions to implement the method according to any one of the second aspect and its possible implementations described above.

[0042] In a seventh aspect, the present application provides a computer-readable storage medium. The computer-readable storage medium stores computer program code, and when it runs on a computer, it causes the computer to execute the method according to any one of the first aspect and its possible implementations described above. These computer-readable storages include but are not limited to one or more of the following: read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), Flash memory, electrically EPROM (EEPROM), hard disk drive (Hard drive).

[0043] In an eighth aspect, the present application provides a computer-readable storage medium. The computer-readable storage medium stores computer program code, and when it runs on a computer, it causes the computer to execute the method according to any one of the second aspect and its possible implementations described above. These computer-readable storages include but are not limited to one or more of the following: read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), Flash memory, electrically EPROM (EEPROM), hard disk drive (Hard drive).

[0044] In a ninth aspect, the present application provides a computer program product containing instructions, which, when running on a computer, causes the computer to execute the method described in any one of the above first aspect and possible implementation manners.

[0045] In a tenth aspect, the present application provides a computer program product containing instructions, which, when running on a computer, causes the computer to execute the method described in any one of the above second aspect and possible implementation manners. Description of the Drawings

[0046] Figure 1 It is a schematic diagram of an application scenario of the method for target tracking provided by an embodiment of the present application.

[0047] Figure 2 It is a schematic diagram of the system architecture of the method for target tracking provided by an embodiment of the present application.

[0048] Figure 3 It is a schematic flowchart of the method for target tracking provided by an embodiment of the present application.

[0049] Figure 4 It is another schematic flowchart of the method for target tracking provided by an embodiment of the present application.

[0050] Figure 5 It is a schematic flowchart of the method for fusing video trajectories and radar trajectories provided by an embodiment of the present application.

[0051] Figure 6 It is a position table of a certain target at different times and different perspectives provided by an embodiment of the present application.

[0052] Figure 7 It is a schematic diagram of the two-dimensional motion trajectories of the original tracking target and other targets provided by an embodiment of the present application.

[0053] Figure 8 It is a schematic diagram of the probability that each other target output by the LSTM neural network is a new target after switching at each moment provided by an embodiment of the present application.

[0054] Figure 9 It is a schematic diagram of the structure of the LSTM unit provided by an embodiment of the present application.

[0055] Figure 10(a) is a schematic diagram of the structure of the LSTM neural network provided by an embodiment of the present application.

[0056] Figure 10(b) is a schematic diagram of the feature pair provided by an embodiment of the present application.

[0057] Figure 11 It is a time period distribution diagram of the training samples for training the LSTM neural network provided by an embodiment of the present application.

[0058] Figure 12 It is a schematic flowchart diagram of a method for secondary judgment using computer vision provided by an embodiment of the present application.

[0059] Figure 13 It is a schematic structural flowchart diagram of a second neural network provided by an embodiment of the present application.

[0060] Figure 14 It is a schematic diagram of a bounding box of an object provided by an embodiment of the present application.

[0061] Figure 15 It is a schematic diagram of a graph model provided by an embodiment of the present application.

[0062] Figure 16 It is a schematic hardware structure diagram of an object tracking device provided by an embodiment of the present application. Specific embodiments

[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0064] Application scenario introduction

[0065] Figure 1 It is a schematic diagram of an application scenario of the object tracking method provided by the present application. Exemplarily, Figure 1There is a set of monitoring devices in front of and diagonally behind the tracking target 16. The front monitoring device 10 includes a monitoring camera 11 and a radar 12, and the diagonally rear monitoring device 13 includes a monitoring camera 14 and a radar 15. The sensing ranges (viewing angles) of the two sets of monitoring devices are different. Each set of monitoring devices will monitor and track the targets that appear within their own sensing ranges, and form the movement trajectory of the target in the world coordinate system. When the monitoring device is a monitoring camera, the sensing range refers to the scene range that the camera can capture; when the monitoring device is a radar, the sensing range refers to the space range that the radar can detect; the viewing angle refers to the direction in which the camera is shooting, and different viewing angles correspond to different sensing ranges. The movement trajectory refers to the position of the target over a period of time and can be represented by coordinates. According to the position calibration of the camera and the radar, the position of the target captured or detected at each moment is projected into the global coordinate system, so that the movement trajectory of the target can be formed. In order to improve the trajectory accuracy in a certain sensing range, in addition to using the monitoring camera to collect video data, a millimeter-wave radar (or lidar) is also installed at a physical position adjacent to the camera. Hereinafter, "radar" is directly used to generally represent "millimeter-wave radar or lidar". By fusing the radar data and video data within the same sensing range, a more accurate movement trajectory can be formed. In addition, in order to reduce the impact of occlusion on tracking and ensure the continuity of the trajectory, it is also necessary to fuse the movement trajectories in multiple sensing ranges, that is, to fuse the trajectory data collected by the monitoring device 10 and the monitoring device 13, so as to obtain the movement trajectory of the target. During the tracking process, the tracking target 16 is about to leave by bus 17. At this moment, since the tracking target 16 gets on the bus and the tracking target disappears, even if the movement trajectories of the tracking target 16 in different sensing ranges (multi-viewpoints) have been fused, it is impossible to continue real-time tracking because the tracking target 16 cannot be found under any lens. Therefore, the solution proposed in this application is: First, based on the movement trajectories of the tracking target 16 and the surrounding targets (such as bus 17 or more other targets), determine whether the tracking target 16 (the original tracking target) has switched behavior (switched means of transportation). If it is confirmed that the tracking target 16 has switched behavior, for example, the tracking target 16 gets on the bus 17, then switch the tracking target to the bus 17 until the original tracking target reappears in a certain monitoring scene picture. Further, after determining the new tracking target through the trajectory, in order to improve the judgment accuracy, video data including the picture where the switching behavior is suspected to occur will also be selected for behavior analysis to determine whether the original tracking target has indeed switched behavior.

[0066] The method for automatic tracking provided by this application can continuously track a target. Even if the target switches to a different means of transportation, the target after the switch can still be monitored and tracked, without missing any time period, greatly improving the convenience and reliability of personnel deployment and tracking.

[0067] System architecture introduction

[0068] Figure 2 It is a schematic diagram of the system architecture of an automatic tracking system provided by this application. As shown in the figure, the system includes a terminal node 21 (Terminal Node, TNode), an edge node 22 (Edge Node, ENode), and a server node 23 (Server Node, SNode). Each node can independently execute computing tasks, and the nodes can also communicate with each other through a network to complete the task distribution and result upload. The network transmission methods include wired transmission and wireless transmission. Among them, the wired transmission method includes data transmission in forms such as Ethernet and optical fiber, and the wireless transmission method includes broadband cellular network transmission methods such as 3G (Third generation), 4G (Fourth generation), or 5G (Fifth generation).

[0069] The terminal node 21 (TNode) can be used to fuse video trajectories and radar trajectories within the same sensing range. The terminal node 21 can be the camera itself or various processor devices with computing capabilities. Data collected by physically adjacent cameras and radars can be directly fused and calculated at the terminal node without network transmission, reducing bandwidth occupancy and latency. The edge node 22 (ENode) can be used to fuse trajectory data under different sensing ranges (viewpoints). The edge node 22 can be an edge computing box, including a switch, a storage unit, a power distribution unit, a computing unit, etc. The server node 23 (SNode) is mainly used to perform cross-target behavior judgment on the tracking target. The server node 23 can be a cloud server, which stores and calculates the data uploaded by the terminal node 21 and the edge node 22 to achieve task allocation. This application does not limit the type of cloud server device and the virtualization management method.

[0070] It should be noted that the tasks executed by each of the above computing nodes are not fixed. For example, the edge-side node 22 can also fuse video trajectories and radar trajectories within the same sensing range, and the server-side node 23 can also fuse motion trajectories under different sensing ranges (multi-view). The specific implementation methods include but are not limited to the following three: (1) Different groups of sensors (cameras and / or radars) directly transmit the collected sensing data to the server-side node 23, and the server-side node 23 performs all calculations and judgments. (2) Different groups of sensors (cameras and / or radars) directly transmit the collected sensing data to the edge-side node 22. The edge-side node 22 first fuses the sensing data of the same group of sensors, and then fuses the trajectory data of different groups to obtain the continuous motion trajectory of the target, and then transmits the calculation result to the server-side node 23. The server-side node 23 calculates and determines the second target based on the trajectory data (which can also be implemented by the edge-side node 22), and performs secondary verification based on the video data. (3) The video and radar data of the same group are first transmitted to the nearest end-side node 21. The end-side node 21 is responsible for fusing the video and radar trajectories within the same sensing range. Then, multiple end-side nodes 21 transmit the fused trajectory data to the edge-side node 22. The edge-side node fuses the trajectories of the same target under different sensing ranges (views) to obtain the continuous motion trajectory of the target. Then, the edge-side node 22 transmits the continuous motion trajectory of the target to the server-side node 23. The server-side node 23 determines the second target based on the trajectory data and performs secondary verification based on the video data. In short, the present invention application does not make specific limitations on which node specifically executes the computing function, which is mainly determined by user habits, network bandwidth, or the computing power of the hardware itself, etc.

[0071] Overall solution

[0072] Next, the overall process of the solution of the present invention will be introduced in combination with Figure 3 The method includes the following steps:

[0073] S31: Obtain the sensing data of the target included in the scene where the first target is located collected by the sensor. Among them, the target includes the first target and other targets, and the first target refers to the original tracking target. Exemplarily, the scene where the first target is located can be an area with a radius of 20 meters centered on the original tracking target. Exemplarily, if the first target is the tracking target 16, then other targets can be vehicles whose distance from the tracking target 16 is less than 20 meters. If the sensor is a surveillance camera, then the sensing data is the video data captured by the surveillance camera; if the sensor is a millimeter-wave radar, then the sensing data is the distance data between the target detected by the millimeter-wave radar and the radar; if the sensor is a lidar, then the sensing data is the position, speed, etc. data of the target detected by the lidar.

[0074] S32: Generate the motion trajectory of each target based on the sensing data collected by the sensor. Each target here refers to the first target and other targets in the scene where the first target is located. Exemplarily, when the sensor is a camera, after pre-calibration, the position of the target in the world coordinate system can be directly obtained from the video data; when the sensor is a radar, the position of the radar itself is fixed and known, and the distance between the measured target and the radar can be used to obtain the position of the target in the world coordinate system; when the sensor is a lidar, the coordinates of the lidar itself are also known, and the three-dimensional dot matrix data of the target can be obtained by emitting laser beams, and information such as the position and speed of the target can be directly obtained. The position of the target at each moment in the world coordinate system constitutes the motion trajectory of the target.

[0075] S33: Determine the second target based on the motion trajectories of the first target and other targets, and use the second target as the switched target. When the first target undergoes a switching behavior (cross-target), the trajectory of the original tracking target and the trajectories of other tracking targets are used to determine the second target, and the second target is used as the new tracking target for continuous tracking.

[0076] This solution takes into account the actual motion of the target during the tracking process, flexibly switches the tracking target, forms a tracking without dead angles and time gaps, and improves the tracking efficiency. Moreover, this solution uses trajectory data to determine the new tracking target. Compared with the existing technology that judges behaviors through computer vision methods throughout the process, the computational complexity of this solution is greatly reduced, and the real-time requirements can be better guaranteed.

[0077] Implementation details of the solution

[0078] After introducing the overall process of the solution, the specific implementation solution of the present invention will be shown in detail below. Exemplarily, the original tracking target is Tracking Target 16, as Figure 4 shown, the complete solution mainly includes five steps (some steps are optional):

[0079] Step S41: For each group of sensors, fuse the sensing data of video and radar in the group of sensors to obtain the motion trajectory of the target within the sensing range of the group of sensors.

[0080] At least two types of sensors can be included in the same group of sensors, and these two types of sensors are in the same orientation. Exemplarily, as Figure 1As shown by camera 11 and radar 12. The same orientation means being physically adjacent, for example, they can be installed on the same utility pole. In the figure, orientation 1 means that camera 11 and radar 12 are in adjacent physical positions, i.e., orientation 1. By performing Kalman fusion on the sensing data of the tracking target 16 collected by the two, the trajectory data of the tracking target 16 sensed and calculated from this position in orientation 1 can be obtained (i.e., the motion trajectory within sensing range 1). Similarly, the positions of camera 13 and radar 14 are also close (both in orientation 2). By fusing the sensing data of the two, the trajectory data of the target sensed from orientation 2 can also be obtained (i.e., the motion trajectory within sensing range 2). It should be noted that for a camera, the sensing range refers to the scene range that the camera can capture; for a radar, it refers to the scene range that the radar can detect. The sensing ranges of sensors in the same orientation (such as camera 11 and radar 12) are approximately the same.

[0081] The targets here include the original tracking target and other targets in the scene where the original tracking target is located. Generating the motion trajectory of the target means generating the corresponding trajectory for each target. Optionally, when collecting video data, in order to improve the accuracy of the video trajectory, a 3D detection method can be used to determine the centroid of the target. In addition, when the target is a vehicle, the local ground-contact feature or the front and rear wheels of the vehicle can also be used to determine the centroid to improve the accuracy of the video trajectory.

[0082] To improve the accuracy of the motion trajectory of the target sensed in a certain orientation, the trajectory data obtained from multiple types of sensors can be fused. For example, the trajectory data obtained through a camera and the trajectory data obtained through a millimeter-wave radar can be fused. In addition, the trajectory data of the camera and the lidar can also be fused, and even the data of the camera, millimeter-wave radar, and lidar can be fused. The types of sensors for fusion in the embodiments of the present application are not limited and depend on the specific situation.

[0083] Taking the motion trajectory of the tracking target 16 within sensing range 1 as an example, the trajectory data of camera 11 and radar 12 are fused to specifically show the fusion method provided by the embodiments of the present application. This method applies the computational idea of Kalman fusion. The trajectory of the tracking target 16 is composed of the positions at each moment. Fusing the trajectory data of the camera and the radar is to fuse the positions of the tracking target 16 provided by these two types of data at each moment to form the motion trajectory of the tracking target 16 within the sensing range. Before fusion, it is necessary to correspond the radar data (position) and video data (position) of the tracking target 16. Exemplarily, the echo image of the radar can be analyzed to generate the contour image of the targets included in the scene where the tracking target 16 is located. Combining the calibration information, the positions of each target monitored by the radar can be distinguished corresponding to which target in the video.

[0084] For the convenience of describing the specific process of fusion, the position of the tracking target 16 obtained by collecting data through the camera is named the video measurement position, and the position of the tracking target 16 obtained by collecting data through the radar is named the radar measurement position. Figure 5 It shows how the video data and radar data of the tracking target 16 are fused at time t. First, there is an optimal estimated position F at time t-1 t-1 , and based on this optimal estimated position F at the previous moment t-1 the predicted position E at time t can be predicted t . The specific prediction formula can be a formula derived based on experience. At time t, the video measurement position V can be obtained according to the video data collected by the camera t . The predicted position E at time t t , and the video measurement position V t are subjected to Kalman fusion to obtain the intermediate optimal estimated position M t . At the same time, the radar measurement position R of the target can be obtained according to the data obtained by the radar at time t t . The intermediate optimal estimated position M t and the radar measurement position R t are subjected to Kalman fusion to obtain the final optimal estimated position F at time t t . The optimal estimated position at each moment is calculated based on the optimal estimated position at the previous moment. The optimal estimated position at the initial moment can be selected as the video measurement position or the radar measurement position at the initial moment, or the position after fusing the two. Thus, the optimal estimated position obtained at each moment constitutes the movement trajectory of the tracking target 16 within the sensing range 1. It should be noted that the above fusion process can also be changed to fuse the radar data first and then the camera data. In short, the embodiments of the present application do not limit the type, quantity of the fused sensors and the order of fusion. In addition to the tracking target 16 (the first target), the movement trajectories of other targets such as the bus 17 are also obtained according to the above method.

[0085] It should be noted that in addition to applying the Kalman idea to fuse the video and radar trajectories, the simplest weighted average algorithm can also be used for fusion. Exemplarily, at time t = 1, the position A of the tracking target 16 can be obtained according to the video data, and the position B of the tracking target can be obtained according to the radar data. The position A and position B are directly subjected to weighted average calculation to obtain the final position of the tracking target 16 within the sensing range 1 at time t = 1. Optionally, the fusion strategy can stipulate that the trajectory data of targets more than 60 meters away from the sensor is mainly based on the radar (higher weight), and the trajectory data of targets within 60 meters of the sensor is mainly based on the video (higher weight).

[0086] The embodiments of the present application adopt an innovative Kalman filtering method to fuse the measurement position data provided by different sensors within the same sensing range, thereby improving the accuracy of the target motion trajectory. It should be noted that the fusion method in the embodiments of the present application does not simply perform Kalman fusion on the video predicted position and the video measurement position, and perform Kalman fusion on the radar predicted position and the radar measurement position, and then fuse the trajectories of these two types of sensors. The key point of the embodiments of the present invention application is to adopt a unified predicted position at the same moment, and then fuse the measurement positions of different sensors. Therefore, an intermediate optimal estimated position will be generated during this process. Such a fusion method enables the final position of the target at each moment to refer to the sensing data of multiple sensors, improving the accuracy of the trajectory.

[0087] It should be noted that step S41 is not necessary in the whole solution. In actual situations, there is no radar (or millimeter-wave radar) installed around the camera. In such a case, the trajectory data within a certain sensing range can be directly obtained from the video data collected by the camera without fusion. In short, step S41 is optional and is mainly determined by the actual situation of the application scenario and the needs of the monitoring personnel, etc.

[0088] Step S42: Fuse the motion trajectories of the target in different sensing ranges.

[0089] The target here includes the original tracking target and other targets in the scene where the original tracking target is located. After obtaining the motion trajectory of the target within a certain sensing range, due to the continuous movement of the target, it may be blocked by foreign objects (billboards, buses), or the target may directly leave the monitoring range, resulting in the interruption of the target's trajectory. In order to obtain the continuous motion trajectory of the target, it is first necessary to associate the same target from different perspectives, and then fuse the trajectories of the same target from different perspectives.

[0090] Because there are multiple targets (tracking target 16, other targets such as bus 17) within each sensing range, it is necessary to first associate the trajectories of the same target in different sensing ranges before fusing the trajectories of the target in different sensing ranges, that is, to ensure that the motion trajectories obtained before fusing the trajectories of multiple sensing ranges belong to the same target. Exemplarily, the specific steps of the association can be as follows: First, obtain the sensing data of each target in different sensing ranges, and extract the features of the target therein (human hair color, clothes color, car color, shape, etc.). Then, pair the target trajectory position P and the target feature C at time t where, represents the trajectory position of target n at time t in view k; Represent the feature information of target n at time t from perspective k, such as the part of the vehicle, the direction of the vehicle head, the contour of the vehicle body, etc. Through a clustering algorithm, cluster and associate the features and trajectory positions detected by each target within different sensing ranges. If they can be clustered into the same category, it is determined that these features and trajectory positions belong to the same object. The clustering algorithm can be a density-based clustering algorithm (DBSCAN, Density-Based Spatial Clustering of Applications with Noise). The embodiments of the present application do not specifically limit the clustering algorithm used to associate the same target. After clustering, if there are these several pair of information in a certain class, then these several pairs of information belong to the same target object. Exemplarily, that is the trajectory positions measured from tracking target 16 at different perspectives (perspectives 1, 2, 4). Thus, at time t, the trajectory positions of the same target from different perspectives have been associated. Next, it is necessary to fuse the trajectory positions measured by the target from multiple perspectives. It should be noted that the "perspective" refers to the shooting direction of the camera, and each perspective corresponds to a sensing range. Different perspectives correspond to different sensing ranges. Cameras in different positions have different perspectives, thus corresponding to different sensing ranges.

[0091] As Figure 6 shown, the entire table is the positions of tracking target 16 at different times and within different sensing ranges after being associated by the clustering algorithm. The time of each row is the same, and the sensing range (perspective) of each column is the same. Here, taking only the camera type of sensors as an example, the details of the method for fusing multi-perspective trajectory data are shown. The embodiments of the present application apply the idea of Kalman fusion to trajectory fusion. The specific fusion method is as follows:

[0092] At time t = 0, which is the initial time, a measurement position at a certain perspective can be randomly selected as the estimated position at the initial time. At time t = 1, there are measurement positions of the tracked target 16 under perspectives 1, 2,..., k. According to principles such as whether these perspectives are approach lanes and the distance between the device corresponding to the perspective and the target, a measurement position under a better perspective is selected as the target measurement position at time t = 1. Predict the position at time t = 1 based on the final position (optimal position) at time t = 0 to obtain the target predicted position at time t = 1. Among them, there are various methods to predict the current time position based on the final position of the previous time. Exemplarily, a fitting method can be used or the position at the current time can be directly calculated based on the speed (including magnitude and direction) and position of the previous time. Perform Kalman fusion on the target predicted position and the selected target measurement position at time t = 1 to obtain the optimal estimated position at time t = 1. Similarly, at time t = 2, a target measurement position under an optimal perspective and the target predicted position at time t = 2 (predicted based on the optimal estimated position at time t = 1) are selected for Kalman fusion to obtain the optimal estimated position at time t = 2. Repeat this process for each time to obtain the continuous motion trajectory of the tracked target 16 after fusing data from different perspectives. Continuous motion trajectories of other surrounding targets can also be obtained according to the above method. It should be noted that in addition to the above method of fusing using the Kalman idea, the most direct weighted average method can also be used to obtain the final position of the target at each time. Exemplarily, at time t = 1, there are perspectives 1, 2, 3,..., k, so there are k positions of the tracked target 16 at time t = 1. Directly perform weighted average calculation on these k positions to obtain the final position of the tracked target 16 at time t = 1.

[0093] In the embodiment of the present application, the idea of Kalman fusion is first applied to the fusion of target motion trajectories under multiple perspectives, and the measurement value at each time is selected from the best perspective, rather than random perspective fusion, which improves the accuracy of trajectory fusion and can effectively solve the problem of target occlusion or target loss during tracking, ensuring the continuity of each target motion trajectory.

[0094] Step S43: Determine the target after the original tracked target switches according to the motion trajectories of the original tracked target and other targets.

[0095] According to step S42, the continuous motion trajectories of the original tracked target and other targets can be obtained. Inputting the continuous motion trajectories of the original tracked target and other targets into a pre-trained neural network model can determine whether the original tracked target has a switching behavior, thereby determining the new tracked target. Exemplarily, a two-dimensional schematic diagram of the trajectory is as Figure 7As shown, the solid line 1 is the movement trajectory of the original tracking target 1 (tracking target 16), and the dotted lines 2, 3, and 4 correspond to the movement trajectories of other targets 2, other target 3, and other target 4. Among them, the other target 2 corresponding to the dotted line 2 can be a bus 17. Optionally, before determining that the original tracking target switches to another target based on the trajectory, the trajectories that are too far from the original tracking target can be filtered. Exemplarily, at time t, when the distance between the tracking target 16 and other targets is greater than 7 meters, then this other target is filtered out. Assuming the position of the original tracking target is (x1, y1) and the position of a certain other target is (x2, y2), exemplarily, the Euclidean distance calculation formula can be used to calculate the distance L between the two: The remaining other targets after filtering can be called candidate targets.

[0096] The process of determining the new target after the original tracking target switches according to the trajectory generally includes the following steps: First, filter the farther other target 4 according to the above method to obtain candidate targets 2 and 3; then establish the trajectory feature space of the original tracking target and the candidate targets, input it into the pre-trained first neural network model, and determine the new target after the original tracking target switches. In the embodiments of the present application, taking the LSTM neural network (the first neural network) as an example, it is introduced how to analyze the trajectory data to determine the new target after switching.

[0097] Before introducing the use of the LSTM neural network for trajectory analysis, first introduce the LSTM neural network.

[0098] The LSTM network is a type of time-recurrent neural network. The LSTM network (as shown in Figure 10(a)) includes LSTM units (such as Figure 9 shown), and the LSTM unit includes three gates: a forget gate, an input gate, and an output gate. The forget gate included in the LSTM unit is used to determine the information to be forgotten, the input gate of the LSTM unit is used to determine the information to be updated, and the output gate of the LSTM unit is used to determine the output value. Exemplarily, Figure 9 shows the LSTM units at three moments. The input of the first LSTM unit is the input at time t - 1, the input of the second LSTM unit is the input at time t, and the input of the third LSTM unit is the input at time t + 1. The structures of the first LSTM unit, the second LSTM unit, and the third LSTM unit are exactly the same. The core of the LSTM network lies in the state of the cell (the cell is the Figure 9 large square box in Figure 9 and the horizontal line crossing in Figure 9 The state of the cell (i.e., t-1 C in t)Like a conveyor belt, it passes through the entire cell and only performs a small amount of linear operations. In this way, information can pass through the entire cell without being changed, and long-term memory retention can be achieved. After each output of the LSTM unit, if there are still trajectory feature pairs of time points that have not been input into the LSTM unit, the output of this time and the trajectory feature pair of the next time point that has not been input into the LSTM unit are input into the LSTM unit until there are no trajectory feature pairs to be input into the LSTM unit after a certain output of the LSTM unit. It should be noted that Figure 9 Only some units of the network are shown. In fact, the entire neural network is formed by combining multiple LSTM units. As shown in Fig. 10(a), the outputs of the LSTM units at each moment will be centrally input into the fully connected layer, and then the classification result (score) is output through the softmax layer. By using the method of regression loss, the specific time point among the multiple input time points is found as the switching time point. Figure 9 And the X in Fig. 10(a) t-1 、X t 、X t+1 Correspond one by one. Exemplarily, in Fig. 10(a), X t-1 、X t 、X t+1 Are a set of feature pairs of target 1 and 2 at times t1, t2, and t3 respectively. As shown in Fig. 10(b), P11 represents the position of target 1 at time t1, and P12 represents the position of target 2 at time t1; V11 represents the speed of target 1 at time t1, V12 represents the speed of target 2 at time t1, and θ1 represents the included angle of the movement directions of target 1 and 2 at time t1. Input the set of feature pairs shown in Fig. 10(b) into the LSTM neural network model shown in Fig. 10(a), and it can be determined whether the set of feature pairs contains a switching behavior (softmax classification is yes or no, and score is a probability value), as well as the time (index).

[0099] In the above method, the first neural network needs to be pre-trained before use. Exemplarily, video data containing the scenario of a person getting into a taxi can be manually searched. Through the video data, the trajectory data of the person and the trajectory data of the taxi the person gets into are obtained. According to the above method (as shown in Fig. 10(b)), the trajectory feature pairs of the person and the taxi are established, that is, the position, speed magnitude, and speed included angle of the person and the taxi at each moment. Then, the feature pairs in different time periods are taken to form different samples and are respectively marked, and then the marked samples are input into the neural network for training. For example, assuming that the trajectory sampling time interval is 1 second, a video is manually found, and the time when the person gets into the car is 11:01:25 ( Figure 11 The 01′25″ inFigure 11 Among them, a, b, c, d, and e are five interleaved time periods, and the length of each time period is 10 seconds. It should be noted that the length of the time period is not fixed. Taking the a time period as an example, its initial moment is 01:10, and its time length is 10 seconds. Select the video data of the person and the taxi in the a time period and convert it into the corresponding trajectory data, and a set of trajectory feature pairs of the person and the taxi at 10 moments (the a time period contains 10 seconds) can be obtained. Take these 10 trajectory feature pairs as a set of trajectory features and mark the classification result as NO (because the person did not get on the taxi within the a time period), and the switching time is none. Similarly, the classification result of a set of feature pairs obtained in the b time period is also marked as NO, and the switching time is none. The initial moment of the e time period is 01'20", and the length is 10s, including the moment when the person gets on the taxi. Therefore, mark the classification result of a set of trajectory feature pairs (10 trajectory feature pairs) obtained in the e time period as YES, and the switching time is 0125. The five sets of feature pairs obtained from the above five time periods can be used as five training samples to train the neural network, and the actual number of training samples is determined by the actual situation or the expected model accuracy. After training the neural network, the real-time trajectory data can be converted into trajectory feature pairs and input into the neural network, and the probability that other targets are the switched targets at each moment can be output.

[0100] Next, the use of the LSTM neural network will be described. Taking Target 1, 2, and 3 as examples, the input to the LSTM neural network is the trajectory feature pairs of Target 1 and Target 2, and the trajectory feature pairs of Target 1 and Target 3. The above trajectory features can be understood as some attributes of the target trajectory, such as position, speed, and angle, as shown below:

[0101] [O1V t , O2V t , O1P t (x,y), O2P t (x, y), θ1 t , t = 0, 1, 2, …, m (1)

[0102] [O1V t , O3V t , O1P t (x, y), O3P t (x, y), θ2 t , t = 0, 1, 2, …, m (2)

[0103] Among them, O1P t (x,y) represents the position of the original Target 1 at time t, and O1V t represents the speed of the original Target 1 at time t; O2P t (x, y) represents the position of the candidate Target 2 at time t, and O2Vt represents the speed of candidate target 2 at time t; O3P t (x, y) represents the position of candidate target 3 at time t, O3V t represents the speed of candidate target 3 at time t; the angle between targets can also be calculated based on the direction of the speed, θ1 t represents the angle between the moving directions of target 1 and target 2 at time t, θ1 t represents the angle between the moving directions of target 1 and target 3 at time t.

[0104] Equation (1) establishes the trajectory feature pairs of the original tracking target 1 and candidate target 2 at each moment, and equation (2) establishes the trajectory feature pairs of the original tracking target 1 and candidate target 3 at each moment. The trajectory feature pairs include the positions, speeds of each target, and the angles between the targets (target 1 and 2, target 1 and 3) (determined by the direction of the speed). The established feature pairs are respectively input into a pre-trained neural network, such as a long short-term memory network (LSTM, Long Short-Term Memory), and the probability that each candidate target is the new target after switching for the original target can be output in real time. As Figure 8 shown, the LSTM neural network can directly output the probability that each candidate target is the new target after switching at each moment. Exemplarily, taking the analysis of candidate target 2 as an example, assume that the current moment is t = 4 and target 1 starts to be tracked. Four feature pairs (a group of feature pairs) of target 1 and 2 from t = 0 to t = 4 established by equation (1) are selected and input into the LSTM neural network shown in Fig. 10(a), and the softmax outputs the probability of switching (as the probability value at time t = 4); when the current moment becomes t = 5, four feature pairs (another group of feature pairs) of target 1 and 2 from t = 1 to t = 5 established by equation (1) are selected and input into the LSTM neural network shown in Fig. 10(a), and the softmax outputs the probability of switching (as the probability value at time t = 5), and so on. The time probability distribution graph of candidate target 2 as the new target after switching can be drawn (as Figure 8 shown by the dashed line 2 in). Figure 8 In, F t is the probability that each candidate target is the new target after switching at each moment, and the horizontal axis, i.e., the time axis, is the current moment. A probability threshold c is preset in advance. At time t0, the probability corresponding to candidate target 2 exceeds the preset threshold, which means that the original target switches to candidate target 2 at time t0.

[0105] It should be noted that after obtaining the trajectory data, in addition to using neural networks for calculation and analysis, traditional machine learning classification models such as support vector machines can also be used for calculation. Or directly judge the obtained trajectory data according to the rules stipulated by humans. The existing research methods mainly use video data to perform intelligent behavior analysis on the target, while the embodiment of the present application uses the movement trajectory of the target to judge whether the target has switched. There is no need to perform a large amount of calculations on the video data, only the change in the target position needs to be considered, which saves computing power and improves the calculation speed while ensuring the real-time nature of tracking; and the calculation is performed using a pre-trained neural network, which ensures the accuracy of the calculation result.

[0106] Step S44: Perform a secondary judgment using computer vision methods.

[0107] Before further judging the switching behavior, a best video needs to be selected first. Through the above embodiments, a moment t0 when the target switches can be obtained. In step S42, a better perspective is selected at each moment. Therefore, the best video can be a video taken at the above best perspective near the moment t0. It should be noted that in addition to only using the video at the best perspective selected in step S42, multiple videos containing the switching scene of the original tracking target can also be selected. The time period of this video is mainly concentrated near the moment t0. Exemplarily, it can be a video of about 6 seconds, and the middle moment of the time period covered by the video is t0. In order to reduce the calculation amount and improve the calculation accuracy, the region of interest (ROI, Region Of Interest) in the video frame can be extracted first, and then the video data after extraction can be input into the pre-trained neural network. Exemplarily, the region of interest may only include the tracked target 16 and the bus 17. The process of performing a secondary judgment using computer vision is as Figure 12 and 13 shown, and includes the following steps:

[0108] Step S121: Input the selected video frame data into a pre-trained convolutional neural network, which can be a Convolutional Neural Networks (CNN), 3D Convolutional Neural Networks (3D CNN), etc. The convolutional neural network is mainly used to extract the features of the target in the video image and generate a bounding box (bbox) corresponding to the target. It needs to be pre-trained before use, and the training set can be images with bounding boxes that have been manually annotated. The convolutional neural network mainly includes an input layer, a convolutional layer, a pooling layer, and a fully connected layer. The input layer is the input of the entire neural network. In a convolutional neural network for processing images, it generally represents the pixel matrix of a picture; each node in the convolutional layer takes only a small block from the previous layer of the neural network, and the size of this small block can be 3*3, 5*5 or other sizes. The convolutional layer is used to perform a more in-depth analysis of each small block in the neural network to obtain features with a higher degree of abstraction; the pooling layer can further reduce the number of nodes in the final fully connected layer, thereby achieving the purpose of reducing the parameters in the entire neural network; the fully connected layer is mainly used to complete the classification task.

[0109] Step S122: The convolutional neural network extracts the features of the target in the video and generates a bounding box. Among them, the features of the target are represented by a series of digital matrices; the position of the target determines the coordinates of the four corners of the bounding box. As Figure 14 shown, both the tracking target 16 and the bus 17 generate a bbox respectively, and the corresponding target classification is generated.

[0110] Step S123: Build a graph model based on the features of the target and the bounding box corresponding to the target. When building the graph model, the targets recognized in the video frame (which can be understood as an ID in the field of target recognition) are used as nodes. The edges connecting the nodes are mainly divided into two categories. The first type of edge has two components: 1) The first part represents the similarity of the target between two consecutive frames. The higher the similarity, the higher the value, and the value range is [0,1]. Obviously, the same person or the same vehicle in consecutive frames has a high similarity, different vehicles and people have a low similarity, and there is basically no similarity between people / vehicles; 2) The second part represents the overlap degree between the bbox of the target in the current frame and the bbox of the target in the next frame, that is, the size of the IoU (intersection of union). If it completely overlaps, the value is 1. The first part and the second part are combined and added according to a preset weight to form the value of the first type of edge. The second type of edge represents the distance between two targets in the actual space. The closer the distance, the larger the value of the edge. Exemplarily, the graph model constructed according to the above method is as Figure 15As shown, the first type of edge can be understood as the horizontal edge in the figure, and the second type of edge can be understood as the vertical edge in the figure.

[0111] Step S124: Input the constructed graph model into the graph convolutional neural network. A general graph convolutional neural network includes a graph convolutional layer, a pooling layer, a fully connected layer, and an output layer. Among them, the graph convolutional layer is similar to the image convolution, performs information transmission inside the graph, and can fully mine the features of the graph; the pooling layer is used for dimensionality reduction; the fully connected layer is used for classification; the output layer outputs the classification result. Exemplarily, the classification result may include getting on the vehicle behavior, getting off the vehicle behavior, human-vehicle approaching behavior, human bending behavior, vehicle door opening and closing behavior, and so on. It should be noted that the output of the neural network is not only a behavior recognition (getting on or getting off the vehicle), but also can include behavior detection, for example, the characteristics of the vehicle that the person gets on can be determined. It should be noted that the switched target determined through step S44 is generally the same as the switched target determined in step S43. In this case, this target can be directly used as the switched target. When the switched targets determined in steps S43 and S44 are different, the switched target determined in step S44 is used as the new tracking target.

[0112] The above graph convolutional neural network also needs to be pre-trained before use. Exemplarily, an artificial video of a person getting on the vehicle is obtained, and the graph models of each target in the video are generated according to steps S121 - S123, and the graph model is marked as the getting on the vehicle behavior, which is used as the training set of the graph convolutional neural network. Therefore, after the graph model generated from the input real-time video data, the graph convolutional neural network can automatically output a classification to determine whether the tracking target has a cross-target behavior.

[0113] It should be noted that step S44 is not necessary. A final new switched target can be directly determined only by step S43, which is mainly determined by the preference of the monitoring personnel or the actual situation. Moreover, the graph convolutional neural network is used for judgment in step S44. It should be noted that in addition to the graph convolutional neural network, other artificial intelligence classification models can also perform behavior recognition on the selected better video.

[0114] The embodiment of the present application uses the method of computer vision to make a secondary judgment on whether the original target has switched, constructs a unique graph model, incorporates more time and space information, and improves the accuracy of the judgment.

[0115] Step S45: Determine the switched target.

[0116] Step S44 is optional. Therefore, there are the following two cases: (1) After executing step S43, directly execute step S45, then use the target determined by step S43 as the switched target; (2) After executing step S43, execute step S44. In step S45, use the switched target determined by step S44 as the final switched target. It should be noted that in the second case (that is, the switched targets determined by step S43 and step S44 are different), in addition to directly using the target determined by step S44 as the switched target, other strategies can also be adopted. For example, comprehensively consider the accuracy of the models used in the two steps and the confidence of the output results, and select one of them as the final switched target.

[0117] Steps S41 - S45 mainly take the scenario of a person getting on the vehicle as an example, and switch the tracking target from tracking target 16 to bus 17. After that, bus 17 can be tracked continuously, and video images of passengers getting off the vehicle are obtained during each parking period to search for tracking target 16. When tracking target 16 reappears in a certain image, switch the tracking target back to tracking target 16, so as to achieve continuous tracking of the target.

[0118] In summary, the automatic cross - target tracking method provided by the embodiments of the present application can be summarized as follows: First, an improved Kalman filtering method is used to fuse video and radar data under the same perspective, which improves the accuracy of the target motion trajectory under a single perspective. Then, the motion trajectories of the same target under multiple perspectives are fused, effectively solving the occlusion problem, and continuous motion trajectories of each target can be obtained. Then, the motion trajectories of the tracking target and other nearby targets are input into a pre - trained neural network to determine whether the original tracking target has switched. If it has switched, the tracking target is replaced with a new target. In addition, to further improve the accuracy of the judgment, on the basis of the previous step, the best video that can capture the image of the switching behavior can be selected for behavior analysis. Input the best video into the neural network to extract features, then construct a unique graph model, and use a pre - trained graph convolutional neural network to make a secondary judgment on whether the target in the video has a cross - target behavior. The above method can be applied not only to tracking suspicious persons, but also to helping find runaway children or lost elderly people, etc. No matter what type the tracked object is, the technical solution provided by the present invention can be applied when tracking it.

[0119] From the perspective of the tracking strategy, the automatic cross-target tracking method provided in this application can automatically switch the tracking target when the original target exhibits cross-target behavior, ensuring the real-time effectiveness of tracking and preventing tracking interruption due to the target taking transportation. From the perspective of the specific implementation of tracking, this application first proposes using the motion trajectory of the target in the world coordinate system to analyze and judge cross-target behavior, improving the accuracy of behavior judgment. In addition, to improve the accuracy of the target's motion trajectory, the present invention also provides a series of measures, including: (1) adopting an original improved Kalman filtering method to fuse the video and radar trajectories in a single perspective; (2) using the Kalman filtering method to fuse the motion trajectories of the same target from multiple perspectives. On this basis, to further improve the judgment accuracy, the present invention also constructs a unique graph model based on the selected best video, and then uses a graph convolutional neural network to perform a secondary judgment on cross-target behavior, incorporating more spatio-temporal information. In summary, the automatic cross-target tracking method provided in this application can achieve real-time, continuous, and accurate tracking of the target, without missing any information in any time period, improving the convenience and reliability of deployment and tracking.

[0120] In addition to the above implementation, this application also provides another variant embodiment for target tracking based on the above embodiments. The method includes the following steps: (1) Obtain the motion trajectory of the tracking target 16 (original tracking target). The trajectory of the original tracking target can be obtained according to the methods described in steps S41 (optional) and S42 in the above embodiments. The trajectory is composed of the positions of the target over a period of time, and the position of the target can be represented by coordinates in a global coordinate system (such as the east-north-up coordinate system). (2) Determine the initial moment when the trajectory disappears as the first moment according to the motion trajectory of the tracking target 16. That is, when the trajectory of the tracking target 16 no longer appears, determine the initial moment when the target disappears as the first moment. (3) Obtain the video data for a period of time before and after the first moment, and this video data includes the images in which the tracking target 16 may be updated to the second target. Exemplarily, the image of the tracking target 16 in this video data is complete and clear. (4) Analyze the video data obtained in the above steps to determine the switched target. Exemplarily, the relevant video data can be input into a pre-trained neural network for analysis. The pre-trained neural network can be the neural network used in step S44, and the switched target is determined by analyzing the video content through the neural network.

[0121] The above method mainly describes the scenario where a person (tracking target 16) gets on a vehicle (the switched tracking target). By determining that the trajectory of the person disappears and then retrieving relevant surrounding videos for analysis, the switched target such as a bus can be obtained, and then the bus is tracked. When the person gets off the vehicle, it is still necessary to search for the tracking target 16 in the surveillance videos along the route of the bus and then continue to track it. In each moment of the whole process of this application embodiment, only the trajectory of one type of target needs to be obtained. Directly retrieving relevant videos for behavior analysis based on the trajectory of the original tracking target reduces the occupancy of computing power to a certain extent and ensures real-time performance and accuracy.

[0122] As described above in connection with Figures 3 - 15 a method for target tracking provided by an embodiment of the present application is described. Next, a target tracking device provided by an embodiment of the present application will be described. It may include an acquisition module and a processing module;

[0123] The acquisition module is configured to acquire sensing data of targets included in the scene where the first target is located. The targets included in the scene where the first target is located include the first target and at least one other target other than the first target, and the first target is the initial tracking target;

[0124] The processing module is configured to generate motion trajectories of the first target and the at least one other target according to the sensing data; and is further configured to determine a second target according to the motion trajectory of the first target and the motion trajectories of the other targets, and use the second target as the switched tracking target.

[0125] Optionally, the processing module is specifically configured to determine a set of candidate targets, where the candidate targets are: the at least one other target, or among the at least one other target, other targets whose distance from the first target is less than a preset threshold; for each candidate target, input the motion trajectory of the first target and the motion trajectory of the candidate target into a pre-trained first neural network to obtain the probability that the candidate target is the second target; determine the second target according to the probabilities that the set of candidate targets are the second target.

[0126] Optionally, the processing module is further configured to detect that a first probability among the probabilities that the at least one candidate target is the second target is higher than a preset threshold, and determine the target corresponding to the first probability as the second target.

[0127] Optionally, the processing module is further configured to, for each candidate target, establish at least one set of trajectory feature pairs based on the movement trajectories of the candidate target and the first target. Each set of trajectory feature pairs includes at least two pairs of trajectory features at consecutive moments. Each pair of trajectory features at each moment includes the position and speed of the first target, the position and speed of the candidate target, and the included angle between the movement directions of the first target and the candidate target. Input the at least one set of trajectory feature pairs into the first neural network to obtain the probability that the candidate target is the second target.

[0128] Optionally, the processing module is further configured to detect that a first probability among the probabilities of the at least one candidate target is higher than a preset threshold, and determine the target corresponding to the first probability as the second target.

[0129] Optionally, setting the moment when the first probability is higher than the preset threshold as the first moment, the processing module is further configured to obtain video data before and after the first moment, where the video data includes the pictures in which the first target may have a switching behavior. Input the video data into the pre-trained second neural network, and determine the third target as the tracking target after switching according to the output result.

[0130] Optionally, the second neural network includes a convolutional neural network and a graph convolutional neural network. Specifically, the processing module is configured to input the selected video frame data into the pre-trained convolutional neural network to output the features and bounding boxes of all targets in the video frame data. Construct a graph model based on the features and bounding boxes. Input the graph model into the pre-trained graph convolutional neural network, and determine the third target as the tracking target after switching according to the output result.

[0131] Optionally, the sensor includes at least two groups of sensors in different orientations. Different groups of sensor modules are in different orientations. For each target in the scene where the first target is located, the processing module is specifically configured to generate at least two movement trajectories of the target respectively according to the sensing data collected by the at least two groups of sensor modules. Fuse the at least two movement trajectories of the target to form the movement trajectory of the target.

[0132] Optionally, each group of sensors includes at least two types of sensors, that is, includes a camera and at least one of the following two types of sensors: millimeter-wave radar and lidar, and the two types of sensors are in the same orientation. For each target included in the scene where the first target is located, the processing module is specifically configured to: generate at least two monitoring trajectories of the target respectively according to the sensing data collected by the at least two types of sensor modules. Fuse the at least two monitoring trajectories of the target to form the movement trajectory of the target.

[0133] The present application further provides another target tracking device, including an acquisition module and a processing module. The acquisition module is configured to acquire sensing data of a first target through a sensor, where the first target is an initial tracking target. The processing module is configured to generate a motion trajectory of the first target according to the sensing data; determine an initial moment when the motion trajectory of the first target disappears as a first moment; acquire video frames for a period of time before and after the first moment, where the video frames include the first target; determine a second target according to the video frames, and use the second target as an updated tracking target.

[0134] Optionally, the processing module is further configured to determine that there is no motion trajectory of the first target after the initial moment, and determine the initial moment as the first moment.

[0135] Optionally, the processing module is further configured to input the video frames into a pre-trained second neural network to determine a second target, and use the second target as an updated tracking target.

[0136] Optionally, the second neural network includes a convolutional neural network and a graph convolutional neural network. The processing module is further configured to: input the video frames into a pre-trained convolutional neural network to output features and bounding boxes of the targets included in the video data; construct a graph model according to the features and bounding boxes of the targets included in the video frames; input the graph model into a pre-trained graph convolutional neural network, determine the second target according to the output result, and use the second target as an updated tracking target.

[0137] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0138] Figure 16 Schematic diagram of the hardware structure of the computing device for target tracking provided by the embodiments of the present application. As Figure 16 shown, the device 160 may include a processor 1601, a communication interface 1602, a memory 1603, and a system bus 1604. The memory 1603 and the communication interface 1602 are connected to the processor 1601 through the system bus 1604 and communicate with each other. The memory 1603 is used to store computer execution instructions, the communication interface 1602 is used to communicate with other devices, and the processor 1601 executes the computer instructions to implement the solutions shown in all the above embodiments.

[0139] Figure 16The system bus mentioned above can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity in illustration, only a thick line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus. The communication interface is used to implement the communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include a Random Access Memory (RAM), and may also include non-volatile memory, such as at least one disk memory.

[0140] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0141] Optionally, an embodiment of the present application further provides a computer-readable storage medium, in which instructions are stored, and when they run on a computer, the computer is caused to execute the method as shown in the above method embodiment.

[0142] Optionally, an embodiment of the present application further provides a chip, which is used to execute the method as shown in the above method embodiment.

[0143] It can be understood that in the embodiments of the present application, the various digital numbers involved are only for the convenience of description and are not used to limit the scope of the embodiments of the present application.

[0144] It can be understood that in the embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for target tracking, characterized in that, The method includes: Obtaining the motion trajectories of the targets included in the scene where the first target is located through a sensor. The targets included in the scene where the first target is located include the first target and at least one other target other than the first target, and the first target is the initial tracking target; Determining a second target according to the motion trajectory of the first target and the motion trajectories of the at least one other target, and using the second target as the updated tracking target; The determining the second target according to the motion trajectory of the first target and the motion trajectories of the at least one other target includes: Determining a set of candidate targets, where the candidate targets are: the at least one other target, or among the at least one other target, the other targets whose distance from the first target is less than a preset threshold; For each of the candidate targets, inputting the motion trajectory of the first target and the motion trajectory of the candidate target into a pre-trained first neural network to obtain the probability that the candidate target is the second target; Determining the second target according to the probabilities that the set of candidate targets are the second target.

2. The method according to claim 1, wherein The determining the second target according to the probabilities that the set of candidate targets are the second target includes: detecting that a first probability among the probabilities that the set of candidate targets are the second target is higher than a preset threshold, and determining the target corresponding to the first probability as the second target.

3. The method according to claim 1 or 2, characterized in that The for each of the candidate targets, inputting the motion trajectory of the first target and the motion trajectory of the candidate target into a pre-trained first neural network to obtain the probability that the candidate target is the second target includes: Establishing at least one set of trajectory feature pairs according to the motion trajectory of the candidate target and the motion trajectory of the first target. Each set of trajectory feature pairs includes at least two trajectory feature pairs at consecutive moments. Each trajectory feature pair at each moment includes the position and speed of the first target at this moment, the position and speed of the candidate target, and the included angle between the motion directions of the first target and the candidate target; Inputting the at least one set of trajectory feature pairs into the first neural network to obtain the probability that the candidate target is the second target.

4. The method according to claim 2, wherein Determining the moment when the first probability is higher than the preset threshold as the first moment. The method further includes: Obtaining video frames for a period of time before and after the first moment, and the video frames include the first target; Inputting the video frames into a pre-trained second neural network, determining a third target according to the output result, and using the third target as the updated tracking target.

5. The method according to claim 4, wherein The second neural network includes a convolutional neural network and a graph convolutional neural network. The inputting the video frames into a pre-trained second neural network, determining the third target according to the output result, and using the third target as the updated tracking target includes: Inputting the video frames into a pre-trained convolutional neural network to output the features and bounding boxes of the targets included in the video frames; Constructing a graph model according to the features and bounding boxes of the targets included in the video frames; Input the graph model into a pre-trained graph convolutional neural network, determine the third target according to the output result, and use the third target as the updated tracking target.

6. The method according to any one of claims 1, 2, 4, and 5, characterized in that The sensor includes at least two groups of sensors in different orientations. For each target included in the scene where the first target is located, the process of obtaining the motion trajectory of the target through the sensor includes: For each group of sensors among the at least two groups of sensors, generate the motion trajectory of the target corresponding to this group of sensors according to the sensing data collected by this group of sensors, so as to obtain at least two motion trajectories of the target, and the at least two motion trajectories of the target are obtained by shooting from different orientations; Fuse the at least two motion trajectories of the target to obtain the fused motion trajectory of the target.

7. The method according to claim 6, wherein Each group of sensors includes at least two types of sensors, and the at least two types of sensors include a camera and at least one of the following two types of sensors: millimeter-wave radar and lidar, and the at least two types of sensors are in the same orientation. For each target included in the scene where the first target is located, for each group of sensors among the at least two groups of sensors, the process of generating the motion trajectory of the target corresponding to this group of sensors according to the sensing data collected by this group of sensors includes: For each type of sensor in this group of sensors, generate the monitoring trajectory of the target corresponding to this type of sensor according to the sensing data collected by this type of sensor, so as to obtain at least two monitoring trajectories of the target; Fuse the at least two monitoring trajectories of the target to obtain the motion trajectory of the target.

8. A method for target tracking, characterized in that, The method includes: Obtain the motion trajectory of the first target through the sensor, and the first target is the initial tracking target; Determine the initial moment when the first target trajectory disappears as the first moment according to the motion trajectory of the first target; Obtain video frames for a period of time before and after the first moment, and the video frames include the first target; Determine the second target according to the video frames, and use the second target as the updated tracking target; The step of determining the initial moment when the first target trajectory disappears as the first moment according to the motion trajectory of the first target includes: Judge that the motion trajectory of the first target does not exist after the initial moment, and determine the initial moment as the first moment; The step of determining the second target according to the video frames and using the second target as the updated tracking target includes: Input the video frames into a pre-trained second neural network, determine the second target according to the output result, and use the second target as the updated tracking target.

9. The method according to claim 8, characterized in that The second neural network includes a convolutional neural network and a graph convolutional neural network. The step of inputting the video frames into a pre-trained second neural network, determining the second target according to the output result, and using the second target as the updated tracking target includes: Input the video frames into a pre-trained convolutional neural network, and output the features and bounding boxes of the targets included in the video frames; Construct a graph model according to the features and bounding boxes of the targets included in the video frames; Input the graph model into a pre-trained graph convolutional neural network, determine the second target according to the output result, and use the second target as the updated tracking target.

10. A device for target tracking, characterized in that, It includes: An acquisition module and a processing module; The acquisition module is used to acquire the sensing data of the targets included in the scene where the first target is located. The targets included in the scene where the first target is located include the first target and at least one other target other than the first target, and the first target is the initial tracking target; The processing module is used to generate the motion trajectories of the first target and the at least one other target according to the sensing data, and determine the second target according to the motion trajectory of the first target and the motion trajectories of the at least one other target, and use the second target as the updated tracking target; Specifically, the processing module is used to determine a set of candidate targets, and the candidate targets are: the at least one other target, or among the at least one other target, the other target whose distance from the first target is less than a preset threshold; For each of the candidate targets, input the motion trajectory of the first target and the motion trajectory of the candidate target into a pre-trained first neural network to obtain the probability that the candidate target is the second target; Determine the second target according to the probabilities that the set of candidate targets are the second target.

11. The device according to claim 10, characterized in that The processing module is further used to detect that a first probability among the probabilities that the set of candidate targets are the second target is higher than a preset threshold, and determine the target corresponding to the first probability as the second target.

12. The device according to claim 10 or 11, characterized in that, The processing module is further used for: For each candidate target, establish at least one set of trajectory feature pairs according to the motion trajectories of the candidate target and the first target. Each set of trajectory feature pairs includes at least two trajectory feature pairs at consecutive moments. Each trajectory feature pair at each moment includes the position and speed of the first target at this moment, the position and speed of the candidate target, and the included angle between the motion directions of the first target and the candidate target; Input the at least one set of trajectory feature pairs into the first neural network to obtain the probability that the candidate target is the second target.

13. The device according to claim 11, wherein The moment when the first probability is higher than the preset threshold is the first moment. The processing module is further used for, Acquire video frames for a period of time before and after the first moment, and the video frames include the first target; Input the video frames into a pre-trained second neural network, determine the third target according to the output result, and use the third target as the updated tracking target.

14. The device according to claim 13, characterized in that, The second neural network includes a convolutional neural network and a graph convolutional neural network. Specifically, the processing module is used for, Input the video frames into a pre-trained convolutional neural network, and output the features and bounding boxes of the targets included in the video frames; Construct a graph model according to the features and bounding boxes of all the targets in the video data; Input the graph model into a pre-trained graph convolutional neural network, determine the third target according to the output result, and use the third target as the updated tracking target.

15. The device according to any one of claims 10, 11, 13, and 14, characterized in that The sensor includes at least two groups of sensors in different orientations. For each target included in the scene where the first target is located, the processing module is specifically configured to, For each group of sensors in the at least two groups of sensors, generate a motion trajectory of the target corresponding to the group of sensors according to the sensing data collected by the group of sensors, so as to obtain at least two motion trajectories of the target. The at least two motion trajectories of the target are obtained by photographing from different orientations; Fuse the at least two motion trajectories of the target to obtain the fused motion trajectory of the target.

16. The device according to claim 15, characterized in that, Each group of sensors includes at least two types of sensors. The at least two types of sensors include a camera and at least one of the following two types of sensors: millimeter-wave radar and lidar, and the at least two types of sensors are in the same orientation. For each target included in the scene where the first target is located, the processing module is specifically configured to, For each type of sensor in the group of sensors, generate a monitoring trajectory of the target corresponding to the type of sensor according to the sensing data collected by the type of sensor, so as to obtain at least two monitoring trajectories of the target; Fuse the at least two monitoring trajectories of the target to obtain the motion trajectory of the target.

17. A device for target tracking, characterized in that, The device includes an acquisition module and a processing module, The acquisition module is configured to obtain sensing data of a first target through a sensor, and the first target is an initial tracking target; The processing module is configured to generate a motion trajectory of the first target according to the sensing data; determine the initial moment when the motion trajectory of the first target disappears as the first moment; obtain video frames for a period of time before and after the first moment, and the video frames include the first target; determine a second target according to the video frames, and use the second target as the updated tracking target; The processing module is further configured to determine that the motion trajectory of the first target does not exist after the initial moment, and determine the initial moment as the first moment; The processing module is further configured to input the video frames into a pre-trained second neural network, determine the second target according to the output result, and use the second target as the updated tracking target.

18. The device according to claim 17, characterized in that, The second neural network includes a convolutional neural network and a graph convolutional neural network. The processing module is further configured to: Input the video frames into a pre-trained convolutional neural network, and output the features and bounding boxes of the targets included in the video data; Construct a graph model according to the features and bounding boxes of the targets included in the video frames; Input the graph model into a pre-trained graph convolutional neural network, determine the second target according to the output result, and use the second target as the updated tracking target.

19. A computing device for target tracking, characterized in that, The computing device includes a processor and a memory, where: Computer instructions are stored in the memory; The processor executes the computer instructions stored in the memory to implement the method according to any one of claims 1-7.

20. A computing device for target tracking, characterized in that, The computing device includes a processor and a memory, where: Computer instructions are stored in the memory; The processor executes the computer instructions stored in the memory to implement the method according to any one of claims 8-9.

21. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program code which, when executed by a computer, causes the computer to execute the method according to any one of claims 1-7.

22. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program code which, when executed by a computer, causes the computer to execute the method according to any one of claims 8-9.

Citation Information

Patent Citations

  • Moving object tracking method and device

    CN104156982A

  • Multi-target tracking processing method and device

    CN108986151A