A multi-sensor target swarm tracking method based on modified DeepSORT
By using Hungarian algorithm and adaptive fusion features to calculate the Mahayana distance in multi-sensing target tracking, combining the least squares method to fit the acceleration prediction curve, and correcting the Kalman filtering, the problem of insufficient prediction of the traditional DeepSORT algorithm when the target suddenly changes, improving the accuracy of target tracking.
Patent Information
- Application Number
- CN202310449562.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-24
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2043-04-24
AI Technical Summary
When the traditional DeepSORT algorithm suddenly accelerates or suddenly stops, the prediction effect of constant-speed linear Kalman filtering is poor, resulting in the failure of target tracking.
The matching method based on Hungarian algorithm and adaptive fusion features are used to calculate the Marshall distance, and the acceleration prediction curve is fitted with the least squares method, the input state of Kalman filtering is corrected, and the accuracy of target matching and tracking is improved.
Correcting Kalman filtering through adaptive fusion features and acceleration prediction improves the accuracy of multi-sensing target tracking, solving the problem of insufficient prediction of traditional DeepSORT algorithm when the target suddenly changes.
Smart Images

Figure CN116543023B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of target tracking and computer vision, and in particular relates to a multi-sensor target swarm intelligence tracking method based on modified DeepSORT. Background Art
[0002] Multi-vision pedestrian tracking algorithms utilize multiple cameras to capture real-time video streams of a scene and automatically detect, track, and identify pedestrians, enabling intelligent and safe monitoring of pedestrians. This algorithm can improve the efficiency of urban safety management, reduce safety hazards and accident rates, and enhance urban security. In modern life, multi-vision pedestrian tracking algorithms are widely used in public security monitoring, intelligent transportation systems, and community security. For example, multi-vision pedestrian tracking algorithms are used in public places such as shopping malls, subway stations, and airports to monitor pedestrians and promptly identify safety hazards such as crowds and congestion. Furthermore, multi-vision pedestrian tracking algorithms can be applied in the transportation sector, helping traffic management departments implement traffic flow monitoring, congestion management, and traffic accident warnings, thereby improving the efficiency and safety of urban transportation.
[0003] Multi-vision pedestrian tracking algorithms can simultaneously utilize multiple cameras for monitoring, thereby improving monitoring efficiency. They can integrate video streams captured by multiple cameras and automatically analyze and track pedestrian movements, reducing the need for manual intervention and lowering monitoring costs. They can also use technologies such as deep learning to improve the accuracy of pedestrian detection and tracking, effectively avoiding missed and false alarms.
[0004] The traditional DeepSORT algorithm uses a constant-speed linear Kalman filter, which is relatively accurate for predicting the state of targets that are moving at a constant speed or with little speed change. However, when the target suddenly accelerates or stops, the prediction effect of the constant-speed linear Kalman filter will be affected, resulting in target tracking failure. Summary of the Invention
[0005] Purpose of the invention: In order to solve the problems existing in the above-mentioned prior art, the present invention provides a multi-sensor target swarm intelligence tracking method based on modified DeepSORT.
[0006] Technical solution: The present invention provides a multi-sensor target swarm tracking method based on modified DeepSORT. The method tracks targets in a task area using multiple cameras, specifically including the following steps:
[0007] Step 1: Use each camera to obtain real-time surveillance video of the task area;
[0008] Step 2: Using the matching method based on the Hungarian algorithm, the matching target of the current video frame between the cameras is obtained and assigned an identity ID;
[0009] Step 3: For each camera, use the DeepSORT method to track the target assigned with the identity ID in the current video frame until the target tracking task is completed.
[0010] Furthermore, the step 2 is specifically as follows:
[0011] Step 2.1: Use YOLOv5 to perform target detection on the current image frames obtained by each camera to obtain the target information in each image frame;
[0012] Step 2.2: Input the target information obtained in step 2.1 into the ResNet residual network to extract the appearance features of the target;
[0013] Step 2.3: Obtain the spatial motion characteristics of the target based on the camera frame rate and the target information obtained in step 2.1;
[0014] Step 2.4: Combine the appearance features of the target in step 2.2 and the spatial motion features of the target in step 2.3 to form an adaptive fusion feature of the target;
[0015] Step 2.5: Using the adaptive fusion features of the target in step 2.4, calculate the Mahalanobis distance and cost matrix;
[0016] Step 2.6: Use the Hungarian algorithm to match the targets between cameras at the same time, obtain the matching targets of the current video frame between the cameras and assign them an identity ID.
[0017] Furthermore, the adaptive fusion feature of the target in step 2.4 is defined as follows:
[0018]
[0019] Among them, ψ represents the set weight, Represents the vector dimension splicing operator, MT l is the spatial motion feature of target l, R l is the appearance feature of target l.
[0020] Furthermore, the expression for setting the weight ψ is as follows:
[0021]
[0022] in, Indicates the distance between the target l and the center of the camera's field of view.
[0023] Furthermore, the expression of the Mahalanobis distance in step 2.5 is as follows:
[0024]
[0025] in, is the lth target p under camera p at time t l Adaptive fusion features and the fth target q under camera q f Adaptive fusion features Mahalanobis distance; for and The covariance matrix between ;
[0026] Furthermore, the cost matrix in step 2.5 is expressed as follows:
[0027]
[0028] Among them, H(p,q) is the cost matrix of camera p and camera q at time t, l∈{1,…,H p}, f∈{1,…,H q}, H p is the number of targets under camera p at time t, H q is the number of targets under camera q at time t.
[0029] Furthermore, the DeepSORT method in step 3 uses the least squares method to fit the acceleration prediction curve, and then corrects the input state detection value of the Kalman filter, specifically including:
[0030] C1: Get the camera at time t k The detection value of the Kalman filter input state of the acquired image frame:
[0031] MT k =(x k ,y k ,w k ,h k ,v xk ,v yk ,v wk ,v hk ) T
[0032] Among them, x k ,y k Respectively represent the horizontal and vertical coordinates of the center point of the target detection frame in the image frame, w k , h k Respectively represent the width and height of the target detection box in the image frame, v xk , v yk , v wk , v hk y k Represents x k ,y k , wk , h k The rate of change;
[0033] C2: Based on the target detection box in C1 at the previous k-1 moments t1,...,t k-1 Acceleration data sequence of acquired image frames Fitting the acceleration prediction curve using the least squares method
[0034]
[0035] Among them, i∈{x,y,w,h}, x,y,w,h are the four parameters of the Kalman filter input state, x,y represent the horizontal and vertical coordinates of the center point of the target detection frame, w,h represent the width and height of the target detection frame respectively; a i1 ,…,a i(k-1) are times t1,...,t k-1 The acceleration data corresponding to the four parameters of the Kalman filter input state, The acceleration fitting prediction curve of the four parameters representing the Kalman filter input state, and is the fitting coefficient of the acceleration fitting prediction curve;
[0036] C3: Using the acceleration prediction curve, the time t k The input state detection value of the Kalman filter of the acquired image frame is corrected to obtain the input state detection value of the Kalman filter after correction:
[0037] MT′ k =(x k ,y k ,w k ,h k ,v′ xk ,v′ yk ,v′ wk ,v′ hk ) T
[0038] Among them, v′ ik is the time t after correction k The change speed of the four parameters corresponding to the Kalman filter input state, Δt is the time difference between two image frames;
[0039] C4: Calculate time t k Kalman filter inputs the acceleration data corresponding to the four parameters of the state and updates the time t k Acceleration data sequence of acquired image frames
[0040]
[0041] Among them, v′ i(k-1) is the time t after correction k-1 The speed of change of the four parameters corresponding to the Kalman filter input state;
[0042] C5: Let k=k+1 and repeat C1 to C5 until the video ends.
[0043] The present invention also provides a multi-sensor target swarm intelligence tracking system based on the modified DeepSORT, which performs target tracking based on the above method, specifically comprising:
[0044] Multiple cameras for real-time surveillance video of the mission area;
[0045] The target matching unit is used to obtain the matching target of the current video frame between cameras and assign an identity ID;
[0046] The target tracking unit is used to track the target assigned with the identity ID in the current video frame using the DeepSORT method for each camera.
[0047] The present invention also provides a computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions that, when executed by a computing device, cause the computing device to perform the method described above.
[0048] The present invention also provides a computing device comprising one or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, and the one or more programs include instructions for executing the method described above.
[0049] Beneficial effects: The present invention discloses a multi-sensor target group intelligence tracking method based on the modified DeepSORT, which tracks the target in the task area through multiple cameras: using each camera to obtain the monitoring video of the task area in real time; using a matching method based on the Hungarian algorithm to obtain the matching target of the current video frame between the cameras and assign an identity ID; for each camera, using the DeepSORT method to track the target assigned the identity ID in the current video frame until the target tracking task is completed. The present invention uses an adaptive fusion feature based on the spatial position of the target to calculate the Mahalanobis distance to measure the target similarity, and finally uses the Hungarian algorithm to achieve target matching, thereby improving the accuracy of target matching. In response to the problem that the uniform linear Kalman filter used in the traditional DeepSORT algorithm has a poor prediction effect when the target suddenly accelerates or stops, the present invention proposes a DeepSORT correction method based on average speed prediction, uses the least squares method to fit the predicted acceleration, and then corrects the input speed parameter of the Kalman filter, thereby improving the accuracy of the Kalman filter prediction of the DeepSORT algorithm, thereby improving the accuracy of target tracking, and has a wide range of application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is the overall flow chart of the present invention;
[0051] Figure 2 A flowchart of DeepSORT after being modified by the DeepSORT modification method based on average velocity prediction of the present invention. DETAILED DESCRIPTION
[0052] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0053] The present invention provides a multi-sensor target group intelligence tracking method based on modified DeepSORT. Figure 1 As shown, the overall steps are as follows:
[0054] Step 1: Multiple cameras simultaneously use YOLOv5 to detect the target in the task area in the current video frame and obtain the target detection information;
[0055] Step 2: Input the targets detected by each camera in step 1 into the ResNet residual network to extract the target's appearance features, and then obtain the target's spatial motion features based on the camera's frame rate and the spatial position of the target detection frame;
[0056] Step 3: Form an adaptive fusion feature based on the appearance features and spatial motion features of the target in step 2;
[0057] Step 4: Propose a matching method based on the Hungarian algorithm:
[0058] The adaptive fusion features in step 3 are used to calculate the Mahalanobis distance, and then the Mahalanobis distance threshold is used to determine the match and calculate the cost matrix. Finally, the cost matrix is input into the Hungarian algorithm to obtain the matching target of the current video frame between the cameras and assign an identity ID.
[0059] Step 5: Each camera tracks the target with the assigned ID using DeepSORT modified by the DeepSORT correction method based on average speed prediction; repeat steps 1 and 5 until the target monitoring and tracking task is completed.
[0060] Combined with attachment Figure 2 It is explained that after DeepSORT is modified by a DeepSORT correction method based on average velocity prediction in the present invention, the steps of DeepSORT for target tracking are as follows:
[0061] E1: The target enters the monitoring area, and the detection generates the target detection frame of the first frame. A detector Detection is initialized for it to record the target's detection frame, appearance features and motion information. At this time, no prediction frame is generated and there is no matching prediction frame. It is a mismatched detector. A new tracker Track is assigned to it to record its prediction frame, appearance features and motion information. For the first frame image, the new tracker Track stores the detection frame as a prediction frame. Then, according to the acceleration sequence in the motion information of the target in Track, the least squares method is used to fit the acceleration prediction curve and calculate the average speed to correct the speed parameter of the Kalman filter input. The Kalman filter predicts the prediction frame of the target in the next frame. At this time, the tracker must be in an unconfirmed state;
[0062] E2: Perform IOU matching on the detection frame of the target in the second frame and the prediction frame generated by the Kalman filter in the previous frame. If they match, the Kalman filter compares and updates the motion information of the detection frame with the motion information of the prediction frame. Then, according to the acceleration sequence in the motion information of the target in Track, the acceleration prediction curve is fitted using the least squares method and the average speed is calculated to correct the speed parameter of the input of the Kalman filter. The Kalman filter generates the prediction frame of the next frame. The detection frame of the Detection of the third frame and the prediction frame of Track are then matched by IOU to calculate the cost matrix. Among them, the tracker of the prediction frame that matches more than 3 times in a row is in the confirmed state (Confirmed), otherwise it is in the unconfirmed state (Unconfirmed);
[0063] E3: Input all the cost matrices calculated in E2 into the Hungarian algorithm to obtain the matching results. There are three matching results at this time. The first is the mismatch tracker, that is, the prediction frame stored by the tracker does not find a matching detection frame. This result is because the target leaves the camera area or is blocked for some reason. The detection frame of the target does not exist, and only the prediction frame of the target in the previous frame remains. If the mismatch tracker is in an unconfirmed state, the mismatch tracker is deleted directly. If the mismatch tracker is in a confirmed state and its survival time is less than t-max, the mismatch tracker Track is retained. If its survival time is greater than t-max, the mismatch tracker is deleted; the second is a mismatch detector, that is, the detection frame of the detector is a newly appeared target, and no matching tracker is found. A new tracker Track is initialized for the mismatch detector; the third is that the tracker and the detector match successfully, indicating that the target tracking of the previous frame and the next frame is successful, and the Kalman filter is updated;
[0064] E4: Repeat E2 to E3 until a confirmed tracker appears, then continue to E5.
[0065] E5: Perform cascade matching operations on the confirmed Track and the target detected by the detector through appearance features and motion information.
[0066] E6: Obtain three results of cascade matching. The first result is that Track and Detection are matched successfully, and the Kalman filter is updated through Track and Detection. The second result is the mismatch detector, and the last result is the mismatch tracker. Then, all the unconfirmed Tracks and mismatched Tracks are matched one by one with the mismatch detector through IOU matching, and then the cost matrix is calculated based on the result of IOU matching (it should be noted that this is the cost matrix of the original DeepSORT method, which is different from the Hungarian algorithm in the previous step 4); then, the mismatch detector is matched with all the unconfirmed Tracks and mismatched Tracks using IOU matching to obtain the matching result, and then the cost matrix is calculated;
[0067] E7: Input the cost matrix obtained in E6 into the Hungarian algorithm to obtain the matching result. There are three matching results at this time. The first result is the mismatch tracker. If the mismatch tracker is in an unconfirmed state, delete the mismatch tracker directly. If the mismatch tracker is in a confirmed state, determine whether its survival time is less than t-max. If the condition is met, retain the mismatch tracker Track; if the condition is not met, delete the mismatch tracker immediately; the second result is the mismatch detector, initialize a new tracker Track for the mismatch detector; the third is that the detector and tracker match successfully, indicating that the target tracking of the previous frame and the next frame is successful, and the Kalman filter is updated;
[0068] E8: Repeat E5 to E7 until the video detection is completed.
[0069] Furthermore, the adaptive fusion feature in step 3 is defined as follows:
[0070] In order to reduce the impact of camera distortion on target matching through appearance features, a weighting function is designed to address the problem that the appearance features of the target are less distorted at the center of the field of view and more distorted at the edge of the field of view:
[0071]
[0072] The maximum value of this function is 1 when ρ=0, and it approaches 0 when ρ→∞, so the value range of this function is (0,1];
[0073] Based on the fact that the target's appearance features have a small degree of distortion at the center of the camera's field of view, and thus contribute more to the target matching task, while the target's appearance features have a large degree of distortion at the edge of the camera's field of view, and thus contribute less to the target matching task, we designed the following feature that adaptively fuses the target's spatial motion features and appearance features based on the distance from the target to the field of view center:
[0074]
[0075] in, Indicates the distance between the target l and the center point of the camera field of view, represents the dimension concatenation operator of vectors, is the normalized spatial motion feature of the target l, is the normalized appearance feature of target l.
[0076] Furthermore, the matching method of step 4 is based on the Hungarian algorithm, which specifically includes steps S1 to S7:
[0077] S1: Use YOLOv5 to detect multiple camera targets and obtain the lth target p under camera p at time tl Spatial motion characteristics and the fth q under camera q f Spatial motion characteristics in, Represents the image target p at time t l The horizontal and vertical coordinates of the center point of the detection frame, Represents the width and height of the detection box respectively, Indicates the change rate of each parameter;
[0078] S2: Use ResNet network to extract target p l Appearance characteristics and target q l Appearance characteristics Among them, p l ∈{1,…,H p},q l ∈{1,…,H q}, κ and κ′ are dimensions, H p is the number of people observed by camera p, H q is the number of people observed by camera q;
[0079] S3: Calculate the target p of camera p l Adaptive fusion features and target q of camera q f Adaptive fusion of features
[0080] S4: Calculate adaptive fusion features and The Mahalanobis distance of:
[0081]
[0082] in, for and The covariance matrix between ;
[0083] S5: According to the Mahalanobis distance threshold th (1) Make a judgment and filter out impossible target matches:
[0084]
[0085] in, A value of 1 indicates that the adaptive fusion feature matching between the target in camera p and the target in other cameras q is successful, and a value of 0 indicates that the matching fails;
[0086] S6: According to Get the cost matrix H(p,q):
[0087]
[0088] S7: Substitute the cost matrix H(p,q) into the Hungarian algorithm to establish target associations between cameras, obtain successfully matched targets, and assign identity IDs.
[0089] Furthermore, the modified DeepSORT method based on average speed prediction in step 5 includes steps C1 to C5:
[0090] C1: Get the camera at time t k The detection value of the Kalman filter input state of the acquired image frame:
[0091] MT k =(x k ,y k ,w k ,h k ,v xk ,v yk ,v wk ,v hk ) T
[0092] Among them, x k ,y k Respectively represent the horizontal and vertical coordinates of the center point of the target detection frame in the image frame, w k , h k Respectively represent the width and height of the target detection box in the image frame, v xk , v yk , v wk , v hk y k Represents x k ,y k , w k , h k The rate of change;
[0093] C2: Based on the target detection box in C1 at the previous k-1 moments t1,...,t k-1 Acceleration data sequence of acquired image frames Fitting the acceleration prediction function using the least squares method related formula
[0094]
[0095] Among them, i∈{x,y,w,h}, x,y,w,h are the four parameters of the Kalman filter input state, x,y represent the horizontal and vertical coordinates of the center point of the target detection frame, w,h represent the width and height of the target detection frame respectively; ai1 ,…,a i(k-1) are times t1,...,t k-1 The acceleration data corresponding to the four parameters of the Kalman filter input state, and is the fitting coefficient of the acceleration prediction function;
[0096] C3: Using the acceleration prediction curve, the time t k The input state detection value of the Kalman filter of the acquired image frame is corrected to obtain the input state detection value of the Kalman filter after correction:
[0097] MT′ k =(x k ,y k ,w k ,h k ,v′ xk ,v′ yk ,v′ wk ,v′ hk ) T
[0098] Among them, v′ ik is the time t after correction k The change speed of the four parameters corresponding to the Kalman filter input state, Δt is the time difference between two image frames;
[0099] C4: Calculate time t k Kalman filter inputs the acceleration data corresponding to the four parameters of the state and updates the time t k The acceleration data sequence of the acquired image frames:
[0100]
[0101] Among them, v′ i(k-1) is the time t after correction k-1 The speed of change of the four parameters corresponding to the Kalman filter input state;
[0102] C5: Let k=k+1 and repeat C1 to C5 until the video ends.
[0103] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any appropriate manner without contradiction. To avoid unnecessary repetition, the present invention will not further describe various possible combinations.
[0104] The present invention also proposes a multi-sensor target swarm intelligence tracking system based on modified DeepSORT, which specifically includes:
[0105] Multiple cameras for real-time surveillance video of the mission area;
[0106] The target matching unit is used to obtain the matching target of the current video frame between cameras and assign an identity ID;
[0107] The target tracking unit is used to track the target assigned with the identity ID in the current video frame using the DeepSORT method for each camera.
[0108] The technical solution of the above-mentioned three-dimensional target detection system is similar to the aforementioned method and will not be repeated here.
[0109] Based on the same technical solution, the present invention also discloses a computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions, which, when executed by a computing device, enable the computing device to execute the above-mentioned multi-sensor target swarm intelligence tracking method based on the modified DeepSORT.
[0110] Based on the same technical solution, the present invention also discloses a computing device, including one or more processors, one or more memories and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing the above-mentioned multi-sensor target swarm intelligence tracking method based on the modified DeepSORT.
[0111] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0112] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0113] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0114] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0115] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.
[0116] The above embodiments are only for illustrating the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the present invention.
Claims
1. A multi-sensor target swarm tracking method based on modified DeepSORT, characterized by: This method uses multiple cameras to track targets in the mission area, and specifically includes the following steps: Step 1: Use each camera to obtain real-time surveillance video of the task area; Step 2: Using the matching method based on the Hungarian algorithm, the matching target of the current video frame between the cameras is obtained and assigned an identity ID; Step 3: For each camera, use the DeepSORT method to track the target assigned to the ID in the current video frame until the target tracking task is completed. The DeepSORT method uses the least squares method to fit the acceleration prediction curve, and then corrects the input state detection value of the Kalman filter. The step 2 is specifically as follows: Step 2.1: Use YOLOv5 to perform target detection on the current image frames obtained by each camera to obtain the target information in each image frame; Step 2.2: Input the target information obtained in step 2.1 into the ResNet residual network to extract the appearance features of the target; Step 2.3: Obtain the spatial motion characteristics of the target based on the camera frame rate and the target information obtained in step 2.1; Step 2.4: Combine the appearance features of the target in step 2.2 and the spatial motion features of the target in step 2.3 to form an adaptive fusion feature of the target; Step 2.5: Using the adaptive fusion features of the target in step 2.4, calculate the Mahalanobis distance and cost matrix; Step 2.6: Use the Hungarian algorithm to match the targets between cameras at the same time, obtain the matching targets of the current video frame between the cameras and assign them an identity ID.
2. The multi-sensor target swarm tracking method based on modified DeepSORT according to claim 1, characterized in that: The adaptive fusion feature of the target in step 2.4 is defined as follows: Among them, ψ represents the set weight, Represents the vector dimension splicing operator, MT l is the spatial motion feature of target l, R l is the appearance feature of target l.
3. The multi-sensor target swarm tracking method based on modified DeepSORT according to claim 2, characterized in that: The expression for setting the weight ψ is as follows: in, Indicates the distance between the target l and the center of the camera's field of view.
4. The multi-sensor target swarm tracking method based on modified DeepSORT according to claim 1, characterized in that: The expression of the Mahalanobis distance in step 2.5 is as follows: in, is the lth target p under camera p at time t l Adaptive fusion features and the fth target q under camera q f Adaptive fusion features Mahalanobis distance; for and The covariance matrix between .
5. The multi-sensor target swarm tracking method based on modified DeepSORT according to claim 4 is characterized in that: The expression of the cost matrix in step 2.5 is as follows: Among them, H(p,q) is the cost matrix of camera p and camera q at time t, l∈{1,…,H p }, f∈{1,…,H q }, H p is the number of targets under camera p at time t, H q is the number of targets under camera q at time t; A value of 1 indicates that the adaptive fusion features of the target in camera p and the targets in other cameras q are successfully matched, and a value of 0 indicates that the match fails.
6. The multi-sensor target swarm tracking method based on modified DeepSORT according to claim 1, characterized in that: The step 3 includes: C1: Get the camera at time t k The detection value of the Kalman filter input state of the acquired image frame: MT k =(x k ,y k ,w k ,h k ,v xk ,v yk ,v wk ,v hk ) T Among them, x k ,y k Respectively represent the horizontal and vertical coordinates of the center point of the target detection frame in the image frame, w k , h k Respectively represent the width and height of the target detection box in the image frame, v xk , v yk , v wk , v hk Represents x k ,y k , w k , h k The rate of change; C2: Based on the target detection box in C1 at the previous k-1 moments t1,...,t k-1 Acceleration data sequence of acquired image frames Fitting the acceleration prediction curve using the least squares method Among them, i∈{x,y,w,h}, x,y,w,h are the four parameters of the Kalman filter input state, x,y represent the horizontal and vertical coordinates of the center point of the target detection frame, w,h represent the width and height of the target detection frame respectively; a i1 ,…,a i(k-1) are times t1,...,t k-1 The acceleration data corresponding to the four parameters of the Kalman filter input state, The acceleration fitting prediction curve of the four parameters representing the Kalman filter input state, and is the fitting coefficient of the acceleration fitting prediction curve; C3: Using the acceleration prediction curve, the time t k The input state detection value of the Kalman filter of the acquired image frame is corrected to obtain the input state detection value of the Kalman filter after correction: MT′ k =(x k ,y k ,w k ,h k ,v′ xk ,v′ yk ,v′ wk ,v′ hk ) T Among them, v′ ik is the time t after correction k The change speed of the four parameters corresponding to the Kalman filter input state, Δt is the time difference between two image frames; C4: Calculate time t k Kalman filter inputs the acceleration data corresponding to the four parameters of the state and updates the time t k Acceleration data sequence of acquired image frames Among them, v′ i(k-1) is the time t after correction k-1 The speed of change of the four parameters corresponding to the Kalman filter input state; C5: Let k=k+1 and repeat C1 to C5 until the video ends.
7. A multi-sensor target swarm tracking system based on modified DeepSORT, characterized by: The system performs target tracking based on the method according to any one of claims 1 to 6, specifically comprising: Multiple cameras for real-time surveillance video of the mission area; The target matching unit is used to obtain the matching target of the current video frame between cameras and assign an identity ID; The target tracking unit is used to track the target assigned with the identity ID in the current video frame using the DeepSORT method for each camera.
8. A computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions, characterized in that: When the instructions are executed by a computing device, the computing device is caused to perform the method according to any one of claims 1 to 6.
9. A computing device, characterized in that The method comprises one or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing the method according to any one of claims 1 to 6.