Target tracking methods, devices, equipment, and media based on multi-view fusion

By using a multi-view fusion method, a global coordinate system is established and a spatiotemporal graph neural network is used to handle the occlusion problem, which solves the problem of incomplete target tracking in existing technologies and achieves global tracking and accurate prediction of human targets in the scene.

CN122023836BActive Publication Date: 2026-07-31ZHONGKE HONGTUO (SUZHOU) INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHONGKE HONGTUO (SUZHOU) INTELLIGENT TECH CO LTD
Filing Date
2026-04-08
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technical solutions cannot fully cover the target to be tracked in the scene, and often the target tracking ID is lost, changed or incorrectly swapped due to two-dimensional occlusion problems.

Method used

A multi-view fusion method is adopted to establish a global coordinate system. Human targets are detected by using two-dimensional images and three-dimensional point clouds from a depth camera. The observation input results are processed by a spatiotemporal graph neural network, and the confidence score is calculated by combining the degree of occlusion, so as to achieve multi-view fusion and global tracking.

Benefits of technology

It alleviates the problem of 2D occlusion, enables global tracking of human targets in the scene, and improves the accuracy of tracking as well as the versatility and scalability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122023836B_ABST
    Figure CN122023836B_ABST
Patent Text Reader

Abstract

This application relates to a target tracking method, apparatus, device, and medium based on multi-view fusion. This application can organically fuse tracking prediction results from multiple camera views at different depths to predict the position and trajectory of each target in the next moment. By modeling human targets in the scene as equivalent to neural network graph nodes in a spatiotemporal manner and calculating graded decay confidence based on the degree of occlusion, the two-dimensional occlusion problem in target tracking is largely alleviated. Based on this, the prediction results from different camera views are organically fused to achieve global tracking of human targets in the scene. Furthermore, this application is compatible with multiple different views, and the system has strong versatility and scalability. More camera views also contribute to more accurate prediction and tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to a target tracking method, apparatus, device, and medium based on multi-view fusion. Background Technology

[0002] Generally speaking, existing technical solutions are limited to a single camera viewpoint and cannot fully cover the target to be tracked in the scene. At the same time, existing technical solutions often encounter two-dimensional occlusion problems when tracking targets, that is, the tracking ID is lost, jumps or incorrectly swapped due to mutual occlusion between the target and obstacles or between the target and other targets in the two-dimensional image. Summary of the Invention

[0003] Based on this, a target tracking method, apparatus, device and medium based on multi-view fusion is provided to solve the technical problem that existing technical solutions cannot cover all the targets to be tracked in the scene well, and the target tracking ID is often lost, changed or incorrectly swapped due to two-dimensional occlusion.

[0004] On the one hand, a target tracking method based on multi-view fusion is provided, the method comprising:

[0005] A global coordinate system is established for the two-dimensional plane where the application scenario is located. A sub-coordinate system is established for each depth camera that acquires two-dimensional images and three-dimensional point clouds in the application scenario. The target position coordinates under each sub-coordinate system are then converted to the global coordinate system.

[0006] Human targets are detected based on the two-dimensional images and three-dimensional point clouds collected by each depth camera, and a global identity tag is bound to the position coordinates of each human target in the global coordinate system.

[0007] For each human target corresponding to a global identity tag, the observation input result and occlusion degree of each human target at the current moment are determined from the camera view of each depth camera based on the two-dimensional images and three-dimensional point clouds collected by multiple depth cameras.

[0008] A spatiotemporal graph neural network is used to process the observation input results of each human target under the camera view of each depth camera to obtain the prediction result of the next moment. The confidence of the prediction result is determined according to the degree of occlusion corresponding to the observation input result. The prediction result is then fused from multiple views to obtain the fused prediction output result.

[0009] The predicted output for each human target is integrated with the observation input, and the visibility of each human body part is deduced by combining the confidence of the predicted result for each depth camera, thereby achieving global tracking of each human target.

[0010] In one embodiment, establishing a global coordinate system for the two-dimensional plane where the application scenario is located, establishing a sub-coordinate system for each depth camera that acquires two-dimensional images and three-dimensional point clouds in the application scenario, and converting the target position coordinates in each sub-coordinate system to the global coordinate system includes:

[0011] A global coordinate system is established for the two-dimensional plane where the application scenario is located. The origin coordinates and X-axis of the global coordinate system are determined. The X-axis is rotated 90 degrees counterclockwise in the world coordinate plane with the origin of the global coordinate system as the rotation center to obtain the Y-axis.

[0012] The position and orientation of each depth camera that acquires 2D images and 3D point clouds in the application scenario are obtained. The viewing angle is determined according to the orientation of the depth camera. A sub-coordinate system is established in the global coordinate system for each depth camera's viewing angle.

[0013] When transforming the target position coordinates in each sub-coordinate system to the global coordinate system, the coordinates of the origin of the sub-coordinate system in the global coordinate system are obtained as follows: , where x O Let y be the x-coordinate of the origin of the sub-coordinate system in the global coordinate system. O Let the ordinate of the origin of the sub-coordinate system be the ordinate of the global coordinate system, and let the unit direction vector corresponding to the positive X-axis of the sub-coordinate system be the corresponding vector in the global coordinate system. , where x X Let y be the x-coordinate of the unit direction vector corresponding to the positive X-axis of the sub-coordinate system in the global coordinate system. X Let X be the ordinate of the unit direction vector corresponding to the positive X-axis of the sub-coordinate system in the global coordinate system. The Y-axis of the sub-coordinate system is obtained by rotating the X-axis counterclockwise by 90 degrees around the origin of the sub-coordinate system in the sub-coordinate plane. The unit direction vector corresponding to the positive Y-axis of the sub-coordinate system is: , It is a rotation matrix;

[0014] If the target position coordinates in the target sub-coordinate system are , where x S Let y be the x-coordinate of the target position in the target sub-coordinate system. S Let be the ordinate of the target position in the target sub-coordinate system, then the target position coordinates correspond to coordinate P in the global coordinate system. W for:

[0015] , where x W Let y be the x-coordinate of the target position in the global coordinate system.W The target position is represented by its ordinate in the global coordinate system.

[0016] In one embodiment, the step of detecting human targets based on the two-dimensional images and three-dimensional point clouds acquired by each depth camera, and binding a global identity tag to the position coordinates of each human target in the global coordinate system, includes:

[0017] For each two-dimensional image captured by a depth camera, a pre-trained deep neural network for human target detection is invoked to detect the key points of each human target and obtain the human targets in the application scenario.

[0018] After detecting a human target, the three-dimensional coordinates of the center of the human target are obtained from the three-dimensional point cloud collected by each depth camera. The vertical axis coordinates of the three-dimensional coordinates are removed and mapped to the two-dimensional plane camera view sub-coordinate system to obtain the position coordinates of the human target in the sub-coordinate system.

[0019] The position coordinates of the human target in the sub-coordinate system are converted to obtain the position coordinates of the human target in the global coordinate system;

[0020] When a new human target is detected in the camera view of any depth camera, the position coordinates of the new human target in the global coordinate system are calculated. Among the human targets detected by each of the other depth cameras, the nearest human target with the shortest Euclidean distance between the two position coordinates and within a preset spacing range is matched.

[0021] If the new human target is matched with its nearest human target in each of the other depth cameras, then a global identity label is bound to the new human target.

[0022] If the new human target fails to find the nearest human target in each of the remaining depth cameras, its global identity label is set to pending, and the pending human target is used as the nearest human target within a preset distance range to match the next newly appearing human target.

[0023] In one embodiment, the step of using a spatiotemporal graph neural network to process the observation input results of each human target from the camera viewpoint of each depth camera to obtain the prediction result for the next moment includes:

[0024] For each depth camera, a spatiotemporal graph neural network is pre-constructed according to the task requirements and trained using the collected target position trajectory data. The spatiotemporal graph neural network is used to predict the future motion trajectory and trend of human targets in the application scenario to obtain the prediction result of the next moment.

[0025] In the spatiotemporal graph neural network corresponding to each depth camera, each graph node is associated with a human target from the camera's viewpoint. Each graph node records a global identity label bound to the corresponding human target. Therefore, the global identity labels for the N human targets are: , , … ;

[0026] The input node features of the spatiotemporal graph neural network are the position coordinates (x, y) and velocity vector (△x, △y) of the human target under the current camera view. It is divided into two parts: {the predicted value P calculated by the spatiotemporal graph neural network at the previous moment for the current moment} and {the observed value S collected and detected by the depth camera at the current moment}. Among them, x is the horizontal coordinate of the human target under the current camera view, y is the vertical coordinate of the human target under the current camera view, △x is the velocity vector of the human target under the horizontal coordinate under the current camera view, and △y is the velocity vector of the human target under the vertical coordinate under the current camera view.

[0027] The input node feature matrix of the spatiotemporal graph neural network has a dimension of T×N×8, where N is the number of target nodes and T is the length of the time window. For a single time t, the dimension of the input feature matrix is ​​N×8, as shown in the following form:

[0028] ;in Let be the predicted value of the Nth target node on the horizontal axis at time t. Let be the predicted value of the Nth target node on the ordinate at time t. Let be the predicted velocity vector of the Nth target node on the horizontal axis at time t. Let be the predicted velocity vector of the Nth target node at time t on the ordinate. Let be the observed value of the Nth target node on the horizontal axis at time t. Let be the observed value of the Nth target node on the ordinate at time t. Let t be the observed velocity vector of the Nth target node on the horizontal axis. Let be the observed velocity vector of the Nth target node at time t on the vertical axis;

[0029] The node adjacency matrix is ​​a symmetric matrix with a dimension of N×N, where the edge weight connecting any two graph nodes is the Euclidean distance between the corresponding two human targets in the sub-coordinate system.

[0030] The output matrix of the spatiotemporal graph neural network has a dimension of 1×N×4. The prediction result output by the spatiotemporal graph neural network is the predicted value P of the position coordinates (x,y) and velocity vectors (△x,△y) of all human targets in the current depth camera's view at the next time t+1.

[0031] In one embodiment, the step of using a spatiotemporal graph neural network to process the observation input results of each human target from the camera viewpoint of each depth camera to obtain the prediction result for the next moment further includes:

[0032] The edge weight connecting any two graph nodes is calculated using the observed value S, and the edge weight connecting any two graph nodes is:

[0033] ;in W represents the edge weight matrix between any two graph nodes in the view of the k-th subsystem camera at the current time t. NN This represents the edge weight between the Nth graph node and the Nth graph node;

[0034] The prediction result output by the spatiotemporal graph neural network is:

[0035] ;in This represents the prediction result of the k-th subsystem at the next time step t+1. Let be the predicted value of the Nth target node on the horizontal axis at time t+1. Let be the predicted value of the Nth target node on the ordinate at time t+1. Let be the predicted velocity vector of the Nth target node on the horizontal axis at time t+1. Let t+1 be the predicted velocity vector of the Nth target node on the ordinate.

[0036] In one embodiment, determining the confidence level of the prediction result based on the occlusion degree corresponding to the observed input result, and performing multi-view fusion on the prediction result to obtain the fused prediction output result includes:

[0037] Obtain the total number K of depth cameras in the application scenario, and treat each depth camera as a subsystem;

[0038] The confidence level of the prediction result is determined based on the degree of occlusion corresponding to the observed input result, and the confidence vector of the prediction result is calculated. This represents the confidence vector of the prediction results of all human targets in the view of the k-th subsystem camera at the current time t, calculated according to their respective occlusion levels.

[0039] The prediction output is obtained by multi-view fusion of the prediction result confidence vector and the prediction result. The prediction output results are described below. This represents the predicted position coordinates (x, y) and velocity vectors (Δx, Δy) of all targets after fusing the prediction results of all subsystems at the next time t+1. This represents the predicted position coordinates (x, y) and velocity vectors (Δx, Δy) of all targets calculated by the k-th subsystem at the next time t+1.

[0040] The fused prediction output for the i-th single human target is represented as follows: ;in, This represents the predicted position coordinates (x, y) and velocity vector (Δx, Δy) of the i-th human target after fusing the prediction outputs of all subsystems at the next time t+1. This represents the confidence level of the predicted output result calculated based on the degree of occlusion of the i-th human target in the view of the k-th subsystem camera at the current time t. This represents the sum of confidence scores of the predicted output results calculated based on the degree of occlusion of the i-th target in the camera view of all subsystems at the current time t.

[0041] In one embodiment, the step of integrating the predicted output result corresponding to each human target with the observation input result, and combining the confidence level of the predicted result corresponding to each depth camera to deduce the visibility of each human body part of the human target, includes:

[0042] For the i-th human target in the view of the k-th subsystem camera at the current time t, obtain the human key points of the human target;

[0043] The confidence level of the prediction result of the human target at the current time t under the camera view of the corresponding subsystem. The following formula applies graded attenuation based on the degree of occlusion: ;in, The value ranges from 0 to 1; For head weight, Weight for the torso; Weighted by left hand Right-hand weight; The value of the Boolean variable is visible in the header; The value of the visible Boolean variable is for the torso. The value of the left-hand visible Boolean variable; The value of the right-hand visible Boolean variable;

[0044] The confidence level of the human body key points is determined based on the confidence level of the prediction results corresponding to each depth camera at the current time t, and the visibility of each part of the human body target is deduced based on the confidence level of the human body key points.

[0045] On the other hand, a target tracking device based on multi-view fusion is provided, the device comprising:

[0046] The coordinate transformation module is used to establish a global coordinate system for the two-dimensional plane where the application scene is located, establish a sub-coordinate system for each depth camera that acquires two-dimensional images and three-dimensional point clouds in the application scene, and convert the target position coordinates under each sub-coordinate system to the global coordinate system.

[0047] The target detection and labeling module is used to detect human targets based on the two-dimensional images and three-dimensional point clouds collected by each depth camera, and bind a global identity label to the position coordinates of each human target in the global coordinate system.

[0048] The observation and occlusion analysis module is used to determine the observation input results and occlusion degree of each human target corresponding to each global identity tag at the current moment from the camera view of each depth camera based on the two-dimensional images and three-dimensional point clouds collected by multiple depth cameras.

[0049] The spatiotemporal graph processing and fusion module is used to process the observation input results of each human target under the camera view of each depth camera using a spatiotemporal graph neural network to obtain the prediction result for the next moment, determine the confidence of the prediction result according to the degree of occlusion corresponding to the observation input result, and perform multi-view fusion on the prediction result to obtain the fused prediction output result.

[0050] The global tracking and visibility estimation module is used to integrate the prediction output results corresponding to each human target with the observation input results, and combine the confidence of the prediction results corresponding to each depth camera to infer the visibility of each human body part of the human target, thereby realizing global tracking of each human target.

[0051] In another aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of a target tracking method based on multi-view fusion.

[0052] In another aspect, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of a target tracking method based on multi-view fusion.

[0053] The aforementioned multi-view fusion-based target tracking method, apparatus, device, and medium can organically integrate tracking and prediction results from multiple camera perspectives at different depths to predict the position and trajectory of each target at the next moment. By modeling human targets in the scene as equivalent to neural network graph nodes in a spatiotemporal manner and calculating graded decay confidence based on the degree of occlusion, the two-dimensional occlusion problem in target tracking is largely alleviated. Based on this, the prediction results from different camera perspectives are organically fused to achieve global tracking of human targets in the scene. Furthermore, this application is compatible with multiple different perspectives, the system has strong versatility and scalability, and more camera perspectives also contribute to more accurate prediction and tracking. Attached Figure Description

[0054] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0055] Figure 1 This is a logic diagram of a target tracking method based on multi-view fusion in one embodiment of this application;

[0056] Figure 2 This is a flowchart illustrating a target tracking method based on multi-view fusion in one embodiment of this application;

[0057] Figure 3 This is a schematic diagram of the input and output flow of a spatiotemporal graph neural network in one embodiment of this application;

[0058] Figure 4 This is a schematic diagram illustrating the conversion of the target position coordinates in each sub-coordinate system to the global coordinate system in one embodiment of this application;

[0059] Figure 5 This is a structural block diagram of a target tracking device based on multi-view fusion in one embodiment of this application;

[0060] Figure 6 This is an internal structural diagram of a computer device in one embodiment of this application. Detailed Implementation

[0061] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0062] In one embodiment, such as Figure 1 , Figure 2, Figure 3 As shown, a target tracking method based on multi-view fusion is provided, including the following steps:

[0063] Step S1: Establish a global coordinate system for the two-dimensional plane where the application scene is located, establish a sub-coordinate system for each depth camera that acquires two-dimensional images and three-dimensional point clouds in the application scene, and convert the target position coordinates under each sub-coordinate system to the global coordinate system.

[0064] Step S2: Detect human targets based on the two-dimensional images and three-dimensional point clouds acquired by each depth camera, and bind a global identity tag to the position coordinates of each human target in the global coordinate system.

[0065] Step S3: For each human target corresponding to a global identity tag, determine the observation input result and occlusion degree of each human target at the current moment from the camera view of each depth camera based on the two-dimensional images and three-dimensional point clouds collected by multiple depth cameras.

[0066] Step S4: The spatiotemporal graph neural network is used to process the observation input results of each human target under the camera view of each depth camera to obtain the prediction result of the next moment. The confidence of the prediction result is determined according to the degree of occlusion corresponding to the observation input result. The prediction result is fused by multiple views to obtain the fused prediction output result.

[0067] Step S5: Integrate the predicted output result corresponding to each human target with the observation input result, and combine the confidence of the prediction result corresponding to each depth camera to deduce the visibility of each human body part of the human target, thereby realizing global tracking of each human target.

[0068] Specifically, this system can organically integrate tracking and prediction results from multiple camera perspectives at different depths to predict the position and trajectory of each target at the next moment. By modeling human targets in the scene as equivalent to nodes in a neural network graph and calculating graded decay confidence based on the degree of occlusion, the two-dimensional occlusion problem in target tracking is largely alleviated. Based on this, the prediction results from different camera perspectives are organically integrated to achieve global tracking of human targets in the scene. Furthermore, this application is compatible with multiple different perspectives, giving the system strong versatility and scalability. More camera perspectives also contribute to more accurate prediction and tracking.

[0069] like Figure 4 As shown, in this embodiment, the step of establishing a global coordinate system for the two-dimensional plane where the application scene is located, establishing a sub-coordinate system for each depth camera that acquires two-dimensional images and three-dimensional point clouds in the application scene, and converting the target position coordinates under each sub-coordinate system to the global coordinate system includes:

[0070] A global coordinate system is established for the two-dimensional plane where the application scenario is located. The origin coordinates and X-axis of the global coordinate system are determined. The X-axis is rotated 90 degrees counterclockwise in the world coordinate plane with the origin of the global coordinate system as the rotation center to obtain the Y-axis.

[0071] The position and orientation of each depth camera that acquires 2D images and 3D point clouds in the application scenario are obtained. The viewing angle is determined according to the orientation of the depth camera. A sub-coordinate system is established in the global coordinate system for each depth camera's viewing angle.

[0072] When transforming the target position coordinates in each sub-coordinate system to the global coordinate system, the coordinates of the origin of the sub-coordinate system in the global coordinate system are obtained as follows: , where x O Let y be the x-coordinate of the origin of the sub-coordinate system in the global coordinate system. O Let the ordinate of the origin of the sub-coordinate system be the ordinate of the global coordinate system, and let the unit direction vector corresponding to the positive X-axis of the sub-coordinate system be the corresponding vector in the global coordinate system. , where x X Let y be the x-coordinate of the unit direction vector corresponding to the positive X-axis of the sub-coordinate system in the global coordinate system. X Let X be the ordinate of the unit direction vector corresponding to the positive X-axis of the sub-coordinate system in the global coordinate system. The Y-axis of the sub-coordinate system is obtained by rotating the X-axis counterclockwise by 90 degrees around the origin of the sub-coordinate system in the sub-coordinate plane. The unit direction vector corresponding to the positive Y-axis of the sub-coordinate system is: , It is a rotation matrix;

[0073] If the target position coordinates in the target sub-coordinate system are , where x S Let y be the x-coordinate of the target position in the target sub-coordinate system. S Let be the ordinate of the target position in the target sub-coordinate system, then the target position coordinates correspond to coordinate P in the global coordinate system. W for:

[0074] , where x W Let y be the x-coordinate of the target position in the global coordinate system. W The target position is represented by its ordinate in the global coordinate system.

[0075] like Figure 4As shown, two subsystems are set up in the application scenario, each equipped with a depth camera to acquire 2D images and 3D point clouds of the scene. Each depth camera has a different position and orientation, with its viewpoint facing the scene horizontally, and they can be parallel or intersecting each other. The known target position coordinates in each camera's viewpoint sub-coordinate system can be converted to the global coordinate system.

[0076] In this embodiment, the step of detecting human targets based on the two-dimensional images and three-dimensional point clouds acquired by each depth camera, and binding a global identity tag to the position coordinates of each human target in the global coordinate system, includes:

[0077] For each two-dimensional image captured by a depth camera, a pre-trained deep neural network for human target detection is invoked to detect the key points of each human target and obtain the human targets in the application scenario.

[0078] After detecting a human target, the three-dimensional coordinates of the center of the human target are obtained from the three-dimensional point cloud collected by each depth camera. The vertical axis coordinates of the three-dimensional coordinates are removed and mapped to the two-dimensional plane camera view sub-coordinate system to obtain the position coordinates of the human target in the sub-coordinate system.

[0079] The position coordinates of the human target in the sub-coordinate system are converted to obtain the position coordinates of the human target in the global coordinate system;

[0080] When a new human target is detected in the camera view of any depth camera, the position coordinates of the new human target in the global coordinate system are calculated. Among the human targets detected by each of the other depth cameras, the nearest human target with the shortest Euclidean distance between the two position coordinates and within a preset spacing range is matched.

[0081] If the new human target is matched with its nearest human target in each of the other depth cameras, then a global identity label is bound to the new human target.

[0082] If the new human target fails to find the nearest human target in each of the remaining depth cameras, its global identity label is set to pending, and the pending human target is used as the nearest human target within a preset distance range to match the next newly appearing human target.

[0083] The target detection network used can be YOLOv10. Based on this, a pre-trained deep neural network for human keypoint detection is invoked to detect the keypoints of each human target. The human keypoint detection network used can be HRNet.

[0084] After detecting a human target in the image, the 3D coordinates of the target's center are obtained from the acquired 3D point cloud. After removing the vertical axis coordinates from the 3D coordinates and mapping them to the 2D plane camera viewpoint sub-coordinate system, the position coordinates of the human target in that sub-coordinate system can be obtained. These coordinates are then further converted to the global coordinate system. It is important to note that each human target in the scene can have its coordinates in the global coordinate system calculated similarly from different subsystem camera views. These coordinates may not be completely identical, but the differences are minor due to some measurement and calculation errors.

[0085] The preset spacing range is preferably within a diameter of 40cm. When binding a global identity tag to a human target, if the global identity tag of the first human target is... The global identity tag for the second human target is The global identity tag for the third human target is And so on.

[0086] During system operation, each subsystem uses this method to detect and calculate the position coordinates of all human targets in the global coordinate system from its own camera's perspective in real time.

[0087] Please see Figure 3 In this embodiment, the step of using a spatiotemporal graph neural network to process the observation input results of each human target under the camera viewpoint of each depth camera to obtain the prediction result for the next moment includes:

[0088] For each depth camera, a spatiotemporal graph neural network is pre-constructed according to the task requirements and trained using the collected target position trajectory data. The spatiotemporal graph neural network is used to predict the future motion trajectory and trend of human targets in the application scenario to obtain the prediction result of the next moment.

[0089] In the spatiotemporal graph neural network corresponding to each depth camera, each graph node is associated with a human target from the camera's viewpoint. Each graph node records a global identity label bound to the corresponding human target. Therefore, the global identity labels for the N human targets are: , , … ;

[0090] The input node features of the spatiotemporal graph neural network are the position coordinates (x, y) and velocity vector (△x, △y) of the human target under the current camera view. It is divided into two parts: {the predicted value P calculated by the spatiotemporal graph neural network at the previous moment for the current moment} and {the observed value S collected and detected by the depth camera at the current moment}. Among them, x is the horizontal coordinate of the human target under the current camera view, y is the vertical coordinate of the human target under the current camera view, △x is the velocity vector of the human target under the horizontal coordinate under the current camera view, and △y is the velocity vector of the human target under the vertical coordinate under the current camera view.

[0091] The input node feature matrix of the spatiotemporal graph neural network has a dimension of T×N×8, where N is the number of target nodes and T is the length of the time window. For a single time t, the dimension of the input feature matrix is ​​N×8, as shown in the following form:

[0092] ;in Let be the predicted value of the Nth target node on the horizontal axis at time t. Let be the predicted value of the Nth target node on the ordinate at time t. Let be the predicted velocity vector of the Nth target node on the horizontal axis at time t. Let be the predicted velocity vector of the Nth target node at time t on the ordinate. Let be the observed value of the Nth target node on the horizontal axis at time t. Let be the observed value of the Nth target node on the ordinate at time t. Let t be the observed velocity vector of the Nth target node on the horizontal axis. Let be the observed velocity vector of the Nth target node at time t on the vertical axis;

[0093] The node adjacency matrix is ​​a symmetric matrix with a dimension of N×N, where the edge weight connecting any two graph nodes is the Euclidean distance between the corresponding two human targets in the sub-coordinate system.

[0094] The output matrix dimension of the spatiotemporal graph neural network is 1×N×(2+2), that is, 1×N×4. The prediction result output by the spatiotemporal graph neural network is the predicted value P of the position coordinates (x,y) and velocity vectors (△x,△y) of all human targets under the current depth camera's camera view at the next time t+1.

[0095] Spatiotemporal graph neural networks typically employ a network structure similar to StemGNN. Their basic principle is to transform the spatiotemporal domain into the frequency domain through Discrete Fourier Transform (DFT) and Graph Fourier Transform (GFT), while simultaneously capturing spatiotemporal dependencies in the frequency domain.

[0096] In this embodiment, the step of using a spatiotemporal graph neural network to process the observation input results of each human target from the camera viewpoint of each depth camera to obtain the prediction result for the next moment further includes:

[0097] The edge weight connecting any two graph nodes is calculated using the observed value S, and the edge weight connecting any two graph nodes is:

[0098] ;in W represents the edge weight matrix between any two graph nodes in the view of the k-th subsystem camera at the current time t. NN This represents the edge weight between the Nth graph node and the Nth graph node;

[0099] The prediction result output by the spatiotemporal graph neural network is:

[0100] ;in This represents the prediction result of the k-th subsystem at the next time step t+1. Let be the predicted value of the Nth target node on the horizontal axis at time t+1. Let be the predicted value of the Nth target node on the ordinate at time t+1. Let be the predicted velocity vector of the Nth target node on the horizontal axis at time t+1. Let t+1 be the predicted velocity vector of the Nth target node on the ordinate.

[0101] Each time the subsystem calls the spatiotemporal graph neural network to complete a forward inference, it can predict the position coordinates and velocity information of all human targets in the current camera view at the next moment. The output prediction value also constitutes the first part of the input node features of the spatiotemporal graph neural network of the subsystem at the next moment.

[0102] In this embodiment, determining the confidence level of the prediction result based on the occlusion degree corresponding to the observed input result, and performing multi-view fusion on the prediction result to obtain the fused prediction output result, includes:

[0103] Obtain the total number K of depth cameras in the application scenario, and treat each depth camera as a subsystem;

[0104] The confidence level of the prediction result is determined based on the degree of occlusion corresponding to the observed input result, and the confidence vector of the prediction result is calculated. This represents the confidence vector of the prediction results of all human targets in the view of the k-th subsystem camera at the current time t, calculated according to their respective occlusion levels.

[0105] The prediction output is obtained by multi-view fusion of the prediction result confidence vector and the prediction result. The prediction output results are described below. This represents the predicted position coordinates (x, y) and velocity vectors (Δx, Δy) of all targets after fusing the prediction results of all subsystems at the next time t+1. This represents the predicted position coordinates (x, y) and velocity vectors (Δx, Δy) of all targets calculated by the k-th subsystem at the next time t+1.

[0106] The fused prediction output for the i-th single human target is represented as follows: ;in, This represents the predicted position coordinates (x, y) and velocity vector (Δx, Δy) of the i-th human target after fusing the prediction outputs of all subsystems at the next time t+1. This represents the confidence level of the predicted output result calculated based on the degree of occlusion of the i-th human target in the view of the k-th subsystem camera at the current time t. This represents the sum of confidence scores of the predicted output results calculated based on the degree of occlusion of the i-th target in the camera view of all subsystems at the current time t.

[0107] In this embodiment, the step of integrating the predicted output result corresponding to each human target with the observation input result, and combining the confidence level of the predicted result corresponding to each depth camera to deduce the visibility of each human body part of the human target, includes:

[0108] For the i-th human target in the view of the k-th subsystem camera at the current time t, obtain the human key points of the human target;

[0109] The confidence level of the prediction result of the human target at the current time t under the camera view of the corresponding subsystem. The following formula applies graded attenuation based on the degree of occlusion: ;in, The value ranges from 0 to 1; For head weight, Weight for the torso; Weighted by left hand Right-hand weight; The value of the Boolean variable is visible in the header; The value of the visible Boolean variable is for the torso. The value of the left-hand visible Boolean variable; The value of the right-hand visible Boolean variable;

[0110] The confidence level of the human body key points is determined based on the confidence level of the prediction results corresponding to each depth camera at the current time t, and the visibility of each part of the human body target is deduced based on the confidence level of the human body key points.

[0111] in, The value ranges from 0 to 1; here, only the visibility of the upper body of the human target is considered, and the following values ​​can be used: head weight. trunk weight Left-hand weight Right-hand weight .

[0112] All are Boolean variables:

[0113] The head is considered visible when at least three out of five keypoints (nose, eyes, ears) within the head region have a confidence level greater than or equal to 0.7. Conversely, it is invisible. ;

[0114] The torso is considered visible when at least two of the four key points within the torso region (both shoulders + both sides of the hips) have a confidence level greater than or equal to 0.6. Conversely, it is invisible. ;

[0115] The left hand is considered visible if at least two of the three key points (left shoulder, left elbow, and left wrist) within the left-hand region have a confidence level greater than or equal to 0.5. Conversely, it is invisible. ;

[0116] The right hand is considered visible if at least two of the three key points (right shoulder, right elbow, and right wrist) within the right-hand region have a confidence level greater than or equal to 0.5. Conversely, it is invisible. .

[0117] Confidence level of prediction results ( This can be used to measure the degree to which the i-th human target and other targets or objects occlude each other in a two-dimensional image from the viewpoint of the k-th subsystem camera: the smaller the occlusion degree, the higher the confidence of the prediction result; the larger the occlusion degree, the lower the confidence of the prediction result.

[0118] According to the above method, confidence weighting is performed based on the degree of occlusion of the same human target under different viewpoints, which can greatly alleviate the two-dimensional occlusion problem in target tracking. Finally, the predicted results of the position coordinates (x, y) and velocity vector (△x, △y) of each human target in the global coordinate system at the next moment are obtained by fusing them together. Combined with the global identity tags that the human targets have been bound to, global tracking of human targets in the scene can be achieved by iterating in this way.

[0119] The aforementioned multi-view fusion-based target tracking method organically integrates tracking prediction results from multiple camera perspectives at different depths to predict the position and trajectory of each target at the next moment. By modeling human targets in the scene as equivalent to neural network graph nodes in a spatiotemporal manner and calculating graded decay confidence based on the degree of occlusion, the two-dimensional occlusion problem in target tracking is largely alleviated. Furthermore, the prediction results from different camera perspectives are organically fused to achieve global tracking of human targets in the scene. In addition, this application is compatible with multiple different perspectives, and the system possesses strong versatility and scalability. More camera perspectives also contribute to more accurate prediction and tracking.

[0120] Specifically, the multi-view fusion target tracking method proposed in this invention models human targets in the world scene as graph nodes in a neural network for spatiotemporal analysis, enabling prediction of the position trajectory of each target at the next moment. It also achieves the organic fusion of perspectives from different subsystems. This method overcomes the shortcomings of traditional techniques by calculating graded attenuation confidence based on the degree of occlusion of the human target, and then weighting and fusing the prediction results from different perspectives, significantly alleviating the two-dimensional occlusion problem in target tracking. Furthermore, this method integrates prediction results from multiple different camera perspectives, enabling global tracking of human targets in the scene. The system is compatible with multiple subsystem perspectives, exhibiting strong scalability, and the increased number of subsystem perspectives also contributes to more accurate prediction and tracking.

[0121] In one embodiment, such as Figure 5 As shown, a target tracking device 10 based on multi-view fusion is provided, including: a coordinate transformation module 1, a target detection and identification module 2, an observation and occlusion analysis module 3, a spatiotemporal map processing and fusion module 4, and a global tracking and visibility estimation module 5.

[0122] The coordinate transformation module 1 is used to establish a global coordinate system for the two-dimensional plane where the application scene is located, establish a sub-coordinate system for each depth camera that acquires two-dimensional images and three-dimensional point clouds in the application scene, and convert the target position coordinates under each sub-coordinate system to the global coordinate system.

[0123] The target detection and identification module 2 is used to detect human targets based on the two-dimensional images and three-dimensional point clouds collected by each depth camera, and bind a global identity tag to the position coordinates of each human target in the global coordinate system.

[0124] The observation and occlusion analysis module 3 is used to determine the observation input results and occlusion degree of each human target at the current moment from the camera view of each depth camera based on the two-dimensional images and three-dimensional point clouds collected by multiple depth cameras for each global identity tag.

[0125] The spatiotemporal graph processing and fusion module 4 is used to process the observation input results of each human target under the camera view of each depth camera using a spatiotemporal graph neural network to obtain the prediction result of the next moment, determine the confidence of the prediction result according to the degree of occlusion corresponding to the observation input result, and perform multi-view fusion on the prediction result to obtain the fused prediction output result.

[0126] The global tracking and visibility estimation module 5 is used to integrate the prediction output result corresponding to each human target with the observation input result, and combine the confidence of the prediction result corresponding to each depth camera to infer the visibility of each human body part of the human target, thereby realizing global tracking of each human target.

[0127] In this embodiment, the step of establishing a global coordinate system for the two-dimensional plane where the application scene is located, establishing a sub-coordinate system for each depth camera that acquires two-dimensional images and three-dimensional point clouds in the application scene, and converting the target position coordinates under each sub-coordinate system to the global coordinate system includes:

[0128] A global coordinate system is established for the two-dimensional plane where the application scenario is located. The origin coordinates and X-axis of the global coordinate system are determined. The X-axis is rotated 90 degrees counterclockwise in the world coordinate plane with the origin of the global coordinate system as the rotation center to obtain the Y-axis.

[0129] The position and orientation of each depth camera that acquires 2D images and 3D point clouds in the application scenario are obtained. The viewing angle is determined according to the orientation of the depth camera. A sub-coordinate system is established in the global coordinate system for each depth camera's viewing angle.

[0130] When transforming the target position coordinates in each sub-coordinate system to the global coordinate system, the coordinates of the origin of the sub-coordinate system in the global coordinate system are obtained as follows: , where x O Let y be the x-coordinate of the origin of the sub-coordinate system in the global coordinate system. O Let the ordinate of the origin of the sub-coordinate system be the ordinate of the global coordinate system, and let the unit direction vector corresponding to the positive X-axis of the sub-coordinate system be the corresponding vector in the global coordinate system. , where x X Let y be the x-coordinate of the unit direction vector corresponding to the positive X-axis of the sub-coordinate system in the global coordinate system. XLet X be the ordinate of the unit direction vector corresponding to the positive X-axis of the sub-coordinate system in the global coordinate system. The Y-axis of the sub-coordinate system is obtained by rotating the X-axis counterclockwise by 90 degrees around the origin of the sub-coordinate system in the sub-coordinate plane. The unit direction vector corresponding to the positive Y-axis of the sub-coordinate system is: , It is a rotation matrix;

[0131] If the target position coordinates in the target sub-coordinate system are , where x S Let y be the x-coordinate of the target position in the target sub-coordinate system. S Let be the ordinate of the target position in the target sub-coordinate system, then the target position coordinates correspond to coordinate P in the global coordinate system. W for:

[0132] , where x W Let y be the x-coordinate of the target position in the global coordinate system. W The target position is represented by its ordinate in the global coordinate system.

[0133] In this embodiment, the step of detecting human targets based on the two-dimensional images and three-dimensional point clouds acquired by each depth camera, and binding a global identity tag to the position coordinates of each human target in the global coordinate system, includes:

[0134] For each two-dimensional image captured by a depth camera, a pre-trained deep neural network for human target detection is invoked to detect the key points of each human target and obtain the human targets in the application scenario.

[0135] After detecting a human target, the three-dimensional coordinates of the center of the human target are obtained from the three-dimensional point cloud collected by each depth camera. The vertical axis coordinates of the three-dimensional coordinates are removed and mapped to the two-dimensional plane camera view sub-coordinate system to obtain the position coordinates of the human target in the sub-coordinate system.

[0136] The position coordinates of the human target in the sub-coordinate system are converted to obtain the position coordinates of the human target in the global coordinate system;

[0137] When a new human target is detected in the camera view of any depth camera, the position coordinates of the new human target in the global coordinate system are calculated. Among the human targets detected by each of the other depth cameras, the nearest human target with the shortest Euclidean distance between the two position coordinates and within a preset spacing range is matched.

[0138] If the new human target is matched with its nearest human target in each of the other depth cameras, then a global identity label is bound to the new human target.

[0139] If the new human target fails to find the nearest human target in each of the remaining depth cameras, its global identity label is set to pending, and the pending human target is used as the nearest human target within a preset distance range to match the next newly appearing human target.

[0140] In this embodiment, the step of using a spatiotemporal graph neural network to process the observation input results of each human target from the camera viewpoint of each depth camera to obtain the prediction result for the next moment includes:

[0141] For each depth camera, a spatiotemporal graph neural network is pre-constructed according to the task requirements and trained using the collected target position trajectory data. The spatiotemporal graph neural network is used to predict the future motion trajectory and trend of human targets in the application scenario to obtain the prediction result of the next moment.

[0142] In the spatiotemporal graph neural network corresponding to each depth camera, each graph node is associated with a human target from the camera's viewpoint. Each graph node records a global identity label bound to the corresponding human target. Therefore, the global identity labels for the N human targets are: , , … ;

[0143] The input node features of the spatiotemporal graph neural network are the position coordinates (x, y) and velocity vector (△x, △y) of the human target under the current camera view. It is divided into two parts: {the predicted value P calculated by the spatiotemporal graph neural network at the previous moment for the current moment} and {the observed value S collected and detected by the depth camera at the current moment}. Among them, x is the horizontal coordinate of the human target under the current camera view, y is the vertical coordinate of the human target under the current camera view, △x is the velocity vector of the human target under the horizontal coordinate under the current camera view, and △y is the velocity vector of the human target under the vertical coordinate under the current camera view.

[0144] The input node feature matrix of the spatiotemporal graph neural network has a dimension of T×N×8, where N is the number of target nodes and T is the length of the time window. For a single time t, the dimension of the input feature matrix is ​​N×8, as shown in the following form:

[0145] ;in Let be the predicted value of the Nth target node on the horizontal axis at time t. Let be the predicted value of the Nth target node on the ordinate at time t. Let be the predicted velocity vector of the Nth target node on the horizontal axis at time t. Let be the predicted velocity vector of the Nth target node at time t on the ordinate. Let be the observed value of the Nth target node on the horizontal axis at time t. Let be the observed value of the Nth target node on the ordinate at time t. Let t be the observed velocity vector of the Nth target node on the horizontal axis. Let be the observed velocity vector of the Nth target node at time t on the vertical axis;

[0146] The node adjacency matrix is ​​a symmetric matrix with a dimension of N×N, where the edge weight connecting any two graph nodes is the Euclidean distance between the corresponding two human targets in the sub-coordinate system.

[0147] The output matrix of the spatiotemporal graph neural network has a dimension of 1×N×4. The prediction result output by the spatiotemporal graph neural network is the predicted value P of the position coordinates (x,y) and velocity vectors (△x,△y) of all human targets in the current depth camera's view at the next time t+1.

[0148] In this embodiment, the step of using a spatiotemporal graph neural network to process the observation input results of each human target from the camera viewpoint of each depth camera to obtain the prediction result for the next moment further includes:

[0149] The edge weight connecting any two graph nodes is calculated using the observed value S, and the edge weight connecting any two graph nodes is:

[0150] ;in W represents the edge weight matrix between any two graph nodes in the view of the k-th subsystem camera at the current time t. NN This represents the edge weight between the Nth graph node and the Nth graph node;

[0151] The prediction result output by the spatiotemporal graph neural network is:

[0152] ;in This represents the prediction result of the k-th subsystem at the next time step t+1. Let be the predicted value of the Nth target node on the horizontal axis at time t+1. Let be the predicted value of the Nth target node on the ordinate at time t+1. Let be the predicted velocity vector of the Nth target node on the horizontal axis at time t+1. Let t+1 be the predicted velocity vector of the Nth target node on the ordinate.

[0153] In this embodiment, determining the confidence level of the prediction result based on the occlusion degree corresponding to the observed input result, and performing multi-view fusion on the prediction result to obtain the fused prediction output result, includes:

[0154] Obtain the total number K of depth cameras in the application scenario, and treat each depth camera as a subsystem;

[0155] The confidence level of the prediction result is determined based on the degree of occlusion corresponding to the observed input result, and the confidence vector of the prediction result is calculated. This represents the confidence vector of the prediction results of all human targets in the view of the k-th subsystem camera at the current time t, calculated according to their respective occlusion levels.

[0156] The prediction output is obtained by multi-view fusion of the prediction result confidence vector and the prediction result. The prediction output results are described below. This represents the predicted position coordinates (x, y) and velocity vectors (Δx, Δy) of all targets after fusing the prediction results of all subsystems at the next time t+1. This represents the predicted position coordinates (x, y) and velocity vectors (Δx, Δy) of all targets calculated by the k-th subsystem at the next time t+1.

[0157] The fused prediction output for the i-th single human target is represented as follows: ;in, This represents the predicted position coordinates (x, y) and velocity vector (Δx, Δy) of the i-th human target after fusing the prediction outputs of all subsystems at the next time t+1. This represents the confidence level of the predicted output result calculated based on the degree of occlusion of the i-th human target in the view of the k-th subsystem camera at the current time t. This represents the sum of confidence scores of the predicted output results calculated based on the degree of occlusion of the i-th target in the camera view of all subsystems at the current time t.

[0158] In this embodiment, the step of integrating the predicted output result corresponding to each human target with the observation input result, and combining the confidence level of the predicted result corresponding to each depth camera to deduce the visibility of each human body part of the human target, includes:

[0159] For the i-th human target in the view of the k-th subsystem camera at the current time t, obtain the human key points of the human target;

[0160] The confidence level of the prediction result of the human target at the current time t under the camera view of the corresponding subsystem. The following formula applies graded attenuation based on the degree of occlusion: ;in, The value ranges from 0 to 1; For head weight, Weight for the torso; Weighted by left hand Right-hand weight; The value of the Boolean variable is visible in the header; The value of the visible Boolean variable is for the torso. The value of the left-hand visible Boolean variable; The value of the right-hand visible Boolean variable;

[0161] The confidence level of the human body key points is determined based on the confidence level of the prediction results corresponding to each depth camera at the current time t, and the visibility of each part of the human body target is deduced based on the confidence level of the human body key points.

[0162] The aforementioned multi-view fusion-based target tracking device can organically integrate tracking prediction results from multiple camera perspectives at different depths to predict the position and trajectory of each target at the next moment. By modeling human targets in the scene as equivalent to neural network graph nodes in a spatiotemporal manner and calculating graded decay confidence based on the degree of occlusion, the two-dimensional occlusion problem in target tracking is largely alleviated. Based on this, the prediction results from different camera perspectives are organically fused to achieve global tracking of human targets in the scene. Furthermore, this application is compatible with multiple different perspectives, and the system has strong versatility and scalability. More camera perspectives also contribute to more accurate prediction and tracking.

[0163] For specific limitations regarding the target tracking device based on multi-view fusion, please refer to the limitations of the target tracking method based on multi-view fusion mentioned above, which will not be repeated here. Each module in the aforementioned target tracking device based on multi-view fusion can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0164] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and the database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores target tracking data based on multi-view fusion. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a target tracking method based on multi-view fusion.

[0165] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0166] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0167] A global coordinate system is established for the two-dimensional plane where the application scenario is located. A sub-coordinate system is established for each depth camera that acquires two-dimensional images and three-dimensional point clouds in the application scenario. The target position coordinates under each sub-coordinate system are then converted to the global coordinate system.

[0168] Human targets are detected based on the two-dimensional images and three-dimensional point clouds collected by each depth camera, and a global identity tag is bound to the position coordinates of each human target in the global coordinate system.

[0169] For each human target corresponding to a global identity tag, the observation input result and occlusion degree of each human target at the current moment are determined from the camera view of each depth camera based on the two-dimensional images and three-dimensional point clouds collected by multiple depth cameras.

[0170] A spatiotemporal graph neural network is used to process the observation input results of each human target under the camera view of each depth camera to obtain the prediction result of the next moment. The confidence of the prediction result is determined according to the degree of occlusion corresponding to the observation input result. The prediction result is then fused from multiple views to obtain the fused prediction output result.

[0171] The predicted output for each human target is integrated with the observation input, and the visibility of each human body part is deduced by combining the confidence of the predicted result for each depth camera, thereby achieving global tracking of each human target.

[0172] For specific limitations on the steps implemented by the processor when executing a computer program, please refer to the limitations on the target tracking method based on multi-view fusion mentioned above, which will not be repeated here.

[0173] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0174] A global coordinate system is established for the two-dimensional plane where the application scenario is located. A sub-coordinate system is established for each depth camera that acquires two-dimensional images and three-dimensional point clouds in the application scenario. The target position coordinates under each sub-coordinate system are then converted to the global coordinate system.

[0175] Human targets are detected based on the two-dimensional images and three-dimensional point clouds collected by each depth camera, and a global identity tag is bound to the position coordinates of each human target in the global coordinate system.

[0176] For each human target corresponding to a global identity tag, the observation input result and occlusion degree of each human target at the current moment are determined from the camera view of each depth camera based on the two-dimensional images and three-dimensional point clouds collected by multiple depth cameras.

[0177] A spatiotemporal graph neural network is used to process the observation input results of each human target under the camera view of each depth camera to obtain the prediction result of the next moment. The confidence of the prediction result is determined according to the degree of occlusion corresponding to the observation input result. The prediction result is then fused from multiple views to obtain the fused prediction output result.

[0178] The predicted output for each human target is integrated with the observation input, and the visibility of each human body part is deduced by combining the confidence of the predicted result for each depth camera, thereby achieving global tracking of each human target.

[0179] For specific limitations on the implementation steps of a computer program when it is executed by a processor, please refer to the limitations on target tracking methods based on multi-view fusion mentioned above, which will not be repeated here.

[0180] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0181] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0182] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A target tracking method based on multi-view fusion, characterized in that, include: A global coordinate system is established for the two-dimensional plane where the application scenario is located. A sub-coordinate system is established for each depth camera that acquires two-dimensional images and three-dimensional point clouds in the application scenario. The target position coordinates under each sub-coordinate system are then converted to the global coordinate system. Human targets are detected based on the two-dimensional images and three-dimensional point clouds collected by each depth camera, and a global identity tag is bound to the position coordinates of each human target in the global coordinate system. For each human target corresponding to a global identity tag, the observation input result and occlusion degree of each human target at the current moment are determined from the camera view of each depth camera based on the two-dimensional images and three-dimensional point clouds collected by multiple depth cameras. A spatiotemporal graph neural network is used to process the observation input results of each human target under the camera view of each depth camera to obtain the prediction result of the next moment. The confidence of the prediction result is determined according to the degree of occlusion corresponding to the observation input result. The prediction result is then fused from multiple views to obtain the fused prediction output result. The predicted output results corresponding to each human target are integrated with the observation input results, and the visibility of each human body part of the human target is deduced by combining the confidence of the prediction results corresponding to each depth camera, so as to realize global tracking of each human target. The process of establishing a global coordinate system for the two-dimensional plane where the application scenario is located, establishing a sub-coordinate system for each depth camera that acquires two-dimensional images and three-dimensional point clouds in the application scenario, and converting the target position coordinates in each sub-coordinate system to the global coordinate system includes: A global coordinate system is established for the two-dimensional plane where the application scenario is located. The origin coordinates and X-axis of the global coordinate system are determined. The X-axis is rotated 90 degrees counterclockwise in the world coordinate plane with the origin of the global coordinate system as the rotation center to obtain the Y-axis. The position and orientation of each depth camera that acquires 2D images and 3D point clouds in the application scenario are obtained. The viewing angle is determined according to the orientation of the depth camera. A sub-coordinate system is established in the global coordinate system for each depth camera's viewing angle. When transforming the target position coordinates in each sub-coordinate system to the global coordinate system, the coordinates of the origin of the sub-coordinate system in the global coordinate system are obtained as follows: , where x O Let y be the x-coordinate of the origin of the sub-coordinate system in the global coordinate system. O Let the ordinate of the origin of the sub-coordinate system be the ordinate of the global coordinate system, and let the unit direction vector corresponding to the positive X-axis of the sub-coordinate system be the corresponding vector in the global coordinate system. , where x X Let y be the x-coordinate of the unit direction vector corresponding to the positive X-axis of the sub-coordinate system in the global coordinate system. X Let X be the ordinate of the unit direction vector corresponding to the positive X-axis of the sub-coordinate system in the global coordinate system. The Y-axis of the sub-coordinate system is obtained by rotating the X-axis counterclockwise by 90 degrees around the origin of the sub-coordinate system in the sub-coordinate plane. The unit direction vector corresponding to the positive Y-axis of the sub-coordinate system is: , It is a rotation matrix; If the target position coordinates in the target sub-coordinate system are , where x S Let y be the x-coordinate of the target position in the target sub-coordinate system. S Let be the ordinate of the target position in the target sub-coordinate system, then the target position coordinates correspond to coordinate P in the global coordinate system. W for: where x W is the horizontal coordinate of the target position in the global coordinate system, y W is the vertical coordinate of the target position in the global coordinate system. 2.The target tracking method based on multi-view fusion according to claim 1, characterized in that, The step of detecting human targets based on the two-dimensional images and three-dimensional point clouds acquired by each depth camera, and binding a global identity tag to the position coordinates of each human target in the global coordinate system, includes: For each two-dimensional image captured by a depth camera, a pre-trained deep neural network for human target detection is invoked to detect the key points of each human target and obtain the human targets in the application scenario. After detecting a human target, the three-dimensional coordinates of the center of the human target are obtained from the three-dimensional point cloud collected by each depth camera. The vertical axis coordinates of the three-dimensional coordinates are removed and mapped to the two-dimensional plane camera view sub-coordinate system to obtain the position coordinates of the human target in the sub-coordinate system. The position coordinates of the human target in the sub-coordinate system are converted to obtain the position coordinates of the human target in the global coordinate system; When a new human target is detected in the camera view of any depth camera, the position coordinates of the new human target in the global coordinate system are calculated. Among the human targets detected by each of the other depth cameras, the nearest human target with the shortest Euclidean distance between the two position coordinates and within a preset spacing range is matched. If the new human target is matched with its nearest human target in each of the other depth cameras, then a global identity label is bound to the new human target. If the new human target fails to find the nearest human target in each of the remaining depth cameras, its global identity label is set to pending, and the pending human target is used as the nearest human target within a preset distance range to match the next newly appearing human target. 3.The target tracking method based on multi-view fusion according to claim 1, characterized in that, The process of using a spatiotemporal graph neural network to process the observation input results of each human target from the camera viewpoint of each depth camera to obtain the prediction result for the next moment includes: For each depth camera, a spatiotemporal graph neural network is pre-constructed according to the task requirements and trained using the collected target position trajectory data. The spatiotemporal graph neural network is used to predict the future motion trajectory and trend of human targets in the application scenario to obtain the prediction result of the next moment. In the spatio-temporal graph neural network corresponding to each depth camera, each graph node is associated with a human target in the camera view of the depth camera, and each graph node records a global identity label bound to the corresponding human target, and the global identity labels of N human targets are , , , ; The input node features of the spatiotemporal graph neural network are the position coordinates (x, y) and velocity vector (△x, △y) of the human target under the current camera view. It is divided into two parts: {the predicted value P calculated by the spatiotemporal graph neural network at the previous moment for the current moment} and {the observed value S collected and detected by the depth camera at the current moment}. Among them, x is the horizontal coordinate of the human target under the current camera view, y is the vertical coordinate of the human target under the current camera view, △x is the velocity vector of the human target under the horizontal coordinate under the current camera view, and △y is the velocity vector of the human target under the vertical coordinate under the current camera view. The input node feature matrix of the spatiotemporal graph neural network has a dimension of T×N×8, where N is the number of target nodes and T is the length of the time window. For a single time t, the dimension of the input feature matrix is ​​N×8, as shown in the following form: ;in Let be the predicted value of the Nth target node on the horizontal axis at time t. Let be the predicted value of the Nth target node on the ordinate at time t. Let be the predicted velocity vector of the Nth target node on the horizontal axis at time t. Let be the predicted velocity vector of the Nth target node at time t on the ordinate. Let be the observed value of the Nth target node on the horizontal axis at time t. Let be the observed value of the Nth target node on the ordinate at time t. Let t be the observed velocity vector of the Nth target node on the horizontal axis. Let be the observed velocity vector of the Nth target node at time t on the vertical axis; The node adjacency matrix is ​​a symmetric matrix with a dimension of N×N, where the edge weight connecting any two graph nodes is the Euclidean distance between the corresponding two human targets in the sub-coordinate system. The output matrix of the spatiotemporal graph neural network has a dimension of 1×N×4. The prediction result output by the spatiotemporal graph neural network is the predicted value P of the position coordinates (x,y) and velocity vectors (△x,△y) of all human targets in the current depth camera's view at the next time t+1. 4.The target tracking method based on multi-view fusion according to claim 3, characterized in that, The method of using a spatiotemporal graph neural network to process the observation input results of each human target from the camera viewpoint of each depth camera to obtain the prediction result for the next moment also includes: The edge weight connecting any two graph nodes is calculated using the observed value S, and the edge weight connecting any two graph nodes is: ; wherein represents the edge weight matrix of any two graph nodes in the kth subsystem camera view at the current time t, W NN represents the edge weight between the Nth graph node and the Nth graph node; The prediction result output by the spatiotemporal graph neural network is: ;in This represents the prediction result of the k-th subsystem at the next time step t+1. Let be the predicted value of the Nth target node on the horizontal axis at time t+1. Let be the predicted value of the Nth target node on the ordinate at time t+1. Let be the predicted velocity vector of the Nth target node on the horizontal axis at time t+1. Let t+1 be the predicted velocity vector of the Nth target node on the ordinate.

5. The multi-view fusion based target tracking method according to claim 4, characterized in that, The step of determining the confidence level of the prediction result based on the occlusion degree corresponding to the observed input result, and performing multi-view fusion on the prediction result to obtain the fused prediction output result includes: Obtain the total number K of depth cameras in the application scenario, and treat each depth camera as a subsystem; According to the occlusion degree corresponding to the observation input result, the confidence of the prediction result is calculated to obtain a prediction result confidence vector, and the prediction result is determined according to the prediction result confidence vector. It is represented that the prediction result confidence vector of all human body targets in the camera view of the kth subsystem at the current time t is calculated according to the respective occlusion degree. The prediction output is obtained by multi-view fusion of the prediction result confidence vector and the prediction result. The prediction output results are described below. This represents the predicted position coordinates (x, y) and velocity vectors (Δx, Δy) of all targets after fusing the prediction results of all subsystems at the next time t+1. This represents the predicted position coordinates (x, y) and velocity vectors (Δx, Δy) of all targets calculated by the k-th subsystem at the next time t+1. The fused prediction output for the i-th single human target is represented as follows: ;in, This represents the predicted position coordinates (x, y) and velocity vector (Δx, Δy) of the i-th human target after fusing the prediction outputs of all subsystems at the next time t+1. This represents the confidence level of the predicted output result calculated based on the degree of occlusion of the i-th human target in the view of the k-th subsystem camera at the current time t. This represents the sum of confidence scores of the predicted output results calculated based on the degree of occlusion of the i-th target in the camera view of all subsystems at the current time t. 6.The target tracking method based on multi-view fusion according to claim 5, characterized in that, The process of integrating the predicted output for each human target with the observation input, and combining the confidence level of the predicted results for each depth camera to deduce the visibility of each human body part of the human target, includes: For the i-th human target in the view of the k-th subsystem camera at the current time t, obtain the human key points of the human target; The confidence level of the prediction result of the human target at the current time t under the camera view of the corresponding subsystem. The following formula applies graded attenuation based on the degree of occlusion: ;in, The value ranges from 0 to 1; For head weight, Weight for the torso; Weighted by left hand Right-hand weight; The value of the Boolean variable is visible in the header; The value of the visible Boolean variable is for the torso. The value of the left-hand visible Boolean variable; The value of the right-hand visible Boolean variable; The confidence level of the human body key points is determined based on the confidence level of the prediction results corresponding to each depth camera at the current time t, and the visibility of each part of the human body target is deduced based on the confidence level of the human body key points.

7. A multi-view fusion based target tracking apparatus for implementing the multi-view fusion based target tracking method according to any one of claims 1 to 6, characterized in that, The device includes: The coordinate transformation module is used to establish a global coordinate system for the two-dimensional plane where the application scene is located, establish a sub-coordinate system for each depth camera that acquires two-dimensional images and three-dimensional point clouds in the application scene, and convert the target position coordinates under each sub-coordinate system to the global coordinate system. The target detection and labeling module is used to detect human targets based on the two-dimensional images and three-dimensional point clouds collected by each depth camera, and bind a global identity label to the position coordinates of each human target in the global coordinate system. The observation and occlusion analysis module is used to determine the observation input results and occlusion degree of each human target corresponding to each global identity tag at the current moment from the camera view of each depth camera based on the two-dimensional images and three-dimensional point clouds collected by multiple depth cameras. The spatiotemporal graph processing and fusion module is used to process the observation input results of each human target under the camera view of each depth camera using a spatiotemporal graph neural network to obtain the prediction result for the next moment, determine the confidence of the prediction result according to the degree of occlusion corresponding to the observation input result, and perform multi-view fusion on the prediction result to obtain the fused prediction output result. The global tracking and visibility estimation module is used to integrate the prediction output results corresponding to each human target with the observation input results, and combine the confidence of the prediction results corresponding to each depth camera to infer the visibility of each human body part of the human target, thereby realizing global tracking of each human target.

8. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.