Pedestrian tracking method and device based on multi-modal data, equipment and storage medium
By introducing an adaptive weighted model of multimodal data in traditional video surveillance systems, combining heterogeneous graph convolution networks and spatiotemporal graph neural networks, the accuracy problem of traditional pedestrian tracking in complex scenarios is solved, and stable and accurate pedestrian tracking in complex environments is achieved.
Patent Information
- Application Number
- CN202510047776.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-13
AI Technical Summary
Traditional video surveillance systems are difficult to provide stable and accurate pedestrian tracking results in occlusion, lighting changes and complex scenarios, especially when pedestrians move quickly or images are blocked.
A pedestrian tracking method based on multimodal data is adopted, and video data, RFID data, Wi-Fi data, Bluetooth data and infrared sensor data are obtained, and the video data is input into the adaptive weighted model is used for weighted fusion, and the movement trajectory and behavioral mode of the first pedestrian are output. This method combines the heterogeneous graph convolution network HGCN and the spatio-temporal graph neural network ST-GNN based on the temporal attention mechanism to accurately track pedestrians in complex scenarios.
It improves the accuracy and stability of pedestrian tracking, can accurately track pedestrians in complex environments (such as video being blocked or blurred, etc.), and predict pedestrian behavior patterns to ensure the reliability of tracking results.
Smart Images

Figure CN120067746A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of pedestrian tracking, and particularly to a pedestrian tracking method, device, equipment and storage medium based on multimodal data. Background Art
[0002] In the construction of modern smart cities, pedestrian tracking technology plays an important role in urban public safety, traffic management and emergency response. Although traditional video surveillance systems have been widely used in these scenarios, in the case of occlusion, light changes and complex scenes, video data often fails to provide stable and accurate pedestrian tracking results.
[0003] Pedestrian tracking methods in related technologies are usually simple filtering algorithms (such as Kalman filtering algorithm, etc.), which are difficult to achieve stable target tracking in complex and dynamically changing environments. When facing complex scenarios such as fast-moving pedestrians, occluded images or changing environments, using a simple filtering algorithm for pedestrian tracking will lead to a decrease in tracking accuracy. Summary of the Invention
[0004] The present disclosure provides a pedestrian tracking method, device, equipment and storage medium based on multimodal data, which can accurately track pedestrians in complex scenarios. The technical solution at least includes the following aspects: In a first aspect, a pedestrian tracking method based on multimodal data is provided, including: obtaining multimodal data of a first area at time k, where the multimodal data includes video data, RFID data, Wi-Fi data, Bluetooth data and infrared sensor data; inputting the multimodal data into an adaptive weighting model, and obtaining a set of first joint points of a first pedestrian at time k output by the adaptive weighting model, where the adaptive weighting model is used to perform weighted fusion on the multimodal data by using an adaptive algorithm and then output the set of first joint points of the first pedestrian at time k; based on the set of first joint points of the first pedestrian at time k, determining the motion trajectory of the first pedestrian by using a heterogeneous graph convolutional network HGCN and a spatio-temporal graph neural network ST-GNN based on a time attention mechanism; and determining a behavior pattern of the first pedestrian based on the motion trajectory of the first pedestrian, where the behavior pattern is used to track the first pedestrian.
[0005] Optionally, the first pedestrian includes n joint points, the multi-modal data includes at least one image modality, the image modality includes the video data and the infrared sensor data, and the adaptive weighting model is used to implement the following method to output the first joint point set of the first pedestrian at time k after weighted fusion of the multi-modal data by using an adaptive algorithm: Based on the multi-modal data of the first region at time k, obtain the second joint point set of the first pedestrian at time k, where the second joint point set includes n joint points of the first pedestrian corresponding to each image modality; Based on the multi-modal data of the first region at time k, determine the adaptive weight of each joint point in the second joint point set; Use the adaptive weight of each joint point in the second joint point set to weight the joint points in the second joint point set to obtain the first joint point set of the first pedestrian at time k, and the first joint point set includes n weighted joint points.
[0006] Optionally, the determining the adaptive weight of each joint point in the second joint point set based on the multi-modal data of the first region at time k includes: using the following formula to determine the adaptive weight of the i-th joint point corresponding to the p-th image modality in the second joint point set:
[0007] where, represents the adaptive weight of the i-th joint point corresponding to the p-th image modality in the second joint point set, represents the confidence of the j-th modality in the i-th joint point at time k, represents the signal strength of the j-th modality in the i-th joint point at time k, the multi-modal data includes data of N modalities, the data of N modalities includes data of P image modalities, and i, j, p, n, N, P are positive integers, the value range of i is from 1 to n, the value range of j is from 1 to N, and the value range of p is from 1 to P.
[0008] Optionally, the method further includes: optimizing the adaptive weight of the i-th joint point corresponding to the p-th image modality in the second joint point set determined by the adaptive weighting model by using deep learning; where, the optimized adaptive weight of the i-th joint point corresponding to the p-th image modality in the second joint point set is represented by the following formula:
[0009] where, represents the optimized adaptive weight of the i-th joint point corresponding to the p-th image modality in the second joint point set, is the adjustment weight for the i-th joint point corresponding to the p-th image modality, and the adjustment weight is obtained by means of deep learning. is the smoothing factor.
[0010] Optionally, determining the motion trajectory of the first pedestrian based on the set of first joint points of the first pedestrian at time k by using a heterogeneous graph convolutional network HGCN and a spatio-temporal graph neural network ST-GNN based on a time attention mechanism includes: constructing a first heterogeneous graph based on the set of first joint points of the first pedestrian at time k, where the first heterogeneous graph includes time edges, spatial edges, and heterogeneous edges; encoding the node features of the first heterogeneous graph by using the HGCN to obtain a plurality of node features; and inputting the plurality of node features into the ST-GNN to obtain the motion trajectory of the first pedestrian output by the ST-GNN.
[0011] Optionally, determining the behavior pattern of the first pedestrian based on the motion trajectory of the first pedestrian includes: establishing the motion process of the pedestrian as a Markov decision process MDP model, where the MDP model is used to predict the behavior pattern of the pedestrian; and inputting the motion trajectory of the first pedestrian into the MDP model to obtain the behavior pattern of the first pedestrian output by the MDP model.
[0012] In a second aspect, there is also provided a pedestrian tracking device based on multimodal data, including: a first acquisition module, configured to acquire multimodal data of a first area at time k, where the multimodal data includes video data, RFID data, Wi-Fi data, Bluetooth data, and infrared sensor data; a second acquisition module, configured to input the multimodal data into an adaptive weighting model and acquire a set of first joint points of the first pedestrian at time k output by the adaptive weighting model, where the adaptive weighting model is used to perform weighted fusion on the multimodal data by using an adaptive algorithm and then output the set of first joint points of the first pedestrian at time k; a trajectory determination module, configured to determine the motion trajectory of the first pedestrian based on the set of first joint points of the first pedestrian at time k by using a heterogeneous graph convolutional network HGCN and a spatio-temporal graph neural network ST-GNN based on a time attention mechanism; and a behavior pattern determination module, configured to determine the behavior pattern of the first pedestrian based on the motion trajectory of the first pedestrian, where the behavior pattern is used to track the first pedestrian.
[0013] Optionally, the first pedestrian includes n joint points, the multimodal data includes at least one image modality, the image modality includes the video data and the infrared sensor data, and the second acquisition module is further configured to implement the weighted fusion of the multimodal data by using an adaptive algorithm and then output the first joint point set of the first pedestrian at the k-th moment in the following manner: based on the multimodal data of the first region at the k-th moment, obtain the second joint point set of the first pedestrian at the k-th moment, where the second joint point set includes n joint points of the first pedestrian corresponding to each image modality; based on the multimodal data of the first region at the k-th moment, determine the adaptive weight of each joint point in the second joint point set; use the adaptive weight of each joint point in the second joint point set to weight the joint points in the second joint point set to obtain the first joint point set of the first pedestrian at the k-th moment, and the first joint point set includes n weighted joint points.
[0014] Optionally, the second acquisition module is further configured to determine the adaptive weight of the i-th joint point corresponding to the p-th image modality in the second joint point set by using the following formula:
[0015] Where represents the adaptive weight of the i-th joint point corresponding to the p-th image modality in the second joint point set, represents the confidence of the j-th modality in the i-th joint point at the k-th moment, represents the signal strength of the j-th modality in the i-th joint point at the k-th moment, the multimodal data includes data of N modalities, the data of the N modalities includes data of P image modalities, and i, j, p, n, N, P are positive integers, and the value range of i is from 1 to n, the value range of j is from 1 to N, and the value range of p is from 1 to P.
[0016] Optionally, the device further includes: an optimization module, configured to optimize the adaptive weight of the i-th joint point corresponding to the p-th image modality in the second joint point set determined by the adaptive weighted model by using deep learning; where the optimized adaptive weight of the i-th joint point corresponding to the p-th image modality in the second joint point set is represented by the following formula:
[0017] Where represents the optimized adaptive weight of the i-th joint point corresponding to the p-th image modality in the second joint point set, is the adjustment weight value of the i-th joint point corresponding to the p-th image modality, and the adjustment weight value is obtained by using deep learning, is the smoothing factor.
[0018] Optionally, the trajectory determination module is further configured to construct a first heterogeneous graph based on the first set of joint points of the first pedestrian at time k, where the first heterogeneous graph includes temporal edges, spatial edges, and heterogeneous edges; use the HGCN to encode the node features of the first heterogeneous graph to obtain multiple node features; input the multiple node features into the ST-GNN to obtain the motion trajectory of the first pedestrian output by the ST-GNN.
[0019] Optionally, the behavior pattern determination module is further configured to establish the motion process of the pedestrian as a Markov decision process MDP model, where the MDP model is used to predict the behavior pattern of the pedestrian; input the motion trajectory of the first pedestrian into the MDP model to obtain the behavior pattern of the first pedestrian output by the MDP model.
[0020] In a third aspect, a computer device is further provided, including: a memory and a processor, where at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to execute the pedestrian tracking method based on multi-modal data in the above embodiments.
[0021] In a fourth aspect, a computer-readable storage medium is further provided, where at least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor to execute the pedestrian tracking method based on multi-modal data in the above embodiments.
[0022] In a fifth aspect, a computer program product is provided, including computer programs / instructions, where the computer programs / instructions, when executed by a processor, implement the method described in the first aspect.
[0023] The beneficial effects brought by the technical solutions provided in the embodiments of the present disclosure at least include: In the embodiments of the present disclosure, by adaptively weighting multi-modal data, the accuracy of each joint point in the obtained first set of joint points can be improved; by inputting the first set of joint points into the HGCN and ST-GNN, the motion trajectory of the first pedestrian can be accurately predicted, and further, the behavior pattern of the first pedestrian can be predicted based on the motion trajectory of the first pedestrian. This behavior pattern can indicate the action trend of the first pedestrian in the next period of time. In this way, even in a complex environment (such as an environment where the video is blocked or the video is blurred), the first pedestrian can be accurately tracked according to the behavior pattern. Description of the Drawings
[0024] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0025] Figure 1 It shows a flowchart of a pedestrian tracking method based on multi-modal data provided by an exemplary embodiment of the present disclosure; Figure 2 It shows a flowchart of a pedestrian tracking method based on multi-modal data provided by another exemplary embodiment of the present disclosure; Figure 3 It is a schematic diagram of a first heterogeneous graph; Figure 4 It shows a schematic structural diagram of a pedestrian tracking device based on multi-modal data provided by an exemplary embodiment of the present disclosure; Figure 5 It is a schematic structural diagram of a computer device provided by an embodiment of the present disclosure. Detailed implementation manners
[0026] Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure belongs. The "first", "second", "third" and similar terms used in the specification and claims of the present patent application do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, the terms such as "a" or "one" do not indicate a quantity limitation, but mean that there is at least one. The terms such as "comprising" or "including" mean that the elements or items appearing before "comprising" or "including" cover the elements or items listed after "comprising" or "including" and their equivalents, and do not exclude other elements or items. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.
[0027] To make the objectives, technical solutions and advantages of the present disclosure clearer, the following further describes the embodiments of the present disclosure in detail with reference to the accompanying drawings.
[0028] Figure 1 It shows a flowchart of a pedestrian tracking method based on multi-modal data provided by an exemplary embodiment of the present disclosure, and this method can be executed by a computer device. Refer to Figure 1 and this method includes: In step 101, obtain the multi-modal data of the first area at time k.
[0029] Multimodal data includes video data, RFID (Radio Frequency Identification) data, Wi-Fi (Wireless Network Communication Technology) data, Bluetooth data, and infrared sensor data.
[0030] Here, the multimodal data includes a first pedestrian, and the first pedestrian is the pedestrian to be tracked. The method in the embodiments of the present disclosure is used to track the first pedestrian. In the case where there are multiple first pedestrians in the multimodal data, the method in the embodiments of the present disclosure can also track these multiple first pedestrians simultaneously.
[0031] Optionally, a high-definition camera is used to obtain video data. The frame rate and resolution of the video are set according to the characteristics of the monitoring area. For example, the resolution and frame rate of the camera close to the road surface (ground) can be relatively low, and the resolution of the camera far from the road surface needs to be relatively high to ensure that the human body image captured by the camera is clearly visible; in addition, the influence of the surrounding scenery can also be collected, and geometric correction is performed in combination with the coordinates of nearby control points to obtain an accurate background reference position, further improving the accuracy of pedestrian positioning.
[0032] Optionally, an RFID reader is used to obtain RFID data. When the mobile device carried by the first pedestrian contains an RFID tag, the RFID reader can read the information in the RFID tag, thereby obtaining the position information of the first pedestrian. This method is particularly suitable for personnel positioning in a large-traffic environment. Since not all mobile devices contain RFID tags, RFID tags can be deployed on the control points with known coordinates in the background stationary objects in the pedestrian monitoring environment. By reading the position information of these fixed tags and applying it to the video signal and infrared signal, pedestrian positioning reference calibration and positioning enhancement are performed.
[0033] Optionally, a Wi-Fi data is obtained through a Wi-Fi signal strength collector. When the mobile device carried by the first pedestrian has Wi-Fi, the Wi-Fi signal strength collector can detect the change in the Wi-Fi signal strength of the mobile device carried by the first pedestrian to achieve rough positioning of the pedestrian.
[0034] Optionally, a Bluetooth data is obtained through a Bluetooth signal strength collector. When the mobile device carried by the first pedestrian has Bluetooth, the Bluetooth signal strength collector can obtain the Bluetooth beacon of the mobile device carried by the first pedestrian for high-precision short-distance positioning, and determine the position of the first pedestrian through the change in the Bluetooth signal strength.
[0035] Optionally, obtaining infrared sensor data is achieved through an infrared sensor. This infrared sensor data can reflect the temperature change of the first pedestrian, which helps to locate the first pedestrian in the case of video occlusion or insufficient lighting.
[0036] It can be seen that the video data and the infrared sensor data exist in the form of images, and both belong to the image modality, which can reflect the posture of the first pedestrian; the RFID data, Wi-Fi data, and Bluetooth data exist in the form of data streams and can reflect the position of the first pedestrian.
[0037] In the case where the multi-modal data includes at least one image modality, that is to say, there is image data. Then, tracking the first pedestrian includes: locating the first pedestrian from the image data (that is, determining the position of the first pedestrian), identifying the first pedestrian from the image data (for example, there may be multiple human bodies in the image data, and determining the first pedestrian from multiple human bodies is also identifying the first pedestrian), and determining the motion trajectory of the first pedestrian based on the image data.
[0038] In the case where the multi-modal data does not have an image modality, tracking the first pedestrian includes: determining the motion trajectory of the first pedestrian based on data of other modalities (such as RFID data, Wi-Fi data, Bluetooth data, etc.).
[0039] In step 102, the multi-modal data is input into the adaptive weighting model, and the set of first joint points of the first pedestrian at time k output by the adaptive weighting model is obtained.
[0040] The adaptive weighting model is used to output the set of first joint points of the first pedestrian at time k after weighted fusion of the multi-modal data by using an adaptive algorithm.
[0041] In step 103, based on the set of first joint points of the first pedestrian at time k, the motion trajectory of the first pedestrian is determined by using HGCN and ST-GNN.
[0042] HGCN (Heterogeneous Graph Convolutional Network) is a new type of neural network architecture that improves the performance of graph neural networks by utilizing the characteristics of hyperbolic space. ST-GNN (Spatio-Temporal Graph Neural Networks) can extract complex spatio-temporal dependencies by integrating graph neural networks and various time learning methods.
[0043] By using HGCN and ST-GNN, the motion trajectory of the first pedestrian can be accurately predicted.
[0044] In step 104, determine the behavior pattern of the first pedestrian based on the movement trajectory of the first pedestrian.
[0045] The behavior pattern is used to track the first pedestrian. Here, the behavior pattern is also the action trend of the first pedestrian in the next period of time. After determining the behavior pattern of the first pedestrian, the movement trajectory of the first pedestrian in the next period of time can be predicted based on the behavior pattern. In this way, even in a complex environment (such as an environment where the video is blocked or the video is blurred, etc.), the first pedestrian can be accurately tracked according to the behavior pattern.
[0046] In the embodiment of the present disclosure, by adaptively weighting the multi-modal data, the accuracy of each joint point in the obtained first set of joint points can be improved; by inputting the first set of joint points into the HGCN and ST-GNN, the movement trajectory of the first pedestrian can be accurately predicted, and then the behavior pattern of the first pedestrian can be predicted based on the movement trajectory of the first pedestrian. This behavior pattern can indicate the action trend of the first pedestrian in the next period of time. In this way, even in a complex environment (such as an environment where the video is blocked or the video is blurred, etc.), the first pedestrian can be accurately tracked according to the behavior pattern.
[0047] Figure 2 The flowchart of the pedestrian tracking method based on multi-modal data provided by another exemplary embodiment of the present disclosure is shown. This method can be executed by a computer device. Refer to Figure 2 , this method includes: In step 201, obtain the multi-modal data of the first area at time k.
[0048] The multi-modal data includes video data, RFID data, Wi-Fi data, Bluetooth data, and infrared sensor data.
[0049] For the content of step 201, refer to the foregoing step 101, which is not elaborated here.
[0050] Optionally, before executing step 202, this method further includes: preprocessing the multi-modal data.
[0051] For the image modal data in the multi-modal data, an image self-supervised denoising algorithm can be used for denoising processing.
[0052] For the RFID data, Wi-Fi data, and Bluetooth data in the multi-modal data, a Kalman filter can be used for denoising processing.
[0053] Regarding the implementation methods of the image self-supervised denoising algorithm and the Kalman filter, there are many in the related technologies, which are not elaborated here.
[0054] Since RFID signals are vulnerable to environmental factors (such as obstacles and reflections), signal strength fluctuations occur. Kalman filtering helps reduce the positioning error caused by signal jitter here by smoothing the signal strength changes. The initial state is set as the initial position and speed estimate of the RFID tag. Subsequently, each time a new RFID signal strength is received, the estimated position of the RFID tag is updated through Kalman filtering, thereby generating stable position information.
[0055] The detection of Wi-Fi signal strength can be applied to indoor positioning such as at subway station security checkpoints. However, the signal strength value is easily affected by wall and equipment interference, resulting in short-term fluctuations. Using Kalman filtering can smooth the time series of signal strength, thereby improving the accuracy of indoor positioning. The signal strength value is used as the observation variable, and the position information of the previous moment is combined for prediction; the current position estimate is updated using the new signal strength value, and jitter and jumps are reduced.
[0056] The signal strength of Bluetooth beacons has high accuracy in short-distance positioning, but it is easily interfered in public areas with a large number of people or in an environment with multiple signal sources, resulting in large fluctuations. Through Kalman filtering, the influence of Bluetooth signal fluctuations can be reduced, and positioning stability can be improved. Set the initial position and error covariance matrix of the Bluetooth signal, and perform prediction and update each time the Bluetooth signal strength is obtained to obtain smooth Bluetooth signal strength and corresponding position information.
[0057] In step 202, the multi-modal data is input into the adaptive weighting model, and the set of first joint points of the first pedestrian at time k output by the adaptive weighting model is obtained.
[0058] The adaptive weighting model is used to output the set of first joint points of the first pedestrian at time k after weighted fusion of multi-modal data using an adaptive algorithm.
[0059] Optionally, the adaptive weighting model implements the weighted fusion of multi-modal data using an adaptive algorithm to output the set of first joint points of the first pedestrian at time k through the following steps a-c.
[0060] Step a, based on the multi-modal data of the first region at time k, obtain the set of second joint points of the first pedestrian at time k.
[0061] The set of second joint points includes n joint points of the first pedestrian corresponding to each image modality.
[0062] For the image modality, the image data collected by different acquisition devices belongs to different modality data. For example, if there are A infrared sensors, the infrared sensor data collected by these A infrared sensors belongs to A infrared sensor data.
[0063] Suppose there are a total of P image modality data, and the first person has n joint points. Then, in the second joint point set at time k, the n joint points corresponding to the p-th image modality can be expressed as . The second joint point set can be expressed as , that is, the second joint point set includes a total of n×P joint points. Among them, i, p, n, and P are all positive integers. The value range of i is from 1 to n, and the value range of p is from 1 to P.
[0064] Step b: Based on the multi-modal data of the first region at time k, determine the adaptive weight of each joint point in the second joint point set.
[0065] Since the multi-modal data comes from the observation of the same environmental scene and there is a coupling phenomenon, the multi-modal data can be weighted by introducing adaptive weights. Optionally, use formula (1) to determine the adaptive weight of the i-th joint point corresponding to the p-th image modality in the second joint point set: (1) In formula (1), represents the adaptive weight of the i-th joint point corresponding to the p-th image modality in the second joint point set, represents the confidence of the j-th modality in the i-th joint point at time k, represents the signal strength of the j-th modality in the i-th joint point at time k. The multi-modal data includes data of N modalities, and among the data of these N modalities, there are data of P image modalities. i, j, p, n, N, and P are positive integers. The value range of i is from 1 to n, the value range of j is from 1 to N, and the value range of p is from 1 to P.
[0066] For different modalities, and are calculated in different ways. The calculation methods of and for different modalities in the multi-modal data are described below.
[0067] Optionally, for the image modality (that is, including video data and infrared sensor data), use formula (2) to calculate , and use formula (3) or formula (4) to calculate .
[0068] (2) In formula (2), m is the total number of joint points detected in the image of the image modality at time k, is the confidence of the r-th joint point; m and r are integers, and the value range of r is from 1 to m. The meanings of other parameters in formula (2) are the same as those in formula (1), and are not elaborated here. Since there may be more than just the first pedestrian in a frame of image, and there are other pedestrians, the total number of joint points m detected in the image at time k is greater than or equal to n.
[0069] Exemplarily, the confidence of the r-th joint point is calculated using a human joint point detection algorithm such as OpenPose. There are many implementation methods of human joint point detection algorithms such as OpenPose in the related art, and are not elaborated here.
[0070] Optionally, the of the video data is calculated using formula (3).
[0071] (3) In formula (3), is the frame rate of the video data, is the resolution of the video data. The meanings of other parameters in formula (3) are the same as those in formula (1) and formula (2), and are not elaborated here.
[0072] Optionally, the of the infrared sensor data is calculated using formula (4).
[0073] (4) In formula (4), is used to indicate the signal-to-noise ratio of the temperature noise. The meanings of other parameters in formula (4) are the same as those in formula (1), and are not elaborated here.
[0074] Optionally, the of Wi-Fi data and Bluetooth data is calculated using formula (5), and the of RFID data is calculated using formula (6).
[0075] (5) In formula (5), is the signal strength, is the expected signal strength, is the standard deviation of the signal fluctuation, is the distance, representing the distance between the signal strength collector and the transmitter, is the expected distance, is the standard deviation of the distance. The meanings of other parameters in formula (5) are the same as those in formula (1), and are not elaborated here.
[0076] (6) In formula (6), is the signal strength, is the minimum value of the signal strength, is the maximum value of the signal strength. The meanings of other parameters in formula (6) are the same as those in formula (1) and are not elaborated here.
[0077] Optionally, the of the RFID data is calculated using formula (7).
[0078] (7) In formula (7), represents the total number of attempts by the RFID tag to read, represents the number of successful reads by the RFID tag. The meanings of other parameters in formula (7) are the same as those in formula (1) and are not elaborated here.
[0079] Optionally, the of the Wi-Fi data is calculated using formula (8).
[0080] (8) In formula (8), represents the total observation time of the first pedestrian, represents the Wi-Fi connection time. The meanings of other parameters in formula (8) are the same as those in formula (1) and are not elaborated here.
[0081] Optionally, the of the Bluetooth data is calculated using formula (9).
[0082] (9) In formula (9), is the signal strength of the Bluetooth data, is the maximum value of the signal strength of the Bluetooth data. The meanings of other parameters in formula (9) are the same as those in formula (1) and are not elaborated here.
[0083] Optionally, after the above is calculated, each needs to be normalized so that the sum of all is 1. The normalization process includes: dividing each calculated by the sum of each that is, dividing by .
[0084] Step c: Weight the joints in the second joint set using the adaptive weights of each joint in the second joint set to obtain the first joint set of the first pedestrian at time k.
[0085] The first set of joint points includes n weighted joint points.
[0086] For the i-th joint point in the first set of joint points, step c can be represented by formula (10).
[0087] (10) In formula (10), is the i-th joint point in the first set of joint points, represents the i-th joint point corresponding to the p-th image modality in the second set of joint points. The meanings of other parameters in formula (10) are the same as those in formula (1), which are omitted here for detailed description.
[0088] The joint points in the first set of joint points obtained by the above weighting method may be different from the true joint points. In order to make the joint points in the first set of joint points as close as possible to the true joint points, in the embodiments of the present disclosure, a deep learning method is used to optimize the weights determined by the adaptive weighting model.
[0089] The goal of this deep learning is to make as small as possible. Among them represents the true value of the i-th joint point at time k.
[0090] Before performing deep learning, a multi-modal data training set needs to be constructed first. The multi-modal data training set includes data of each modality and the true values of the first pedestrian joint points in each image modality .
[0091] During the deep learning process, the adaptive weight of the i-th joint point corresponding to the p-th image modality in the optimized second set of joint points is represented by formula (11).
[0092] (11) In formula (11), represents the adaptive weight of the i-th joint point corresponding to the p-th image modality in the optimized second set of joint points, is the adjusted weight value of the i-th joint point corresponding to the p-th image modality, and the adjusted weight value is obtained by deep learning, is the smoothing factor, .
[0093] Here, the process of deep learning is also the process of learning the adjusted weight value. After the deep learning is completed, the optimal adjusted weight value can be determined. During the deep learning process, after the end of the previous round of training, all should be retained and used for the next round of training. After multiple rounds of training, the model converges. When the model converges, That is the final .
[0094] After optimizing the weights determined by the adaptive weighted model in the way of deep learning, formula (10) in step c becomes the form of formula (12).
[0095] (12) The meanings of the parameters in formula (12) are the same as those in formula (10) and formula (11), and the detailed description is omitted here.
[0096] In step 203, based on the set of first joint points of the first pedestrian at time k, the motion trajectory of the first pedestrian is determined by using HGCN and ST-GNN.
[0097] Optionally, step 203 includes the following steps d-f.
[0098] Step d, based on the set of first joint points of the first pedestrian at time k, construct a first heterogeneous graph.
[0099] The first heterogeneous graph includes temporal edges, spatial edges and heterogeneous edges. Figure 3 is a schematic diagram of the first heterogeneous graph. Figure 3 Part (a) of is a schematic diagram of the temporal edge, Figure 3 Part (b) of is a schematic diagram of the spatial edge, Figure 3 Part (c) of is a schematic diagram of the heterogeneous edge.
[0100] Next, in combination with Figure 3 Brief descriptions of the temporal edge, spatial edge and heterogeneous edge are given. The temporal edge is the edge obtained by connecting the same joint points in two adjacent frames of images; the spatial edge refers to the edge that connects the joint points into the shape of a human body in a certain frame of image; the heterogeneous edge refers to the edge obtained by connecting the same joint points in the images at the same moment in different modalities.
[0101] The first heterogeneous graph can be expressed as , where is the first heterogeneous graph, is the set of first joint points; is the set of edges, and the set of edges includes temporal edges, spatial edges and heterogeneous edges; is to map the joint point to the node type of a specific modality, is to map the edge to the edge type of the specific relationship. In the heterogeneous graph, the node feature matrix X and the adjacency matrix A containing edge weights are initialized. In step e, HGCN will encode each node feature based on the node feature matrix X.
[0102] Step e: Use HGCN to encode the node features of the first heterogeneous graph to obtain multiple node features.
[0103] Optionally, for node , the process of its feature update is represented by formula (13).
[0104] (13) In formula (13), is the feature of the updated node ; is the set of edge types, including time correlation edges, space correlation edges, etc.; is the neighbor node set connected to node through edge type ; is the normalization coefficient used to balance the influence of different types of edges; is the neighbor node set related to edge type ; is the feature of neighbor node at the th layer; is the activation function, for example, ReLU (Rectified Linear Unit).
[0105] For different node types in the heterogeneous graph, a type weight matrix is used for type-specific feature transformation and type-level aggregation. The process of this feature update is represented by formula (14).
[0106] (14) In formula (14), is the type weight matrix, is the feature of node at the th layer. The meanings of other parameters in formula (14) are the same as those in formula (13) and are not elaborated here.
[0107] After encoding the node features of the first heterogeneous graph using HGCN, the multiple node features output by HGCN are in a unified embedding representation form , where is the number of layers of HGCN.
[0108] Step f: Input the multiple node features into ST-GNN to obtain the motion trajectory of the first pedestrian output by ST-GNN.
[0109] The ST-GNN includes an input layer, a graph convolutional layer, an activation function layer, a spatio-temporal graph neural network layer, and a decoding layer that are connected in sequence.
[0110] The input layer is used to receive multiple node features from the HGCN , and the multiple node features include the high-level semantic features of the nodes. Optionally, the input layer is also used to receive the initialized adjacency matrix A, which is used to define the connection relationships between the nodes in the heterogeneous graph.
[0111] The graph convolutional layer is used to transmit information through the adjacency matrix and update the node features.
[0112] The activation function layer is used to perform a non-linear transformation on the node features after graph convolution. Activation functions such as ReLU, Sigmoid, and Tanh can be used in the activation function layer.
[0113] The spatio-temporal graph neural network layer is a graph convolutional layer that combines spatio-temporal relationships and is used to capture the dynamic associations in space and time. The features generated through temporal attention will be combined with the spatio-temporal graph structure to further update the node features.
[0114] The decoding layer is used to map the spatio-temporal GNN trajectory output to corresponding feature points.
[0115] The HGCN inputs the node features into the ST-GNN. After the ST-GNN processes the node features, the trajectory points can be output in the decoding layer, and these trajectory points are the motion trajectories of the first pedestrian.
[0116] In step 204, the motion process of the pedestrian is established as an MDP model.
[0117] The MDP (Markov Decision Process) model is used to predict the behavior patterns of pedestrians. The basic components of the MDP are: state space, action space, transition probability, and reward function.
[0118] In the embodiments of the present disclosure, the state space is used to describe information such as the position, speed, and direction of each joint point of the pedestrian.
[0119] The position is the current position of each joint point of a pedestrian, which is a three-dimensional coordinate (x, y, z). The speed refers to the speed of each joint point of each pedestrian. The behavior refers to the current behavior pattern of the pedestrian (such as walking, staying, running, etc.). The time information refers to the current time, that is, the time of the frame where the current pedestrian monitoring video is located. For the convenience of calculation, the trajectory-based time series information can be discretized, and the continuous time, position, speed and other information can be discretized into a finite state set.
[0120] The action space is used to describe the set of actions of a pedestrian such as "forward", "backward", "stay", "turn", "accelerate or decelerate", etc. Forward means that the pedestrian moves forward a certain distance. Backward means that the pedestrian moves backward a certain distance. Stay means that the pedestrian stays at the current position. Turn means that the pedestrian changes the direction of travel. Accelerate or decelerate means that the pedestrian changes the moving speed.
[0121] The action space needs to be adjusted according to the target task. If the task is to detect abnormal behaviors, "abnormal behaviors" can be included as special actions, such as crowding, pushing or even trampling among pedestrians, etc.
[0122] The transition probability is used to describe the possibility of a pedestrian changing from one action to another. The probability of a pedestrian changing from the forward behavior to turning, the probability of a pedestrian changing from staying to moving forward, and so on. When initially constructing the MDP model, the transition probability can be obtained by statistical analysis of historical data, and the transition probability can be continuously updated during the subsequent reinforcement training process of the model.
[0123] The reward function is a mapping function used in the training process of the pedestrian behavior prediction model based on trajectory points. When designing the reward function, factors such as the correctness of the behavior pattern, spatial rationality, and discount factor etc.
[0124] Among them, the correctness of the behavior pattern refers to giving a reward based on the matching degree between the predicted behavior according to the trajectory and the actual behavior. If the trajectory matches the predicted behavior pattern, a positive reward is given; if it deviates from the expected trajectory or behavior, a negative reward is given.
[0125] Spatial rationality means that if the trajectory of a pedestrian conforms to the actual situation in space (for example, there are no abnormalities such as passing through walls), a higher reward is given. For spatial rationality, corresponding knowledge needs to be input into the model. That is to say, this MDP model is data-knowledge dual-driven.
[0126] Discount factor is a parameter between 0 and 1 during the training process of this MDP model, which is used to measure the importance of the reward value returned by the above reward function. A higher discount factor indicates that this reward has a greater impact on the decision-making of the model.
[0127] After the MDP model is established, some historical pedestrian trajectory point data can be used to train the MDP model, enabling the MDP model to learn how to select the optimal actions to obtain the maximum cumulative reward. After the MDP model is trained, it can judge the behavior pattern of a pedestrian based on the input motion trajectory of the pedestrian. The behavior patterns include "walking", "staying", or "running fast", etc.
[0128] In step 205, the motion trajectory of the first pedestrian is input into the MDP model, and the behavior pattern of the first pedestrian output by the MDP model is obtained.
[0129] By analyzing the behavior pattern, it is possible to predict the future motion trend of the pedestrian, thus playing a predictive and supplementary role in tracking interruptions caused by occlusion or perception loss.
[0130] In the case where the recognized behavior pattern has abnormal behaviors (such as wandering or ramming), an alarm for abnormal behaviors can also be given.
[0131] Optionally, when there is a first pedestrian in the image modality at the k-th moment and there is no first pedestrian in the image modality at the (k + 1)-th moment, the data of other modalities at the (k + 1)-th moment and the behavior pattern of the first pedestrian at the k-th moment are used to track the pedestrian. The data of other modalities can determine the position of the first pedestrian, and the behavior pattern can predict the action trend of the first pedestrian. In the case where there is no first pedestrian in the image modality, the behavior pattern can play an auxiliary correction role in the position of the first pedestrian determined by using the data of other modalities.
[0132] For example, if the behavior pattern of the first pedestrian predicted at the k-th moment is a fast running pattern, the image modality of the first pedestrian is lost at the (k + 1)-th moment, and the data of other modalities at the (k + 1)-th moment indicates that the first pedestrian is in a stopped state, it means that the data of other modalities may be incorrect. At this time, the first pedestrian can be continuously tracked according to the fast running pattern predicted at the k-th moment.
[0133] The following is an apparatus embodiment of the present application. For details not described in detail in the apparatus embodiment, reference can be made to the above method embodiment.
[0134] Figure 4 The structural schematic diagram of a pedestrian tracking apparatus based on multi-modal data provided by an exemplary embodiment of the present disclosure is shown. Refer to Figure 4 The pedestrian tracking apparatus 400 based on multi-modal data includes: a first acquisition module 401, a second acquisition module 402, a trajectory determination module 403, and a behavior pattern determination module 404.
[0135] The first acquisition module 401 is used to acquire multimodal data of the first area at time k, and the multimodal data includes video data, RFID data, Wi-Fi data, Bluetooth data, and infrared sensor data.
[0136] The second acquisition module 402 is used to input the multimodal data into the adaptive weighting model and obtain the set of first joint points of the first pedestrian at time k output by the adaptive weighting model. The adaptive weighting model is used to output the set of first joint points of the first pedestrian at time k after weighted fusion of the multimodal data using an adaptive algorithm.
[0137] The trajectory determination module 403 is used to determine the motion trajectory of the first pedestrian based on the set of first joint points of the first pedestrian at time k, using the heterogeneous graph convolutional network HGCN and the spatio-temporal graph neural network ST-GNN based on the time attention mechanism.
[0138] The behavior pattern determination module 404 is used to determine the behavior pattern of the first pedestrian based on the motion trajectory of the first pedestrian, and the behavior pattern is used to track the first pedestrian.
[0139] Optionally, the first pedestrian includes n joint points, the multimodal data includes at least one image modality, the image modality includes video data and infrared sensor data, and the second acquisition module 402 is further used to implement the output of the set of first joint points of the first pedestrian at time k after weighted fusion of the multimodal data using an adaptive algorithm in the following manner: based on the multimodal data of the first area at time k, obtain the set of second joint points of the first pedestrian at time k, and the set of second joint points includes n joint points of the first pedestrian corresponding to each image modality; based on the multimodal data of the first area at time k, determine the adaptive weight of each joint point in the set of second joint points; use the adaptive weight of each joint point in the set of second joint points to weight the joint points in the set of second joint points to obtain the set of first joint points of the first pedestrian at time k, and the set of first joint points includes n weighted joint points.
[0140] Optionally, the second acquisition module 402 is further used to determine the adaptive weight of the i-th joint point corresponding to the p-th image modality in the set of second joint points using the following formula:
[0141] where represents the adaptive weight of the i-th joint point corresponding to the p-th image modality in the set of second joint points, represents the confidence of the j-th modality at time k for the i-th joint point, It represents the signal strength of the j-th modality for the i-th joint point at time k. The multi-modal data includes data of N modalities, and the data of N modalities includes data of P image modalities. i, j, p, n, N, and P are positive integers. The value range of i is from 1 to n, the value range of j is from 1 to N, and the value range of p is from 1 to P.
[0142] Optionally, the device further includes: an optimization module 405, configured to optimize the adaptive weight of the i-th joint point corresponding to the p-th image modality in the second joint point set determined by the adaptive weighted model by means of deep learning; wherein, the adaptive weight of the i-th joint point corresponding to the p-th image modality in the optimized second joint point set is represented by the following formula:
[0143] wherein, It represents the adaptive weight of the i-th joint point corresponding to the p-th image modality in the optimized second joint point set, is the adjusted weight value of the i-th joint point corresponding to the p-th image modality, and the adjusted weight value is obtained by means of deep learning, is the smoothing factor.
[0144] Optionally, the trajectory determination module 403 is further configured to construct a first heterogeneous graph based on the first joint point set of the first pedestrian at time k. The first heterogeneous graph includes a temporal edge, a spatial edge, and a heterogeneous edge; use HGCN to encode the node features of the first heterogeneous graph to obtain multiple node features; input the multiple node features into ST-GNN to obtain the motion trajectory of the first pedestrian output by ST-GNN.
[0145] Optionally, the behavior pattern determination module 404 is further configured to establish the motion process of the pedestrian as a Markov decision process MDP model, and the MDP model is used to predict the behavior pattern of the pedestrian; input the motion trajectory of the first pedestrian into the MDP model to obtain the behavior pattern of the first pedestrian output by the MDP model.
[0146] It should be noted that: when the above-mentioned pedestrian tracking device based on multi-modal data performs pedestrian tracking, only the above-mentioned division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the above-mentioned pedestrian tracking device based on multi-modal data and the embodiment of the pedestrian tracking method based on multi-modal data belong to the same concept. For the specific implementation process, please refer to the method embodiment, which will not be elaborated here.
[0147] In the embodiments of the present disclosure, the division of modules is illustrative. It is only a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present disclosure, each functional module may be integrated in a processor, may exist independently physically, or two or more modules may be integrated into one module. The above integrated module may be implemented in the form of hardware or in the form of a software functional module.
[0148] If the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a terminal device (which may be a personal computer, a mobile phone, or a communication device, etc.) or a processor to execute all or part of the steps of the method in each embodiment of the present disclosure. The aforementioned storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0149] Figure 5 It is a schematic structural diagram of a computer device provided by an embodiment of the present disclosure. As Figure 5 shown, the computer device 500 includes: a processor 501 and a memory 502.
[0150] The processor 501 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 501 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 501 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 501 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 501 may further include an AI (Artificial Intelligence) processor, which is used to process computational operations related to machine learning.
[0151] The memory 502 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 502 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 502 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 501 to implement the multi-modal data-based pedestrian tracking method provided in the embodiments of the present disclosure.
[0152] Those skilled in the art can understand that Figure 5 the structure shown in
[0153] does not constitute a limitation on the computer device 500, and may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component layout.
[0154] The embodiments of the present disclosure also provide a non-temporary computer-readable storage medium. When the instructions in the storage medium are executed by the processor of the computer device, the computer device can execute the multi-modal data-based pedestrian tracking method provided in the embodiments of the present disclosure.
[0155] The above are only alternative embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A pedestrian tracking method based on multimodal data, characterized in that: The method comprises: Acquire multimodal data of the first area at time k, where the multimodal data includes video data, RFID data, Wi-Fi data, Bluetooth data, and infrared sensor data; The multimodal data is input into an adaptive weighted model, and a first joint point set of the first pedestrian at time k output by the adaptive weighted model is obtained, wherein the adaptive weighted model is used to perform weighted fusion on the multimodal data using an adaptive algorithm and then output the first joint point set of the first pedestrian at time k; Based on the first joint point set of the first pedestrian at time k, a heterogeneous graph convolutional network HGCN and a spatiotemporal graph neural network ST-GNN based on a temporal attention mechanism are used to determine the motion trajectory of the first pedestrian; A behavior pattern of the first pedestrian is determined based on the motion trajectory of the first pedestrian, and the behavior pattern is used to track the first pedestrian.
2. The method according to claim 1, characterized in that The first pedestrian includes n joint points, the multimodal data includes at least one image modality, and the image modality includes the video data and the infrared sensor data. The adaptive weighted model is used to implement the weighted fusion of the multimodal data using an adaptive algorithm and then output the first joint point set of the first pedestrian at time k in the following manner: Based on the multimodal data of the first area at time k, obtaining a second joint point set of the first pedestrian at time k, where the second joint point set includes n joint points of the first pedestrian corresponding to each of the image modalities; Determining an adaptive weight of each joint point in the second joint point set based on the multimodal data of the first area at time k; The joint points in the second joint point set are weighted using the adaptive weight of each joint point in the second joint point set to obtain a first joint point set of the first pedestrian at time k, wherein the first joint point set includes n weighted joint points.
3. The method according to claim 2, characterized in that The step of determining the adaptive weight of each joint point in the second joint point set based on the multimodal data of the first region at time k includes: The adaptive weight of the i-th joint point corresponding to the p-th image modality in the second joint point set is determined by the following formula: in, represents the adaptive weight of the i-th joint point corresponding to the p-th image modality in the second joint point set, represents the confidence of the j-th mode at time k for the i-th joint point, It represents the signal strength of the j-th modality for the i-th joint point at time k, the multimodal data includes data of N modalities, the N modal data includes data of P image modalities, i, j, p, n, N, P are positive integers, the value range of i is 1 to n, the value range of j is 1 to N, and the value range of p is 1 to P.
4. The method according to claim 3, characterized in that: The method further comprises: Optimizing the adaptive weight of the i-th joint point corresponding to the p-th image modality in the second joint point set determined by the adaptive weighted model by using a deep learning method; The adaptive weight of the i-th joint point corresponding to the p-th image modality in the optimized second joint point set is expressed by the following formula: in, represents the adaptive weight of the i-th joint point corresponding to the p-th image modality in the optimized second joint point set, is the adjustment weight of the i-th joint point corresponding to the p-th image modality, and the adjustment weight is obtained by deep learning. is the smoothing factor.
5. The method according to any one of claims 1 to 4, characterized in that: The method of determining the motion trajectory of the first pedestrian based on the first joint point set of the first pedestrian at time k by using a heterogeneous graph convolutional network HGCN and a spatiotemporal graph neural network ST-GNN based on a temporal attention mechanism includes: Based on the first joint point set of the first pedestrian at time k, construct a first heterogeneous graph, wherein the first heterogeneous graph includes a time edge, a space edge, and a heterogeneous edge; Encoding the node features of the first heterogeneous graph using the HGCN to obtain a plurality of node features; The multiple node features are input into the ST-GNN to obtain the motion trajectory of the first pedestrian output by the ST-GNN.
6. The method according to any one of claims 1 to 4, characterized in that: The determining the behavior pattern of the first pedestrian based on the motion trajectory of the first pedestrian includes: The movement process of pedestrians is established as a Markov decision process (MDP) model, and the MDP model is used to predict the behavior pattern of pedestrians; The motion trajectory of the first pedestrian is input into the MDP model to obtain a behavior pattern of the first pedestrian output by the MDP model.
7. A pedestrian tracking device based on multimodal data, characterized in that: The device comprises: A first acquisition module is used to acquire multimodal data of the first area at time k, where the multimodal data includes video data, RFID data, Wi-Fi data, Bluetooth data, and infrared sensor data; a second acquisition module, configured to input the multimodal data into an adaptive weighted model, and to obtain a first joint point set of the first pedestrian at time k output by the adaptive weighted model, wherein the adaptive weighted model is configured to perform weighted fusion on the multimodal data using an adaptive algorithm and then output the first joint point set of the first pedestrian at time k; A trajectory determination module, used to determine the motion trajectory of the first pedestrian based on the first joint point set of the first pedestrian at time k, using a heterogeneous graph convolutional network HGCN and a spatiotemporal graph neural network ST-GNN based on a temporal attention mechanism; A behavior pattern determination module is used to determine the behavior pattern of the first pedestrian based on the motion trajectory of the first pedestrian, and the behavior pattern is used to track the first pedestrian.
8. A computer device, characterized in that: The computer device comprises: a memory and a processor, wherein at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by a processor to implement the method according to any one of claims 1 to 6.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
End-to-end multi-target identification, tracking and prediction method
CN114169241A
Human body motion track extraction method and system based on vision algorithm
CN117474948A
Intelligent traffic monitoring method and device based on multi-modal data fusion and graph neural network, and electronic equipment
CN118366311A
Personnel attitude estimation method and system based on multi-model graph neural network fusion
CN118942153A
Human body action recognition method, human body action recognition system, and device
WO2022000420A1
Cited By
Multi-modal network global routing optimization method and device, electronic equipment and medium
CN120358190A