Pedestrian tracking method and device based on multi-modal data, equipment and storage medium
By combining adaptive weighting of multimodal data with HGCN and ST-GNN, the accuracy problem of traditional pedestrian tracking in complex environments is solved, and accurate tracking and behavior prediction are achieved even when there is occlusion or changes in lighting.
Patent Information
- Application Number
- CN202510047776.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-01-13
AI Technical Summary
Traditional video surveillance systems struggle to achieve stable pedestrian tracking in complex and dynamically changing environments, especially when there are obstructions or changes in lighting conditions, which can reduce tracking accuracy.
Adaptive weighted fusion of multimodal data (video data, RFID data, Wi-Fi data, Bluetooth data, and infrared sensor data) is used, combined with a heterogeneous graph convolutional network (HGCN) and a spatiotemporal graph neural network (ST-GNN) based on a time attention mechanism, to predict pedestrian movement trajectories and determine behavioral patterns.
Accurately tracking pedestrians in complex environments improves the accuracy of motion trajectory prediction, enabling accurate tracking and prediction of future movement trends even when video is obscured or blurry.
Smart Images

Figure CN120067746B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of pedestrian tracking, and in particular relates to a pedestrian tracking method and device based on multi-modal data, equipment and a storage medium. BACKGROUND
[0002] In modern smart city construction, pedestrian tracking technology plays an important role in urban public safety, traffic management and emergency response. Although traditional video surveillance systems have been widely used in these scenarios, video data often fails to provide stable and accurate pedestrian tracking results in the presence of occlusion, changes in lighting and complex scenes.
[0003] The pedestrian tracking method in the related art is usually a simple filtering algorithm (such as Kalman filtering algorithm), which is difficult to achieve stable target tracking in complex and dynamic environments. When facing complex scenes such as fast movement of pedestrians, image occlusion or environmental changes, using a simple filtering algorithm for pedestrian tracking will result in a decrease in tracking accuracy. SUMMARY
[0004] The present disclosure provides a pedestrian tracking method and device based on multi-modal data, equipment and a storage medium, which can accurately track pedestrians in complex scenes. The technical solution at least includes the following solutions:
[0005] In a first aspect, a pedestrian tracking method based on multi-modal data is provided, comprising: obtaining multi-modal data of a first region at time k, the multi-modal data comprising video data, RFID data, Wi-Fi data, Bluetooth data and infrared sensor data; inputting the multi-modal data into an adaptive weighting model, and obtaining a first set of junctions of the first pedestrian at time k output by the adaptive weighting model, the adaptive weighting model being configured to output the first set of junctions of the first pedestrian at time k after weighting and fusing the multi-modal data using an adaptive algorithm; determining a motion trajectory of the first pedestrian based on the first set of junctions of the first pedestrian at time k using a heterogeneous graph convolution network (HGCN) and a spatio-temporal graph neural network (ST-GNN) based on a time attention mechanism; and determining a behavior pattern of the first pedestrian based on the motion trajectory of the first pedestrian, the behavior pattern being used to track the first pedestrian.
[0006] Optionally, the first pedestrian includes n joints, the multi-modal data includes at least one image modality, the image modality includes the video data and the infrared sensor data, and the adaptive weighting model is configured to output the first joint set of the first pedestrian at the k-th moment by weighting and fusing the multi-modal data using the adaptive algorithm in the following manner: based on the multi-modal data of the first region at the k-th moment, obtaining a second joint set of the first pedestrian at the k-th moment, the second joint set including n joints of the first pedestrian corresponding to each image modality; based on the multi-modal data of the first region at the k-th moment, determining an adaptive weight of each joint in the second joint set; weighting the joints in the second joint set using the adaptive weight of each joint in the second joint set to obtain the first joint set of the first pedestrian at the k-th moment, the first joint set including n weighted joints.
[0007] Optionally, the method further comprises: determining the adaptive weight of the i-th joint corresponding to the p-th image modality in the second joint set using the adaptive weighting model in a deep learning manner; and wherein the adaptive weight of the i-th joint corresponding to the p-th image modality in the second joint set after optimization is represented by the following formula:
[0008]
[0009] wherein, represents the adaptive weight of the i-th joint corresponding to the p-th image modality in the second joint set, represents the confidence of the j-th modality for the i-th joint at the k-th moment, represents the signal strength of the j-th modality for the i-th joint at the k-th moment, the multi-modal data includes data of N modalities, the data of the N modalities includes data of P image modalities, i, j, p, n, N, and P are positive integers, the value range of i is 1 to n, the value range of j is 1 to N, and the value range of p is 1 to P.
[0010] Optionally, the method further comprises: determining the adaptive weight of the i-th joint corresponding to the p-th image modality in the second joint set using the adaptive weighting model in a deep learning manner; and wherein the adaptive weight of the i-th joint corresponding to the p-th image modality in the second joint set after optimization is represented by the following formula:
[0011]
[0012] wherein, represents the adaptive weight of the i-th joint corresponding to the p-th image modality in the second joint set after optimization, an adjustment weight of an i th joint node corresponding to a p th image modality, the adjustment weight being obtained by using deep learning, a smoothing factor.
[0013] Optionally, the motion trajectory of the first pedestrian is determined based on the first set of joint nodes of the first pedestrian at the k th moment by using a heterogeneous graph convolution network (HGCN) and a spatio-temporal graph neural network (ST-GNN) based on a time attention mechanism, including: a first heterogeneous graph is constructed based on the first set of joint nodes of the first pedestrian at the k th moment, the first heterogeneous graph including time edges, space edges and heterogeneous edges; node features of the first heterogeneous graph are encoded by using the HGCN to obtain a plurality of node features; the plurality of node features are input into the ST-GNN to obtain the motion trajectory of the first pedestrian output by the ST-GNN.
[0014] Optionally, the behavior pattern of the first pedestrian is determined based on the motion trajectory of the first pedestrian, including: a motion process of a pedestrian is established as a Markov decision process (MDP) model, the MDP model being used to predict a behavior pattern of a pedestrian; the motion trajectory of the first pedestrian is input into the MDP model to obtain the behavior pattern of the first pedestrian output by the MDP model.
[0015] In a second aspect, a pedestrian tracking device based on multi-modal data is also provided, including: a first acquisition module configured to acquire multi-modal data of a first region at a k th moment, the multi-modal data including video data, RFID data, Wi-Fi data, Bluetooth data and infrared sensor data; a second acquisition module configured to input the multi-modal data into an adaptive weighting model and acquire a first set of joint nodes of the first pedestrian at the k th moment output by the adaptive weighting model, the adaptive weighting model being configured to output the first set of joint nodes of the first pedestrian at the k th moment after weighting and fusing the multi-modal data by using an adaptive algorithm; a trajectory determination module configured to determine a motion trajectory of the first pedestrian based on the first set of joint nodes of the first pedestrian at the k th moment by using a heterogeneous graph convolution network (HGCN) and a spatio-temporal graph neural network (ST-GNN) based on a time attention mechanism; and a behavior pattern determination module configured to determine a behavior pattern of the first pedestrian based on the motion trajectory of the first pedestrian, the behavior pattern being used to track the first pedestrian.
[0016] Optionally, the first pedestrian includes n joints, the multi-modal data includes at least one image modality, the image modality includes the video data and the infrared sensor data, and the second acquisition module is further configured to: acquire a second joint set of the first pedestrian at the k th moment based on the multi-modal data of the first region at the k th moment, the second joint set including n joints of the first pedestrian corresponding to each image modality; determine an adaptive weight of each joint in the second joint set based on the multi-modal data of the first region at the k th moment; and weight the joints in the second joint set based on the adaptive weight of each joint in the second joint set, to obtain a first joint set of the first pedestrian at the k th moment, the first joint set including n weighted joints.
[0017] Optionally, the second acquisition module is further configured to determine the adaptive weight of the i th joint corresponding to the p th image modality in the second joint set according to the following formula:
[0018]
[0019] wherein, represents the adaptive weight of the i th joint corresponding to the p th image modality in the second joint set, represents the confidence of the j th modality for the i th joint at the k th moment, represents the signal strength of the j th modality for the i th joint at the k th moment, the multi-modal data including data of N modalities, the data of the N modalities including data of P image modalities, i, j, p, n, N, and P being positive integers, the value range of i being 1 to n, the value range of j being 1 to N, and the value range of p being 1 to P.
[0020] Optionally, the device further includes an optimization module configured to optimize the adaptive weight of the i th joint corresponding to the p th image modality in the second joint set determined by the adaptive weighting model in a deep learning manner, wherein the adaptive weight of the i th joint corresponding to the p th image modality in the optimized second joint set is represented by the following formula:
[0021]
[0022] wherein, represents the adaptive weight of the i th joint corresponding to the p th image modality in the optimized second joint set, The adjustment weights are the values for the i-th joint corresponding to the p-th image modality, and these adjustment weights are obtained using deep learning. This is a smoothing factor.
[0023] Optionally, the trajectory determination module is further configured to construct a first heterogeneous graph based on the first set of first joint points of the first pedestrian at time k, the first heterogeneous graph including time edges, spatial edges and heterogeneous edges; encode the node features of the first heterogeneous graph using the HGCN to obtain multiple node features; input the multiple node features into the ST-GNN to obtain the motion trajectory of the first pedestrian output by the ST-GNN.
[0024] Optionally, the behavior pattern determination module is further configured to establish the pedestrian's movement process as a Markov Decision Process (MDP) model, the MDP model being used to predict the pedestrian's behavior pattern; inputting the first pedestrian's movement trajectory into the MDP model to obtain the first pedestrian's behavior pattern output by the MDP model.
[0025] Thirdly, a computer device is also provided, comprising: a memory and a processor, wherein the memory stores at least one computer program, the at least one computer program being loaded and executed by the processor to perform the pedestrian tracking method based on multimodal data described in the above embodiments.
[0026] Fourthly, a computer-readable storage medium is also provided, wherein at least one computer program is stored in the computer-readable storage medium, the at least one computer program being loaded and executed by a processor to perform the pedestrian tracking method based on multimodal data described in the above embodiments.
[0027] Fifthly, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the method described in the first aspect.
[0028] The beneficial effects of the technical solutions provided in this disclosure include at least the following:
[0029] In this embodiment, by adaptively weighting the multimodal data, the accuracy of each joint in the obtained first joint set can be improved. By inputting the first joint set into HGCN and ST-GNN, the movement trajectory of the first pedestrian can be accurately predicted, and then the pedestrian's behavior pattern can be predicted based on the movement trajectory. This behavior pattern can indicate the pedestrian's movement trend over a future period of time. Thus, even in complex environments (such as environments where the video is occluded or blurry), the pedestrian can be accurately tracked based on the behavior pattern. Attached Figure Description
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings needed to be used in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without any creative effort based on these drawings.
[0031] Figure 1 A flow chart of a pedestrian tracking method based on multi-modal data provided by an example embodiment of the present disclosure is shown;
[0032] Figure 2 A flow chart of a pedestrian tracking method based on multi-modal data provided by another example embodiment of the present disclosure is shown;
[0033] Figure 3 is a schematic diagram of a first heterogeneous graph;
[0034] Figure 4 A structural schematic diagram of a pedestrian tracking device based on multi-modal data provided by an example embodiment of the present disclosure is shown;
[0035] Figure 5 is a structural schematic diagram of a computer device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0036] Unless otherwise defined, technical terms or scientific terms used herein should be understood as their common meanings to those skilled in the art to which the present disclosure pertains. The terms “first”, “second”, “third” and the like used in the specification and claims of the present patent application do not denote any order, quantity or importance, but are only used to distinguish different components. Similarly, the terms “one” or “a” and the like do not denote quantity limitation, but mean that there is at least one. The terms “include” or “contain” and the like mean that the elements or objects appearing before “include” or “contain” cover the elements or objects listed after “include” or “contain” and their equivalents, and do not exclude other elements or objects. The terms “connect” or “connected” and the like are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect.
[0037] In order to make the purposes, technical solutions and advantages of the present disclosure clearer, the embodiments of the present disclosure will be described in further detail below with reference to the drawings.
[0038] Figure 1 A flow chart of a pedestrian tracking method based on multi-modal data provided by an example embodiment of the present disclosure is shown, which can be executed by a computer device. Referring to Figure 1 , the method comprises:
[0039] In step 101, multi-modal data of the first region at the k-th moment is acquired.
[0040] The multi-modal data includes video data, RFID (Radio Frequency Identification) data, Wi-Fi data, Bluetooth data, and infrared sensor data.
[0041] Here, the multi-modal data includes a first pedestrian, which is a pedestrian that needs to be tracked. The method in the present disclosure is used to track the first pedestrian, and in the case where multiple first pedestrians exist in the multi-modal data, the method in the present disclosure can also track the multiple first pedestrians simultaneously.
[0042] Optionally, a high-definition camera is used to acquire the video data. The frame rate and resolution of the video are set according to the characteristics of the monitoring region, for example, the resolution and frame rate of a camera close to the road surface (ground) can be lower, and the resolution of a camera far from the road surface needs to be higher to ensure that the human body image captured by the camera is clear and visible; in addition, the influence of the surrounding scenery can be collected, and geometric correction is performed in combination with the coordinates of nearby control points to obtain accurate background reference positions, further improving the accuracy of pedestrian positioning.
[0043] Optionally, an RFID reader is used to acquire the RFID data. In the case where the mobile device carried by the first pedestrian contains an RFID tag, the RFID reader can read the information in the RFID tag, thereby acquiring the position information of the first pedestrian. This method is particularly suitable for personnel positioning in high-traffic environments. Since not all mobile devices contain RFID tags, RFID tags can be arranged on control points with known coordinates in the background static objects in the pedestrian monitoring environment, and the position information of these fixed tags is read and applied to the video signal and infrared signal to perform pedestrian positioning reference calibration and positioning enhancement.
[0044] Optionally, a Wi-Fi signal strength collector is used to acquire the Wi-Fi data. In the case where the mobile device carried by the first pedestrian has Wi-Fi, the Wi-Fi signal strength collector can detect the change in the Wi-Fi signal strength of the mobile device carried by the first pedestrian, thereby achieving rough positioning of the pedestrian.
[0045] Optionally, a Bluetooth signal strength collector is used to acquire the Bluetooth data. In the case where the mobile device carried by the first pedestrian has Bluetooth, the Bluetooth signal strength collector can acquire the Bluetooth beacon of the mobile device carried by the first pedestrian to perform high-precision short-distance positioning, and determine the position of the first pedestrian through the change in the Bluetooth signal strength.
[0046] Optionally, the infrared sensor data is obtained by an infrared sensor. The infrared sensor data can reflect the temperature change of the first pedestrian, and is helpful for locating the first pedestrian in the case of video occlusion or insufficient light.
[0047] As can be seen, the video data and the infrared sensor data are in the form of images, both of which belong to image modalities and can reflect the posture of the first pedestrian; the RFID data, the Wi-Fi data and the Bluetooth data are in the form of data streams and can reflect the position of the first pedestrian.
[0048] In the case where the multi-modal data includes at least one image modality, that is, there is image data. The tracking of the first pedestrian includes: locating the first pedestrian from the image data (that is, determining the position of the first pedestrian), identifying the first pedestrian from the image data (for example, there can be multiple human bodies in the image data, and determining the first pedestrian from the multiple human bodies is to identify the first pedestrian), and determining the motion trajectory of the first pedestrian based on the image data.
[0049] In the case where the multi-modal data does not include an image modality, the tracking of the first pedestrian includes: determining the motion trajectory of the first pedestrian based on the data of other modalities (such as RFID data, Wi-Fi data, Bluetooth data, etc.).
[0050] In step 102, the multi-modal data is input into the adaptive weighting model, and the first set of key nodes of the first pedestrian at time k output by the adaptive weighting model is obtained.
[0051] The adaptive weighting model is used to output the first set of key nodes of the first pedestrian at time k after weighting and fusing the multi-modal data by using an adaptive algorithm.
[0052] In step 103, based on the first set of key nodes of the first pedestrian at time k, the motion trajectory of the first pedestrian is determined by using HGCN and ST-GNN.
[0053] HGCN (Heterogeneous Graph Convolutional Network) is a new type of neural network architecture that improves the performance of graph neural networks by utilizing the properties of hyperbolic space. ST-GNN (Spatio-Temporal Graph Neural Networks) can extract complex spatio-temporal dependencies by integrating graph neural networks and various temporal learning methods.
[0054] By using HGCN and ST-GNN, the motion trajectory of the first pedestrian can be accurately predicted.
[0055] In step 104, the behavior pattern of the first pedestrian is determined based on the motion trajectory of the first pedestrian.
[0056] The behavior pattern is used to track the first pedestrian. Here, the behavior pattern is the action trend of the first pedestrian in a future period of time. After the behavior pattern of the first pedestrian is determined, the motion trajectory of the first pedestrian in the future period of time can be predicted based on the behavior pattern. In this way, even in a complex environment (such as an environment in which a video is blocked or a video is blurred), the first pedestrian can be accurately tracked according to the behavior pattern.
[0057] In the embodiments of the present disclosure, by adaptively weighting the multi-modal data, the accuracy of each joint node in the obtained first joint node set can be improved; by inputting the first joint node set to the HGCN and the ST-GNN, the motion trajectory of the first pedestrian can be accurately predicted, and then the behavior pattern of the first pedestrian can be predicted based on the motion trajectory of the first pedestrian. The behavior pattern can indicate the action trend of the first pedestrian in a future period of time. In this way, even in a complex environment (such as an environment in which a video is blocked or a video is blurred), the first pedestrian can be accurately tracked according to the behavior pattern.
[0058] Figure 2 A flowchart of a method for pedestrian tracking based on multi-modal data provided by another example embodiment of the present disclosure is shown. The method can be executed by a computer device. Referring to FIG. 2, Figure 2 The method comprises:
[0059] In step 201, multi-modal data of a first region at a k-th time point is obtained.
[0060] The multi-modal data comprises video data, RFID data, Wi-Fi data, Bluetooth data, and infrared sensor data.
[0061] The content of step 201 is described above in step 101, and is not described here in detail.
[0062] Optionally, before step 202 is executed, the method further comprises: pre-processing the multi-modal data.
[0063] For image modal data in the multi-modal data, an image self-supervised denoising algorithm can be used for denoising processing.
[0064] For RFID data, Wi-Fi data, and Bluetooth data in the multi-modal data, a Kalman filter can be used for denoising processing.
[0065] The implementation of the image self-supervised denoising algorithm and the Kalman filter is relatively common in the related art, and is not described here in detail.
[0066] Since RFID signal is easily affected by environmental factors (such as obstacles, reflections), leading to signal strength fluctuation. Kalman filter is used to smooth the signal strength variation, helping to reduce the positioning error caused by signal jitter. The initial state is set as the initial position and speed estimation of the RFID tag, and then the estimated RFID tag position is updated by Kalman filter every time a new RFID signal strength is received, thereby generating stable position information.
[0067] The detection of Wi-Fi signal strength can be applied to indoor positioning in subway station security check and the like, but the signal strength value is easily affected by walls and equipment interference, resulting in short-term fluctuations. Using Kalman filter can smooth the time series of signal strength, thereby improving the accuracy of indoor positioning. The signal strength value is used as an observation variable, combined with the position information at the previous moment for prediction; the current position estimation is updated using the new signal strength value, and jitter and jumping are reduced.
[0068] The signal strength of the Bluetooth beacon has high accuracy in short-distance positioning, but it is easily disturbed in public areas with large flow or in multi-signal source environment, resulting in large fluctuations. Through Kalman filter, the influence of Bluetooth signal fluctuation can be reduced, and the positioning stability can be improved. The initial position and error covariance matrix of the Bluetooth signal are set, and prediction and update are performed after each Bluetooth signal strength is obtained, to obtain smooth Bluetooth signal strength and corresponding position information.
[0069] In step 202, the multi-modal data is input into the adaptive weighting model, and a first set of key nodes of the first pedestrian at the k moment output by the adaptive weighting model is obtained.
[0070] The adaptive weighting model is used to output the first set of key nodes of the first pedestrian at the k moment by using an adaptive algorithm to weight and fuse the multi-modal data.
[0071] Optionally, the adaptive weighting model uses steps a-c as follows to output the first set of key nodes of the first pedestrian at the k moment by using an adaptive algorithm to weight and fuse the multi-modal data.
[0072] Step a, based on the multi-modal data of the first area at the k moment, a second set of key nodes of the first pedestrian at the k moment is obtained.
[0073] The second set of key nodes includes n key nodes of the first pedestrian corresponding to each image modality.
[0074] For image modalities, image data collected by different acquisition devices belongs to different modal data. For example, there are A infrared sensors, and the infrared sensor data collected by the A infrared sensors belongs to A infrared sensor data.
[0075] Suppose that there are P image modalities in total and n joints in common for the first pedestrian, the n joints corresponding to the pth image modality in the second joint set at the kth time can be expressed as The second joint set can be expressed as That is, the second joint set includes n*P joints in total. Wherein i, p, n, P are positive integers, the value range of i is 1 to n, and the value range of p is 1 to P.
[0076] Step b, determining the adaptive weight of each joint in the second joint set based on the multi-modal data of the first region at the kth time.
[0077] Since the multi-modal data comes from the observation of the same environment scene, there is a coupling phenomenon, so the adaptive weight can be introduced to realize the weighting of the multi-modal data. Alternatively, formula (1) is used to determine the adaptive weight of the ith joint corresponding to the pth image modality in the second joint set:
[0078] (1)
[0079] In formula (1), represents the adaptive weight of the ith joint corresponding to the pth image modality in the second joint set, represents the confidence of the jth modality for the ith joint at the kth time, represents the signal strength of the jth modality for the ith joint at the kth time, the multi-modal data includes data of N modalities, and the data of the P image modalities among the N modalities, i, j, p, n, N, P are positive integers, the value range of i is 1 to n, the value range of j is 1 to N, and the value range of p is 1 to P.
[0080] For different modalities, and are calculated in different ways. The calculation methods of and for different modalities in multi-modal data are described below.
[0081] Alternatively, for image modalities (that is, including video data and infrared sensor data), formula (2) is used to calculate , and formula (3) or formula (4) is used to calculate .
[0082] (2)
[0083] In formula (2), m is the total number of joints detected in the image of the image modality at the kth time, is the confidence of the rth joint node; m, r are integers, and the value range of r is 1 to m. The meanings of other parameters in formula (2) are the same as those in formula (1), and details are omitted here. Since there can be more than one pedestrian in a frame of image, the total number of joint nodes detected in the image at time k is greater than or equal to n.
[0084] Exemplarily, the confidence of the rth joint node is calculated by using a human joint node detection algorithm such as OpenPose. The implementation of the human joint node detection algorithm such as OpenPose is relatively common in the related art, and details are omitted here.
[0085] Optionally, the confidence of the video data is calculated by using formula (3). Formula (3) is used for calculation.
[0086] (3)
[0087] In formula (3), is the frame rate of the video data, is the resolution of the video data. The meanings of other parameters in formula (3) are the same as those in formula (1) and formula (2), and details are omitted here.
[0088] Optionally, the confidence of the infrared sensor data is calculated by using formula (4). Formula (4) is used for calculation.
[0089] (4)
[0090] In formula (4), is the signal-to-noise ratio used to indicate temperature noise. The meanings of other parameters in formula (4) are the same as those in formula (1), and details are omitted here.
[0091] Optionally, the confidence of the Wi-Fi data and the Bluetooth data is calculated by using formula (5). Formula (5) is used for calculation, and the confidence of the RFID data is calculated by using formula (6). Formula (6) is used for calculation.
[0092] (5)
[0093] In formula (5), is the signal strength, is the expected signal strength, is the standard deviation of signal fluctuation, is the distance, representing the distance between the signal strength collector and the transmitter, is the expected distance, is the standard deviation of the distance. The meanings of other parameters in formula (5) are the same as those in formula (1), and details are omitted here.
[0094] (6)
[0095] In formula (6), is the signal strength, is the minimum value of the signal strength, is the maximum value of the signal strength. The meanings of other parameters in formula (6) are the same as those in formula (1), and are omitted here.
[0096] Optionally, the is calculated by using formula (7).
[0097] (7)
[0098] In formula (7), represents the total number of attempts of reading by the RFID tag, represents the number of successful reading by the RFID tag. The meanings of other parameters in formula (7) are the same as those in formula (1), and are omitted here.
[0099] Optionally, the is calculated by using formula (8).
[0100] (8)
[0101] In formula (8), represents the total observation time of the first pedestrian, represents the Wi-Fi connection time. The meanings of other parameters in formula (8) are the same as those in formula (1), and are omitted here.
[0102] Optionally, the is calculated by using formula (9).
[0103] (9)
[0104] In formula (9), is the signal strength of the Bluetooth data, is the maximum value of the signal strength of the Bluetooth data. The meanings of other parameters in formula (9) are the same as those in formula (1), and are omitted here.
[0105] Optionally, the above After the calculation, each needs to be normalized so that the sum of all is 1. The normalization processing includes: dividing each calculated by the sum of the various that is, by .
[0106] Step c, the adaptive weight of each joint in the second set of joints is used to weight the joints in the second set of joints, to obtain the first set of joints of the first pedestrian at the k moment.
[0107] The first set of joints includes n weighted joints.
[0108] For the i th joint in the first set of joints, step c can be represented by formula (10).
[0109] (10)
[0110] In formula (10), is the i th joint in the first set of joints, represents the i th joint corresponding to the p th image modality in the second set of joints. The meanings of other parameters in formula (10) are the same as in formula (1), and details are omitted here.
[0111] The joints in the first set of joints obtained by using the above weighting method may be different from the true joints. In order to make the joints in the first set of joints as close to the true joints as possible, the weight determined by the adaptive weighting model is optimized in the embodiment of the disclosure by using deep learning.
[0112] The target of deep learning is to make as small as possible. Wherein represents the true value of the i th joint at the k moment.
[0113] Before deep learning, a multi-modal data training set is first constructed, which includes data of each modality and the true value of the first pedestrian joint in each image modality .
[0114] In the deep learning process, the adaptive weight of the i th joint corresponding to the p th image modality in the optimized second set of joints is represented by formula (11).
[0115] (11)
[0116] In formula (11), represents the adaptive weight of the i th joint corresponding to the p th image modality in the optimized second set of joints, is the adjustment weight of the i th joint corresponding to the p th image modality, and the adjustment weight is obtained by using deep learning, is a smoothing factor, .
[0117] Here, the process of deep learning is also the process of learning to adjust the weights, and the optimal adjustment weight can be determined after the completion of deep learning. In the deep learning process, after the end of the previous round of training, all the generated should be retained and used for the next round of training. After multiple rounds of training, the model converges, and the model at the time of convergence is the final .
[0118] After the weights determined by the self-adaptive weighting model optimized in the manner of deep learning, the formula (10) in step c becomes the form of formula (12).
[0119] (12)
[0120] The meanings of the parameters in formula (12) are the same as those in formula (10) and formula (11), and detailed description is omitted here.
[0121] In step 203, based on the first node set of the first pedestrian at time k, the motion trajectory of the first pedestrian is determined by using HGCN and ST-GNN.
[0122] Optionally, step 203 includes the following steps d-f.
[0123] Step d, based on the first node set of the first pedestrian at time k, a first heterogeneous graph is constructed.
[0124] The first heterogeneous graph includes time edges, space edges and heterogeneous edges. Figure 3 is a schematic diagram of the first heterogeneous graph. Figure 3 Part (a) of is a schematic diagram of the time edge, Figure 3 Part (b) of is a schematic diagram of the space edge, Figure 3 Part (c) of is a schematic diagram of the heterogeneous edge.
[0125] The following will be described in conjunction with Figure 3 The time edge, the space edge and the heterogeneous edge are briefly described. The time edge is the edge obtained by connecting the same node in the adjacent two frames of images; the space edge refers to the edge connecting the nodes to the shape of the human body in a certain frame of image; the heterogeneous edge refers to the edge obtained by connecting the same node in the images at the same time under different modalities.
[0126] The first heterogeneous graph can be represented as , wherein is the first heterogeneous graph, is the first node set; is the set of edges, which includes time edges, space edges and heterogeneous edges; is the mapping of the node to the node type of a specific modality, is the mapping of the edge The type of the edge mapped to the specific relationship. In the heterogeneous graph, the node feature matrix X and the adjacency matrix A containing the edge weight are initialized, and in step e, the HGCN encodes each node feature based on the node feature matrix X.
[0127] Step e, the node features of the first heterogeneous graph are encoded by the HGCN to obtain a plurality of node features.
[0128] Optionally, for the node , the feature updating process is represented by formula (13).
[0129] (13)
[0130] In formula (13), is the feature of the updated node ; is a set of edge types, including time-related edges, space-related edges, etc. is a set of neighbor nodes connected to the node through the edge type ; is a normalization coefficient for balancing the influence of different types of edges. is a set of neighbor nodes related to the edge type ; is the feature of the neighbor node in the first layer; is an activation function, for example, ReLU (Rectified Linear Unit, linear rectifier function).
[0131] For different node types in the heterogeneous graph, type-specific feature conversion is performed using the type weight matrix , and type-level aggregation is performed, and the feature updating process is represented by formula (14).
[0132] (14)
[0133] In formula (14), is the type weight matrix, is the feature of the node in the first layer. The meanings of other parameters in formula (14) are the same as in formula (13), and are omitted here.
[0134] After the node features of the first heterogeneous graph are encoded by the HGCN, the plurality of node features output by the HGCN are unified embedding representations , wherein is the number of layers of the HGCN.
[0135] In step f, the plurality of node features are input into the ST-GNN to obtain a motion trajectory of the first pedestrian output by the ST-GNN.
[0136] The ST-GNN comprises an input layer, a graph convolutional layer, an activation function layer, a spatio-temporal graph neural network layer, and a decoding layer connected in sequence.
[0137] The input layer is configured to receive the plurality of node features from the HGCN The plurality of node features contain high-level semantic features of the nodes. Optionally, the input layer is also configured to receive an initialized adjacency matrix A, which is used to define the connection relationship between the nodes in the heterogeneous graph.
[0138] The graph convolutional layer is configured to update the node features by passing information through the adjacency matrix.
[0139] The activation function layer is configured to perform a nonlinear transformation on the node features after graph convolution. The activation function layer can use ReLU, Sigmoid, Tanh, or other activation functions.
[0140] The spatio-temporal graph neural network layer is a graph convolutional layer that combines spatial and temporal relationships, and is configured to capture the dynamic association between space and time. The features generated by the time attention are combined with the spatio-temporal graph structure to further update the node features.
[0141] The decoding layer is configured to map the corresponding features of the spatio-temporal GNN trajectory output to trajectory points.
[0142] The HGCN inputs the node features to the ST-GNN, and after the ST-GNN processes the node features, the trajectory points are output in the decoding layer, which are the motion trajectory of the first pedestrian.
[0143] In step 204, the motion process of the pedestrian is established as an MDP model.
[0144] The MDP (Markov Decision Process) model is used to predict the behavior pattern of the pedestrian. The basic elements of the MDP include state space, action space, transition probability, and reward function.
[0145] In the embodiments of the present disclosure, the state space is used to describe the position, speed, direction, and other information of each joint node of the pedestrian.
[0146] The position is the current position of each joint of the pedestrian, which is a three-dimensional coordinate (x, y, z). The velocity refers to the velocity of each joint of each pedestrian. The behavior refers to the current behavior mode of the pedestrian (such as walking, staying, running, etc.). The time information refers to the current time, that is, the time of the frame in which the current pedestrian monitoring video is located. In order to facilitate calculation, the time sequence information based on the trajectory can be discretized, and the continuous time, position, velocity and other information can be discretized into a limited state set.
[0147] The action space is used to describe the set of actions of the pedestrian, such as "forward", "backward", "stay", "turn", "accelerate or decelerate", etc. Forward refers to the pedestrian advancing a certain distance. Backward refers to the pedestrian retreating a certain distance. Stay refers to the pedestrian staying at the current position. Turning refers to the pedestrian changing the direction of travel. Acceleration or deceleration refers to the pedestrian changing the speed of movement.
[0148] The action space needs to be adjusted according to the target task. If the task is to detect abnormal behavior, "abnormal behavior" can be included as a special action, such as crowding, pushing and even trampling between pedestrians.
[0149] The transition probability is used to describe the possibility of the pedestrian changing from one action to another. The probability of the pedestrian changing from forward to turning, the probability of the pedestrian changing from staying to forward, etc. The transition probability can be obtained by statistical analysis of historical data when the MDP model is initially constructed, and the transition probability can be continuously updated during the reinforcement training process of the model afterwards.
[0150] The reward function is a mapping function used in the process of training the pedestrian behavior prediction model based on trajectory points. When designing the reward function, the correctness of the behavior mode, the spatial rationality, the discount factor , etc. need to be considered.
[0151] The correctness of the behavior mode refers to the matching degree between the predicted behavior according to the trajectory and the actual behavior. If the trajectory matches the predicted behavior mode, a positive reward is given. If it deviates from the expected trajectory or behavior, a negative reward is given.
[0152] The spatial rationality refers to giving a higher reward if the trajectory of the pedestrian is in line with the actual situation in space (for example, there is no abnormality such as wall penetration). For spatial rationality, the model needs to be input with corresponding knowledge. That is, the MDP model is driven by data and knowledge.
[0153] The discount factor is a parameter between 0 and 1 in the process of training the MDP model, which is used to measure the importance of the reward value returned by the above reward function. A higher discount factor indicates that the reward has a greater impact on the decision of the model.
[0154] After the MDP model is established, some historical pedestrian trajectory point data can be used to train the MDP model, so that the MDP model learns how to select the optimal action to obtain the maximum cumulative reward. After the MDP model is trained, the behavior mode of the pedestrian can be determined according to the input motion trajectory of the pedestrian. The behavior mode includes "walking", "staying" or "fast running", etc.
[0155] In step 205, the motion trajectory of the first pedestrian is input into the MDP model to obtain the behavior mode of the first pedestrian output by the MDP model.
[0156] By analyzing the behavior mode, the future motion trend of the pedestrian can be predicted, thereby playing a predictive supplement role for the tracking interruption caused by occlusion or perception loss.
[0157] In the case that the recognized behavior mode has abnormal behavior (such as wandering or rushing), an abnormal behavior alarm can also be performed.
[0158] Optionally, in the case that the first pedestrian exists in the image modality at the k th moment and the first pedestrian does not exist in the image modality at the k+1 th moment, the pedestrian is tracked by using the data of other modalities at the k+1 th moment and the behavior mode of the first pedestrian at the k th moment. The data of other modalities can determine the position of the first pedestrian, and the behavior mode can predict the action trend of the first pedestrian. In the case that the first pedestrian does not exist in the image modality, the behavior mode can play a role of auxiliary correction to the position of the first pedestrian determined by using the data of other modalities.
[0159] For example, the behavior mode of the first pedestrian predicted at the k th moment is a fast running mode, the image modality of the first pedestrian is lost at the k+1 th moment, and the data of other modalities at the k+1 th moment indicates that the first pedestrian is in a stop state. It is indicated that the data of other modalities may be wrong, and the first pedestrian can be tracked according to the fast running mode predicted at the k th moment.
[0160] The following is an apparatus embodiment of the present application. For details not described in detail in the apparatus embodiment, reference can be made to the above method embodiments.
[0161] Figure 4 A structure schematic diagram of a pedestrian tracking apparatus based on multi-modal data provided by one example embodiment of the present disclosure is shown. Referring to Figure 4 The pedestrian tracking apparatus 400 based on multi-modal data includes a first acquisition module 401, a second acquisition module 402, a trajectory determination module 403 and a behavior mode determination module 404.
[0162] The first acquisition module 401 is configured to acquire multi-modal data of the first region at the k-th moment, the multi-modal data comprising video data, RFID data, Wi-Fi data, Bluetooth data and infrared sensor data.
[0163] The second acquisition module 402 is configured to input the multi-modal data into an adaptive weighting model, and acquire a first set of key nodes of the first pedestrian at the k-th moment output by the adaptive weighting model, the adaptive weighting model being configured to output the first set of key nodes of the first pedestrian at the k-th moment after weighted fusion of the multi-modal data by using an adaptive algorithm.
[0164] The trajectory determination module 403 is configured to determine a motion trajectory of the first pedestrian based on the first set of key nodes of the first pedestrian at the k-th moment, by using a heterogeneous graph convolution network (HGCN) and a spatio-temporal graph neural network (ST-GNN) based on a time attention mechanism.
[0165] The behavior pattern determination module 404 is configured to determine a behavior pattern of the first pedestrian based on the motion trajectory of the first pedestrian, the behavior pattern being configured to track the first pedestrian.
[0166] Optionally, the first pedestrian comprises n key nodes, the multi-modal data comprises at least one image modality, the image modality comprising the video data and the infrared sensor data, and the second acquisition module 402 is further configured to output the first set of key nodes of the first pedestrian at the k-th moment after weighted fusion of the multi-modal data by using an adaptive algorithm in the following manner: acquiring a second set of key nodes of the first pedestrian at the k-th moment based on the multi-modal data of the first region at the k-th moment, the second set of key nodes comprising n key nodes of the first pedestrian corresponding to each image modality; determining an adaptive weight of each key node in the second set of key nodes based on the multi-modal data of the first region at the k-th moment; and weighting the key nodes in the second set of key nodes by using the adaptive weight of each key node in the second set of key nodes, to obtain the first set of key nodes of the first pedestrian at the k-th moment, the first set of key nodes comprising n weighted key nodes.
[0167] Optionally, the second acquisition module 402 is further configured to determine the adaptive weight of the i-th key node corresponding to the p-th image modality in the second set of key nodes by using the following formula:
[0168]
[0169] wherein, represents the adaptive weight of the i-th key node corresponding to the p-th image modality in the second set of key nodes, represents the confidence of the j-th modality at the k-th moment with respect to the i-th key node, represents the signal intensity of the jth modality for the ith joint at the kth moment, the multi-modal data includes N modal data, the N modal data includes P image modal data, i, j, p, n, N, P are positive integers, the value range of i is 1 to n, the value range of j is 1 to N, and the value range of p is 1 to P.
[0170] Optionally, the device further comprises an optimization module 405, configured to optimize the adaptive weight of the ith joint corresponding to the pth image modality in the second joint set determined by the adaptive weighting model in a deep learning manner; wherein the adaptive weight of the ith joint corresponding to the pth image modality in the optimized second joint set is represented by the following formula:
[0171]
[0172] wherein, represents the adaptive weight of the ith joint corresponding to the pth image modality in the optimized second joint set, is the adjustment weight of the ith joint corresponding to the pth image modality, and the adjustment weight is obtained in a deep learning manner, is a smoothing factor.
[0173] Optionally, the trajectory determination module 403 is further configured to construct a first heterogeneous graph based on the first joint set of the first pedestrian at the kth moment, the first heterogeneous graph comprising time edges, space edges and heterogeneous edges; encode the node features of the first heterogeneous graph using HGCN to obtain a plurality of node features; and input the plurality of node features into the ST-GNN to obtain the motion trajectory of the first pedestrian output by the ST-GNN.
[0174] Optionally, the behavior pattern determination module 404 is further configured to establish the motion process of the pedestrian as a Markov decision process MDP model, the MDP model being used to predict the behavior pattern of the pedestrian; and input the motion trajectory of the first pedestrian into the MDP model to obtain the behavior pattern of the first pedestrian output by the MDP model.
[0175] It should be noted that: when the pedestrian tracking device based on multi-modal data provided by the above embodiment performs pedestrian tracking, only the division of the above functional modules is exemplified, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the pedestrian tracking device based on multi-modal data provided by the above embodiment and the pedestrian tracking method based on multi-modal data belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0176] The division of the modules in the embodiments of the present disclosure is illustrative, and is merely a logical function division. In actual implementation, another division manner can be used. In addition, each function module in each embodiment of the present disclosure can be integrated in one processor, or can be physically separated, or two or more modules can be integrated in one module. The integrated module can be implemented in the form of hardware or in the form of a software function module.
[0177] When the integrated module is implemented in the form of a software function module and sold or used as an independent product, the integrated module can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present disclosure, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing an end device (which can be a personal computer, a mobile phone, or a communication device, etc.) or a processor to execute all or part of the steps of the method of each embodiment of the present disclosure. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0178] Figure 5 is a structural schematic diagram of a computer device provided by the embodiments of the present disclosure. As shown in Figure 5 the computer device 500 includes a processor 501 and a memory 502.
[0179] The processor 501 can include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor 501 can be implemented in at least one of a hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), a PLA (Programmable Logic Array). The processor 501 can also include a main processor and a coprocessor, the main processor being a processor for processing data in an awake state, also referred to as a CPU (Central Processing Unit), and the coprocessor being a low-power processor for processing data in a standby state. In some embodiments, the processor 501 can be integrated with a GPU (Graphics Processing Unit) that is responsible for rendering and drawing content required to be displayed by the display screen. In some embodiments, the processor 501 can further include an AI (Artificial Intelligence) processor for processing computing operations related to machine learning.
[0180] The memory 502 can include one or more computer-readable storage media that can be non-transitory. The memory 502 can also include a high-speed random access memory, and a nonvolatile memory such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 502 is used to store at least one instruction for being executed by the processor 501 to implement the pedestrian tracking method based on multi-modal data provided in the embodiments of the present disclosure.
[0181] Those skilled in the art can understand that the structure shown in the figure does not constitute a limitation on the computer device 500, and can include more or fewer components than those shown, or combine certain components, or adopt different component arrangements. Figure 5
[0182] The embodiments of the present disclosure also provide a non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of a computer device, the computer device is enabled to perform the pedestrian tracking method based on multi-modal data provided in the embodiments of the present disclosure.
[0183] The embodiments of the present disclosure also provide a computer program product, including computer programs / instructions, which, when executed by a processor, implement the pedestrian tracking method based on multi-modal data provided in the embodiments of the present disclosure.
[0184] The above merely provides the optional embodiments of the present disclosure, but does not intend to limit the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A pedestrian tracking method based on multi-modal data, characterized in that, The method comprises: acquiring multi-modal data of a first region at a k moment, the multi-modal data comprising video data, RFID data, Wi-Fi data, Bluetooth data and infrared sensor data; inputting the multi-modal data into an adaptive weighting model, and acquiring a first joint set of a first pedestrian at the k moment output by the adaptive weighting model, the adaptive weighting model being configured to output the first joint set of the first pedestrian at the k moment after weighted fusion of the multi-modal data by using an adaptive algorithm; determining a motion trajectory of the first pedestrian based on the first joint set of the first pedestrian at the k moment by using a heterogeneous graph convolution network (HGCN) and a spatio-temporal graph neural network (ST-GNN) based on a time attention mechanism; determining a behavior mode of the first pedestrian based on the motion trajectory of the first pedestrian, the behavior mode being used for tracking the first pedestrian; the first pedestrian comprises n joints, the multi-modal data comprises at least one image modality, the image modality comprising the video data and the infrared sensor data, and the adaptive weighting model is configured to output the first joint set of the first pedestrian at the k moment after weighted fusion of the multi-modal data by using an adaptive algorithm in the following manner: acquiring a second joint set of the first pedestrian at the k moment based on the multi-modal data of the first region at the k moment, the second joint set comprising n joints of the first pedestrian corresponding to each image modality; determining an adaptive weight of each joint in the second joint set based on the multi-modal data of the first region at the k moment; weighting the joints in the second joint set by using the adaptive weight of each joint in the second joint set to obtain the first joint set of the first pedestrian at the k moment, the first joint set comprising n weighted joints; the method further comprises: optimizing the adaptive weight of the i-th joint corresponding to the p-th image modality in the second joint set determined by the adaptive weighting model in a deep learning manner; wherein, represents an adaptive weight of the ith joint node corresponding to the pth image modality in the second joint node set, represents a confidence of the jth modality for the ith joint node at the kth time, represents a signal strength of the jth modality for the ith joint node at the kth time, the multi-modal data including data of N modalities, the data of N modalities including data of P image modalities, i, j, p, n, N, P being positive integers, the value range of i being 1 to n, the value range of j being 1 to N, the value range of p being 1 to P.
2. The method of claim 1, wherein, wherein the adaptive weight of the i-th joint corresponding to the p-th image modality in the second joint set after optimization is represented by the following formula: the method further comprises: optimizing the adaptive weight of the i-th joint corresponding to the p-th image modality in the second joint set determined by the adaptive weighting model in a deep learning manner; wherein, represents the adaptive weight of the ith node corresponding to the pth image modality in the optimized second node set, is the adjustment weight of the ith node corresponding to the pth image modality, and the adjustment weight is obtained in a deep learning manner, is a smoothing factor.
3. The method according to claim 1 or 2, characterized in that, wherein the adaptive weight of the i-th joint corresponding to the p-th image modality in the second joint set after optimization is represented by the following formula: the method further comprises: optimizing the adaptive weight of the i-th joint corresponding to the p-th image modality in the second joint set determined by the adaptive weighting model in a deep learning manner; wherein the adaptive weight of the i-th joint corresponding to the p-th image modality in the second joint set after optimization is represented by the following formula: the method further comprises: optimizing the adaptive weight of the i-th joint corresponding to the p-th image modality in the second joint set determined by the adaptive weighting model in a deep learning manner; wherein the adaptive weight of the i-th joint corresponding to the p-th image modality in the second joint set after optimization is represented by the following formula: input the plurality of node features into the ST-GNN to obtain a motion trajectory of the first pedestrian output by the ST-GNN.
4. The method according to claim 1 or 2, characterized in that, The method further includes: establishing a motion process of a pedestrian as a Markov decision process (MDP) model, the MDP model being used to predict a behavior pattern of the pedestrian; inputting the motion trajectory of the first pedestrian into the MDP model to obtain a behavior pattern of the first pedestrian output by the MDP model.
5. A pedestrian tracking apparatus based on multi-modal data, characterized by, The device includes: a first obtaining module configured to obtain multi-modal data of a first region at a k-th time point, the multi-modal data including video data, RFID data, Wi-Fi data, Bluetooth data, and infrared sensor data; a second obtaining module configured to input the multi-modal data into an adaptive weighting model, and obtain a first node set of a first pedestrian at the k-th time point output by the adaptive weighting model, the adaptive weighting model being used to output the first node set of the first pedestrian at the k-th time point after weighted fusion of the multi-modal data by using an adaptive algorithm; a trajectory determining module configured to determine a motion trajectory of the first pedestrian based on the first node set of the first pedestrian at the k-th time point by using a heterogeneous graph convolution network (HGCN) and a spatio-temporal graph neural network (ST-GNN) based on a time attention mechanism; a behavior pattern determining module configured to determine a behavior pattern of the first pedestrian based on the motion trajectory of the first pedestrian, the behavior pattern being used to track the first pedestrian; the first pedestrian includes n nodes, and the multi-modal data includes at least one image modality, the image modality including the video data and the infrared sensor data, and the second obtaining module is further configured to output the first node set of the first pedestrian at the k-th time point after weighted fusion of the multi-modal data by using an adaptive algorithm in the following manner: obtain a second node set of the first pedestrian at the k-th time point based on the multi-modal data of the first region at the k-th time point, the second node set including n nodes of the first pedestrian corresponding to each image modality; determine an adaptive weight of each node in the second node set based on the multi-modal data of the first region at the k-th time point; weight the nodes in the second node set by using the adaptive weights of the nodes in the second node set to obtain the first node set of the first pedestrian at the k-th time point, the first node set including n weighted nodes; the determination of the adaptive weight of each node in the second node set based on the multi-modal data of the first region at the k-th time point includes: the adaptive weight of an i-th node corresponding to a p-th image modality in the second node set is determined by using the following formula: wherein, represents an adaptive weight of the ith joint node corresponding to the pth image modality in the second joint node set, represents a confidence of the jth modality for the ith joint node at the kth time, represents a signal strength of the jth modality for the ith joint node at the kth time, the multi-modal data including data of N modalities, the data of N modalities including data of P image modalities, i, j, p, n, N, P being positive integers, the value range of i being 1 to n, the value range of j being 1 to N, the value range of p being 1 to P.
6. A computer device, comprising: The computer device includes a memory and a processor, the memory stores at least one computer program, the at least one computer program is loaded and executed by the processor to implement the method in any one of claims 1 to 4.
7. A computer readable storage medium characterized in that, The computer readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the method in any one of claims 1 to 4.
8. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instruction is executed by the processor to implement the method in any one of claims 1 to 4.
Citation Information
Patent Citations
Intelligent traffic monitoring method and device based on multi-modal data fusion and graph neural network, and electronic equipment
CN118366311A
Personnel attitude estimation method and system based on multi-model graph neural network fusion
CN118942153A