Kan-based full attention interaction and rnn conditional vae for pedestrian trajectory prediction

Through the method of KAN-based full-attention interaction and RNN conditional VAE, the efficiency and accuracy issues of pedestrian trajectory prediction in open environments are solved. By constructing data lists, designing encoders and optimizers, the social intention characteristics of pedestrians are captured, achieving efficient and accurate trajectory prediction.

CN120031917BActive Publication Date: 2025-10-21WUXI HUIHANG INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510190924.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-10-21
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

Existing technologies have difficulty in achieving efficient and accurate pedestrian trajectory prediction in open environments, especially when faced with pedestrians' high autonomy and flexibility, dynamic interactive behaviors, and high computational costs.

Method used

A method based on KAN full attention interaction and RNN conditional VAE is adopted. By constructing a data list containing the location information of all pedestrians in the scene, an encoder including a historical reasoning module and a future approximation module is designed. The RCVAE network is used for encoding and decoding, and a parallel prediction optimizer is combined for global optimization to capture the social intention characteristics between pedestrians.

Benefits of technology

It achieves more efficient and accurate pedestrian trajectory prediction, can more accurately capture the social intention characteristics between pedestrians in spatial scenes, and improves the success rate and accuracy of trajectory prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031917B_ABST
    Figure CN120031917B_ABST
Patent Text Reader

Abstract

The application discloses a pedestrian trajectory prediction method based on KAN full attention interaction and RNN conditional VAE, relates to the technical field of unmanned driving, and can improve the accuracy and success rate of pedestrian trajectory prediction.The application comprises the following steps: extracting data based on the tracking results of consecutive frames of pedestrians, and constructing a data list containing the position information of all pedestrians in a scene; selecting a target agent and neighbors to model the motion state through the data list, and constructing the input tensor of the network; designing an encoder containing a historical reasoning module and a future approximation module to encode the input tensor of the network; designing an RCVAE network to decode the processing results of the encoder, and obtaining the trajectory representation of the agent in the scene; and designing a parallel prediction optimizer to globally optimize the decoded trajectory, so that efficient and accurate pedestrian trajectory prediction is achieved.The application is suitable for pedestrian trajectory prediction in an open environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of unmanned driving technology, and in particular to an efficient pedestrian trajectory prediction method based on full attention interaction of KAN and RNN conditional VAE. Background Art

[0002] The core of pedestrian trajectory prediction technology lies in combining historical data with the current environment to determine a pedestrian's future movement trends. In most cases, predicting pedestrian trajectories is extremely difficult. Pedestrians are highly autonomous and flexible; they perceive their neighbors' potential behavior and adjust their paths accordingly. This dynamic interaction complicates trajectory prediction. When faced with emergencies, pedestrians make instinctive decisions, and this unpredictable behavior increases the difficulty of prediction. Individual differences in the same scenario place higher demands on the model's adaptability and generalization capabilities. Therefore, an efficient pedestrian trajectory prediction method based on the full-attention interaction of a KAN and a RNN conditional VAE is urgently needed to achieve efficient and accurate pedestrian trajectory prediction in open environments.

[0003] Early methods relied on fixed rules and parameter settings, making these models difficult to flexibly adjust and adapt to changes in the environment and pedestrian behavior. Existing LSTM-based methods introduce a social pooling layer to analyze interactions between pedestrians within a grid, but ignore the social information that may exist between grid boundaries. Although Transformer-based methods allow for long-term dependency modeling and large-scale parallel training, when faced with constantly changing pedestrians, these methods incur high computational costs and provide minimal benefits. VAE-based methods study pedestrian trajectory prediction from an uncertainty perspective. However, these methods do not directly intervene in the update of latent variables through the generated target distribution, which leads to the problem of pedestrian trajectory collapse. Therefore, how to achieve accurate and efficient pedestrian trajectory prediction has become a problem that needs to be studied and solved. Summary of the Invention

[0004] In response to the above technical problems, an embodiment of the present invention provides an efficient pedestrian trajectory prediction method based on KAN's full attention interaction and RNN conditional VAE, which can improve the success rate and accuracy of pedestrian trajectory prediction.

[0005] To achieve the above objectives, the embodiments of the present invention adopt the following technical solutions:

[0006] A pedestrian trajectory prediction method based on KAN full attention interaction and RNN conditional VAE, the prediction method includes:

[0007] S1. Extract data based on the tracking results of pedestrians in consecutive frames and build a data list containing the location information of all pedestrians in the scene;

[0008] S2. Selecting a target agent and neighbors through the data list to perform motion state modeling and construct an input tensor of the network;

[0009] S3. Designing an encoder including a history reasoning module and a future approximation module to encode the input tensor of the network;

[0010] S4. Design an RCVAE network to decode the processing results of the encoder to obtain the trajectory representation of the agent in the scene;

[0011] S5. Design a parallel prediction optimizer to perform global optimization processing on the decoded trajectory to achieve efficient and accurate pedestrian trajectory prediction.

[0012] Furthermore, the S1 includes:

[0013] For each image frame within a historical moment, the location information of all pedestrians in the image is extracted and a data list is constructed. The first column is the selected moment, the second column is the agent ID, and the third and fourth columns are the position coordinates of the agent in two-dimensional space at the current moment. The target agent is selected for investigation in turn, and the other agents are temporarily designated as global neighbors.

[0014] Furthermore, through the data list, the target agent and neighbors are selected to perform motion state modeling and construct the input tensor of the network, specifically including:

[0015] Model the motion state of the selected target agent and its neighbors in the global historical period;

[0016] Analyze the motion state of the target agent and construct the absolute motion state tensor A of the target agent i4 , the relative motion state tensor R related to the number of neighbors corresponding to the target agent i4 ;

[0017] Analyze the social status of the target agent and its neighbors, and construct a social status tensor S related to the number of neighbors corresponding to the target agent. i3 .

[0018] Furthermore, an encoder including a historical reasoning module and a future approximation module is designed to encode the input tensor of the network, specifically including:

[0019] The historical reasoning module adopts the agent-neighbor interaction attention mechanism to achieve efficient extraction of social intent;

[0020] The future approximation module is used to enhance the real motion state.

[0021] Furthermore, an RCVAE network is designed to decode the processing results of the corresponding encoder and obtain the trajectory representation of the agent in the scene, including:

[0022] The RCVAE network receives social intention information and augmented reality motion state information from the encoder to approximate the distribution of latent variables;

[0023] At each time step, the latent variables are fused with the generated trajectory to complete the update of the latent variables, and then the latent variables are decoded to obtain the trajectory representation of the target agent.

[0024] Furthermore, a parallel prediction optimizer is designed to perform global optimization processing on the decoding trajectory, which includes:

[0025] Process the decoded trajectory data as the initial input information of the network;

[0026] A global optimization process is performed on the initial input information.

[0027] Furthermore, the target agent and its neighbors are selected within the global historical period to perform motion state modeling, including:

[0028] According to the relationship between the target agent and the neighbor group, the target agent i and neighbor j are selected for motion state decomposition, and the perception information of the target agent is incorporated to prepare for constructing the input tensor of the network.

[0029] Furthermore, the target agent's motion state is analyzed to construct the target agent's absolute motion state tensor A. i4 , the relative motion state tensor R related to the number of neighbors corresponding to the target agent i4 , specifically including:

[0030] Define the position of the selected agent i in the scene at time t as The tangential velocity is expressed as The normal velocity is expressed as

[0031] In constructing the agent's absolute motion state tensor In the example, the feature dimension is set to Where N represents the number of agents processed at the current time t; about the relative motion state tensor The feature dimension is set to N n Represents the number of neighbors, which is determined by the number of pedestrians in the scene.

[0032] Furthermore, the social status of the target agent and its neighbors is analyzed, and a social status tensor S related to the number of neighbors corresponding to the target agent is constructed. i3 , which includes:

[0033] The distance between target agent i and neighbor j at time t It is the key factor to characterize the current state interaction between them, as shown in formula (1). In order to enhance the understanding of social behavior and better characterize the social motion state, their velocity angle is calculated. As shown in formula (2). Consider the time period from the current moment to the end of the historical period (T H -t) The time γ required for the target agent i and neighbor j to meet at the current speed, select a smaller time period as the time for them to continue traveling along the tangential direction while maintaining the tangential speed, and associate it with the distance at the current moment to obtain the final distance in the historical period. As shown in formula (3), (4). Finally, the motion perception state of neighbor j at time t is obtained Throughout the historical period H The inner neighbor state tensor is

[0034] Furthermore, the historical reasoning module involves an agent-neighbor interaction attention mechanism to achieve efficient extraction of social intent, including:

[0035] For the absolute motion state tensor A i4 , relative motion state tensor R i4 , social state tensor S i3 Processing is performed to project the feature dimensions of each tensor into a high-dimensional space, such as equations (5), (6), and (7):

[0036]

[0037] Where f S , f A , f R For embedded neural networks, d A ,d R ,d S , respectively represent A i , R i , S i Feature dimensions in high-dimensional space;

[0038] The hidden units at each time step Using embedded neural network f E Perform feature enhancement to obtain the query tensor As shown in formula (8);

[0039]

[0040] By querying the tensor Query the social status tensor at time t The information is aggregated to obtain the attention weight;

[0041] Apply a distance mask to the attention weight, set to be greater than r i The non-neighbors are set to 0 in the perception mask, otherwise they are regarded as neighbors and set to 1, resulting in a sparse matrix As shown in formula (9):

[0042]

[0043] r i is the observation radius;

[0044] Apply attention weights to relative motion states Get attention score As shown in formula (10):

[0045]

[0046] Will Acts as a query tensor on Obtain the attention weight of the distance mask, assign it to the first step of attention calculation, and substitute it into formula (10) to obtain the new attention score Thus, the attention score calculated in the second step is obtained As shown in formula (11):

[0047]

[0048] Attention score The number of neighbors N n The dimensions are averaged to complete the social intention extraction. It is integrated with the motion state information of the target agent as the RNN network f HISE Input, such as formula (12):

[0049]

[0050] When reaching the historical period T H At the last moment within, we get the hidden unit The hidden unit As the initial unit of the RCVAE hidden unit, through the embedded neural network f C Initialize it to get the initial forward state of the decoder As shown in formula (13):

[0051]

[0052] Furthermore, the future approximation module is used to enhance the real motion state, including:

[0053] For the motion state tensor A i and the relative motion state tensor R i The time dimension is flipped and then the embedded neural network f AR For the motion state tensor A i and the relative motion state tensor R i Perform feature enhancement, as shown in formula (14):

[0054]

[0055] Through the reverse recurrent network f FASE renew Forming a backward state As shown in formula (15):

[0056]

[0057] Furthermore, the RCVAE network receives social intent information and augmented reality motion state information from the encoder to approximate the distribution of latent variables, specifically including:

[0058] According to the chain rule in probability theory, the probability distribution of the target trajectory is modeled as formula (16):

[0059]

[0060] There are two ways to express the latent variables of target agent i at time step t in the future period: in training mode In test mode, It is fused with the trajectory information at the current time step, and the zg After feature enhancement, it is passed into the RNN network f de Hidden state Update as shown in formula (17):

[0061]

[0062] Introduce a noise variable ∈~N(0,1) sampled from the standard normal distribution to perform a reparameterization operation on the latent variable to repair the gradient information. In training mode, it is formula (18), and in test mode, it is formula (19):

[0063]

[0064] Where β and δ are the neural network parameters that need to be optimized during the training process.

[0065] Furthermore, at each time step, the latent variables are fused with the generated trajectory to complete the update of the latent variables, and then the latent variables are decoded to obtain the trajectory representation of the target agent, specifically including:

[0066] The latent variables and social intention information are integrated to decode the target trajectory, as shown in formula (20):

[0067]

[0068] Where α is the neural network parameter that needs to be optimized during the training process;

[0069] After introducing the latent variables, Equation (16) is rewritten as Equation (21) in training mode and Equation (22) in test mode:

[0070]

[0071] Get the target agent i in the future period T P Random prediction of inner spatial position.

[0072] Furthermore, the data processing of the decoded trajectory, as the initial input information of the network, specifically includes:

[0073] The trajectory information decoded by the decoder during the prediction period Compression in the same dimension Where 0 represents the initial input information of the optimizer, the depth of the optimizer is set to L layers, m represents the number of nodes in the l layer, n represents the number of nodes in the l+1 layer, Φ l is the activation function matrix corresponding to the l-layer nodes.

[0074] Furthermore, a global optimization process is performed on the initial input information, including:

[0075] The optimization process from the optimizer layer l to the layer l+1 is expressed as formula (23):

[0076]

[0077] Finally, the trajectory generated by the optimizer will be restored to the original dimensional space, and the final trajectory representation for the entire prediction period will be obtained at one time, as shown in formula (24):

[0078]

[0079] Beneficial effects:

[0080] The embodiment of the present invention provides an efficient pedestrian trajectory prediction method based on KAN full attention interaction and RNN conditional VAE. Through data preprocessing, encoding-decoding structure and global optimizer processing, it more accurately captures the social intention characteristics between pedestrians in spatial scenes and achieves more efficient and accurate trajectory prediction.

[0081] Based on the tracking results of pedestrians in consecutive frames, data is extracted to construct a data list containing the location information of all pedestrians in the scene; through the data list, the target agent and neighbors are selected for motion state modeling to construct the network input tensor; the input tensor of the network is encoded by an encoder including a historical reasoning module and a future approximation module; an RCVAE network is designed to decode the processing results of the encoder to obtain the trajectory representation of the agent in the scene; and a parallel prediction optimizer is designed to perform global optimization processing on the decoded trajectory to achieve efficient and accurate pedestrian trajectory prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0083] Figure 1 A schematic diagram of the overall framework flow provided for an embodiment of the present invention;

[0084] Figure 2 A schematic diagram of a data extraction module provided in an embodiment of the present invention;

[0085] Figure 3 A schematic diagram of trajectory analysis provided by an embodiment of the present invention;

[0086] Figure 4 A schematic diagram of a network structure provided by an embodiment of the present invention;

[0087] Figure 5 A schematic diagram of an optimizer provided in an embodiment of the present invention;

[0088] Figure 6 A schematic diagram of complex scene trajectory visualization provided by an embodiment of the present invention;

[0089] Figure 7 A schematic diagram of prediction results using an NBA dataset provided by an embodiment of the present invention;

[0090] Figure 8 A schematic diagram of multiple trajectory prediction results under the SDD dataset provided by an embodiment of the present invention;

[0091] Figure 9 The present invention provides a flow chart of the method. DETAILED DESCRIPTION

[0092] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. The embodiments of the present invention will be described in detail below, with examples of the embodiments illustrated in the accompanying drawings. Throughout, identical or similar reference numerals represent identical or similar elements or elements having identical or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and intended only to explain the present invention and are not to be construed as limiting the present invention. Those skilled in the art will appreciate that, unless otherwise stated, the singular forms "a," "an," "said," and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" as used in the description of the present invention refers to the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when an element is referred to as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or intervening elements may be present. Furthermore, "connected" or "coupled" as used herein may include wireless connections or couplings. The term "and / or" as used herein includes any and all combinations of one or more associated listed items. It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art in the art to which the present invention belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless defined as such herein.

[0093] The embodiment of the present invention provides a pedestrian trajectory prediction method based on full attention interaction of KAN and RNN conditional VAE, such as Figure 9 Shown, including:

[0094] S1. Extract data based on the tracking results of pedestrians in consecutive frames and build a data list containing the location information of all pedestrians in the scene;

[0095] S2. Selecting a target agent and neighbors through the data list to perform motion state modeling and construct an input tensor of the network;

[0096] S3. Designing an encoder including a history reasoning module and a future approximation module to encode the input tensor of the network;

[0097] S4. Design the RCVAE network to decode the processing results of the encoder and obtain the trajectory representation of the agent in the scene;

[0098] S5. Design a parallel prediction optimizer to perform global optimization processing on the decoded trajectory to achieve efficient and accurate pedestrian trajectory prediction.

[0099] In this embodiment, in S1, data is extracted based on the pedestrian tracking results of consecutive frames to construct a data list containing the position information of all pedestrians in the scene, including:

[0100] like Figure 2 As shown in the figure, for each historical moment in the image, the location information of all pedestrians is extracted and a data list is constructed. The first column is the selected moment, the second column is the agent ID, and the third and fourth columns are the position coordinates of the agent at the current moment in two-dimensional space. The target agent is selected for investigation in sequence, and the other agents are temporarily designated as global neighbors.

[0101] In this embodiment, in S2, the data list is used to focus on the target agent and its neighbors in the global historical period to perform motion state modeling; the motion state of the target agent is analyzed to construct the absolute motion state tensor A of the target agent. i3 , the relative motion state tensor R related to the number of neighbors corresponding to the target agent i4 Analyze the social status of the target agent and its neighbors, and construct a social status tensor S related to the number of neighbors corresponding to the target agent i3 .

[0102] Specifically, the relationship between the target agent and its neighbor group is modeled. The target agent i and neighbor j are selected for motion state decomposition, and the target agent's perception information is incorporated to prepare for the construction of the network input tensor.

[0103] like Figure 3 , define the position of the selected agent i in the scene at time t as The tangential velocity is expressed as The normal velocity is expressed as During the process of moving, the agent tends to move in a relatively stable state of motion. Under this quasi-steady state, although acceleration or relative acceleration exists, their changes are very small or occur on a very long time scale. Therefore, when constructing the agent's absolute motion state tensor In the example, the feature dimension is set to Where N represents the number of agents processed at the current time t. The feature dimension is set to N n Represents the number of neighbors, which is determined by the number of pedestrians in the scene.

[0104] The distance between target agent i and neighbor j at time t It is the key factor to characterize the current state interaction between them, as shown in formula (1). In order to enhance the understanding of social behavior and better characterize the social motion state, their velocity angle is calculated. As shown in formula (2). Consider the time period from the current moment to the end of the historical period (T H -t) The time γ required for the target agent i and neighbor j to meet at the current speed, select a smaller time period as the time for them to continue traveling along the tangential direction while maintaining the tangential speed, and associate it with the distance at the current moment to obtain the final distance in the historical period. As shown in formula (3), (4). Finally, the motion perception state of neighbor j at time t is obtained Throughout the historical period H The inner neighbor state tensor is

[0105]

[0106] In this embodiment, in S3, an encoder including a historical reasoning module and a future approximation module is designed to encode the input tensor of the network. The historical reasoning module involves an agent-neighbor interaction attention mechanism to achieve efficient extraction of social intent; the future approximation module is used to enhance the real motion state.

[0107] In order to efficiently utilize limited computing resources, improve the model's expressiveness, and better capture more complex patterns and relationships in the data. Figure 4 , for the A i4 , R i4 , S i3 Process them and project their feature dimensions into a high-dimensional space, as shown in Equations (5), (6), and (7).

[0108]

[0109]

[0110] Where f S , f A , f R It is an embedded neural network. d A ,d R ,d S , respectively represent A i , R i , S i Feature dimensions in high-dimensional space.

[0111] In the historical reasoning module, we use RNN as the main architecture to model historical time series data. In order to enrich the input information, improve the expressiveness of features, and efficiently and accurately extract social intent for network learning. For each time step, the hidden unit Embedded neural network fE is used for feature enhancement to obtain query tensor It is used as the query tensor, as shown in formula (8). Query the social status tensor The information is aggregated to obtain the attention weight. This process achieves the n The linear computational complexity is O(N n ). When calculating the attention weights, we consider that the influence of distant neighbors on the selected agent is very weak based on social norms, and the improvement brought by including them in the calculation process is far less than the higher computational cost. Therefore, we apply a distance mask to the attention weights. H The distance between agent i and each neighbor j in the period With the observation radius r i We set it to be greater than r i The non-neighbors are set to 0 in the perception mask, otherwise they are regarded as neighbors and set to 1. Then we can get a sparse matrix As shown in formula (9), the attention weight is applied to the relative motion state. Get attention score As shown in formula (10).

[0112]

[0113]

[0114]

[0115] Next, Acts as a query tensor on Get the attention weight of the distance mask. Assign it to the attention weight obtained in the first step of attention calculation Thus, the attention score calculated in the second step is obtained As shown in formula (11). Similarly, this process achieves the n The linear computational complexity is O(N n ). Considering the errors caused by individual extreme values ​​in the neighborhood, we The number of neighbors N n The dimensions are averaged to complete the social intention extraction. It is integrated with the motion state information of the target agent as the RNN network f HISEInput, as shown in formula (12).

[0116]

[0117] When the last moment in the historical period is reached, the hidden unit is obtained It will serve as the initial unit of the RCVAE hidden unit. Through the embedded neural network f C Initialize it to get the initial forward state of the decoder As shown in formula (13), we embed a module that can set the number of trajectories K. This module will Create a new dimension and stack it K times to achieve diversified reasoning in the test mode.

[0118]

[0119] In order to achieve a better approximation effect on the distribution of latent variables in the decoder, the absolute motion state tensor and the relative motion state tensor of the entire future prediction period are fused to obtain the augmented reality motion state tensor B i .like Figure 4 , in order to meet the requirements of the posterior distribution in the decoder, A i and R i The time dimension is flipped, and then f AR To A i and R i Perform feature enhancement as shown in formula (14).

[0120]

[0121] for Through the reverse recurrent network f FASE Update, forming a backward state As shown in formula (15).

[0122]

[0123] In this embodiment, in S4, the RCVAE network is designed to receive social intention information and augmented reality motion state information from the encoder to approximate the distribution of latent variables; at each time step, the latent variables are fused with the generated trajectory to complete the update of the latent variables, and then the latent variables are decoded to obtain the trajectory representation of the target agent.

[0124] Specifically, such as Figure 4 As shown, according to the chain rule in probability theory, the probability distribution of the target trajectory can be modeled as formula (16).

[0125]

[0126] There are two ways to express the potential variables of target agent i at time step t in the future period. In test mode, It is fused with the trajectory information at the current time step, and the zg After feature enhancement, it is passed into the RNN network f de Hidden state Update as shown in formula (17).

[0127]

[0128] The updated hidden state contains the information of the latent variable, which is then combined with the backward state Approximate the latent variable to achieve the update of the latent variable. To compensate for the damage to the gradient information caused by the latent variable during the sampling process, we introduce a noise variable ∈~N(0,1) sampled from the standard normal distribution to perform a reparameterization operation on the latent variable to repair the gradient information. In training mode, it is Equation (18), and in test mode, it is Equation (19).

[0129]

[0130] Where β and δ are the neural network parameters that need to be optimized during the training process.

[0131] In order to prevent the loss of social intention information, we fuse the latent variables and social intention information to decode the target trajectory, as shown in Equation (20):

[0132]

[0133] Where α is the neural network parameter that needs to be optimized during the training process.

[0134] According to the above content, after introducing the latent variables, Equation (16) can be rewritten as Equation (21) in the training mode and Equation (22) in the test mode.

[0135]

[0136] So far, we have obtained the target agent i in the future period T P Random prediction of inner spatial position.

[0137] In this embodiment, in S5, a parallel prediction optimizer is designed to perform global optimization processing on the decoding trajectory.

[0138] Specifically, such as Figure 5 As shown, for We compress its trajectory information into the same dimension to obtain Where 0 represents the initial input information of the optimizer. The depth of the optimizer is set to L layers. Set m to represent the number of nodes in layer l, n to represent the number of nodes in layer l+1, Φ l is the activation function matrix corresponding to the l-layer nodes.

[0139] The optimization process for layer l→l+1 is expressed as formula (23).

[0140]

[0141] Finally, the trajectory generated by the optimizer will be restored back to the original dimensional space, and the final trajectory representation for the entire prediction period will be obtained at one time, as shown in Equation (24).

[0142]

[0143] The method proposed in this invention is tested on the public datasets ETH / UCY, SDD, and NBA. Figure 6 As shown in Figure 3, the social information is richer. Neighbors are represented by gray tracks, and the agent's attention score to the neighbor is characterized by its transparency. Lower transparency means higher attention score, and higher transparency means lower attention score. Figure 7 The figure is a schematic diagram of trajectory prediction results under the NBA dataset. In order to better demonstrate the advantages of the present invention, a variety of historical trajectories are selected and their prediction results are displayed, and the prediction results are compared with related methods. According to the visualization results, it can be found that even if the historical trajectory is turning throughout the entire period, the present invention can achieve good prediction of the agent's movement direction in the future period. Figure 8 Figure 2 shows the visualization of multiple trajectory predictions for the SDD dataset. Not only does the present invention accurately predict the results, but the remaining 19 predictions also reflect the diversity of agent motion. Even with the agent's historical trajectory in row 1 and column 2, where speed and direction are constantly changing, the present invention still accurately predicts and presents a variety of possible outcomes.

Claims

1. A pedestrian trajectory prediction method based on KAN full attention interaction and RNN conditional VAE, characterized by: The prediction method comprises: S1. Extract data based on the tracking results of pedestrians in consecutive frames and build a data list containing the location information of all pedestrians in the scene; S2. Using the data list, select the target agent and neighbors to perform motion state modeling and construct the network input tensor; specifically, the following steps are performed: The target agent and its neighbors are selected in the global historical period to model their motion state. Based on the relationship between the target agent and its neighbor group, the target agent i and its neighbor j are selected for motion state decomposition, and the target agent's perception information is incorporated to prepare for the construction of the network's input tensor. Analyze the motion state of the target agent and construct the absolute motion state tensor A of the target agent i4 , the relative motion state tensor R related to the number of neighbors corresponding to the target agent i4 ; Define the position of the selected agent i in the scene at time t as The tangential velocity is expressed as The normal velocity is expressed as In constructing the agent's absolute motion state tensor In the example, the feature dimension is set to Where N represents the number of agents processed at the current time t; about the relative motion state tensor The feature dimension is set to N n represents the number of neighbors, which is determined by the number of pedestrians in the scene; Analyze the social status of the target agent and its neighbors, and construct a social status tensor S related to the number of neighbors corresponding to the target agent. i3 ; The distance between target agent i and neighbor j at time t It is the key factor to characterize the current state interaction between them, as shown in formula (1); in order to enhance the understanding of social behavior and better characterize the social motion state, their velocity angle is calculated As shown in formula (2); Examine the time period from the present moment to the end of the historical period (T H -t) The time γ required for the target agent i and neighbor j to meet at the current speed, select a smaller time period as the time for them to continue traveling along the tangential direction while maintaining the tangential speed, and associate it with the distance at the current moment to obtain the final distance in the historical period. As shown in formula (3), (4); finally, the motion perception state of neighbor j at time t is obtained Throughout the historical period H The inner neighbor state tensor is S3. Designing an encoder including a history reasoning module and a future approximation module to encode the input tensor of the network; S4. Design an RCVAE network to decode the processing results of the encoder to obtain the trajectory representation of the agent in the scene; S5. Design a parallel prediction optimizer to perform global optimization processing on the decoded trajectory to achieve efficient and accurate pedestrian trajectory prediction.

2. The method according to claim 1, characterized in that Said S1 comprises: For each image frame within a historical moment, the location information of all pedestrians in the image is extracted and a data list is constructed. The first column is the selected moment, the second column is the agent ID, and the third and fourth columns are the position coordinates of the agent in two-dimensional space at the current moment. The target agent is selected for investigation in turn, and the other agents are temporarily designated as global neighbors.

3. The method according to claim 1, characterized in that For the input tensor of the network, an encoder including a history reasoning module and a future approximation module is designed to encode the network, specifically including: The historical reasoning module adopts the agent-neighbor interaction attention mechanism to achieve efficient extraction of social intent; For the absolute motion state tensor A i4 , relative motion state tensor R i4 , social state tensor S i3 Processing is performed to project the feature dimensions of each tensor into a high-dimensional space, such as equations (5), (6), and (7): Where f S , f A , f R For embedded neural networks, d A ,d R ,d S , respectively represent A i , R i , S i Feature dimensions in high-dimensional space; The hidden units at each time step Using embedded neural network f E Perform feature enhancement to obtain the query tensor As shown in formula (8); By querying the tensor Query the social status tensor at time t The information is aggregated to obtain the attention weight; Apply a distance mask to the attention weight, set to be greater than r i The non-neighbors are set to 0 in the perception mask, otherwise they are regarded as neighbors and set to 1, resulting in a sparse matrix As shown in formula (9): r i is the observation radius; Apply attention weights to relative motion states Get attention score As shown in formula (10): Will Acts as a query tensor on Obtain the attention weight of the distance mask, assign it to the first step of attention calculation, and substitute it into formula (10) to obtain the new attention score Thus, the attention score calculated in the second step is obtained As shown in formula (11): Attention score The number of neighbors N n The dimensions are averaged to complete the social intention extraction. It is integrated with the motion state information of the target agent as the RNN network f HISE Input, such as formula (12): When reaching the historical period T H At the last moment within, we get the hidden unit The hidden unit As the initial unit of the RCVAE hidden unit, through the embedded neural network f C Initialize it to get the initial forward state of the decoder As shown in formula (13): The future approximation module is used to enhance the real motion state and to calculate the motion state tensor A. i and the relative motion state tensor R i The time dimension is flipped and then the embedded neural network f AR For the motion state tensor A i and the relative motion state tensor R i Perform feature enhancement, as shown in formula (14): Through the reverse recurrent network f FASE renew Forming a backward state As shown in formula (15):

4. The method according to claim 1, wherein Design an RCVAE network to decode the processing results of the corresponding encoder and obtain the trajectory representation of the agent in the scene, including: The RCVAE network receives social intention information and augmented reality motion state information from the encoder to approximate the distribution of latent variables; According to the chain rule in probability theory, the probability distribution of the target trajectory is modeled as formula (16): There are two ways to express the latent variables of target agent i at time step t in the future period: in training mode In test mode, It is fused with the trajectory information at the current time step, and the zg After feature enhancement, it is passed into the RNN network f de Hidden state Update as shown in formula (17): Introduce a noise variable ∈~N(0,1) sampled from the standard normal distribution to perform a reparameterization operation on the latent variable to repair the gradient information. In training mode, it is formula (18), and in test mode, it is formula (19): In the formula, β and δ are the neural network parameters that need to be optimized during the training process; At each time step, the latent variables are fused with the generated trajectory to complete the update of the latent variables, and then the latent variables are decoded to obtain the trajectory representation of the target agent; the latent variables and social intention information are fused to decode the target trajectory, as shown in Equation (20): Where α is the neural network parameter that needs to be optimized during the training process; After introducing the latent variables, Equation (16) is rewritten as Equation (21) in training mode and Equation (22) in test mode: Get the target agent i in the future period T P Random prediction of inner spatial position.

5. The method according to claim 1, wherein For the decoding trajectory, a parallel prediction optimizer is designed to perform global optimization processing, which includes: The data processing of the decoded trajectory is used as the initial input information of the network; the trajectory information decoded by the decoder during the prediction period is used as the initial input information of the network; Compression in the same dimension Where 0 represents the initial input information of the optimizer, the depth of the optimizer is set to L layers, m represents the number of nodes in the l layer, n represents the number of nodes in the l+1 layer, Φ l is the activation function matrix corresponding to the l-layer nodes; A global optimization process is performed on the initial input information; the optimization process of the optimizer layer l to layer l+1→layer l+1 is expressed as formula (23): Finally, the trajectory generated by the optimizer will be restored to the original dimensional space, and the final trajectory representation for the entire prediction period will be obtained at one time, as shown in formula (24):

Citation Information

Patent Citations

  • Pedestrian trajectory prediction method based on global dynamic scene information depth modeling

    CN113538506A

  • Urban scene-oriented pedestrian trajectory prediction method, model and storage medium

    CN115071762A