A pedestrian trajectory prediction method based on social intention interaction

By adopting a social intention interaction method in pedestrian trajectory prediction, using video data and state tensors for intention interaction and trajectory optimization, the problem of accuracy of pedestrian trajectory prediction in complex scenarios in the prior art is solved, and higher prediction accuracy and reliability are achieved.

CN119559705BActive Publication Date: 2025-06-17NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510117046.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-17
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

Existing pedestrian trajectory prediction methods are difficult to accurately capture the social norms and intrinsic connections between pedestrians in complex scenarios, resulting in the increase in errors as the prediction period increases.

Method used

The pedestrian trajectory prediction method based on social intent interaction is adopted. By extracting pedestrian position information from the video data, constructing a data list, performing kinematic analysis, obtaining the state tensors of subjects and neighbor groups, performing intention interaction and decoding, optimizing trajectory representation, and finally DBSCAN clustering processing is performed.

Benefits of technology

It improves the accuracy of pedestrian trajectory prediction, and can more accurately capture social norms between pedestrians in spatial scenes, reduce errors, and achieve more reliable trajectory prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559705B_ABST
    Figure CN119559705B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a pedestrian trajectory prediction method based on social intention interaction, which relates to the technical field of intelligent manufacturing and can improve the accuracy and success rate of pedestrian trajectory prediction. The present invention includes: extracting pedestrian position information in a continuous frame scene to create a data list; designing a motion state analysis module through the data list to perform kinematic analysis on the pedestrian, and respectively constructing state tensors of the main body and the neighbor group; building a network to encode the state tensors and extract intention information to achieve intention interaction; decoding the interaction result to obtain the trajectory representation of the main body in the scene; and optimizing the trajectory representation by using the designed trajectory optimizer to achieve complete and accurate pedestrian trajectory prediction. The present invention is applicable to pedestrian trajectory prediction in a dense pedestrian scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the branch of prediction technology in general image data processing or generation technology, and is mainly used in the intelligent manufacturing scenario involving the new generation of information technology. In particular, it relates to a pedestrian trajectory prediction method based on social intention interaction. Background Art

[0002] Pedestrian trajectory prediction refers to predicting the possible movement trajectory of a pedestrian in a future period when given the historical data of the pedestrian's position within a certain period of time. This technology has wide applications in strong human-machine interaction fields such as autonomous driving and intelligent manufacturing. However, the complex situations in actual scenarios pose great challenges to the pedestrian trajectory prediction task. On the one hand, pedestrian movement has a high degree of autonomy and flexibility. Especially in a social environment, the walking route of a pedestrian is affected by various complex factors. On the other hand, the number of neighbors of the main pedestrian in the spatial scene is constantly changing. How to process them and analyze the interaction of social intentions between them and the main body is crucial for the final trajectory prediction result. Therefore, there is an urgent need for a pedestrian trajectory prediction method based on social intention interaction to achieve effective and accurate pedestrian trajectory prediction in the spatial field.

[0003] Early methods were difficult to capture complex environments and potential state information, and they have been replaced by data-driven deep learning methods. Existing methods have solved some challenges. The method based on LSTM analyzes the social interaction behavior between the main body and neighbors and treats the interaction information equally, unable to mine potential key information at future moments. Although the method based on the attention mechanism no longer treats the interaction information between the main body and neighbors equally, it is only limited to the external interaction between attention mechanisms and cannot capture the internal connection between the main body and neighbors. At the same time, it wastes more computing resources in dealing with pedestrians at a long distance. The method based on VAE uses latent variables generated over time to carry out trajectory prediction work, but the social analysis of neighborhood relationships is not deep enough. There is a common problem with existing methods that the error gradually increases as the prediction period grows. Therefore, how to improve the accuracy of pedestrian trajectory prediction in complex scenarios to achieve reliable pedestrian trajectory prediction has become a problem that needs to be studied and solved. Summary of the Invention

[0004] Embodiments of the present invention provide a pedestrian trajectory prediction method based on social intention interaction, which can improve the accuracy of pedestrian trajectory prediction.

[0005] To achieve the above object, the embodiments of the present invention adopt the following technical solutions:

[0006] A pedestrian trajectory prediction method based on social intention interaction, comprising:

[0007] S1. Extract the pedestrian position information in the scene from the consecutive frames of the captured video data, and create a data list according to the pedestrian position information;

[0008] S2. Use the data list to perform kinematic analysis on the pedestrians, and obtain the state tensors of the main body and the neighbor group;

[0009] S3. Use the state tensors to perform intention interaction and obtain the interaction result;

[0010] S4. Decode the interaction result and obtain the trajectory representation of the main body in the scene;

[0011] S5. After optimizing the trajectory representation, perform DBSCAN clustering on the finally obtained trajectory.

[0012] Among them, in the created data list, the first column is the selected observation time, the second column is the pedestrian id selected as the main body, and the third and fourth columns are the abscissa information and ordinate information of the pedestrian as the main body in the two-dimensional plane. For example: for each frame of the observed image, extract the current positions of all pedestrians in the image and construct a data list. The first column is the selected observation time, the second column is the pedestrian id selected as the main body, and the third and fourth columns are the positions of the main body at the current moment in the two-dimensional plane. A target can be selected in turn for investigation, and other pedestrians are temporarily regarded as global neighbors.

[0013] Furthermore, it is possible to focus on the motion states of all pedestrians under global observation and design a motion state analysis module; specifically, introduce the perception information of pedestrians to design the motion state module and model according to the relationship between the selected main body and the neighbor group. Decompose the motion states of the selected main body i and the neighbor j, and integrate the perception information to prepare for further constructing the state tensors of the main body and the neighbor group. S2 specifically includes: performing motion state analysis on the current main body to obtain the state tensor s of the main body; performing motion state analysis on the neighbor group of the current main body to obtain the state tensor o of the neighbor group. Among them, the performing motion state analysis on the current main body to obtain the state tensor S of the main body includes: establishing the state tensor of the main body within the entire observation period of the continuous frames T ob wherein the state tensor of the main body within the observation period of the continuous frames is , N represents the batch of main bodies to be processed, represents the real number field; the motion state of the main body i at time t is ([[]] x t i , y t i , v t i , θ ti ), x t i , y t i respectively represent the abscissa and ordinate of the position of entity i in the scene at time t, v t i , θ t i respectively represent the tangential velocity variable and the normal velocity variable of entity i in the scene at time t; , , the tangential velocity of entity i is expressed as v t ix , and the normal velocity of entity i is expressed as v t iy .

[0014] After constructing the state tensor of entity i, the motion state of each entity in the neighbor group can be decomposed in the same way as analyzing the state tensor of the entity, and the state tensor of each entity in the neighbor group can be obtained, so as to further describe the interaction information between the entity and the neighbor group. The motion state analysis of the neighbor group of the current entity to obtain the state tensor O of the neighbor group includes: establishing the entire observation period of the continuous frames T ob in which the state tensor of the neighbor group is , N n represents the number of neighbors, and the motion perception state of neighbor j at time t is expressed as , represents the distance between entity i and neighbor j at time t, θ t ij represents the included angle between the velocities of entity i and neighbor j, , , x t j , y t j respectively represent the abscissa and ordinate of the position of neighbor j in the scene at time t, v t j represents the tangential velocity variable of neighbor j in the scene at time t, → is the vector symbol, ob t ij represents the corresponding T ob final distance; where, , , v t ij represents the velocity of agent i relative to neighbor j at time t. v t jx represents the tangential velocity of neighbor j, and t represents the current time. ξ represents the time period from the previous time to the end of the observation period ( T ob - t ) and the time required for agent i and neighbor j to meet at the current velocity.

[0015] S3 of this embodiment includes: preprocessing the obtained state tensor, and then obtaining attention weights using the intention information of the neighbor group; determining the result of intention interaction using the attention weights.

[0016] The preprocessing of the obtained state tensor includes: performing linear transformation processing and position encoding on the state tensor of the neighbor group, and obtaining the preprocessing result for the neighbor group , where d model represents the feature dimension of O pe The vector projected onto Query is Q pe h o The vector projected onto Key is K h o The vector projected onto Value is V h o , where , h = 1, …, ε , d ε is the feature dimension after ε projections; performing linear transformation processing and position encoding on the state tensor of the agent, and obtaining the preprocessing result for the agent , where the Query vector Q is obtained after performing three linear transformations on Spe s , the Key vector K s and the Value vector V s , and .

[0017] The obtaining of attention weights using the intention information of the neighbor group includes: obtaining the attention weight A of each head in the neighbor group h o , , M o represents the perception mask. When is greater than​r i If not, it is determined as a non - neighbor. r i is the observation radius, and the input of the interaction result decoder is MultiHead(Q h o , K h o , V h o ). The input of the main body intention extraction module is .

[0018] Among them, the distance between the main body i and each global neighbor j during the entire T ob period is examined, and the relationship between the distance and the observation radius and r i is set. Those greater than r i are regarded as non - neighbors and set to 0 in the perception mask, while the others are regarded as neighbors and set to 1. Thus, a sparse perception mask can be obtained. As shown in Eqs. (8) and (9), the designed neighbor perception intention extraction module has two outputs, namely MultiHead(Q h o , K h o , V h o ) and . The former is used as the input of the interaction result decoder, and the latter is used as the input of the main body intention extraction module.

[0019] In the formula, the attention weight , the perception mask , the attention score , the projection matrix , and '.' is used to concatenate the results of the attention functions of h indexed ε heads together.

[0020] The method of using the attention weight to determine the result of intention interaction includes: averaging in the number of neighbors in the N n dimension, where , T ob represents the last moment of the observation period; then expanding the dimension of V S to obtain the expanded result S of V, and expanding the dimension of A S to obtain the expanded result S of Aof A , A S is the attention weight in the main intention extraction module, ; then, the influence of the neighbor group attention weight is adjusted by the scaling factor λ to obtain the final attention score , W S is the learned linear transformation matrix.

[0021] S4 in this embodiment includes: decoding the interaction result between the main body and the neighbor group to obtain the trajectory representation of the main body; decoding the attention score, and the decoding result is used as the input of the neighbor perception intention extraction module in the next layer of the network. Specifically, decoding the interaction result to obtain the trajectory representation of the main body in the scene, including: decoding the interaction result between the main body and the neighbor group to obtain the trajectory representation of the main body; decoding the attention score result and using it as the input of the neighbor perception intention extraction module in the next layer of the network. Decoding the interaction result between the main body and the neighbor group to obtain the trajectory representation of the main body; decoding the attention score result and using it as the input of the neighbor perception intention extraction module in the next layer of the network, including: decoding the interaction result and the attention score result, and respectively inputting them into the corresponding decoders for processing. To alleviate the problems of gradient disappearance and explosion and obtain smoother gradients, the first part of the encoder is designed as a residual connection. It consists of a dropout function, an add function, and layer normalization. After processing the input, two output results C s and C o are obtained, as shown in Equations (12) and (13).

[0022] To improve the embedding quality, C s and C o are sent to a multi-layer perceptron for processing, as shown in Equations (14) and (15). In the equations, C ’ o , C ’ s are the processing results of C o , C s after passing through the multi-layer perceptron, respectively. To slow down model degradation, a layer of residual connection is stacked after the multi-layer perceptron. The final processing results C ” o , C ” s are obtained, as shown in Equations (16) and (17). At the last layer of the decoder, the output C ” s on the s branch will be processed through a fully connected layer to obtain the trajectory representation , where l represents the number of predicted trajectories, T pred represents the number of periods of the predicted trajectories.

[0023] Optionally, after optimizing the trajectory representation using the designed trajectory optimizer, it includes: using a loss function for multi-dimensional information fusion to supervise the prediction results; performing DBSCAN clustering on the finally obtained multiple trajectories to automatically determine the number of clusters. This method is more flexible than manually setting the number of categories and is more in line with the pedestrian autonomous navigation strategy. The loss function for multi-dimensional information fusion supervises the prediction results, including: comparing the predicted position with the real position for each period. Calculating their Euclidean distance to obtain the distance prediction error for each period L dis , as shown in Equation (18). In the equation p t igt is the real position of the main body i at time t, T pred represents the number of periods to be predicted, p t i represents the two-dimensional spatial coordinates of the main body i at time step t.

[0024] For the processing of angles, introduce the last moment of the observation period T ob . Establish a connection between each prediction moment and the previous moment to construct a direction vector, and then examine the included angle with the real trajectory direction vector, as shown in Equation (19). The objective loss function of the model L As shown in Equation (20), by backpropagating the distance error and angle error between each prediction result and the real result to the network, the objective loss function is minimized L result.

[0025] The pedestrian trajectory prediction method based on social intention interaction provided by the embodiment of the present invention combines the trajectory prediction requirements of pedestrians in the space field with the functions in the field of computer deep learning, more accurately captures the social norms among pedestrians in the space scene, and realizes a more real and accurate trajectory prediction. The present invention includes: extracting the position information of pedestrians in the scene in continuous frames and creating a data list; through the data list, designing a motion state analysis module to perform kinematic analysis on pedestrians, and respectively constructing state tensors of the main body and the neighbor group; building a network to encode and extract intention information from the state tensors to achieve intention interaction; decoding the interaction results to obtain the trajectory representation of the main body in the scene; using the designed trajectory optimizer to optimize the trajectory representation, thereby improving the accuracy of pedestrian trajectory prediction.

[0026] Formulas (1) to (20) in this embodiment include:

[0027]

[0028] Description of the Drawings

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0030] Figure 1 Schematic diagram of the overall framework process provided by the embodiment of the present invention;

[0031] Figure 2 Schematic diagram of the data extraction module provided by the embodiment of the present invention;

[0032] Figure 3 Schematic diagram of the trajectory analysis provided by the embodiment of the present invention;

[0033] Figure 4 Schematic diagram of the network structure provided by the embodiment of the present invention;

[0034] Figure 5 Schematic diagram of the perceptual mask attention provided by the embodiment of the present invention;

[0035] Figure 6 Heat map of the social intention interaction attention weight provided by the embodiment of the present invention;

[0036] Figure 7 Schematic diagram of the loss function of the multi-dimensional information fusion provided by the embodiment of the present invention;

[0037] Figure 8 Schematic diagram of the simple scenario trajectory visualization provided by the embodiment of the present invention;

[0038] Figure 9 Schematic diagram of the complex scenario trajectory visualization provided by the embodiment of the present invention;

[0039] Figure 10 Schematic diagram of the attention map within the observation range provided by the embodiment of the present invention;

[0040] Figure 11 Schematic diagram of the method flow provided by the present invention. Detailed Embodiments

[0041] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and cannot be construed as a limitation of the present invention. Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used here may include wireless connection or coupling. The term "and / or" used here includes any unit and all combinations of one or more related listed items. Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used here have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless defined as here.

[0042] An embodiment of the present invention provides a pedestrian trajectory prediction method based on social intention interaction, as Figure 11 shown, including:

[0043] S1. Extract the pedestrian position information in the scene from the consecutive frames of the captured video data, and create a data list according to the pedestrian position information;

[0044] S2. Use the data list to perform kinematic analysis on the pedestrians, and obtain the state tensors of the main body and the neighbor group;

[0045] S3. Use the state tensors to perform intention interaction and obtain an interaction result;

[0046] S4. Decode the interaction result and obtain the trajectory representation of the main body in the scene;

[0047] S5. After optimizing the trajectory representation, perform DBSCAN clustering processing on the finally obtained trajectory.

[0048] In this embodiment, in S1, in the scenario of continuous frame extraction, the pedestrian position information is extracted to create a data list, which includes:

[0049] As Figure 2 shown, for each frame of the image at the observation moment, the current positions of all pedestrians in the image are extracted to construct a data list. The first column is the selected observation moment, the second column is the pedestrian id selected as the subject, and the third and fourth columns are the positions of the subject in the two-dimensional plane at the current moment. One target is selected for investigation in turn, and other pedestrians are temporarily defined as global neighbors.

[0050] In this embodiment, in S2, through the data list, a motion state analysis module is designed to perform kinematic analysis on pedestrians, and the state tensors of the subject and the neighbor group are constructed respectively, including: focusing on the motion states of all pedestrians under global observation, designing a motion state analysis module; performing motion state analysis on the current subject to construct the state tensor s of the subject; performing motion state analysis on the neighbor group of the current subject to construct the state tensor o of the neighbor group.

[0051] Specifically, the perception information of pedestrians is introduced to design the motion state module, and modeling is performed according to the relationship between the selected subject and the neighbor group. As Figure 3 shown, the selected subject i and neighbor j are decomposed in terms of motion state, and the perception information is incorporated to prepare for further constructing the state tensors of the subject and the neighbor group.

[0052] Define the position of a certain subject i in the scenario at time t as ( x t i , y t i ), the tangential velocity is expressed as v t ix , and the normal velocity is expressed as v t iy . The motion state of the subject is described based on the position and velocity vector of the subject, as shown in equations (1) and (2). Finally, the motion state of the currently selected subject i at time t is expressed as ( x t i , y t i , v t i , θ t i ), then the state tensor of the subject during the entire observation period T ob is , where NRepresents the main batch to be processed.

[0053] After constructing the state tensor of agent i, the motion states of each agent in the neighbor group are decomposed in the same way. Thus, the interaction information between the agent and the neighbor group is further described. The distance between agent i and neighbor j at time t is a key factor to characterize the current intended interaction between them, as shown in Equation (3). To enhance the understanding of social behavior and better characterize the social motion state, their velocity angle θ t ij is introduced, as shown in Equation (4).

[0054] Integrating the perception of potential impacts at future times, throughout the entire time domain, assuming that pedestrians are not affected by sudden factors during the movement process, and examining the time period from the current time to the end of the observation period ( T ob - t ) and the time ξ required for agent i and neighbor j to meet at the current speed. The smaller time period is selected as the time for them to continue moving along the tangent direction with the tangential speed, which is associated with the distance at the current time, and the final distance under the observation period ob t ij is obtained, as shown in Equations (5) and (6). Finally, the motion perception state of neighbor j at time t is represented as , then the entire observation period T ob the state tensor of the neighbor group within is , where N n represents the number of neighbors, and the size of the tensor o is related to the value of the number of neighbors N n .

[0055] In this embodiment, in S3, a network is built to encode the state tensor and extract intention information to achieve intention interaction, including: designing a neighbor perception intention extraction module to extract the intention information of the neighbor group to obtain the attention score result and attention weight; designing a main body intention extraction module to receive the attention weight output from the neighbor perception intention extraction module to achieve intention interaction.

[0056] Specifically, as Figure 4 shown, the neighbor group state tensor o is obtained after linear transformation and position encoding and projected to the Query vector Q ε times respectively with different, learned linear projections h o , and the Key vector K ho and the Value vector V h o , where , h = 1, …, ε , d ε is the feature dimension after ε times of projection. According to formula (7), the attention weight A of each head in the neighbor group is obtained h o . Examine the distance T ob between the subject i and each global neighbor j during the entire and the observation radius r i size relationship. Set those greater than r i as non-neighbors and set them to 0 in the perception mask. Otherwise, they are regarded as neighbors and set to 1. Furthermore, a sparse perception mask can be obtained, as shown in Figure 5 . As shown in equations (8) and (9), the designed neighbor perception intention extraction module has two outputs, namely MultiHead(Q h o , K h o , V h o ) and . The former is used as the input to the interaction result decoder, and the latter is used as the input to the subject intention extraction module.

[0057] In the formula, the attention weight , the perception mask , the attention score , the projection matrix , and ‘.’ is used to concatenate the results of the attention functions of the h indexed ε heads together.

[0058] The subject state tensor s is obtained through linear transformation and position encoding to get . After three linear transformations, the Query vector Q S , the Key vector K s and the Value vector V S are obtained, where . To eliminate the influence of random errors and outliers in the neighbor group, average processing is performed on the number of neighbors in the attention weight Nn dimension to obtain . Considering the match with , for V SPerform dimensional expansion to obtain . Meanwhile, as shown in Equation (10), after obtaining the attention scores, A S is also subjected to dimensional expansion to obtain . A scaling factor λ (this parameter is ultimately obtained through training) is introduced to adjust the influence of the attention weights of the introduced neighbor group, and the final interaction result is obtained through a fully connected layer, as shown in Equation (11). As Figure 6 shown, the interaction result is visualized, and both the abscissa and ordinate represent historical moments. Among them, row (a) represents the heat map of the influence weights between the subject states at each historical moment, row (b) represents the heat map of the weights between the neighbor group states at each historical moment after being processed by the perception mask, and row (c) represents the heat map of the interaction influence weights after (b) is superimposed on (a).

[0059] In the formula is the attention weight in the subject intention extraction module, is the learned linear transformation matrix.

[0060] In this embodiment, in S4, the interaction result is decoded to obtain the trajectory representation of the subject in the scene, including: decoding the interaction result between the subject and the neighbor group to obtain the trajectory representation of the subject; decoding the attention score result as the input of the neighbor perception intention extraction module in the next layer of the network.

[0061] Specifically, as Figure 4 shown, the interaction result Attention(Q s , K s , V s ) and the attention score result MultiHead(Q h o , K h o , V h o ) are decoded and respectively input into the corresponding decoders for processing. To alleviate the problems of gradient disappearance and explosion and obtain smoother gradients, the first part of the encoder is designed as a residual connection. It consists of a dropout function, an add function, and layer normalization. After processing the input, C s and C o are obtained respectively, as shown in Equations (12) and (13).

[0062] To improve the embedding quality, C s and C o are sent to a multi-layer perceptron for processing, as shown in Equations (14) and (15). In the formula, C ’ o , C ’ sThey are C respectively o ,C s The processing results after passing through the multi-layer perceptron. To slow down model degradation, a residual connection layer is stacked after the multi-layer perceptron. The final processing result C is obtained ” o ,C ” s ,as shown in Eqs. (16) and (17). At the last layer of the decoder, the branch where s is located outputs C ” s After that, it will be processed through a fully connected layer to obtain the trajectory representation ,where l represents the number of predicted trajectories T pred represents the number of periods of the predicted trajectory

[0063] In this embodiment, in S5, the designed trajectory optimizer is used to optimize the trajectory representation to achieve complete and accurate pedestrian trajectory prediction, including designing a loss function for multi-dimensional information fusion to supervise the prediction results; performing DBSCAN clustering on the finally obtained multiple trajectories to automatically determine the number of classifications

[0064] Specifically, as Figure 7 shown, the predicted position of each period is compared with the real position. Calculate their Euclidean distance to obtain the distance prediction error of each period, as shown in Eq. (18). In the formula p t igt is the real position of the subject i at time t

[0065] For the processing of angles, the last moment of the observation period T ob is introduced. A direction vector is constructed by establishing a connection between each predicted moment and the previous moment, and then the included angle with the real trajectory direction vector is examined, as shown in Eq. (19). The objective loss function of the model is shown in Eq. (20). By backpropagating the distance error and angle error between each prediction result and the real result to the network, the objective loss function is minimized L As a result. Performing DBSCAN clustering on the finally obtained multiple trajectories to automatically determine the number of classifications is more flexible than manually setting the number of categories and more in line with the pedestrian autonomous navigation strategy

[0066] The method proposed in the present invention is tested on the public datasets ETH, UCY, and SDD. As Figure 8 shown, it shows the trajectory prediction results of the agent itself in a simple scenario and the neighbors increasing from 1 to 4 in turn. The method proposed in the present invention makes better predictions than the SOTA model SocialVAE. As Figure 9As shown, a scenario with richer social information and a larger number of pedestrians is presented. Neighbors are represented by gray trajectories, and the attention score is characterized by its transparency. The lower the transparency, the higher the attention score, and the higher the transparency, the lower the attention score. As Figure 10 shown, it is the attention map within the observation range, where HOTEL is in the first row, UNIV is in the second row, and ZARA is in the third row. The yellow circle is the observation area of the agent, the red dot is the position of the agent itself in the current scenario, the red line segment represents the historical trajectory of the agent in the current scenario, the blue line segment represents the true trajectory, and the orange line segment represents the predicted trajectory.

[0067] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments. The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A pedestrian trajectory prediction method based on social intention interaction, characterized in that: include: S1, extracting pedestrian position information in a scene from consecutive frames of captured video data, and creating a data list according to the pedestrian position information; S2. Analyze the kinematics of the pedestrians using the data list, and obtain state tensors of the subject and the neighbor group; S3. Use the state tensor to perform intention interaction and obtain an interaction result; S4, decoding the interaction result and obtaining a trajectory representation of the subject in the scene; S5, after optimizing the trajectory representation, performing DBSCAN clustering processing on the finally obtained trajectory; In the created data list, the first column is the selected observation time, the second column is the pedestrian ID selected as the subject, and the third and fourth columns are the horizontal and vertical coordinate information of the pedestrian as the subject in the two-dimensional plane; S2 includes: analyzing the motion state of the current subject to obtain the state tensor S of the subject; analyzing the motion state of the neighbor group of the current subject to obtain the state tensor O of the neighbor group; The motion state analysis of the current subject is performed to obtain the subject's state tensor S, including: obtaining the entire observation period of the continuous frame T ob The state tensor of the inner body , N represents the main batch that needs to be processed, represents the field of real numbers; Among them, the motion state of subject i at time t is ( x t i , y t i , v t i , θ t i ), x t i, y t i They represent the horizontal and vertical coordinates of the position of subject i in the scene at time t, v t i, θ t i They represent the tangential velocity variable and normal velocity variable of subject i in the scene at time t respectively; , , the tangential velocity of body i is expressed as v t ix , the normal velocity of body i is expressed as v t iy .

2. The method according to claim 1, characterized in that The motion state analysis of the neighbor group of the current subject is performed to obtain the state tensor O of the neighbor group, including: The entire observation period of the consecutive frames is established T ob The state tensor of the inner neighbor group is , N n represents the number of neighbors, and the motion perception state of neighbor j at time t is expressed as , represents the distance between subject i and neighbor j at time t, θ t ij represents the velocity angle between subject i and neighbor j, , , x t j , y t j They represent the horizontal and vertical coordinates of the position of neighbor j in the scene at time t, v t j represents the tangential velocity variable of neighbor j in the scene at time t, → is a vector symbol, ob t ij Indicates the corresponding T ob The final distance in, , , v t ij represents the velocity of subject i relative to neighbor j at time t, v t jx represents the tangential velocity of neighbor j, t represents the current moment, ξ Represents the time period from the previous moment to the end of the observation period ( T ob -t ) is the time required for the subject i to meet its neighbor j at the current speed.

3. The method according to claim 1, characterized in that S3 include: The obtained state tensor is preprocessed, and then the attention weight is obtained using the intention information of the neighbor group; The attention weights are used to determine the outcome of the intended interaction.

4. The method according to claim 3, characterized in that The preprocessing of the obtained state tensor includes: performing linear transformation processing and position encoding on the state tensor of the neighbor group, and obtaining a preprocessing result for the neighbor group. ,in, d model Indicates O pe The characteristic dimension, O pe The vector projected to Query is Q h o , the vector projected to Key is K h o , the vector projected to Value is V h o , ,h=1,…, ε , d ε For passing ε Feature dimension after projection; Perform linear transformation and position encoding on the subject's state tensor, and obtain the preprocessing results for the subject , where S pe After three linear transformations, we get the Query vector Q s , Key vector K s and Value vector V s ,and .

5. The method according to claim 4, characterized in that The method of using the intention information of the neighbor group to obtain the attention weight includes: obtaining the attention weight A of each head in the neighbor group h o , , M o represents the perceptual mask, when Greater than r i When it is determined to be a non-neighbor, r i is the observation radius, and the input of the interaction result decoder is , the input of the subject intention extraction module is .

6. The method according to claim 5, characterized in that The using the attention weight to determine the result of the intention interaction includes: The number of neighbors N n The average processing result is ,in , T ob represents the last moment of the observation period; Then V S Expand the dimension and get V S The expansion result of , and, for A S Expand the dimension to get A S The expansion result of , A S is the attention weight in the subject intention extraction module, ; Then the influence of the neighbor group attention weight is adjusted by the proportional factor λ to obtain the final attention score , W S is the learned linear transformation matrix.

7. The method according to claim 1, characterized in that S4 include: Decode the interaction results between the subject and the neighbor group to obtain the trajectory representation of the subject; The attention scores are decoded and used as input to the neighbor-aware intent extraction module in the next layer of the network.

Citation Information

Patent Citations

  • Pedestrian trajectory prediction method based on global dynamic scene information depth modeling

    CN113538506A