Intelligent body motion mode prediction method and device for open scene
By using multi-scale historical aggregation encoder and efficient autoregressive decoder in the agent motion mode prediction technology, the multimodal features are deeply optimized, which solves the problem of poor feature fusion effect in the existing technology and achieves more efficient agent motion mode prediction.
Patent Information
- Application Number
- CN202510250109.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The existing agent motion mode prediction technology has poor feature fusion effect when processing multimodal features and cannot fully capture the dynamic interaction between multimodal features.
Multimodal features are processed using a multi-scale historical aggregation encoder and an efficient autoregressive decoder to generate multi-scale features and initial prediction results, and the target agent motion mode prediction results are further generated through a binary Gaussian hybrid model.
Through deep optimization of feature fusion in the time dimension, the effect of multimodal feature fusion is improved, and the accuracy of the prediction of the agent's motion mode is significantly improved.
Smart Images

Figure CN120105339A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and device for predicting intelligent body motion patterns in open scenarios. Background Art
[0002] Motion pattern recognition technology for multiple intelligent entities such as crowds and motor vehicles is widely used in intelligent monitoring, robot navigation, and autonomous driving.
[0003] Motion pattern prediction for open scenarios refers to accurately identifying the motion behavior patterns of all intelligent agents in open scenarios such as urban roads and public places through real-time data collection and intelligent analysis, thereby predicting possible future motion trends, thereby providing dynamic warnings of collision risks for system decisions of mobile robots or self-driving cars, and making safer path planning.
[0004] Most existing intelligent body motion pattern prediction technologies use autoregressive generation methods, that is, using deep learning methods to gradually predict future motion patterns. However, when dealing with multimodal features, this method usually simply connects features of different modalities together. This simple feature connection method cannot fully capture the dynamic interaction relationship between multimodal features, resulting in poor feature fusion effects. Summary of the invention
[0005] The present invention provides an open scene-oriented intelligent body motion mode prediction method and device, which are used to solve the technical problem that the existing intelligent body motion mode prediction technology leads to poor feature fusion effect.
[0006] The first aspect of the present invention provides a method for predicting motion patterns of an intelligent agent in an open scene, comprising:
[0007] Acquire the agent's historical motion pattern sequence, the agent's future motion planning sequence, a surveillance video dataset, and a city road map, and perform map encoding on the city road map to generate a map feature vector;
[0008] Generate a plurality of two-dimensional spatial position coordinates according to the original annotation file of the surveillance video data set, and perform homogeneous coordinate transformation and normalization on the plurality of two-dimensional spatial position coordinates to generate a plurality of motion state information;
[0009] Generate a time dimension edge sequence according to the plurality of motion state information and the plurality of two-dimensional space position coordinates using a preset spatial edge rule;
[0010] The agent's historical motion pattern sequence, the agent's future motion planning sequence and the time dimension edge sequence are respectively embedded to generate a historical motion time step embedding corresponding to the agent's historical motion pattern sequence, a future motion time step embedding corresponding to the agent's future motion planning sequence and an edge time step embedding corresponding to the time dimension edge sequence;
[0011] Using a preset multi-scale history aggregation encoder to encode the historical motion time step embedding, the future motion time step embedding and the edge time step embedding respectively, to generate multi-scale historical motion features corresponding to the historical motion time step embedding, multi-scale future motion features corresponding to the future motion time step embedding and multi-scale edge features corresponding to the edge time step embedding;
[0012] Outputting an initial prediction result according to the multi-scale historical motion features, the multi-scale future motion features, the multi-scale edge features, the map feature vector and the historical motion pattern sequence of the intelligent agent through a plurality of preset efficient autoregressive decoders;
[0013] A preset binary Gaussian mixture model is used to generate a target agent motion mode prediction result based on the initial prediction result.
[0014] Optionally, the performing map encoding on the city road map to generate a map feature vector includes:
[0015] Dividing the city road map and outputting a plurality of map sub-blocks;
[0016] Transforming each of the map sub-blocks to generate a sub-block vector corresponding to each of the map sub-blocks;
[0017] Performing a linear transformation on each of the sub-block vectors, and outputting a linear embedding corresponding to each of the sub-block vectors;
[0018] Adding the preset position coding matrix to each of the linear embeddings respectively to generate an intermediate embedding corresponding to each of the linear embeddings;
[0019] Based on the multi-head self-attention mechanism, the attention weight of each map sub-block is calculated respectively to determine the attention weight corresponding to each map sub-block;
[0020] Adding each of the attention weights to the intermediate embedding corresponding to each of the attention weights to generate a target embedding corresponding to each of the intermediate embeddings;
[0021] Performing layer normalization on each of the target embeddings to determine a normalized embedding corresponding to each of the target embeddings;
[0022] A multi-layer perceptron is used to generate a map feature vector according to the multiple normalized embeddings.
[0023] Optionally, the adopting a preset spatial edge rule to generate a time dimension edge sequence according to the plurality of motion state information and the plurality of two-dimensional space position coordinates includes:
[0024] Smoothing the plurality of motion state information and the plurality of two-dimensional space position coordinates to generate a plurality of smoothed motion state information and a plurality of smoothed two-dimensional space position coordinates;
[0025] Constructing a node set based on the plurality of smooth motion state information and the plurality of smooth two-dimensional space position coordinates;
[0026] Using preset spatial edge rules to construct edge sets according to multiple nodes in the node set;
[0027] Based on the edge set, a time dimension edge sequence is determined.
[0028] Optionally, the embedding processing is performed on the agent historical motion pattern sequence, the agent future motion planning sequence and the time dimension edge sequence respectively to generate the historical motion time step embedding corresponding to the agent historical motion pattern sequence, the future motion time step embedding corresponding to the agent future motion planning sequence and the edge time step embedding corresponding to the time dimension edge sequence, including:
[0029] Perform one-dimensional convolution on the agent's historical motion pattern sequence, the agent's future motion planning sequence, and the time dimension edge sequence respectively, and output the historical motion convolution features corresponding to the agent's historical motion pattern sequence, the future motion convolution features corresponding to the agent's future motion planning sequence, and the edge convolution features corresponding to the time dimension edge sequence;
[0030] Nonlinearly mapping the historical motion convolution feature, the future motion convolution feature and the edge convolution feature respectively to generate a historical motion label embedding corresponding to the historical motion convolution feature, a future motion label embedding corresponding to the future motion convolution feature and an edge label embedding corresponding to the edge convolution feature;
[0031] Performing position embedding processing on the agent's historical motion pattern sequence, the agent's future motion planning sequence, and the time dimension edge sequence respectively, and outputting the historical motion position embedding corresponding to the agent's historical motion pattern sequence, the future motion position embedding corresponding to the agent's future motion planning sequence, and the edge position embedding corresponding to the time dimension edge sequence;
[0032] Performing a weighted summation of the historical motion mark embedding and the historical motion position embedding to determine a historical motion time step embedding;
[0033] Performing a weighted summation on the future motion tag embedding and the future motion position embedding to determine a future motion time step embedding;
[0034] A weighted sum is performed on the edge label embedding and the edge position embedding to determine an edge time step embedding.
[0035] Optionally, the preset multi-scale history aggregation encoder includes a self-attention submodule, a first multi-scale branch module, a second multi-scale branch module, and a third multi-scale branch module; the data processing process of the preset multi-scale history aggregation encoder is specifically as follows:
[0036] Slicing the input time step embedding input to the preset multi-scale history aggregation encoder to generate a first sub-feature, a second sub-feature, and a third sub-feature;
[0037] Embedding the input time step as an input to a self-attention submodule, and outputting a first self-attention feature;
[0038] Using a first multi-scale branch module to perform multi-scale feature extraction on the input time step embedding to generate a first multi-scale feature;
[0039] Performing multi-scale feature extraction on the first sub-feature by a second multi-scale branch module to generate a second multi-scale feature;
[0040] Inputting the second sub-feature into a third multi-scale branch module for multi-scale feature extraction to generate a third multi-scale feature;
[0041] Using the third sub-feature as the input of the self-attention sub-module, and outputting a second self-attention feature;
[0042] Respectively perform layer normalization on the first self-attention feature, the first multi-scale feature, the second multi-scale feature, the third multi-scale feature, and the second self-attention feature, and output a first self-attention normalized feature, a first multi-scale normalized feature, a second multi-scale normalized feature, a third multi-scale normalized feature, and a second self-attention normalized feature;
[0043] The first self-attention normalized feature, the first multi-scale normalized feature, the second multi-scale normalized feature, the third multi-scale normalized feature, and the second self-attention normalized feature are concatenated to generate an output multi-scale feature.
[0044] Optionally, embedding the input time step as an input of a self-attention submodule and outputting a first self-attention feature comprises:
[0045] Based on the multi-head self-attention mechanism, performing attention calculation on the input time step embedding to determine the multi-head self-attention corresponding to the input time step embedding;
[0046] Add the multi-head self-attention and the input time step embedding element by element, and output a multi-head self-attention time step feature;
[0047] Performing layer normalization on the multi-head self-attention time step features to generate multi-head self-attention normalized features;
[0048] Performing one-dimensional convolution on the multi-head self-attention normalized features, and outputting multi-head self-attention one-dimensional convolution features;
[0049] Performing nonlinear mapping on the multi-head self-attention one-dimensional convolutional features to generate multi-head self-attention nonlinear features;
[0050] Perform one-dimensional convolution on the multi-head self-attention nonlinear features and output the first self-attention feature.
[0051] Optionally, the third multi-scale branch module includes a self-attention submodule and a self-attention distillation downsampling submodule; the step of inputting the second sub-feature into the third multi-scale branch module for multi-scale feature extraction to generate a third multi-scale feature includes:
[0052] Using the second sub-feature as the input of the self-attention sub-module, and outputting the self-attention feature corresponding to the second sub-feature;
[0053] Using the self-attention distillation downsampling submodule to distill and downsample the self-attention feature corresponding to the second sub-feature to generate a maximum pooling feature of the second sub-feature;
[0054] The second sub-feature maximum pooling feature is used as the input of the self-attention sub-module, and the third multi-scale feature is output.
[0055] Optionally, the adopting the self-attention distillation downsampling submodule to distill and downsample the self-attention feature corresponding to the second sub-feature to generate a maximum pooling feature of the second sub-feature includes:
[0056] Perform a one-dimensional convolution on the self-attention feature corresponding to the second sub-feature, and output a convolution feature of the second sub-feature;
[0057] performing batch normalization on the second sub-feature convolutional features to generate second sub-feature normalized features;
[0058] Downsampling the second sub-feature normalized feature, and outputting the second sub-feature downsampled feature;
[0059] Perform maximum pooling on the second sub-feature downsampling feature to generate a second sub-feature maximum pooling feature.
[0060] Optionally, the data processing process of the preset efficient autoregressive decoder is specifically as follows:
[0061] A multi-layer perceptron is used to perform feature mapping on the input historical motion pattern sequence input to the preset efficient autoregressive decoder to generate high-dimensional features of historical motion;
[0062] Splicing the historical motion high-dimensional features with the input edge features, the input map features, and the input future motion features input to the preset efficient autoregressive decoder to generate a spliced feature;
[0063] Performing a global adaptive one-dimensional average pooling operation on the concatenated features to generate a global feature representation;
[0064] Performing temporal excitation on the global feature representation and outputting temporal excitation features;
[0065] Multiplying the time excitation feature and the concatenation feature element by element to generate a time series feature representation;
[0066] A multi-layer perceptron is used to perform feature mapping on the time series feature representation, and output high-dimensional features of the time series;
[0067] Performing layer normalization and nonlinear mapping on the high-dimensional features of the time series, and outputting nonlinear features of the time series;
[0068] Adding the nonlinear characteristics of the time series and the historical moment prediction results element by element to generate the current moment prediction result;
[0069] splicing the current moment prediction result and the target subsequence in the input historical motion pattern sequence to generate a new input historical motion pattern sequence;
[0070] Performing time step embedding and position encoding on the new input historical motion pattern sequence, and outputting a coding feature sequence;
[0071] Based on a single-layer masked multi-head probabilistic sparse self-attention mechanism, an attention operation is performed on the encoded feature sequence to output a probabilistic sparse self-attention;
[0072] Add the probabilistic sparse self-attention and the encoding feature sequence element by element, and output a probabilistic sparse self-attention feature sequence;
[0073] Performing layer normalization on the probabilistic sparse self-attention feature sequence to determine a query matrix;
[0074] Performing linear projection on the input historical motion features input to the preset efficient autoregressive decoder to generate a key matrix and a value matrix;
[0075] generating an efficient decoding feature sequence according to the query matrix, the key matrix and the value matrix;
[0076] Performing layer normalization on the efficient decoding feature sequence, and outputting an efficient decoding normalized feature sequence;
[0077] A multi-layer perceptron is used to perform feature mapping on the efficient decoding normalized feature sequence to generate an output prediction result.
[0078] A second aspect of the present invention provides an open scene-oriented intelligent body motion pattern prediction device, comprising:
[0079] An acquisition module is used to acquire the historical motion pattern sequence of the intelligent agent, the future motion planning sequence of the intelligent agent, the monitoring video data set and the urban road map, and map encode the urban road map to generate a map feature vector;
[0080] According to the module, it is used to generate a plurality of two-dimensional space position coordinates according to the original annotation file of the monitoring video data set, and perform homogeneous coordinate transformation and normalization on the plurality of the two-dimensional space position coordinates to generate a plurality of motion state information;
[0081] An adopting module, used for generating a time dimension edge sequence according to a plurality of the motion state information and a plurality of the two-dimensional space position coordinates by adopting a preset space edge rule;
[0082] An embedding module is used to embed the agent's historical motion pattern sequence, the agent's future motion planning sequence and the time dimension edge sequence respectively, to generate a historical motion time step embedding corresponding to the agent's historical motion pattern sequence, a future motion time step embedding corresponding to the agent's future motion planning sequence and an edge time step embedding corresponding to the time dimension edge sequence;
[0083] an encoding module, configured to respectively encode the historical motion time step embedding, the future motion time step embedding and the edge time step embedding using a preset multi-scale history aggregation encoder, to generate multi-scale historical motion features corresponding to the historical motion time step embedding, multi-scale future motion features corresponding to the future motion time step embedding and multi-scale edge features corresponding to the edge time step embedding;
[0084] A decoding module, configured to output an initial prediction result according to the multi-scale historical motion features, the multi-scale future motion features, the multi-scale edge features, the map feature vector and the agent historical motion pattern sequence through a plurality of preset efficient autoregressive decoders;
[0085] The output result module is used to generate a target intelligent body motion mode prediction result based on the initial prediction result by using a preset binary Gaussian mixture model.
[0086] A third aspect of the present invention provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method for predicting the motion pattern of an intelligent body for open scenes as described in any one of the above items.
[0087] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed, implements the steps of the method for predicting motion patterns of an intelligent body for open scenes as described in any one of the above items.
[0088] A fifth aspect of the present invention provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer executes the steps of the method for predicting intelligent body motion patterns for open scenes as described in any one of the above items.
[0089] It can be seen from the above technical solutions that the present invention has the following advantages:
[0090] The above technical scheme of the present invention provides a method for predicting the motion pattern of an intelligent agent for open scenes. First, the intelligent agent's historical motion pattern sequence, the intelligent agent's future motion planning sequence, a monitoring video data set and a city road map are obtained, and the city road map is map-encoded to generate a map feature vector; then, multiple two-dimensional spatial position coordinates are generated according to the original annotation file of the monitoring video data set, and the multiple two-dimensional spatial position coordinates are homogeneously transformed and normalized to generate multiple motion state information; a preset spatial edge rule is used to generate a time dimension edge sequence according to the multiple motion state information and the multiple two-dimensional spatial position coordinates; the intelligent agent's historical motion pattern sequence, the intelligent agent's future motion planning sequence and the time dimension edge sequence are respectively embedded to generate the historical motion time step embedding corresponding to the intelligent agent's historical motion pattern sequence, the future motion time step embedding corresponding to the intelligent agent's future motion planning sequence and the edge time step embedding corresponding to the time dimension edge sequence; a preset multi-scale historical aggregation encoder is used to embed the historical motion time step embedding, the future motion time step embedding and the edge time step embedding respectively. The method comprises the following steps: encoding by embedding the historical motion time step to generate multi-scale historical motion features corresponding to the historical motion time step embedding, multi-scale future motion features corresponding to the future motion time step embedding, and multi-scale edge features corresponding to the edge time step embedding; outputting the initial prediction result according to the multi-scale historical motion features, multi-scale future motion features, multi-scale edge features, map feature vectors and the agent historical motion pattern sequence through multiple preset efficient autoregressive decoders; finally, generating the target agent motion pattern prediction result according to the initial prediction result by using the preset binary Gaussian mixture model; based on the above scheme, based on the preset multi-scale historical aggregation encoder and multiple preset efficient autoregressive decoders, processing the multimodal features, i.e., the acquired agent historical motion pattern sequence, agent future motion planning sequence, monitoring video data set and urban road map, outputting the initial prediction result, and then generating the target agent motion pattern prediction result according to the initial prediction result by combining the preset binary Gaussian mixture model, the present invention can deeply optimize the feature fusion in the time dimension, thereby improving the effect of multimodal feature fusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0092] Figure 1 A flowchart of a method for predicting motion patterns of an intelligent agent in an open scene provided in Embodiment 1 of the present invention;
[0093] Figure 2 A schematic diagram of the structure of a preset multi-scale history aggregation encoder provided in Embodiment 1 of the present invention;
[0094] Figure 3 The overall framework diagram of the intelligent body motion mode prediction method for open scenes provided in the first embodiment of the present invention;
[0095] Figure 4 A schematic diagram of the time compression and excitation (TSE module) provided in the first embodiment of the present invention;
[0096] Figure 5 A training flow chart of a preset binary Gaussian mixture model provided in the second embodiment of the present invention;
[0097] Figure 6 A schematic diagram of a flow chart of motion mode prediction using a preset binary Gaussian mixture model provided in the second embodiment of the present invention;
[0098] Figure 7 This is a structural block diagram of an intelligent body motion pattern prediction device for open scenes provided in Example 3 of the present invention. DETAILED DESCRIPTION
[0099] The embodiments of the present invention provide a method and device for predicting the motion pattern of an intelligent body in an open scene, which are used to solve the technical problem that the existing intelligent body motion pattern prediction technology leads to poor feature fusion effect.
[0100] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0101] See also Figure 1 , Figure 1 A flowchart of the steps of a method for predicting the motion pattern of an intelligent body in an open scene provided in Example 1 of the present invention.
[0102] The present invention provides an open scene-oriented intelligent body motion mode prediction method, comprising:
[0103] Step 101: Obtain the agent's historical motion pattern sequence, the agent's future motion planning sequence, a surveillance video dataset, and a city road map, and perform map encoding on the city road map to generate a map feature vector.
[0104] It should be noted that the present invention also supports multiple usage scenarios such as intelligent monitoring, autonomous driving, and mobile robots. Users can select one of the following two types of data sets for model training or model prediction according to their needs: a surveillance video data set from a rooftop bird's-eye view, and a large-scale data set that provides complete sensor suite data for autonomous vehicles (such as camera / lidar / millimeter-wave radar / inertial measurement unit / GPS data).
[0105] Furthermore, the input data sources for modeling and feature extraction required for model training are determined based on the factors that have a key impact on the motion pattern in the multi-agent scene of the dataset. When using a third-person perspective video dataset, the input sources for modeling and feature extraction required for model training include historical motion patterns and social interactions between agents in the video scene. When using a dataset of mobile agents with a complete sensor suite, the input sources for modeling and feature extraction in model training include: historical motion patterns, past interactions between the self and other agents; city road maps (such as grid maps); and future motion planning of the self.
[0106] Furthermore, the kinematic properties required for motion pattern modeling are calculated based on the original annotations of the dataset. Since the original annotation data of large-scale autonomous driving datasets is very large and the packaging is complex, the present invention is described in conjunction with a surveillance video type dataset as an example.
[0107] Furthermore, for the physical modeling of the motion pattern, the present invention defines the time step and the motion state at different moments in the motion for different types of agents. Agents, forming a set The current time is recorded as .
[0108] The historical time step, represented by a set of discrete time steps, is defined as: .in, Indicates a specific moment in the historical data, ranging from -H (the starting point of the observation window) to 0 (the most recent observation moment).
[0109] Future time steps are also represented by a set of discrete time steps, defined as: .in, It represents a prediction of a future moment, ranging from 1 (the first time step after the most recent observation moment) to T (the end point of the prediction window).
[0110] Historical state definition: For the first Agents ( ), which at the historical moment The actual state is defined as: .in, It is the dimension of the agent state, and its value varies according to the specific application scenario. For example, the state dimension of a pedestrian is 3 (position, speed, acceleration); the state dimension of a vehicle is 4 (position, speed, acceleration, heading angle).
[0111] Future State Definition: At a future moment , No. The predicted state of an agent only contains position information and is defined as: The self-agent at the future moment The actual state of the object includes position, velocity, acceleration, and heading angle, which are defined as: .
[0112] Based on the above foundation, the physical representation of different types of motion pattern data is generated: the historical motion pattern (the sequence of historical motion patterns of the intelligent agent) is ,express The state of an agent at a historical time step; the predicted future motion mode is , represents the predicted state of each agent in the future time step; the future motion plan of the self-agent (the sequence of future motion plans of the agent) is .
[0113] Furthermore, the process of performing map encoding on the urban road map and generating a map feature vector can be achieved by executing the following sub-steps S11 to S18:
[0114] Step S11, dividing the city road map and outputting a plurality of map sub-blocks;
[0115] Step S12: transform each map sub-block to generate a sub-block vector corresponding to each map sub-block;
[0116] Step S13, performing a linear transformation on each sub-block vector, and outputting a linear embedding corresponding to each sub-block vector;
[0117] Step S14, adding the preset position coding matrix to each linear embedding respectively to generate an intermediate embedding corresponding to each linear embedding;
[0118] Step S15: Based on the multi-head self-attention mechanism, the attention weight of each map sub-block is calculated to determine the attention weight corresponding to each map sub-block;
[0119] Step S16: adding each attention weight to the intermediate embedding corresponding to each attention weight to generate a target embedding corresponding to each intermediate embedding;
[0120] Step S17: perform layer normalization on each target embedding to determine a normalized embedding corresponding to each target embedding;
[0121] Step S18: using a multi-layer perceptron to generate a map feature vector based on multiple normalized embeddings.
[0122] It should be noted that, first, the input map image (city road map) is divided into sub-blocks of a fixed size of 25×25 pixels, and all the image sub-blocks are converted into corresponding sub-block vectors. Each sub-block vector will be linearly embedded to match the specified dimension, that is, multiple linear embeddings are obtained. The preset position encoding matrix is added to the linear embedding to obtain multiple intermediate embeddings, and then the multi-head self-attention mechanism is applied to calculate the attention weight of each patch (map sub-block) relative to other patches. The output of the self-attention corresponding to each map sub-block is added to its corresponding intermediate embedding, and then layer normalization is performed. Finally, a multi-layer perceptron (MLP) is used to process the embedding of each sub-block and map it to the specified output size, which is recorded as the map feature vector. .
[0123] Step 102: Generate multiple two-dimensional spatial position coordinates according to the original annotation file of the monitoring video data set, and perform homogeneous coordinate transformation and normalization on the multiple two-dimensional spatial position coordinates to generate multiple motion state information.
[0124] The motion state information includes velocity components and acceleration.
[0125] It should be noted that the two-dimensional spatial position coordinates of multiple agents in each frame are extracted through the original annotation file of the monitoring video dataset, which is recorded as and , representing the horizontal and vertical coordinates respectively.
[0126] Furthermore, if the coordinate units are inconsistent, unit conversion and calibration are required to ensure the uniformity of the model input. In some cases, the original coordinates may be pixel coordinates (e.g., image plane units) rather than world coordinates (e.g., actual geographic units). Since the present invention is based on the world coordinate system, homogeneous coordinate transformation is required at this time. The homogeneous coordinate transformation formula is: ,in, represents the homogeneous form of pixel coordinates, It is the transformation matrix from pixel coordinates to world coordinates.
[0127] Furthermore, the converted world coordinates need to be normalized, that is: ,in, is the scaling factor in homogeneous coordinates; Represents the normalized world coordinates.
[0128] Furthermore, the normalized world coordinates are used to calculate the motion state information. Assuming the frame rate is fixed at , the calculation formula of velocity component (including lateral velocity component and longitudinal velocity component) is:
[0129] , ;
[0130] in, and Indicates the horizontal and vertical coordinates of the next frame; and Indicates the horizontal and vertical coordinates of the current frame; Indicates the time interval between frames (fixed according to the data frame rate).
[0131] Furthermore, based on the known speed calculation formula, the acceleration (including lateral acceleration and longitudinal acceleration) is calculated as follows:
[0132] , ;
[0133] in, and Indicates the horizontal and vertical speeds of the next frame; and Indicates the horizontal and vertical speed of the current frame.
[0134] Step 103: Generate a time dimension edge sequence according to a plurality of motion state information and a plurality of two-dimensional space position coordinates using a preset spatial edge rule.
[0135] Specifically, step 103 may include the following sub-steps S31-S34:
[0136] Step S31, performing smoothing processing on the plurality of motion state information and the plurality of two-dimensional space position coordinates to generate a plurality of smoothed motion state information and a plurality of smoothed two-dimensional space position coordinates;
[0137] Step S32: constructing a node set based on the multiple smooth motion state information and the multiple smooth two-dimensional space position coordinates;
[0138] Step S33: construct an edge set according to multiple nodes in the node set using a preset spatial edge rule;
[0139] Step S34: Determine the edge sequence in the time dimension based on the edge set.
[0140] It should be noted that for the noise in the position or velocity data (such as manual annotation errors), smoothing processing is required, such as using a low-pass filter or moving average method. After the smoothing processing is completed, The scene is represented as a spatial graph For a node collection: , each node Representing an Agent exist The position and motion state information at each moment. According to the node type, each node is classified as a vehicle or a pedestrian. For the edge set: , represents a directed edge in the spatial graph. Each edge Representatives in Time Node For Node impact.
[0141] Furthermore, in order to capture the asymmetric interactions between agents and enhance the generalization ability of the spatial graph, the edge definition rules are as follows, that is, the preset spatial edge rules are as follows: The conditions for the existence of directed edges: when the node and The two-dimensional space position distance is less than When the perception scope of the semantic type belongs to The perception range is obtained through experience and adjusted according to the actual scenario. Edge weight: If there is a directed edge between nodes, then ;otherwise .
[0142] Furthermore, we construct a set of adjacent edges: for each node , the edge set between it and its neighbor nodes is defined as: Adjacent edges with the same semantic relationship (e.g., pedestrian→pedestrian, pedestrian→vehicle) are combined into a shared weight vector and processed uniformly.
[0143] Based on the above foundation, the edge sequence of the time dimension (time dimension edge sequence) is generated. Specifically, all processed spatial edges constitute a time sequence in the time dimension, which is defined as: Time Series All nodes representing historical states will be used as input and sent to the encoder for feature encoding.
[0144] Step 104, respectively embed the agent's historical motion pattern sequence, the agent's future motion planning sequence and the time dimension edge sequence to generate the historical motion time step embedding corresponding to the agent's historical motion pattern sequence, the future motion time step embedding corresponding to the agent's future motion planning sequence and the edge time step embedding corresponding to the time dimension edge sequence.
[0145] It should be noted that the HD map is used as a static context, and the time series of dynamic motion patterns and interactions in the scene are vectorized through time step embedding and used as the input of the encoder. A combination of label embedding and position embedding is used to embed the time step.
[0146] Specifically, step 104 may include the following sub-steps S41-S46:
[0147] Step S41, performing one-dimensional convolution on the agent's historical motion pattern sequence, the agent's future motion planning sequence, and the time dimension edge sequence respectively, and outputting the historical motion convolution features corresponding to the agent's historical motion pattern sequence, the future motion convolution features corresponding to the agent's future motion planning sequence, and the edge convolution features corresponding to the time dimension edge sequence;
[0148] Step S42, nonlinearly mapping the historical motion convolution features, the future motion convolution features and the edge convolution features respectively to generate historical motion label embeddings corresponding to the historical motion convolution features, future motion label embeddings corresponding to the future motion convolution features and edge label embeddings corresponding to the edge convolution features;
[0149] Step S43, respectively perform position embedding processing on the agent's historical motion pattern sequence, the agent's future motion planning sequence, and the time dimension edge sequence, and output the historical motion position embedding corresponding to the agent's historical motion pattern sequence, the future motion position embedding corresponding to the agent's future motion planning sequence, and the edge position embedding corresponding to the time dimension edge sequence;
[0150] Step S44, performing weighted summation on the historical motion mark embedding and the historical motion position embedding to determine the historical motion time step embedding;
[0151] Step S45, performing weighted summation on the future motion mark embedding and the future motion position embedding to determine the future motion time step embedding;
[0152] Step S46: Perform weighted summation on the edge label embedding and the edge position embedding to determine the edge time step embedding.
[0153] It should be noted that tag embedding is the process of mapping the time series of multiple input sources into a high-dimensional continuous vector space, which is conducive to capturing the semantic similarity between tags. A one-dimensional convolution filter (kernel width 3, step size 1) and a nonlinear activation function Leaky ReLU are used for tag embedding. Tag embedding is represented as a fixed dimension size of Specifically, the time series of multiple input sources are the agent's historical motion pattern sequence, the agent's future motion planning sequence, and the time dimension edge sequence. A one-dimensional convolution filter is used to perform one-dimensional convolution on these three inputs respectively to obtain the convolution features corresponding to the three inputs. Then, a nonlinear activation function Leaky ReLU is used to perform nonlinear mapping on the convolution features corresponding to the three inputs to obtain the label embedding corresponding to the agent's historical motion pattern sequence, the agent's future motion planning sequence, and the time dimension edge sequence.
[0154] Furthermore, in order to retain the timing information in the motion pattern, position embedding based on sine and cosine functions is introduced to generate time step embedding, that is, the position embedding of the agent's historical motion pattern sequence, the agent's future motion planning sequence, and the time dimension edge sequence is processed based on the position embedding of sine and cosine functions to obtain the position embedding corresponding to the agent's historical motion pattern sequence, the agent's future motion planning sequence, and the time dimension edge sequence. A position code is added to each marker to indicate the relative position of the marker in the motion pattern sequence, thereby helping the model to distinguish the motion states at different time points or different steps. The calculation method of position embedding is:
[0155] ;
[0156] ;
[0157] in, , refers to the d-th feature of the C-dimensional vector space, is the position embedding vector (position embedding). Finally, the result of the time step embedding It is represented as (i.e., the weighted sum of position embedding and tag embedding):
[0158] ;
[0159] in, is a scaling factor used to keep the magnitude of the token embedding and position embedding consistent. If the input sequence is normalized, it is recommended to set .
[0160] Step 105: Use a preset multi-scale history aggregation encoder to encode the historical motion time step embedding, the future motion time step embedding, and the edge time step embedding respectively to generate multi-scale historical motion features corresponding to the historical motion time step embedding, multi-scale future motion features corresponding to the future motion time step embedding, and multi-scale edge features corresponding to the edge time step embedding.
[0161] It should be noted that the present invention designs a multi-scale history aggregation encoder (MHAE), that is, a multi-scale history aggregation encoder is preset to perform feature encoding on time step embedding. The time series of several data sources that have completed time step embedding are input into the multi-scale history aggregation encoder for feature encoding. MHAE can generate three-dimensional feature tensors at multiple time scales, and finally output a representation that aggregates features at different time scales.
[0162] For further information, see Figure 2The preset multi-scale history aggregation encoder consists of two self-attention sub-modules, a first multi-scale branch module, a second multi-scale branch module, and a third multi-scale branch module. The size of the input data of the first self-attention sub-module is , the size of the input data of the first self-attention submodule is ; The first multi-scale branch module consists of four self-attention submodules and three self-attention distillation downsampling submodules, and its input data size is The second multi-scale branch module consists of three self-attention submodules and two self-attention distillation downsampling submodules, and its input data size is ; The third multi-scale branch module consists of two self-attention submodules and one self-attention distillation downsampling submodule, and its input data size is .
[0163] Optionally, the data processing process of presetting the multi-scale history aggregation encoder may include the following sub-steps S51-S58:
[0164] Step S51, slicing the input time step embedding input to the preset multi-scale history aggregation encoder to generate a first sub-feature, a second sub-feature and a third sub-feature;
[0165] Step S52: embed the input time step as the input of the self-attention submodule, and output the first self-attention feature;
[0166] Specifically, step S52 may include the following sub-steps S521-S526:
[0167] Step S521: Based on the multi-head self-attention mechanism, perform attention calculation on the input time step embedding to determine the multi-head self-attention corresponding to the input time step embedding;
[0168] Step S522, add the multi-head self-attention and input time step embedding element by element, and output the multi-head self-attention time step feature;
[0169] Step S523, performing layer normalization on the multi-head self-attention time step features to generate multi-head self-attention normalized features;
[0170] Step S524, performing one-dimensional convolution on the multi-head self-attention normalized features, and outputting multi-head self-attention one-dimensional convolution features;
[0171] Step S525, performing nonlinear mapping on the multi-head self-attention one-dimensional convolutional features to generate multi-head self-attention nonlinear features;
[0172] Step S526: perform one-dimensional convolution on the multi-head self-attention nonlinear features and output the first self-attention feature.
[0173] It should be noted that a self-attention submodule is constructed in the feature branch of each scale. The feature extraction of MHAE relies on the multi-head self-attention mechanism, where the self-attention formula is:
[0174] ;
[0175] The multi-head self-attention formula is:
[0176] ;
[0177] ;
[0178] Where Q, K, V are query, key and value matrices, , , and are the weight matrices corresponding to the i-th head respectively. yes and The dimension of is the scaling factor. The scaling factor is 5, and the number of attention heads Serves 8.
[0179] Furthermore, a self-attention block is constructed as the unit module for MHAE feature extraction, and its output is recorded as ,The data processing principle of the self-attention submodule can be expressed as:
[0180] ;
[0181] Among them, the symbol Represents the input data, is element-by-element addition, is the activation function. and Both use a one-dimensional convolution filter with a kernel size of 1, but The number of output channels is four times the number of input channels, and vice versa. The number of output channels is 1 / 4 of the number of input channels; LayerNorm is layer normalization.
[0182] It is worth mentioning that the data processing principle of the self-attention submodule in the first multi-scale branch module, the second multi-scale branch module, and the third multi-scale branch module is consistent with the data processing principle of the self-attention submodule described here, and the present invention will not go into details.
[0183] Step S53: using a first multi-scale branch module to perform multi-scale feature extraction on the input time step embedding to generate a first multi-scale feature;
[0184] Step S54: performing multi-scale feature extraction on the first sub-feature through a second multi-scale branch module to generate a second multi-scale feature;
[0185] Step S55: input the second sub-feature into the third multi-scale branch module to perform multi-scale feature extraction to generate a third multi-scale feature;
[0186] The third multi-scale branch module includes a self-attention sub-module and a self-attention distillation downsampling sub-module.
[0187] Specifically, step S55 may include the following sub-steps S551-S553:
[0188] Step S551: taking the second sub-feature as the input of the self-attention sub-module, and outputting the self-attention feature corresponding to the second sub-feature;
[0189] Step S552: Use the self-attention distillation downsampling submodule to distill and downsample the self-attention feature corresponding to the second sub-feature to generate a maximum pooling feature of the second sub-feature;
[0190] Specifically, step S552 may include the following sub-steps S5521-S5524:
[0191] Step S5521, performing one-dimensional convolution on the self-attention feature corresponding to the second sub-feature, and outputting the convolution feature of the second sub-feature;
[0192] Step S5522: performing batch normalization on the second sub-feature convolution feature to generate a second sub-feature normalized feature;
[0193] Step S5523, downsample the normalized feature of the second sub-feature, and output the downsampled feature of the second sub-feature;
[0194] Step S5524: perform maximum pooling on the second sub-feature downsampling feature to generate the second sub-feature maximum pooling feature.
[0195] Step S553: Use the second sub-feature maximum pooling feature as the input of the self-attention sub-module, and output the third multi-scale feature.
[0196] It should be noted that information distillation and time dimension downsampling are performed after each self-attention block. On the one hand, distillation can reduce the number of parameters and calculations in the model by reducing the feature dimension, thereby reducing the computational cost of the model. On the other hand, information distillation can provide a regularization mechanism. It improves the generalization ability of the model by introducing lower-dimensional feature representations, thereby preventing the model from overfitting the training data. Assume that the output feature of any layer of self-attention distillation (except the last layer) is , then from Layer to Layer, the forward function of the distillation operation, that is, the data processing principle of the self-attention distillation downsampling submodule, can be expressed as:
[0197] ;
[0198] in, It is the input of the self-attention distillation downsampling submodule; The output of the self-attention distillation downsampling submodule; is a one-dimensional convolution filter with a kernel size of 3 and the same number of input and output channels, The kernel size is 3 and the stride is 2, and BN is the batch normalization of the samples. (Maximum pooling) operation, each distillation layer can reduce the size of the time dimension of the feature tensor It is drastically pruned to half of its original size, thereby reducing the redundant information contained in the sequence; ELU is downsampling.
[0199] It is worth mentioning that the data processing principle of the self-attention distillation downsampling submodule in the first multi-scale branch module, the second multi-scale branch module, and the third multi-scale branch module is consistent with the data processing principle of the self-attention distillation downsampling submodule described here, and the present invention will not go into details.
[0200] Step S56: using the third sub-feature as the input of the self-attention sub-module, and outputting the second self-attention feature;
[0201] Step S57, respectively perform layer normalization on the first self-attention feature, the first multi-scale feature, the second multi-scale feature, the third multi-scale feature, and the second self-attention feature, and output the first self-attention normalized feature, the first multi-scale normalized feature, the second multi-scale normalized feature, the third multi-scale normalized feature, and the second self-attention normalized feature;
[0202] Step S58: concatenate the first self-attention normalized feature, the first multi-scale normalized feature, the second multi-scale normalized feature, the third multi-scale normalized feature, and the second self-attention normalized feature to generate an output multi-scale feature.
[0203] The input time step embedding is feature data of the input preset multi-scale history aggregation encoder. It can be understood that the input time step embedding can correspond to any feature data input to the preset multi-scale history aggregation encoder for data processing during model training or model application;
[0204] The data of the first sub-feature, the second sub-feature, and the third sub-feature are intermediate feature data generated in the preset multi-scale history aggregation encoder;
[0205] The output multi-scale feature is feature data output by a preset multi-scale history aggregation encoder. It can be understood that it can correspond to any feature data output after the preset multi-scale history aggregation encoder performs data processing during model training or model application.
[0206] It should be noted that the splicing and aggregation of the features of each scale branch are completed by stacking different numbers of self-attention modules and self-attention distillation layers. Different stacked layers have different numbers of self-attention distillation layers, indicating that the time dimension distillation multiples of different stacked layers in the multi-scale branch are different. Layers with lower distillation ratios are used for slices with smaller time dimensions, while layers with higher distillation ratios are used for slices with larger time dimensions. Specifically, the ratio of information distillation matches the ratio of the time dimension of the input slices, and the output dimensions of the multi-scale branches after the stacked layers are all Finally, the feature map of the multi-scale branch is concatenated with the feature map of the constraint branch in the time dimension to obtain the expression of the multi-scale historical feature map, that is, the processing process of the multi-scale feature, which can be expressed as:
[0207] ;
[0208] in, represents the splicing operation along the time dimension, , They represent the feature maps of the complete copy, 1 / 2 slice, 1 / 4 slice and 1 / 8 slice of the input sequence respectively, namely the first multi-scale normalized feature, the second multi-scale normalized feature, the third multi-scale normalized feature, and the second self-attention normalized feature. The feature map representing the original input sequence (i.e., the first self-attention normalized features). (Multi-scale features) are used as simulated motion pattern features of a certain length, added before the earliest observation time of the original input features, and are composed of slice features of different historical lengths. Within these slice sequences, the motion pattern data conforms to the law of change of the real motion pattern over time. The addition of long multi-scale motion pattern features enables MHAE to provide multi-head attention for the subsequent decoder. and It covers the important characteristic information of motion pattern data at different time scales.
[0209] Furthermore, the processing process of inputting historical motion time step embedding, future motion time step embedding and edge time step embedding into the preset multi-scale history aggregation encoder for encoding can refer to the process of processing the input time step embedding by the preset multi-scale history aggregation encoder, so as to obtain the multi-scale historical motion features corresponding to the historical motion time step embedding, the multi-scale future motion features corresponding to the future motion time step embedding and the multi-scale edge features corresponding to the edge time step embedding, and the present invention will not go into details.
[0210] It is worth mentioning that the present invention uses the original input tensor (input time step embedding) as a separate constraint branch of MHAE, and constructs four copies of the input for feature extraction at different time scales. In the branch used as a constraint, only one self-attention block is applied to encode the original input features without changing their time dimension. In the multi-scale branch, the four copies of the original input are processed separately. One copy retains the full time dimension, and the other three copies extract subsequences of 1 / 2, 1 / 4, and 1 / 8 of the time dimension closest to the current time step by slicing the input features. Figure 2 The sizes of the three slices are , , ,in Represents the batch size. Multi-time scale feature aggregation is achieved by encoding these history slices of different lengths.
[0211] Furthermore, motion pattern data of different scales may have different temporal variation patterns, such as periodic changes or rapid changes. By extracting features at different scales, MHAE can focus on both the global and local structures of motion pattern data and better adapt to data with different temporal variation patterns. Features with larger temporal dimensions can capture the long-term dependencies and overall trends of motion pattern sequences, while features with smaller temporal dimensions can capture local details and short-term changes. In addition, the collected data may be affected by noise interference. Features with smaller scales can help the model capture local changes in a short period of time more sensitively, making the model more adaptable and robust to noise.
[0212] Step 106: Output an initial prediction result through multiple preset efficient autoregressive decoders according to multi-scale historical motion features, multi-scale future motion features, multi-scale edge features, map feature vectors and the historical motion pattern sequence of the intelligent agent.
[0213] Please note that Figure 3, the present invention designs an efficient autoregressive decoder (EAD) for decoding. EAD is an efficient decoder, and its contributions mainly include two parts. On the one hand, the real-time calculation time of the model is reduced by using the probabilistic sparse self-attention mechanism; on the other hand, the prediction accuracy is improved by combining the autoregressive decoding of the historical labels, which is different from the traditional autoregressive decoder. The specific structure of EAD is shown in the attached Figure 3 The autoregressive decoding part on the right extracts part of the output it generates as the input of the multi-layer perceptron, and combines it with the preset binary Gaussian mixture model trained based on re-parameterization to generate the target agent's motion pattern prediction result.
[0214] It is worth mentioning that the present invention decodes according to multi-scale historical motion features, multi-scale future motion features, multi-scale edge features, map feature vectors and agent historical motion pattern sequences through seven preset efficient autoregressive decoders to output initial prediction results; among them, the input of the first preset efficient autoregressive decoder is multi-scale historical motion features, multi-scale future motion features, multi-scale edge features, map feature vectors and agent historical motion pattern sequences, and the input of the remaining preset efficient autoregressive decoders is multi-scale historical motion features, multi-scale future motion features, multi-scale edge features, map feature vectors and new agent historical motion pattern sequences, and the output prediction results generated by the previous preset efficient autoregressive decoder; the output prediction result generated by the last preset efficient autoregressive decoder is the initial prediction result.
[0215] Optionally, the data processing process of presetting the efficient autoregressive decoder may include the following sub-step S61:
[0216] Step S61, using a multi-layer perceptron to perform feature mapping on the input historical motion pattern sequence input to a preset high-efficiency autoregressive decoder to generate historical motion high-dimensional features;
[0217] Step S62, concatenating the historical motion high-dimensional features, the input edge features, the input map features, and the input future motion features input to the preset efficient autoregressive decoder to generate concatenated features;
[0218] Step S63, performing a global adaptive one-dimensional average pooling operation on the concatenated features to generate a global feature representation;
[0219] Step S64, performing time excitation on the global feature representation and outputting the time excitation feature;
[0220] Please note that Figure 4 , a time compression and excitation (TSE module) mechanism is designed for feature fusion and time dimension expansion. The decoder further combines social interaction features (i.e., multi-scale edge features) , grid map features (i.e. map feature vector) , ego-vehicle motion planning features (i.e. multi-scale future motion features) , through time dimension splicing and TSE optimization feature expression, a high-dimensional nonlinear prediction result is formed. The input of the TSE module is multiple input sources at any time The feature concatenation of Time compression and time excitation are two key steps of the TSE module. They work together to change the representation and extraction of time series features by learning and adjusting the weights of channels. The structural diagram of the TSE module is shown in Figure 4 shown.
[0221] Furthermore, for time compression, in the time compression step, the TSE module performs a global adaptive one-dimensional average pooling operation (GAP) on the concatenated multimodal sequence features, with the aim of compressing the time dimension of the features while retaining important information. The global adaptive one-dimensional average pooling operation can be expressed as:
[0222] ;
[0223] in, It is the global feature representation; It is a global adaptive one-dimensional average pooling.
[0224] Furthermore, for each time channel, the TSE module calculates the average of all modes in the channel. In this way, the TSE module helps the model better understand the overall time series characteristics and reduce dependence on a specific mode, thereby enhancing the robustness of the model.
[0225] Furthermore, for time excitation, the goal of time excitation is to learn weights for each channel and adjust the input features according to their importance. The calculation process (i.e., the processing process of time excitation) is as follows:
[0226] ;
[0227] in, represents the global feature representation obtained through the time compression step. It is a multi-layer perceptron that maps the pooled global features to a very low dimension (1 / 16 of the original dimension) to focus on important features. The introduction of the ReLU function enhances the nonlinearity of the module, enabling the TSE module to learn more complex weight distributions. Multi-layer perceptron The features are mapped back to the original high dimension for more fine-grained adjustments and weighting. The activation output is normalized to the range of [0,1] through the Sigmoid function to obtain the weight of each channel. These weights represent the importance of each channel, where a weight value close to 1 indicates that the channel contributes more to the feature representation, while a weight value close to 0 indicates a smaller contribution.
[0228] Step S65, multiplying the time excitation feature and the splicing feature element by element to generate a time series feature representation;
[0229] Based on the above foundation, using these channel weights, the TSE module reweights the original input features. For each time channel, the original feature is multiplied by the corresponding channel weight to obtain the reweighted feature. In this way, features from important time channels are amplified, while features from unimportant time channels are suppressed. The overall forward function expression of the TSE module, that is, the processing process of time series feature representation, can be expressed as:
[0230] ;
[0231] in, Represents the element-by-element multiplication operation. This step completes the re-weighting of the features and generates the final time series feature representation .
[0232] Step S66, using a multi-layer perceptron to perform feature mapping on the time series feature representation, and outputting high-dimensional features of the time series;
[0233] Step S67, performing layer normalization and nonlinear mapping on the high-dimensional features of the time series, and outputting the nonlinear features of the time series;
[0234] Step S68, adding the time series nonlinear characteristics and the historical moment prediction results element by element to generate the current moment prediction result;
[0235] Step S69, concatenating the current prediction result with the target subsequence in the input historical motion pattern sequence to generate a new input historical motion pattern sequence;
[0236] The target subsequence in the input historical motion pattern sequence is the feature data corresponding to the historical motion state pattern label (-H to -1).
[0237] It should be noted that after the time-step embedding operation (TE) is performed on the decoder input, it is processed by the multi-layer perceptron, layer normalization, and LeakyReLU activation function in sequence, and then added element by element with the historical moment prediction result (initialized as the historical motion pattern sequence of the agent input to the first preset efficient autoregressive decoder) to obtain the current moment prediction result. This process can be expressed as:
[0238] in, Predict the outcome for the current moment; predicting outcomes for historical moments; is the input historical motion pattern sequence.
[0239] Furthermore, the newly generated prediction results are spliced into the historical input to form a new input sequence, that is, the prediction results at the current moment and the target subsequence in the input history motion pattern sequence Splice to generate a new input history motion pattern sequence , the process can be expressed as:
[0240] ;
[0241] It is worth mentioning that the generated new input historical motion pattern sequence is used as the input of the next preset efficient autoregressive decoder.
[0242] Step S610, performing time step embedding and position encoding on the new input historical motion pattern sequence, and outputting a coding feature sequence;
[0243] It should be noted that the processing principle of time step embedding can refer to the embedding process of the above-mentioned intelligent agent historical motion pattern sequence, the intelligent agent future motion planning sequence and the time dimension edge sequence, and the present invention will not go into details.
[0244] Furthermore, the position encoding is added to the new input historical motion pattern sequence after the time step embedding process by using a preset position encoding matrix.
[0245] Step S611: Based on the single-layer mask multi-head probabilistic sparse self-attention mechanism, perform attention operation on the encoded feature sequence and output probabilistic sparse self-attention;
[0246] Step S612, adding the probabilistic sparse self-attention and the encoding feature sequence element by element, and outputting the probabilistic sparse self-attention feature sequence;
[0247] Step S613, performing layer normalization on the probabilistic sparse self-attention feature sequence to determine a query matrix;
[0248] Step S614, linearly projecting the input historical motion features input to the preset efficient autoregressive decoder to generate a key matrix and a value matrix;
[0249] Step S615, generating an efficient decoding feature sequence according to the query matrix, the key matrix and the value matrix;
[0250] Step S616, performing layer normalization on the efficient decoding feature sequence, and outputting an efficient decoding normalized feature sequence;
[0251] Step S617: Use a multi-layer perceptron to perform feature mapping on the efficient decoding normalized feature sequence to generate an output prediction result.
[0252] The input historical motion pattern sequence, the input edge feature, the input map feature, the input future motion feature, and the input historical motion feature are feature data for inputting a preset efficient autoregressive decoder. It can be understood that the input historical motion pattern sequence, the input edge feature, the input map feature, the input future motion feature, and the input historical motion feature can correspond to any feature data inputted into a preset efficient autoregressive decoder for data processing during model training or model application.
[0253] Data such as historical motion high-dimensional features and time excitation features are intermediate feature data generated in a preset efficient autoregressive decoder;
[0254] The output prediction result is the feature data output by the preset efficient autoregressive decoder. It can be understood that it can correspond to any feature data output after the preset efficient autoregressive decoder performs data processing during model training or model application.
[0255] It should be noted that the present invention constructs an efficient decoding module (EDB) to optimize real-time computing performance. Real-time performance is the key to the successful deployment of motion pattern prediction models in autonomous driving systems. Unlike the encoding process, the decoder usually only needs to focus on the relevant parts of the generated sequence without capturing the semantic and spatial information of the complete input sequence. Traditional self-attention models calculate weights for all elements, resulting in high computational costs and low efficiency. The probabilistic sparse self-attention mechanism solves this problem by limiting the attention mechanism to only act on key query interactions that meet the sparsity criterion. The probabilistic sparse self-attention mechanism is represented as follows:
[0256] ;
[0257] in, is a sparse matrix associated only with the primary query that satisfies the sparsity criterion. The sparsity criterion is defined as:
[0258] ;
[0259] in, For query q i The maximum mean value on the key set K, ; is the square root of the dimension of the input data feature; is the transpose of the jth key; is the i-th query vector; L is the number of keys, usually expressed as L K or L Q ; is a part of the attention weight, which represents the similarity between a given query and a specific key.
[0260] Furthermore, by focusing on the most critical data elements, probabilistic sparse self-attention reduces unnecessary calculations, thereby speeding up processing time. Based on the single-layer masked multi-head probabilistic sparse self-attention (MMPSA), the present invention constructs an efficient decoding module (EDB), whose forward function is used to calculate the query matrix and applies mask operations in the time dimension to improve computational efficiency. Its structure is as follows (i.e., the processing process of the query matrix):
[0261] ;
[0262] Furthermore, efficient decoding of feature sequences The processing process can be expressed as:
[0263] in, is the query matrix; is the input data; K and V are the input historical motion features The corresponding key matrix and value matrix.
[0264] It is worth mentioning that the present invention introduces historical labels to improve the autoregressive decoding structure. Another core feature of EAD is the uniqueness of its autoregressive decoding structure. Different from conventional methods, the input data of EAD includes To the current time The position data of all agents is defined as:
[0265] ;
[0266] Among them, for , the data is the real motion mode label; and for , where the data is the prediction result generated by the model at the previous moment. By introducing the joint input of the true label and the historical prediction result, the decoder can achieve stronger model robustness.
[0267] Furthermore, the output prediction results generated by the first six preset efficient autoregressive decoders replace the input historical motion pattern sequence as the input, and are concatenated with the input edge features, input map features, and input future motion features after being processed by the multi-layer perceptron to perform subsequent processing, and the output prediction result generated by the last preset efficient autoregressive decoder is used as the initial prediction result.
[0268] Step 107: Use a preset binary Gaussian mixture model to generate a target agent motion mode prediction result based on the initial prediction result.
[0269] It should be noted that the prediction part is extracted from the initial prediction result , use it as the input of the multi-layer perceptron, map these prediction vectors into intermediate feature vectors through the multi-layer perceptron (MLP), and then input them into the preset binary Gaussian mixture model to output a variety of possible motion mode distributions, perform maximum probability sampling or random sampling on the multiple possible motion mode distributions, and generate the target agent motion mode prediction results.
[0270] As a comparison of technical effects, it can be combined with existing technologies for reference. Motion pattern recognition technology is widely used in intelligent monitoring, robot navigation, autonomous driving and other fields. Open scene dynamic warning accurately identifies the motion pattern of the intelligent body and predicts future motion trends through real-time data collection and intelligent analysis, provides dynamic collision risk warning, and supports path planning. Traditional methods have limited prediction capabilities in complex scenes. Neural network methods have made certain progress, but there are still problems such as insufficient time domain feature mining, high decoding complexity, and insufficient feature fusion. To this end, the present invention proposes a method for predicting the motion pattern of an intelligent body for open scenes, aiming to solve the shortcomings of the existing technology in time domain feature exploration, decoder computational complexity and multimodal feature fusion. By introducing a multi-scale historical aggregation encoder and an efficient autoregressive decoder, combined with a binary Gaussian mixture model, the prediction accuracy is improved, and the motion patterns of multiple intelligent bodies (such as crowds, motor vehicles, etc.) can be more accurately predicted, providing more reliable dynamic warnings and path planning support for intelligent monitoring, robot navigation, autonomous driving and other fields.
[0271] In an embodiment of the present invention, the present invention provides a method for predicting the motion pattern of an intelligent agent for open scenes. First, the intelligent agent's historical motion pattern sequence, the intelligent agent's future motion planning sequence, a monitoring video data set and a city road map are obtained, and the city road map is map-encoded to generate a map feature vector; then, multiple two-dimensional spatial position coordinates are generated according to the original annotation file of the monitoring video data set, and the multiple two-dimensional spatial position coordinates are homogeneously transformed and normalized to generate multiple motion state information; a preset spatial edge rule is used to generate a time dimension edge sequence according to the multiple motion state information and the multiple two-dimensional spatial position coordinates; the intelligent agent's historical motion pattern sequence, the intelligent agent's future motion planning sequence and the time dimension edge sequence are respectively embedded to generate the historical motion time step embedding corresponding to the intelligent agent's historical motion pattern sequence, the future motion time step embedding corresponding to the intelligent agent's future motion planning sequence and the edge time step embedding corresponding to the time dimension edge sequence; a preset multi-scale historical aggregation encoder is used to embed the historical motion time step embedding, the future motion time step embedding and the edge The time step embedding is used for encoding to generate multi-scale historical motion features corresponding to the historical motion time step embedding, multi-scale future motion features corresponding to the future motion time step embedding, and multi-scale edge features corresponding to the edge time step embedding; an initial prediction result is output according to the multi-scale historical motion features, multi-scale future motion features, multi-scale edge features, map feature vectors and the agent historical motion pattern sequence through multiple preset efficient autoregressive decoders; finally, a preset binary Gaussian mixture model is used to generate a target agent motion pattern prediction result according to the initial prediction result; based on the above scheme, based on a preset multi-scale historical aggregation encoder and multiple preset efficient autoregressive decoders, multimodal features, i.e., the acquired agent historical motion pattern sequence, agent future motion planning sequence, monitoring video data set and urban road map are processed to output the initial prediction result, and then combined with the preset binary Gaussian mixture model to generate the target agent motion pattern prediction result according to the initial prediction result, the present invention can deeply optimize the feature fusion in the time dimension, thereby improving the effect of multimodal feature fusion.
[0272] For better explanation, refer to Figure 5 , showing Figure 5 The training flow chart of the preset binary Gaussian mixture model provided in the second embodiment of the present invention may include the following steps:
[0273] 1. Obtain the dataset for model training; 2. Calculate the kinematic properties required for modeling based on the original annotations of the dataset; 3. Physical modeling of movement patterns; 4. Spatial graph representation of social interactions of intelligent agents, time dimension series representation and feature preprocessing; 4. Timestamp embedding and city map encoding of multi-source time series; 5. Multi-scale history aggregation encoder (MHAE) for feature encoding of timestamp embedding; 6. Efficient autoregressive decoder (EAD) for decoding; 7. Time compression and excitation (TSE) mechanism for feature fusion and time dimension expansion; 7. Model training.
[0274] It should be noted that for the training of the preset binary Gaussian mixture model, in the autoregressive decoder, the initial output of the last step of decoding is ; Next, extract the prediction part from the initial output , which is the final output of the decoder: ,in, Represents the predicted spatial coordinate position of the agent at the time step, that is, the agent position at each predicted time step can be expressed as a two-dimensional position vector, is the two-dimensional position vector of agent i at any time step t in the time range of t=-H to t=T.
[0275] Further, these prediction vectors are mapped into intermediate feature vectors through multi-layer perceptrons (MLPs). Then four relatively simple predefined submodules are designed to convert the intermediate features into Gaussian mixture model (GMM) parameters. Each submodule contains only linear layers to map the hidden dimensions of the intermediate features into GMM parameter spaces with different dimensional sizes.
[0276] Specifically, the outputs of these four submodules include a binary Gaussian mixture model Four types of parameters: mean , logarithmic standard deviation , correlation coefficient , the logarithm of the mixing coefficient .for components, with the following constraints: .in, , and Through MLP initialization, The Tanh activation function and MLP are used to initialize the covariance matrix and probability density function of GMM.
[0277] Represents the time step The covariance matrix of the position is calculated as:
[0278] ;
[0279] in, is the probability density function of the bivariate Gaussian distribution, calculated using the following formula:
[0280] in, is the logarithmic standard deviation in the transverse direction; is the logarithmic standard deviation in the longitudinal direction.
[0281] Furthermore, since sampling directly from GMM will result in the inability of gradient propagation, the present invention uses a reparameterization technique. First, from the standard normal distribution Sampling Standardized 2D Random Variables Through affine transformation, the standard normal distribution samples are converted into samples of the target distribution. For the bivariate Gaussian distribution, the affine transformation is as follows:
[0282] ;
[0283] In order to minimize the gap between the predicted output and the actual motion pattern, the present invention supervises the model by using regression loss The final output is:
[0284] ;
[0285] in, represents the total number of agents in a batch of samples, and Represents any agent in the training sample At time step The predicted motion patterns and ground truth.
[0286] During the validation phase of training, the model is supervised by minimizing the negative log-likelihood loss Learning of all trainable parameters:
[0287] ;
[0288] Based on the above foundation, the present invention can obtain a motion pattern prediction model with better prediction effect and smaller deviation, that is, a preset binary Gaussian mixture model.
[0289] For example, see Figure 6, according to the usage scenario, the input data required by the model is obtained in real time; the trained preset binary Gaussian mixture model is used for model reasoning and future motion pattern generation. Based on the GMM parameters of the trained preset binary Gaussian mixture model, possible future motion patterns are generated by sampling from the GMM. The two optional sampling methods include: 1) selecting the maximum probability mode of the GMM to generate the most likely motion pattern; 2) randomly sampling several motion patterns from the GMM to capture the diversity and uncertainty of future motion patterns. By switching the sampling strategy, the most likely future motion pattern and the motion pattern with greater randomness can be simulated.
[0290] Furthermore, the predicted motion pattern generated by the model is based on the world coordinate system. If it is in actual needs such as visualization, it needs to be converted to other coordinate systems (such as pixel coordinate system). The formula for homogeneous coordinate conversion from the world coordinate system to the pixel coordinate system is: ;
[0291] in, Represents the homogeneous form of world coordinates; is the transformation matrix from world coordinates to pixel coordinates. After conversion, the coordinates need to be normalized, that is: , is the scaling factor in homogeneous coordinates.
[0292] In an embodiment of the present invention, since sampling directly from GMM will cause the gradient to be unable to be transferred, a reparameterization technique is used. The present invention first randomly samples a two-dimensional random variable from a standard normal distribution, converts it into a sample of the target distribution through an affine transformation, and generates a training sample of a Gaussian distribution. The training of the model is supervised using regression loss, and these samples are used to calculate the regression loss of the model to minimize the gap between the predicted motion pattern and the actual motion pattern. In the verification stage of training, the learning of all trainable parameters of the model is supervised by minimizing the negative log-likelihood loss.
[0293] See also Figure 7 , Figure 7 This is a structural block diagram of an intelligent body motion pattern prediction device for open scenes provided in Example 3 of the present invention.
[0294] The present invention provides an open scene-oriented intelligent body motion pattern prediction device, comprising:
[0295] The acquisition module 701 is used to acquire the historical motion pattern sequence of the intelligent agent, the future motion planning sequence of the intelligent agent, the monitoring video data set and the urban road map, and perform map encoding on the urban road map to generate a map feature vector;
[0296] According to module 702, it is used to generate multiple two-dimensional space position coordinates according to the original annotation file of the monitoring video data set, and perform homogeneous coordinate transformation and normalization on the multiple two-dimensional space position coordinates to generate multiple motion state information;
[0297] Adopting module 703, for generating a time dimension edge sequence according to a plurality of motion state information and a plurality of two-dimensional space position coordinates by adopting a preset spatial edge rule;
[0298] The embedding module 704 is used to embed the agent's historical motion pattern sequence, the agent's future motion planning sequence and the time dimension edge sequence respectively, and generate the historical motion time step embedding corresponding to the agent's historical motion pattern sequence, the future motion time step embedding corresponding to the agent's future motion planning sequence and the edge time step embedding corresponding to the time dimension edge sequence;
[0299] The encoding module 705 is used to respectively encode the historical motion time step embedding, the future motion time step embedding and the edge time step embedding using a preset multi-scale history aggregation encoder to generate multi-scale historical motion features corresponding to the historical motion time step embedding, multi-scale future motion features corresponding to the future motion time step embedding and multi-scale edge features corresponding to the edge time step embedding;
[0300] A decoding module 706, configured to output an initial prediction result according to multi-scale historical motion features, multi-scale future motion features, multi-scale edge features, map feature vectors and the agent historical motion pattern sequence through a plurality of preset efficient autoregressive decoders;
[0301] The output result module 707 is used to generate a target intelligent body motion mode prediction result based on the initial prediction result by using a preset binary Gaussian mixture model.
[0302] Furthermore, the acquisition module 701 is specifically used for:
[0303] Divide the city road map and output multiple map sub-blocks;
[0304] Transform each map sub-block to generate a sub-block vector corresponding to each map sub-block;
[0305] Perform a linear transformation on each sub-block vector and output the linear embedding corresponding to each sub-block vector;
[0306] Add the preset position encoding matrix to each linear embedding respectively to generate an intermediate embedding corresponding to each linear embedding;
[0307] Based on the multi-head self-attention mechanism, the attention weight of each map sub-block is calculated separately to determine the attention weight corresponding to each map sub-block;
[0308] Add each attention weight to the intermediate embedding corresponding to each attention weight to generate the target embedding corresponding to each intermediate embedding;
[0309] Perform layer normalization on each target embedding to determine the normalized embedding corresponding to each target embedding;
[0310] A multi-layer perceptron is used to generate a map feature vector based on multiple normalized embeddings.
[0311] Further, module 703 is used to:
[0312] Smoothing the multiple motion state information and the multiple two-dimensional space position coordinates to generate multiple smoothed motion state information and multiple smoothed two-dimensional space position coordinates;
[0313] Constructing a node set based on multiple smooth motion state information and multiple smooth two-dimensional spatial position coordinates;
[0314] Using preset spatial edge rules to construct edge sets according to multiple nodes in the node set;
[0315] Based on the edge set, determine the edge sequence in the time dimension.
[0316] Furthermore, the embedding module 704 is specifically used for:
[0317] Perform one-dimensional convolution on the agent's historical motion pattern sequence, the agent's future motion planning sequence, and the time dimension edge sequence respectively, and output the historical motion convolution features corresponding to the agent's historical motion pattern sequence, the future motion convolution features corresponding to the agent's future motion planning sequence, and the edge convolution features corresponding to the time dimension edge sequence;
[0318] Nonlinear mapping is performed on the historical motion convolution features, the future motion convolution features and the edge convolution features respectively to generate the historical motion label embedding corresponding to the historical motion convolution features, the future motion label embedding corresponding to the future motion convolution features and the edge label embedding corresponding to the edge convolution features;
[0319] Perform position embedding processing on the agent's historical motion pattern sequence, the agent's future motion planning sequence, and the time dimension edge sequence respectively, and output the historical motion position embedding corresponding to the agent's historical motion pattern sequence, the future motion position embedding corresponding to the agent's future motion planning sequence, and the edge position embedding corresponding to the time dimension edge sequence;
[0320] Perform weighted summation of historical motion tag embedding and historical motion position embedding to determine historical motion time step embedding;
[0321] Perform a weighted summation of the future motion tag embedding and the future motion position embedding to determine the future motion time step embedding;
[0322] The edge time step embedding is determined by taking a weighted sum of the edge label embedding and the edge position embedding.
[0323] Furthermore, the preset multi-scale history aggregation encoder includes a self-attention submodule, a first multi-scale branch module, a second multi-scale branch module, and a third multi-scale branch module; the data processing process of the preset multi-scale history aggregation encoder is specifically as follows:
[0324] Slicing the input time step embedding input to the preset multi-scale history aggregation encoder to generate a first sub-feature, a second sub-feature, and a third sub-feature;
[0325] The input time step embedding is used as the input of the self-attention submodule, and the first self-attention feature is output;
[0326] Using a first multi-scale branch module to perform multi-scale feature extraction on the input time step embedding to generate a first multi-scale feature;
[0327] Performing multi-scale feature extraction on the first sub-feature through a second multi-scale branch module to generate a second multi-scale feature;
[0328] Inputting the second sub-feature into the third multi-scale branch module for multi-scale feature extraction to generate a third multi-scale feature;
[0329] The third sub-feature is used as the input of the self-attention sub-module and the second self-attention feature is output;
[0330] Respectively perform layer normalization on the first self-attention feature, the first multi-scale feature, the second multi-scale feature, the third multi-scale feature, and the second self-attention feature, and output the first self-attention normalized feature, the first multi-scale normalized feature, the second multi-scale normalized feature, the third multi-scale normalized feature, and the second self-attention normalized feature;
[0331] The first self-attention normalized feature, the first multi-scale normalized feature, the second multi-scale normalized feature, the third multi-scale normalized feature, and the second self-attention normalized feature are concatenated to generate an output multi-scale feature.
[0332] Furthermore, the input time step embedding is used as the input of the self-attention submodule, and the first self-attention feature is output, including:
[0333] Based on the multi-head self-attention mechanism, the attention calculation is performed on the input time step embedding to determine the multi-head self-attention corresponding to the input time step embedding;
[0334] Add the multi-head self-attention and input time step embedding element by element, and output the multi-head self-attention time step features;
[0335] Perform layer normalization on the multi-head self-attention time step features to generate multi-head self-attention normalized features;
[0336] Perform one-dimensional convolution on the normalized features of multi-head self-attention and output one-dimensional convolution features of multi-head self-attention;
[0337] Perform nonlinear mapping on the one-dimensional convolution features of multi-head self-attention to generate multi-head self-attention nonlinear features;
[0338] Perform one-dimensional convolution on the multi-head self-attention nonlinear features and output the first self-attention feature.
[0339] Furthermore, the third multi-scale branch module includes a self-attention sub-module and a self-attention distillation downsampling sub-module; the second sub-feature is input into the third multi-scale branch module for multi-scale feature extraction to generate a third multi-scale feature, including:
[0340] The second sub-feature is used as the input of the self-attention sub-module, and the self-attention feature corresponding to the second sub-feature is output;
[0341] The self-attention distillation and downsampling submodule is used to distill and downsample the self-attention feature corresponding to the second sub-feature to generate the maximum pooling feature of the second sub-feature;
[0342] The second sub-feature maximum pooling feature is used as the input of the self-attention sub-module and the third multi-scale feature is output.
[0343] Furthermore, the self-attention distillation downsampling submodule is used to distill and downsample the self-attention feature corresponding to the second sub-feature to generate the maximum pooling feature of the second sub-feature, including:
[0344] Perform one-dimensional convolution on the self-attention feature corresponding to the second sub-feature, and output the convolution feature of the second sub-feature;
[0345] Batch normalize the second sub-feature convolution features to generate second sub-feature normalized features;
[0346] Downsampling the normalized feature of the second sub-feature, and outputting the downsampled feature of the second sub-feature;
[0347] The second sub-feature downsampling feature is subjected to maximum pooling to generate the second sub-feature maximum pooling feature.
[0348] Furthermore, the data processing process of the preset efficient autoregressive decoder is specifically as follows:
[0349] A multi-layer perceptron is used to perform feature mapping on the input historical motion pattern sequence input to a preset efficient autoregressive decoder to generate high-dimensional features of historical motion;
[0350] The historical motion high-dimensional features are concatenated with the input edge features, the input map features, and the input future motion features input to the preset efficient autoregressive decoder to generate concatenated features;
[0351] Perform a global adaptive one-dimensional average pooling operation on the concatenated features to generate a global feature representation;
[0352] Perform temporal excitation on the global feature representation and output the temporal excitation feature;
[0353] Multiply the time excitation feature and the concatenation feature element by element to generate a time series feature representation;
[0354] A multi-layer perceptron is used to perform feature mapping on the time series feature representation and output high-dimensional features of the time series;
[0355] Perform layer normalization and nonlinear mapping on the high-dimensional features of the time series, and output the nonlinear features of the time series;
[0356] The nonlinear characteristics of the time series and the prediction results of the historical moments are added element by element to generate the prediction results of the current moment;
[0357] The prediction result at the current moment is concatenated with the target subsequence in the input historical motion pattern sequence to generate a new input historical motion pattern sequence;
[0358] Perform time step embedding and position encoding on the new input historical motion pattern sequence, and output the encoded feature sequence;
[0359] Based on the single-layer masked multi-head probabilistic sparse self-attention mechanism, the attention operation is performed on the encoded feature sequence, and the probabilistic sparse self-attention is output;
[0360] Add the probabilistic sparse self-attention and the encoded feature sequence element by element, and output the probabilistic sparse self-attention feature sequence;
[0361] Perform layer normalization on the probabilistic sparse self-attention feature sequence to determine the query matrix;
[0362] Linearly project the input historical motion features input to a preset efficient autoregressive decoder to generate a key matrix and a value matrix;
[0363] Generate efficient decoding feature sequence according to query matrix, key matrix and value matrix;
[0364] Perform layer normalization on the efficient decoding feature sequence and output an efficient decoding normalized feature sequence;
[0365] A multi-layer perceptron is used to perform feature mapping on the efficient decoding normalized feature sequence to generate output prediction results.
[0366] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0367] An embodiment of the present invention also provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory; when the computer program is executed by the processor, the processor executes the steps of a method for predicting motion patterns of an intelligent body for open scenes as described in the first embodiment above.
[0368] An embodiment of the present invention also provides a computer-readable storage medium having a computer program / instruction stored thereon. When the computer program / instruction is executed by a processor, the steps of a method for predicting motion patterns of an intelligent body for open scenes as described in the first embodiment above are implemented.
[0369] An embodiment of the present invention also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of a method for predicting motion patterns of an intelligent body for open scenes as described in the first embodiment above.
[0370] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0371] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0372] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for predicting motion patterns of intelligent agents in open scenarios, characterized in that: include: Acquire the agent's historical motion pattern sequence, the agent's future motion planning sequence, a surveillance video dataset, and a city road map, and perform map encoding on the city road map to generate a map feature vector; Generate a plurality of two-dimensional spatial position coordinates according to the original annotation file of the surveillance video data set, and perform homogeneous coordinate transformation and normalization on the plurality of two-dimensional spatial position coordinates to generate a plurality of motion state information; Generate a time dimension edge sequence according to the plurality of motion state information and the plurality of two-dimensional space position coordinates using a preset spatial edge rule; The agent's historical motion pattern sequence, the agent's future motion planning sequence and the time dimension edge sequence are respectively embedded to generate a historical motion time step embedding corresponding to the agent's historical motion pattern sequence, a future motion time step embedding corresponding to the agent's future motion planning sequence and an edge time step embedding corresponding to the time dimension edge sequence; Using a preset multi-scale history aggregation encoder to encode the historical motion time step embedding, the future motion time step embedding and the edge time step embedding respectively, to generate multi-scale historical motion features corresponding to the historical motion time step embedding, multi-scale future motion features corresponding to the future motion time step embedding and multi-scale edge features corresponding to the edge time step embedding; Outputting an initial prediction result according to the multi-scale historical motion features, the multi-scale future motion features, the multi-scale edge features, the map feature vector and the historical motion pattern sequence of the intelligent agent through a plurality of preset efficient autoregressive decoders; A preset binary Gaussian mixture model is used to generate a target agent motion mode prediction result based on the initial prediction result.
2. The method for predicting the motion pattern of an intelligent agent in an open scene according to claim 1, characterized in that: The step of performing map encoding on the city road map to generate a map feature vector includes: Dividing the city road map and outputting a plurality of map sub-blocks; Transforming each of the map sub-blocks to generate a sub-block vector corresponding to each of the map sub-blocks; Performing a linear transformation on each of the sub-block vectors, and outputting a linear embedding corresponding to each of the sub-block vectors; Adding the preset position coding matrix to each of the linear embeddings respectively to generate an intermediate embedding corresponding to each of the linear embeddings; Based on the multi-head self-attention mechanism, the attention weight of each map sub-block is calculated respectively to determine the attention weight corresponding to each map sub-block; Adding each of the attention weights to the intermediate embedding corresponding to each of the attention weights to generate a target embedding corresponding to each of the intermediate embeddings; Performing layer normalization on each of the target embeddings to determine a normalized embedding corresponding to each of the target embeddings; A multi-layer perceptron is used to generate a map feature vector according to the multiple normalized embeddings.
3. The method for predicting the motion pattern of an intelligent agent in an open scene according to claim 1, characterized in that: The step of using a preset spatial edge rule to generate a time dimension edge sequence according to the plurality of motion state information and the plurality of two-dimensional spatial position coordinates includes: Smoothing the plurality of motion state information and the plurality of two-dimensional space position coordinates to generate a plurality of smoothed motion state information and a plurality of smoothed two-dimensional space position coordinates; Constructing a node set based on the plurality of smooth motion state information and the plurality of smooth two-dimensional space position coordinates; Using preset spatial edge rules to construct edge sets according to multiple nodes in the node set; Based on the edge set, a time dimension edge sequence is determined.
4. The method for predicting the motion pattern of an intelligent agent in an open scene according to claim 1, characterized in that: The embedding process is respectively performed on the agent historical motion pattern sequence, the agent future motion planning sequence and the time dimension edge sequence to generate the historical motion time step embedding corresponding to the agent historical motion pattern sequence, the future motion time step embedding corresponding to the agent future motion planning sequence and the edge time step embedding corresponding to the time dimension edge sequence, including: Perform one-dimensional convolution on the agent's historical motion pattern sequence, the agent's future motion planning sequence, and the time dimension edge sequence respectively, and output the historical motion convolution features corresponding to the agent's historical motion pattern sequence, the future motion convolution features corresponding to the agent's future motion planning sequence, and the edge convolution features corresponding to the time dimension edge sequence; Nonlinearly mapping the historical motion convolution feature, the future motion convolution feature and the edge convolution feature respectively to generate a historical motion label embedding corresponding to the historical motion convolution feature, a future motion label embedding corresponding to the future motion convolution feature and an edge label embedding corresponding to the edge convolution feature; Performing position embedding processing on the agent's historical motion pattern sequence, the agent's future motion planning sequence, and the time dimension edge sequence respectively, and outputting the historical motion position embedding corresponding to the agent's historical motion pattern sequence, the future motion position embedding corresponding to the agent's future motion planning sequence, and the edge position embedding corresponding to the time dimension edge sequence; Performing a weighted summation of the historical motion mark embedding and the historical motion position embedding to determine a historical motion time step embedding; Performing a weighted summation on the future motion tag embedding and the future motion position embedding to determine a future motion time step embedding; A weighted sum is performed on the edge label embedding and the edge position embedding to determine an edge time step embedding.
5. The method for predicting motion patterns of intelligent agents in open scenarios according to claim 1, characterized in that: The preset multi-scale history aggregation encoder includes a self-attention submodule, a first multi-scale branch module, a second multi-scale branch module, and a third multi-scale branch module; the data processing process of the preset multi-scale history aggregation encoder is specifically as follows: Slicing the input time step embedding input to the preset multi-scale history aggregation encoder to generate a first sub-feature, a second sub-feature, and a third sub-feature; Embedding the input time step as an input to a self-attention submodule, and outputting a first self-attention feature; Using a first multi-scale branch module to perform multi-scale feature extraction on the input time step embedding to generate a first multi-scale feature; Performing multi-scale feature extraction on the first sub-feature by a second multi-scale branch module to generate a second multi-scale feature; Inputting the second sub-feature into a third multi-scale branch module for multi-scale feature extraction to generate a third multi-scale feature; Using the third sub-feature as the input of the self-attention sub-module, and outputting a second self-attention feature; Respectively perform layer normalization on the first self-attention feature, the first multi-scale feature, the second multi-scale feature, the third multi-scale feature, and the second self-attention feature, and output a first self-attention normalized feature, a first multi-scale normalized feature, a second multi-scale normalized feature, a third multi-scale normalized feature, and a second self-attention normalized feature; The first self-attention normalized feature, the first multi-scale normalized feature, the second multi-scale normalized feature, the third multi-scale normalized feature, and the second self-attention normalized feature are concatenated to generate an output multi-scale feature.
6. The method for predicting the motion pattern of an intelligent agent in an open scene according to claim 5, characterized in that: The step of embedding the input time step as an input of a self-attention submodule and outputting a first self-attention feature comprises: Based on the multi-head self-attention mechanism, performing attention calculation on the input time step embedding to determine the multi-head self-attention corresponding to the input time step embedding; Add the multi-head self-attention and the input time step embedding element by element, and output a multi-head self-attention time step feature; Performing layer normalization on the multi-head self-attention time step features to generate multi-head self-attention normalized features; Performing one-dimensional convolution on the multi-head self-attention normalized features, and outputting multi-head self-attention one-dimensional convolution features; Performing nonlinear mapping on the multi-head self-attention one-dimensional convolutional features to generate multi-head self-attention nonlinear features; Perform one-dimensional convolution on the multi-head self-attention nonlinear features and output the first self-attention feature.
7. The method for predicting the motion pattern of an intelligent agent in an open scene according to claim 5, characterized in that: The third multi-scale branch module includes a self-attention submodule and a self-attention distillation downsampling submodule; the second sub-feature is input into the third multi-scale branch module for multi-scale feature extraction to generate a third multi-scale feature, including: Using the second sub-feature as the input of the self-attention sub-module, and outputting the self-attention feature corresponding to the second sub-feature; Using the self-attention distillation downsampling submodule to distill and downsample the self-attention feature corresponding to the second sub-feature to generate a maximum pooling feature of the second sub-feature; The second sub-feature maximum pooling feature is used as the input of the self-attention sub-module, and the third multi-scale feature is output.
8. The method for predicting the motion pattern of an intelligent agent in an open scene according to claim 7, characterized in that: The self-attention distillation and downsampling submodule is used to distill and downsample the self-attention feature corresponding to the second sub-feature to generate a maximum pooling feature of the second sub-feature, including: Perform a one-dimensional convolution on the self-attention feature corresponding to the second sub-feature, and output a convolution feature of the second sub-feature; performing batch normalization on the second sub-feature convolutional features to generate second sub-feature normalized features; Downsampling the second sub-feature normalized feature, and outputting the second sub-feature downsampled feature; Perform maximum pooling on the second sub-feature downsampling feature to generate a second sub-feature maximum pooling feature.
9. The method for predicting the motion pattern of an intelligent agent in an open scene according to claim 1, characterized in that: The data processing process of the preset efficient autoregressive decoder is specifically as follows: A multi-layer perceptron is used to perform feature mapping on the input historical motion pattern sequence input to the preset efficient autoregressive decoder to generate high-dimensional features of historical motion; Splicing the historical motion high-dimensional features with the input edge features, the input map features, and the input future motion features input to the preset efficient autoregressive decoder to generate a spliced feature; Performing a global adaptive one-dimensional average pooling operation on the concatenated features to generate a global feature representation; Performing temporal excitation on the global feature representation and outputting temporal excitation features; Multiplying the time excitation feature and the concatenation feature element by element to generate a time series feature representation; A multi-layer perceptron is used to perform feature mapping on the time series feature representation, and output high-dimensional features of the time series; Performing layer normalization and nonlinear mapping on the high-dimensional features of the time series, and outputting nonlinear features of the time series; Adding the nonlinear characteristics of the time series and the historical moment prediction results element by element to generate the current moment prediction result; splicing the current moment prediction result and the target subsequence in the input historical motion pattern sequence to generate a new input historical motion pattern sequence; Performing time step embedding and position encoding on the new input historical motion pattern sequence, and outputting a coding feature sequence; Based on a single-layer masked multi-head probabilistic sparse self-attention mechanism, an attention operation is performed on the encoded feature sequence to output a probabilistic sparse self-attention; Add the probabilistic sparse self-attention and the encoding feature sequence element by element, and output a probabilistic sparse self-attention feature sequence; Performing layer normalization on the probabilistic sparse self-attention feature sequence to determine a query matrix; Performing linear projection on the input historical motion features input to the preset efficient autoregressive decoder to generate a key matrix and a value matrix; generating an efficient decoding feature sequence according to the query matrix, the key matrix and the value matrix; Performing layer normalization on the efficient decoding feature sequence, and outputting an efficient decoding normalized feature sequence; A multi-layer perceptron is used to perform feature mapping on the efficient decoding normalized feature sequence to generate an output prediction result.
10. An intelligent body motion pattern prediction device for open scenes, characterized in that: include: An acquisition module is used to acquire the historical motion pattern sequence of the intelligent agent, the future motion planning sequence of the intelligent agent, the monitoring video data set and the urban road map, and map encode the urban road map to generate a map feature vector; According to the module, it is used to generate a plurality of two-dimensional space position coordinates according to the original annotation file of the monitoring video data set, and perform homogeneous coordinate transformation and normalization on the plurality of the two-dimensional space position coordinates to generate a plurality of motion state information; An adopting module, used for generating a time dimension edge sequence according to a plurality of the motion state information and a plurality of the two-dimensional space position coordinates by adopting a preset space edge rule; An embedding module is used to embed the agent's historical motion pattern sequence, the agent's future motion planning sequence and the time dimension edge sequence respectively, to generate a historical motion time step embedding corresponding to the agent's historical motion pattern sequence, a future motion time step embedding corresponding to the agent's future motion planning sequence and an edge time step embedding corresponding to the time dimension edge sequence; an encoding module, configured to respectively encode the historical motion time step embedding, the future motion time step embedding and the edge time step embedding using a preset multi-scale history aggregation encoder, to generate multi-scale historical motion features corresponding to the historical motion time step embedding, multi-scale future motion features corresponding to the future motion time step embedding and multi-scale edge features corresponding to the edge time step embedding; A decoding module, configured to output an initial prediction result according to the multi-scale historical motion features, the multi-scale future motion features, the multi-scale edge features, the map feature vector and the agent historical motion pattern sequence through a plurality of preset efficient autoregressive decoders; The output result module is used to generate a target intelligent body motion mode prediction result based on the initial prediction result by using a preset binary Gaussian mixture model.
Citation Information
Patent Citations
Vehicle trajectory prediction method based on lane point future trajectory offset auxiliary supervision
CN116403176A
Multi-modal space-time model for accurate motion prediction based on visual fusion
CN117315603A
Vehicle trajectory prediction method and model based on improved Transform model and target point guidance, and electronic equipment
CN118823731A
Hand motion generation method and device for musical instrument playing and medium
CN119091905A
Intelligent agent action prediction method based on multi-scale spatial perception
CN119399570A