Semantic enhancement pre-training method for intention understanding
Through fine-grained masking strategy and SimSiam comparison learning, the prediction ability of the intention-understanding model in the long-tail sample scenario is improved, and the problem of insufficient semantic consistency learning of long-tail samples is solved, and more efficient intention prediction is achieved.
Patent Information
- Application Number
- CN202510527673.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
AI Technical Summary
The existing intent-to-understand model performs poorly in long-tail sample scenarios and relies on large-scale annotation of data, resulting in insufficient prediction capabilities in complex scenarios, especially in insufficient semantic consistency learning of long-tail samples.
The semantic enhancement pre-training framework is adopted to improve the feature learning ability of long-tail samples through fine-grained sequence reconstruction tasks and coarse-grained intention comparison tasks. Combined with SimSiam's stop gradient mechanism, masking strategies and multimodal future decoder are designed to improve the model's feature learning ability of long-tail samples.
The model's intention prediction performance in different scenarios is improved, especially in long-tail scenarios, and the feature learning and prediction accuracy of long-tail samples are enhanced.
Smart Images

Figure CN120449973A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computers, and more particularly, to a semantic enhancement pre-training method for intent understanding. Background Art
[0002] Trajectory encoding is used to learn the characteristics of the agent's historical trajectory. Drawing on time series modeling methods such as natural language processing, recurrent neural networks
[24] (RNN) and its variants, long short-term memory networks (LSTM), gated recurrent units (GRU), etc. are often used to extract the time domain characteristics of trajectory sequences. Usually, each agent's historical trajectory is processed separately using a recurrent neural network or a one-dimensional convolutional neural network to extract its trajectory representation. Some scholars
[30] form a holistic representation of the agent trajectory in the scene, use a two-dimensional convolutional neural network to directly extract the deep features of the scene as a whole, learn the spatiotemporal correlation of individual movements, and then obtain the corresponding representation of each individual by indexing the overall representation of the scene.
[0003] Trajectory decoding focuses on how to generate the agent's final trajectory. While the use of models such as recurrent neural networks (RNNs) mentioned above focuses on improving the model's ability to represent input data, trajectory decoding research focuses on aspects such as output data representation, multimodal trajectory output, and loss function design. Output representations can be categorized into two types based on the type of output trajectory: output trajectory sequence and output trajectory probability distribution. The output trajectory sequence simply and intuitively utilizes linear transformations or multilayer perceptrons (MLPs) to directly convert the intermediate representations extracted by the network into a future trajectory sequence, or simply converts the trajectory point positions at the next moment and uses recursion to continuously derive future trajectories. Loss functions commonly used include mean squared error (MSE) and smoothL1. Other methods convert trajectory point sequences in Euclidean space into other spaces, such as converting trajectory data to the Frenet coordinate system. One method for outputting trajectory probability distributions is to construct a parameterized trajectory distribution, utilizing the parameters of the network output distribution. Another type of method outputs a non-parametric trajectory distribution. Generally, a fully convolutional network is used to output a heat map of the same scale as the current scene, indicating the distribution of future position coordinates in the scene. Alternatively, the possible trajectory endpoints are densely pre-set in the scene space and refined through secondary evaluation, clustering, and other methods to obtain a more accurate trajectory output.
[0004] Due to the influence of factors such as individual will, future trajectories are difficult to accurately infer solely from historical trajectories. Compared to assuming a single-mode trajectory output, a range of methods consider constructing multimodal future trajectory distributions to improve the trajectory prediction model's ability to describe future trajectory uncertainty. Most of these methods rely on the introduction of additional variables. A fully data-driven approach does not incorporate any semantic information into these variables. For example, they use mixture density networks to model future trajectory distributions and construct multimodal distributions of future trajectories based on mixture Gaussian models. Alternatively, they train multiple trajectory decoders simultaneously to output multiple trajectories and then score or rank the quality of these multiple predictions. Other methods capture the multimodality of the data by introducing perturbations such as noise, employing methods such as conditional variational autoencoders (CVAEs) and generative adversarial networks (GANs). Since the true value of a trajectory sample corresponds to multiple predicted trajectories output by the network during each optimization, these methods often require optimization strategies to ensure diversity in the predicted trajectories, such as the Winner-Takes-All (WTA) loss function. Other studies introduce semantic information for such variables in a top-down manner. These variables can be described from small to large according to the semantic scale: the short-term "actions" that the predicted individual may take, the medium- to long-term "decisions" or "intentions", and the long-term "behavioral preferences".
[0005] Current mainstream approaches to intention understanding are primarily based on supervised learning, with their core architecture typically employing an encoder-decoder paradigm. This paradigm first fuses the target agent's dynamic state information with environmental contextual information. The encoder module then extracts spatiotemporal and interaction features, ultimately outputting a multimodal distribution of the target agent's future potential intentions. This architectural design fully considers the complex spatiotemporal dependencies and interactions in traffic scenarios, effectively capturing the uncertainty and multi-possibility of driving behavior. In encoder design, researchers have primarily employed three representative approaches: methods based on convolutional neural networks (CNNs) excel at processing regular, gridded environmental representations; methods based on graph neural networks (GNNs) are more suitable for modeling dynamic interactions; and attention mechanisms (particularly self-attention modules) demonstrate unique advantages in modeling long-range dependencies. Representative work, such as VectorNet, innovatively employs a layered graph network architecture, first locally encoding individual traffic elements and then modeling high-order interactions between them through a global graph network. LaneGCN, on the other hand, specifically designs a graph convolution operator for lane networks, accurately encoding complex lane topologies through a multi-hop message passing mechanism. Decoder design is primarily categorized into two paradigms: anchor-based and anchor-free. Anchor-based methods (such as MultiPath and TNT) rely on prior knowledge to predefine a set of typical motion pattern anchors, predicting trajectory distributions through a dual-branch classification + regression architecture. Their advantage lies in their strong physical interpretability. Anchor-free methods (such as DenseTNT and MotionTransformer) use an end-to-end approach to directly predict future trajectory points from the original feature space, avoiding the bias inherent in anchor design but requiring a larger amount of training data. Recent research trends indicate that the two paradigms are mutually integrative, with hybrid approaches such as learnable anchors or probabilistic anchors being used to balance model performance and generalization.
[0006] These methods face two major challenges in practical applications. First, their performance relies heavily on supervised training using large amounts of labeled data, which not only increases data collection and annotation costs but also limits the model's applicability in scenarios where labels are scarce. Second, because training data is often concentrated on common scenarios, the model's transferability to unseen scenarios or cross-domain data is poor, resulting in significantly reduced generalization performance in complex real-world environments. For example, when transferring knowledge from understanding autonomous driving to understanding sports, the discrepancy between driving behavior and sports rules often significantly reduces the model's ability to predict or understand individual intentions. We further analyzed long-tail samples (i.e., low-frequency, edge cases) and found significant differences between their historical intentions (past behavior patterns) and future intentions (actual subsequent behavior). This dynamic variation makes it difficult for traditional models to capture their global patterns, resulting in significant errors in understanding individual behavioral intentions. This phenomenon reveals a key flaw in mainstream models: their over-reliance on the distribution of head samples during training leads to insufficient learning of the semantic consistency of long-tail samples (i.e., the correlation between past behavior and future intentions). Specifically, the model tends to force long-tail samples into common patterns, ignoring their unique dynamic characteristics. This bias not only reduces prediction accuracy but can also lead to safety risks in real-world applications (such as misjudging emergency lane changes or unexpected obstacles, and misinterpreting user intent).
[0007] In intent understanding pre-training tasks, the long-tail sample problem is prevalent. This refers to a small number of trajectory sequence points that are temporally and spatially distributed far from the mainstream data. Existing methods often ignore or improperly handle these long-tail samples, resulting in poor model performance on small numbers of samples. Traditional pre-training methods typically select 30% of individuals and mask their entire history, while masking the entire future trajectory sequences of the remaining 70%. While simple, this approach has significant limitations:
[0008] (1) Information loss: Blocking all historical or future trajectories at once causes the model to lose a large amount of contextual information, affecting the learning of individual behavior patterns.
[0009] (2) Long-tail sample neglect: Long-tail samples are more likely to be neglected in this rough masking strategy due to their small amount of data, and it is difficult for the model to effectively learn the characteristics of these samples. Summary of the Invention
[0010] The purpose of the embodiments of the present disclosure is to provide a semantic enhancement pre-training method for intent understanding. The present invention aims to address the shortcomings of existing supervised learning-based models in intent prediction when faced with a large amount of unlabeled data, and the problem that the MAE model based on self-supervised learning is insufficient in predicting the subject's intention / behavior in complex scenarios because it does not consider the semantic consistency of the subject's intention / behavior when dealing with long-tail problems.
[0011] In general, a semantic enhancement pre-training method for intent understanding is provided, including a semantic enhancement pre-training framework and an intent understanding fine-tuning framework, for training a trajectory prediction intent understanding model for a model that processes trajectory prediction information:
[0012] The trajectory prediction information model first needs to rotate and translate the entire scene element around the predicted target;
[0013] The semantic enhancement pre-training framework first needs to design a fine-grained sequence reconstruction task and a coarse-grained intent comparison task;
[0014] The fine-grained sequence reconstruction task adopts a masking strategy in the time dimension. It randomly selects specific time steps for masking in the historical and future intention sequence information of the object's motion trajectory to retain more historical and future patterns. The masked results are processed using a standard Transformer block as an encoder. A complete token set is generated by concatenating a learnable initialization matrix. The masked state is predicted by the cross-self decoder and the fine-grained reconstruction loss is calculated.
[0015] The coarse-grained intent comparison task adds a similarity-based loss and uses SimSiam's stop-gradient mechanism to allow the model to learn the relationship between historical and future behavior patterns. The pre-training loss consists of a fine-grained reconstruction loss and a coarse-grained comparison loss.
[0016] The intent understanding fine-tuning framework uses a pre-trained semantic enhancement encoder to initialize the encoder of the prediction network, uses a multimodal future decoder to generate the multimodal predicted intent of the target subject, and uses an independent prediction network to generate a confidence score; the fine-tuning loss is decomposed into trajectory regression loss and confidence classification loss, and the Winner-Takes-All strategy is used to optimize the downstream fine-tuning task loss.
[0017] The specific method of randomly selecting a specific time step for masking in the historical and future intention sequence information is:
[0018]
[0019] One second consists of 10 time steps, where the meaning of each variable is: Represents the historical information of the predicted target, Represents the historical information of surrounding participants, Represents the future information of the predicted target, represents the future information of surrounding participants, Represents visible historical information, Represents visible future information.
[0020] The specific method of generating a complete token set by connecting the learnable initialization matrix is: the output of the encoder is connected to the learnable initialization matrix X' hmask and X' fmask Concatenated, this matrix has the same shape as the mask tokens:
[0021]
[0022] This process generates two complete sets of tokens X i h’ and X i f’ , corresponding to the original input history future trajectory sequence; then, they are fed into the cross self decoder to predict the values of the masked history state and future state respectively, where the meaning of each variable is: Represents the visible historical representation of the predicted target after passing through the encoder, Represents the visible future representation of the predicted target after passing through the encoder.
[0023] The fine-grained reconstruction loss is as follows:
[0024]
[0025] The meaning of each variable is: represents the historical representation of the initial visible target, Represents the historical representation of the predicted target after passing through the autoencoder, represents the future representation of the initially visible target being predicted, Represents the future representation of the predicted target after passing through the autoencoder.
[0026] The coarse-grained intent comparison task is implemented as follows: first, a prediction function is performed using the encoder output:
[0027]
[0028] Then minimize the negative cosine similarity between past and future representations:
[0029]
[0030] Where ‖·‖ is l2 normalization;
[0031] Finally, a stop gradient mechanism is added to the similarity to form a coarse-grained contrastive loss:
[0032]
[0033] The meaning of each variable is: Represents the visible historical representation of the predicted target after passing through the encoder, Represents the future representation of the predicted target after the prediction function, Represents the visible future representation of the predicted target after passing through the encoder, Represents the historical representation of the predicted target after passing through the prediction function.
[0034] The specific implementation of the intent understanding fine-tuning framework is as follows: input the historical trajectory of the object in the current scene:
[0035]
[0036] A multimodal future decoder Dtrato is used to generate the predicted multimodal trajectory of the target agent. This decoder is implemented using MLP:
[0037]
[0038] To evaluate the reliability of each predicted trajectory, a separate prediction network is used to generate a confidence score:
[0039]
[0040] Decompose the loss into separate trajectory regression and confidence classification losses with equal weights:
[0041]
[0042] A winner-takes-all strategy is also used to optimize the loss of downstream fine-tuning tasks to ensure the diversity of results and further optimize the model:
[0043]
[0044] The meaning of each variable is: Expressing true future intentions, Represents the best mode among the predicted K modes.
[0045] The technical effects to be achieved by the embodiments of the present invention are:
[0046] A semantically enhanced trajectory pre-training framework is proposed to improve the model's prediction performance of the subject's future intentions in different scenarios, especially in long-tail scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The above and other objects and features of the present disclosure will become more apparent from the following description in conjunction with the accompanying drawings.
[0048] Figure 1 1 is a schematic diagram illustrating an architecture diagram of a semantic enhancement pre-training method for intent understanding according to an embodiment of the present disclosure;
[0049] Figure 2 is a schematic diagram illustrating an intent understanding fine-tuning framework according to an embodiment of the present disclosure;
[0050] Figure 3 is a schematic diagram showing quantitative results according to an embodiment of the present disclosure;
[0051] Figure 4 is a schematic diagram showing qualitative results according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0052] The following detailed description is provided to help the reader gain a comprehensive understanding of the methods, devices and / or systems described herein. However, various changes, modifications and equivalents of the methods, devices and / or systems described herein will be clear after understanding the disclosure of the present application. For example, the order of operations described herein is merely an example and is not limited to those orders set forth herein, but can be changed as will be clear after understanding the disclosure of the present application, except for operations that must occur in a specific order. In addition, for greater clarity and conciseness, descriptions of features known in the art may be omitted.
[0053] The features described herein can be implemented in different forms and should not be construed as limited to the examples described herein. Rather, the examples described herein are provided to illustrate only some of the many possible ways to implement the methods, devices, and / or systems described herein, which will become clear after understanding the disclosure of this application.
[0054] As used herein, the term "and / or" includes any one of the associated listed items and any combination of any two or more.
[0055] Although terms such as "first," "second," and "third" may be used herein to describe various members, components, regions, layers, or portions, these members, components, regions, layers, or portions should not be limited by these terms. Instead, these terms are used solely to distinguish one member, component, region, layer, or portion from another member, component, region, layer, or portion. Thus, what is referred to as a first member, first component, first region, first layer, or first portion in the examples described herein may also be referred to as a second member, second component, second region, second layer, or second portion without departing from the teachings of the examples.
[0056] In the specification, when an element (such as a layer, region, or substrate) is described as being “on,” “connected to,” or “coupled to” another element, the element may be directly “on,” “connected to,” or “coupled to” the other element, or one or more other elements may be present therebetween. Conversely, when an element is described as being “directly on,” “directly connected to,” or “directly coupled to” another element, there may be no other elements present therebetween.
[0057] The terms used herein are intended only to describe various examples and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular is intended to include the plural. The terms "comprise," "include," and "have" indicate the presence of the recited features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.
[0058] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure pertains after understanding the present disclosure. Unless expressly defined otherwise herein, terms (such as those defined in general dictionaries) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and should not be interpreted in an idealized or overly formal manner.
[0059] Furthermore, in describing the examples, when it is deemed that a detailed description of well-known related structures or functions would cause ambiguous interpretation of the present disclosure, such detailed description will be omitted.
[0060] Figure 1 is a schematic diagram illustrating a semantic enhancement pre-training method for intent understanding according to an embodiment of the present disclosure.
[0061] In order to achieve the above-mentioned purpose of the invention, the technical framework adopted by the present invention is as follows Figure 1 shown.
[0062] Randomly select some moments and mask the trajectory points at those moments. Compared to traditional methods, this fine-grained masking strategy preserves more contextual information, helping the model better understand and model temporal dependencies in trajectory data during pre-training. Especially for long-tail samples, the fine-grained masking strategy helps retain more effective information and improve learning performance for these few samples. Furthermore, we introduce the SimSiam method to further enhance the model's feature learning capabilities. SimSiam has excellent performance in contrastive learning, and we apply it to historical and future trajectory sequences that have been randomly masked over time. This allows the model to more effectively capture dynamic changes in time series data, enhance feature learning for long-tail samples, and achieve superior performance in downstream tasks. Overall, our method, through a fine-grained temporal masking strategy and SimSiam contrastive learning, not only addresses the information loss issue of existing methods but also enhances the model's representation of intent data. Especially for long-tail samples, these improvements help preserve and effectively utilize more effective information, enabling the model to better learn the features of these few samples and provide more powerful feature support for downstream tasks.
[0063] The present invention mainly includes the following two frameworks:
[0064] (1) Semantic enhancement pre-training framework, such as Figure 1 As shown in the figure, a fine-grained sequence reconstruction task and a coarse-grained intent comparison task are designed. Fine-grained reconstruction learning adopts a masking strategy in the time dimension, randomly selecting specific time steps in the historical and future intent sequence information for masking, preserving more historical and future patterns. The masked results are processed using a standard Transformer block as an encoder, and a complete token set is generated by connecting a learnable initialization matrix. The masked state is predicted by the cross-self-decoder, and the fine-grained reconstruction loss is calculated. Coarse-grained contrastive learning designs a comparison task based on the historical and future levels. By adding a similarity-based loss and utilizing SimSiam's stopping gradient mechanism, the model learns the relationship between historical and future behavior patterns. The pre-training loss consists of the fine-grained reconstruction loss and the coarse-grained contrastive loss.
[0065] (2) Intention understanding fine-tuning framework, such as Figure 2 As shown, the prediction network's encoder is initialized with a pretrained semantic enhancement encoder. A multimodal future decoder is used to generate the target subject's multimodal predicted intent. A separate prediction network is then used to generate confidence scores. The fine-tuning loss is decomposed into a trajectory regression loss (Huber loss) and a confidence classification loss (cross-entropy loss). A winner-takes-all strategy is used to optimize the downstream fine-tuning task loss. Ultimately, this improves the model's intent understanding capabilities.
[0066] Experimental verification:
[0067] (1) Quantitative results
[0068] like Figure 3 As shown, in an evaluation on the Argoverse 2 (AV2) dataset, our model achieves the best performance in both minADE and minFDE metrics in a map-free setting. In a map-based setting, minADE increases by only 0.04 and minFDE by only 0.10 compared to the state-of-the-art model, demonstrating that it can learn more robust representations of agent behavior intentions. Compared to a model trained from scratch, our model improves minADE by approximately 4.5% and minFDE by approximately 5.1%. For long-tail scene samples, our method significantly improves prediction performance compared to the baseline model without sacrificing overall performance. For example, on the top 1% of long-tail samples, minFDE decreases from 13.03 to 11.75.
[0069] (2) Qualitative results:
[0070] The multi-scene visualization results are shown in the figure Figure 4 Figure 2: Qualitative results on the AV2 validation set, showing trajectory predictions for the target subject (orange square) and surrounding subjects (blue squares). Green arrows indicate true values, red arrows indicate multimodal predictions (K=3), and color shading indicates probability. These results cover turning, straight driving, roadside parking, and lane changing scenarios.
[0071] While some embodiments of the present disclosure have been shown and described, it will be appreciated by those skilled in the art that changes may be made to these embodiments without departing from the principles and spirit of the disclosure, the scope of which is defined by the claims and their equivalents.
Claims
1. A semantic enhancement pre-training method for intent understanding, characterized in that: It includes a semantic enhancement pre-training framework and an intent understanding fine-tuning framework, which trains the trajectory prediction intent understanding model for models that process trajectory prediction information; The trajectory prediction information model first needs to rotate and translate the entire scene element around the predicted target; The semantic enhancement pre-training framework first needs to design a fine-grained sequence reconstruction task and a coarse-grained intent comparison task; The fine-grained sequence reconstruction task adopts a masking strategy in the time dimension. It randomly selects specific time steps for masking in the historical and future intention sequence information of the object's motion trajectory to retain more historical and future patterns. The masked results are processed using a standard Transformer block as an encoder. A complete token set is generated by concatenating a learnable initialization matrix. The masked state is predicted by the cross-self decoder and the fine-grained reconstruction loss is calculated. The coarse-grained intent comparison task adds a similarity-based loss and uses SimSiam's stop-gradient mechanism to allow the model to learn the relationship between historical and future behavior patterns. The pre-training loss consists of a fine-grained reconstruction loss and a coarse-grained comparison loss. The intent understanding fine-tuning framework uses a pre-trained semantic enhancement encoder to initialize the encoder of the prediction network, uses a multimodal future decoder to generate the multimodal predicted intent of the target subject, and uses an independent prediction network to generate a confidence score; the fine-tuning loss is decomposed into trajectory regression loss and confidence classification loss, and the Winner-Takes-All strategy is used to optimize the downstream fine-tuning task loss.
2. A semantic enhancement pre-training method for intent understanding according to claim 1, characterized in that: The specific method of randomly selecting a specific time step for masking in the historical and future intention sequence information is: One second consists of 10 time steps, where the meaning of each variable is: Represents the historical information of the predicted target, Represents the historical information of surrounding participants, Represents the future information of the predicted target, represents the future information of surrounding participants, Represents visible historical information, Represents visible future information.
3. A semantic enhancement pre-training method for intent understanding according to claim 2, characterized in that: The specific method of generating a complete token set by connecting the learnable initialization matrix is: the output of the encoder is connected to the learnable initialization matrix X' hmask and X' fmask Concatenated, this matrix has the same shape as the mask tokens: This process generates two complete sets of tokens X i h’ and X i f’ , corresponding to the original input history future trajectory sequence; then, they are fed into the cross self decoder to predict the values of the masked history state and future state respectively, where the meaning of each variable is: Represents the visible historical representation of the predicted target after passing through the encoder, Represents the visible future representation of the predicted target after passing through the encoder.
4. A semantic enhancement pre-training method for intent understanding according to claim 3, characterized in that: The fine-grained reconstruction loss is as follows: The meaning of each variable is: represents the historical representation of the initial visible target, Represents the historical representation of the predicted target after passing through the autoencoder, represents the future representation of the initially visible target being predicted, Represents the future representation of the predicted target after passing through the autoencoder.
5. A semantic enhancement pre-training method for intent understanding according to claim 4, characterized in that: The coarse-grained intent comparison task is implemented as follows: first, a prediction function is performed using the encoder output: Then minimize the negative cosine similarity between past and future representations: Where ‖·‖ is l2 normalization; Finally, a stop gradient mechanism is added to the similarity to form a coarse-grained contrastive loss: The meaning of each variable is: Represents the visible historical representation of the predicted target after passing through the encoder, Represents the future representation of the predicted target after the prediction function, Represents the visible future representation of the predicted target after passing through the encoder, Represents the historical representation of the predicted target after passing through the prediction function.
6. A semantic enhancement pre-training method for intent understanding according to claim 5, characterized in that: The specific implementation of the intent understanding fine-tuning framework is as follows: input the historical trajectory of the object in the current scene: A multimodal future decoder Dtrato is used to generate the predicted multimodal trajectory of the target agent. This decoder is implemented using MLP: To evaluate the reliability of each predicted trajectory, a separate prediction network is used to generate a confidence score: Decompose the loss into separate trajectory regression and confidence classification losses with equal weights: A winner-takes-all strategy is also used to optimize the loss of downstream fine-tuning tasks to ensure the diversity of results and further optimize the model: The meaning of each variable is: Expressing true future intentions, Represents the best mode among the predicted K modes.
Citation Information
Cited By
End-to-end automatic driving long tail scene data acquisition and automatic labeling method
CN121614891A
Power battery life prediction method based on mechanism perception sparse coding and WTA prediction
CN122286689A