Method and device for predicting short-time target point of other vehicle through automatic driving

By conducting top3 intention predictions for traffic participants in the autonomous driving system for the next 2s, combining the encoder-decoder model and high-precision map, the problem of difficult to accurately predict short-term target points of traffic participants in the prior art is solved, and more accurate and efficient prediction results are achieved.

CN120067718APending Publication Date: 2025-05-30COWA TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510153498.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing autonomous driving prediction module is difficult to accurately predict the short-term target points of traffic participants in complex road conditions, resulting in untimely responses and may cause collisions. The long-term prediction line ignores short-term accuracy when ensuring the accuracy of the end point, resulting in inefficient traffic efficiency.

Method used

By predicting the top3 intentions of each traffic participant in the next 2s, switching to behavioral predictions, using the encoder-decoder structure to build a model, combining high-precision maps and traffic participants' characteristics and historical motion information, distributed training is carried out to generate short-term prediction points.

Benefits of technology

It realizes accurate prediction of short-term target points of traffic participants in complex road conditions, improves the response speed and accuracy of downstream planning control, and ensures safety and redundancy while taking into account traffic efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067718A_ABST
    Figure CN120067718A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for predicting a short-time target point of other vehicles through automatic driving, and the method comprises the steps: S1, data processing: carrying out the data cleaning and scene pre-recognition marking of original data, making an exclusive training data set of an intention network, and balancing the distribution of all intention scenes; s2, anchor point selection: clustering of goal points, namely target end points, is carried out on obstacles of each category, and fixed goal point distribution of each category is obtained; s3, anchor point classification: carrying out the automatic marking of a training data set based on the goal point distribution, i.e., the short-time behavior of each type of obstacle finally falls on a goal point distribution diagram; s4, model building: building a basic model by using an encoder-decoder structure; s5, model training: training a training data set by using a distributed training method, loading the goal point data as priori knowledge before training and taking the goal point data as an original inquiry matrix, and inquiring a final prediction result; and S6, model deployment: converting the trained model file into a static reasonable file.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving, and particularly to a method and device for predicting short-term target points of other vehicles in autonomous driving. Background Art

[0002] With the development of autonomous driving technology over the years, it has evolved from the initial simple lane-keeping function to the complex road condition urban automatic pilot function. In the complex road conditions of the city, the cooperation of multiple modules is particularly important. The upstream perception, as the eyes of the vehicle, provides all the vehicles, pedestrians, and obstacles around the host vehicle at the current moment. The downstream planning and control are the steering wheel, accelerator, and brakes, replacing humans to drive safely and smoothly. The prediction module, as the intermediate bridge between perception and planning and control, predicts the future behaviors of all traffic participants perceived and tells the planning and control the future drivable space and possible risks.

[0003] Currently, the autonomous driving prediction module generally predicts the trajectory points of traffic participants for the next 8 seconds. Each point represents 0.1 second, for a total of 80 points. There are mainly two methods for generating trajectory points. One is to generate them according to rules (see S. Pellegrini, A. Ess, and L. Van Gool, “Predicting pedestrian trajectories,” in Visual Analysis of Humans. Springer, 2011, pp. 473–491), that is, to generate a smooth curve using kinematic constraints based on the motion state of other vehicles and the lane information where they are located. The other is to use a neural network model to learn historical data and infer a smooth curve for other vehicles in real time (see J. Gao, C. Sun, H. Zhao, Y. Shen, D. Anguelov, C. Li, and C. Schmid, “Vectornet: Encoding hd maps and agent dynamics from vectorized representation,” in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 11525–11533; C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 652–660).

[0004] However, the existing prediction line has only one modality, that is, it only outputs a definite prediction line at each moment. Moreover, in order to fit the fixed 8-second length of the prediction line, the prediction line often has only one speed change trend, that is, it is always accelerating or always decelerating, which will cause the downstream to react in a timely manner. In order to fit the prediction line, the prediction lines at each moment are often smoothed, to ensure the overall rationality from the 0th second to the 8th second. However, the side effect of this is that short-term predictions will often deviate from the true prediction position due to this smoothing.

[0005] In the automatic driving in complex urban road conditions, the primary guarantee is safe driving, that is, the prediction of the future potential intentions of other traffic participants on the road should be timely and accurate. In the past, when using prediction lines, they often jumped continuously in relatively complex road conditions. Since the long-term prediction line needs to ensure the accuracy of the end point, it often ignores the short-term accuracy, which may lead to untimely responses in downstream planning and control and the possibility of collisions. While ensuring safety, it is also necessary to maintain traffic efficiency. For example, if all vehicles and pedestrians' future traffic predictions cannot be given in real time at a large crossroads during the evening rush hour, the vehicle itself will be stagnant, causing serious traffic jams. In such congested scenarios, the long-term prediction line often occupies the entire road, resulting in low traffic efficiency. Summary of the Invention

[0006] What the present invention aims to solve is safety redundancy and timely response. By predicting the top 3 intentions of each traffic participant in the next 2 seconds, the trajectory prediction is switched to behavior-level prediction, enabling the downstream to make more anthropomorphic judgments. It achieves safety with redundancy while taking into account traffic efficiency.

[0007] To achieve the above object, the technical solution of the present invention provides a method for predicting the short-term target points of other vehicles in automatic driving, which includes the following steps: S1 Data processing: By performing data cleaning and scene pre-identification and labeling on the original data, a dedicated training data set for the intention network is made to balance the distribution of each intention scene; S2 Anchor point selection: Cluster the goal points, that is, the target end points, for each category of obstacles to obtain the fixed goal point distribution for each category; S3 Anchor point classification: Automatically label the training data set based on the goal point distribution, that is, the short-term behaviors of each type of obstacle will ultimately fall on the goal point distribution map; S4 Model construction: Use the encoder-decoder structure to construct the basic model; S5 Model training: Use the distributed training method to train the training data set. Among them, the goal point data is loaded as prior knowledge before training and used as the original query matrix to query the final prediction result; S6 Model deployment: Convert the trained model file into a static inferable file.

[0008] Further, step S1 specifically includes: collecting operation data in the real vehicle's autonomous driving state, extracting data offline, and the target length of the extracted data is N frames, where the N-frame data includes historical frame data, current frame data, and future frame data, and the value range of N is 60-110 frames; cleaning all the collected data; performing scene recognition and labeling on the cleaned data, where the lane successor relationship of the high-precision map at the intersection is involved in scene judgment; after scene recognition, labeling each piece of data, and then performing data distribution statistics on all the labeled data, deleting the data with a high distribution according to the statistical results or collecting and supplementing the data with a low distribution, and finally achieving relative balance of the data.

[0009] Further, in step S4, encode the upstream information, calculate the attention for all the encoded elements and prior anchor points, and finally decode the calculation result into the intention point result, where the upstream information includes the characteristics of traffic participants and their historical movement information, and the high-precision map near the vehicle's position.

[0010] Further, in step S5, when training the model, the loss function uses the position deviation and classification loss. The position deviation is the distance between the true end position of the obstacle and the predicted intention point position, and the classification loss uses the winner-takes-all strategy, that is, the difference between the probability of the predicted intention point with the highest probability and the probability of the intention point closest to the true end point among all the predicted intention points.

[0011] Further, in step S6, based on the converted model, perform inference on all obstacles simultaneously and issue the results. The inference result is the short-term prediction point of the current obstacle.

[0012] The technical solution of the present invention also provides a device for predicting the short-term target point of other vehicles in autonomous driving, which includes the following modules: Data processing module: By cleaning the original data and performing pre-recognition and labeling of the scene, make a dedicated training dataset for the intention network and balance the distribution of each intention scene; Anchor point selection module: Cluster the goal points, that is, the target end points, for each category of obstacles to obtain the fixed goal point distribution of each category; Anchor point classification module: Automatically label the training dataset based on the goal point distribution, that is, the short-term behavior of each type of obstacle will finally fall on the goal point distribution map; Model construction module: Use the encoder-decoder structure to build the basic model; Model training module: Use the distributed training method to train the training dataset. Among them, load the goal point data as prior knowledge before training and use it as the original query matrix to query the final prediction result; Model deployment module: Convert the trained model file into a static inferable file.

[0013] Further, in the data processing module, operation data in the real vehicle's autonomous driving state is collected, and data extraction is performed offline. The target length of the extracted data is N frames, where the N-frame data includes historical frame data, current frame data, and future frame data, and the value range of N is 60 - 110 frames. All the collected data is subjected to data cleaning. The cleaned data is used for scene identification and labeling, where the lane successor relationship of the high-precision map at intersections is involved in scene judgment. After scene identification, each piece of data is labeled, and then the data distribution of all the labeled data is statistically analyzed. According to the statistical results, data with a relatively high distribution of labels is deleted, or data with a relatively low distribution of labels is collected and supplemented, ultimately achieving relative data balance.

[0014] Further, in the model construction module, upstream information is encoded, and attention calculations are performed on all the encoded elements and prior anchor points. Finally, the calculation results are decoded into intention point results, where the upstream information includes the characteristics of traffic participants and their historical movement information, as well as the high-precision map near the vehicle's position.

[0015] Further, in the model training module, when training the model, the loss function uses position deviation and classification loss. The position deviation is the distance between the true end position of the obstacle and the predicted intention point position, and the classification loss uses the winner-takes-all strategy, that is, the difference between the probability of the predicted intention point with the highest probability and the probability of the intention point closest to the true end point among all the predicted intention points.

[0016] Further, in the model deployment module, based on the converted model, all obstacles are simultaneously inferred and distributed. The inference result is the short-term prediction point of the current obstacle. Description of the Drawings

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is a flowchart of the method for predicting the short-term target point of other vehicles in autonomous driving of the present invention;

[0019] Figure 2 is a schematic diagram of the short-term prediction result based on intention of the present invention;

[0020] Figure 3 is a schematic diagram of the lane-changing scene including left lane change and right lane change of the present invention;

[0021] Figure 4 It is a schematic diagram of the distribution of goal points for three categories of motor vehicles, non-motor vehicles, and pedestrians in the present invention. Specific embodiments

[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.

[0023] For some special scenarios and obstacles, the results of trajectory prediction will have unclear problems (such as a vehicle just about to come out of a community, a pedestrian crossing the road). Most of these scenarios are due to less data, and the trajectory prediction model cannot fit well. Therefore, short-term intention prediction is proposed to provide a short-term prediction result based on intention for the downstream planning and control module.

[0024] Upstream input:

[0025] All obstacles are based on the historical state characteristics of map positioning and perception fusion, including but not limited to a series of features such as the position, speed, orientation angle, turn signal, traffic light, etc. of the obstacle that can be iteratively trained for intention in the forward direction.

[0026] Downstream output (as Figure 2 shown):

[0027] Motor vehicle: The position P(x, y) of the obstacle in the next 2s

[0028] Pedestrian: The position P(x, y) of the obstacle in the next 2s

[0029] Non-motor vehicle: The position P(x, y) of the obstacle in the next 2s

[0030] As Figure 1 shown, the method of the present invention mainly includes the following aspects:

[0031] Data processing:

[0032] Since intention training has higher requirements for the original data, by data cleaning, scene recognition and labeling, a dedicated data set for the intention network is made to balance the distribution of each intention scene, so that the scene distribution of the training data received by the model contains all the scenes, and the proportion of all scenes cannot be too small, otherwise the model may selectively ignore the scenes with too small proportions. At the same time, it should also conform to the approximate proportion of the scene distribution during actual vehicle operation.

[0033] First, collect the operation data of a real vehicle in the autonomous driving state and extract the data offline. The target length of the extracted data is N frames (60 - 110 frames). When it is, for example, 110 frames, it can include 29 frames of historical frame data, 1 frame of current frame data, and 80 frames of future frame data.

[0034] After the collection is completed, perform data cleaning on all the collected data. The cleaning criteria are as follows: within the lifespan of 34 frames of an obstacle, the inter-frame movement distance of a motor vehicle shall not exceed 3m, that of a non-motor vehicle shall not exceed 2m, and that of a pedestrian shall not exceed 1m. The inter-frame speed change of a pedestrian, motor vehicle, and non-motor vehicle shall not exceed 3m / s, and the inter-frame orientation change shall not exceed 0.5 radians.

[0035] See Figure 3 , for the cleaned data, perform scene recognition and labeling. For the intention network, scenarios such as intersections, non-intersections, motor vehicles, non-motor vehicles, pedestrians, low speed, medium speed, high speed, lane change, and straight are designed. Among them, for scenarios involving intersections, the lane successor relationship of the high-precision map is used to judge the scene.

[0036] After performing scene recognition on all the cleaned data and labeling each piece of data, one piece of data can contain several scene labels. Then, perform data distribution statistics on all the labeled data. According to the statistical results, delete the data with a high distribution of labels or collect and supplement the data with a low distribution, ultimately achieving relative balance of the data (which needs to be consistent with the specific business scenario, not an equal balance).

[0037] Anchor point selection:

[0038] Cluster the goal points, that is, the target end points, for each category of obstacles to obtain the fixed goal point distribution for each category.

[0039] Anchor point classification:

[0040] Based on this distribution, perform automatic annotation of the training data, that is, the short-term behavior of each category of obstacles will ultimately fall on this goal point distribution map, as Figure 4 shown.

[0041] Model construction:

[0042] The basic model uses an encoder-decoder structure to encode the upstream information including the characteristics of traffic participants and historical movement information, and all the road lines of the high-precision map near the location of the ego vehicle. Then, perform attention calculation on all the encoded elements and prior anchor points, and finally decode the calculation results into intention point results.

[0043] Model training:

[0044] Use the distributed training method to train the above 1 million cleaned and balanced training data. Among them, the goal point data of the three categories (motor vehicles, non-motor vehicles, and pedestrians) are loaded as prior knowledge before training and used as the original query matrix to query the final prediction results. This training process has a total of 30 rounds. The loss function uses the position deviation and classification loss. The position deviation is the L2 distance between the true end position of the obstacle and the predicted intention point position. The classification loss uses the winner-takes-all strategy, that is, the difference between the probability of the predicted intention point with the highest probability and the probability of the intention point closest to the true end point among all predicted intention points.

[0045] Model deployment:

[0046] Use libtorch as the deployment framework to convert the trained model file into a static inferable file.

[0047] According to the converted model, perform inference on all obstacles simultaneously. The inference result is the short-term prediction point of the current obstacle and is sent down.

[0048] In an embodiment of the present invention, a method for predicting the short-term target point of other vehicles in autonomous driving is provided, which includes the following steps: S1 Data processing: By performing data cleaning and scene pre-identification and labeling on the original data, a dedicated training dataset for the intention network is made to balance the distribution of each intention scene; S2 Anchor point selection: Cluster the goal points, that is, the target end points, for obstacles of each category to obtain the fixed goal point distribution of each category; S3 Anchor point classification: Automatically label the training dataset based on the goal point distribution, that is, the short-term behavior of obstacles of each type will finally fall on the goal point distribution map; S4 Model construction: Use the encoder-decoder structure to construct the basic model; S5 Model training: Use the distributed training method to train the training dataset. Among them, the goal point data is loaded as prior knowledge before training and used as the original query matrix to query the final prediction results; S6 Model deployment: Convert the trained model file into a static inferable file.

[0049] Further, step S1 specifically includes: collecting operation data in the real vehicle's autonomous driving state, extracting data offline, with the target length of the extracted data being N frames, where the N-frame data includes historical frame data, current frame data, and future frame data; cleaning all the collected data; performing scene recognition and labeling on the cleaned data, where the lane successor relationship of the high-precision map at intersections is involved in scene judgment; after scene recognition, labeling each piece of data, then statistically analyzing the data distribution of all labeled data, deleting data with a high distribution of labels or collecting and supplementing data with a low distribution of labels according to the statistical results, and finally achieving relative balance of the data.

[0050] Further, in step S4, encode the upstream information, calculate the attention for all encoded elements and prior anchor points, and finally decode the calculation result into the intention point result, where the upstream information includes the characteristics of traffic participants (motor vehicles, non-motor vehicles, pedestrians, ego vehicles, and general static obstacles) and their historical movement information, and the high-precision map near the ego vehicle's position (all road centerline and road boundary line information).

[0051] Further, in step S5, when training the model, the loss function uses the position deviation and classification loss. The position deviation is the distance between the true end position of the obstacle and the predicted intention point position, and the classification loss uses the winner-takes-all strategy, that is, the difference between the probability of the predicted intention point with the highest probability and the probability of the intention point closest to the true end point among all predicted intention points.

[0052] Further, in step S6, based on the converted model, perform inference on all obstacles simultaneously and issue the results. The inference result is the short-term prediction point of the current obstacle.

[0053] In an embodiment of the present invention, there is also provided a device for predicting the short-term target points of other vehicles in autonomous driving, which includes the following modules: Data processing module: By performing data cleaning and scene pre-identification and labeling on the original data, a dedicated training data set for the intention network is made to balance the distribution of each intention scene; Anchor point selection module: Cluster the goal points, i.e., the target end points, for obstacles of each category to obtain the fixed goal point distribution of each category; Anchor point classification module: Automatically label the training data set based on the goal point distribution, that is, the short-term behaviors of obstacles of each type will ultimately fall on the goal point distribution map; Model construction module: Use the encoder-decoder structure to construct the basic model; Model training module: Use the distributed training method to train the training data set. Among them, the goal point data is loaded as prior knowledge before training and used as the original query matrix to query the final prediction result; Model deployment module: Convert the trained model file into a static inferable file.

[0054] Further, in the data processing module, the operation data in the real vehicle autonomous driving state is collected, and data extraction is performed offline. The target length of the extracted data is N frames, where the N-frame data includes historical frame data, current frame data, and future frame data; All the collected data is subjected to data cleaning; The cleaned data is subjected to scene identification and labeling, and the lane successor relationship of the high-precision map is used to judge the scene when it comes to intersections; After scene identification, each piece of data is labeled, and then the data distribution of all the labeled data is statistically analyzed. According to the statistical results, the data with a high distribution label is deleted or the data with a low distribution label is collected and supplemented, and finally the relative balance of the data is achieved.

[0055] Further, in the model construction module, the upstream information is encoded, and attention calculation is performed on all the encoded elements and the prior anchor points. Finally, the calculation result is decoded into the intention point result, where the upstream information includes the characteristics of traffic participants (motor vehicles, non-motor vehicles, pedestrians, the ego vehicle, and general static obstacles) and their historical motion information, and the high-precision map near the ego vehicle position (all road centerline and road boundary line information).

[0056] Further, in the model training module, when training the model, the loss function uses the position deviation and the classification loss. The position deviation is the distance between the actual end position of the obstacle and the predicted intention point position, and the classification loss uses the winner-takes-all strategy, that is, the difference between the probability of the predicted intention point with the highest probability and the probability of the intention point closest to the actual end point among all the predicted intention points.

[0057] Furthermore, in the model deployment module, based on the converted model, inferences are performed on all obstacles simultaneously and the results are sent out. The inference result is the short-term prediction point of the current obstacle.

[0058] Advantageous technical effects of the present invention:

[0059] The intention prediction of traffic participants is the basis and an important input variable for downstream planning and control, and is used in obstacle judgment, screening, dynamic speed planning, and final decision-making. The data-driven intention point output helps the downstream decider module replace the original complex rule-based logical judgment. The original rule code was highly complex and increasingly difficult to update. The intention point model can directly provide the determined intention point and the corresponding probability, effectively reducing the rule code maintenance work of the decider module.

[0060] Since the intention model was launched, according to the feedback from downstream, the intention can effectively improve the situation that could not be judged by the original rules, such as unprotected left turns and obstacle vehicle ramp merges. Especially in intersection scenarios, the intention has shown great advantages, and there has been no case of obvious deterioration since its use, that is, a smooth transition from rules to models and an improvement in benefits in some scenarios have been achieved.

[0061] The data-driven intention judgment is more friendly for iterative updates. The intention model can be updated by batch cleaning and processing of the latest data, and this process is imperceptible to downstream users.

[0062] Based on case-by-case data, the intention point has an intention determination situation no worse than the current internal cut-in rule of the decision-making. Using this information, a relatively stable interactive intention determination can be given, which can be used as a replacement for the current rule.

[0063] Existing statistical data shows that the optimized case ratio is 97.2%, the difficult-to-process ratio is 2.8%, and the deterioration ratio is 0%.

[0064] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent replacements or changes, should be covered by the protection scope of the present invention.

Claims

1. A method for predicting short-term target points of other vehicles by automatic driving, characterized in that: The steps include: S1 Data Processing: By cleaning the raw data and pre-identifying and labeling the scenes, we can create a dedicated training data set for the intent network and balance the distribution of various intent scenes. S2 Anchor point selection: cluster the goal points (i.e., target endpoints) of each category of obstacles to obtain a fixed goal point distribution for each category; S3 anchor point classification: Automatically annotate the training data set based on the goal point distribution, that is, the short-term behavior of each category of obstacles will eventually fall on the goal point distribution map; S4 model building: Use the encoder-decoder structure to build a basic model; S5 model training: Use the distributed training method to train the training data set, where the goal point data is loaded as prior knowledge before training and used as the original query matrix to query the final prediction results; S6 model deployment: Convert the trained model file into a static inferable file.

2. The method according to claim 1, characterized in that Step S1 specifically includes: For the operational data collection of the real vehicle in the state of autonomous driving, data extraction is performed offline, and the target length of the extracted data is N frames, where the N frames of data include historical frame data, current frame data, and future frame data; Perform data cleaning on all collected data; The cleaned data is used for scene recognition and labeling, which involves using the lane succession relationship of high-precision maps at intersections for scene judgment; After scene recognition, each piece of data is labeled, and then data distribution statistics are performed on all labeled data. Based on the statistical results, data of labels with high distribution is deleted or data of labels with low distribution is collected and supplemented, ultimately achieving a relative balance of data.

3. The method according to claim 2, characterized in that In step S4, the upstream information is encoded, and attention calculation is performed on all encoded elements and prior anchor points. Finally, the calculation results are decoded into intention point results. The upstream information includes the characteristics of traffic participants and their historical movement information, and a high-precision map near the vehicle's position.

4. The method according to claim 3, characterized in that: In step S5, when training the model, the loss function uses position deviation and classification loss. The position deviation is the distance between the actual end position of the obstacle and the predicted intention point position. The classification loss uses a winner-takes-all strategy, that is, the difference between the probability of the predicted intention point with the highest probability and the probability of the intention point closest to the actual end point among all predicted intention points.

5. The method according to claim 4, characterized in that In step S6, all obstacles are inferred and issued simultaneously according to the converted model, and the inference result is the short-term prediction point of the current obstacle.

6. A device for predicting short-term target points of other vehicles by automatic driving, characterized in that: Includes the following modules: Data processing module: This module cleans the raw data and pre-identifies and labels the scenes to create a dedicated training data set for the intent network and balance the distribution of various intent scenes. Anchor point selection module: cluster the goal points of each category of obstacles, i.e. the target end points, to obtain the fixed goal point distribution of each category; Anchor point classification module: Automatically annotate the training data set based on the goal point distribution, that is, the short-term behavior of each category of obstacles will eventually fall on the goal point distribution map; Model building module: Use the encoder-decoder structure to build a basic model; Model training module: Use the distributed training method to train the training data set, where the goal point data is loaded as prior knowledge before training and used as the original query matrix to query the final prediction results; Model deployment module: converts the trained model files into static inferable files.

7. The device according to claim 6, characterized in that In the data processing module, For the operational data collection of the real vehicle in the state of autonomous driving, data extraction is performed offline, and the target length of the extracted data is N frames, where the N frames of data include historical frame data, current frame data, and future frame data; Perform data cleaning on all collected data; The cleaned data is used for scene recognition and labeling, which involves using the lane succession relationship of high-precision maps at intersections for scene judgment; After scene recognition, each piece of data is labeled, and then data distribution statistics are performed on all labeled data. Based on the statistical results, data of labels with high distribution is deleted or data of labels with low distribution is collected and supplemented, ultimately achieving a relative balance of data.

8. The device according to claim 7, characterized in that In the model building module, the upstream information is encoded, and attention calculation is performed on all encoded elements and prior anchor points. Finally, the calculation results are decoded into intention point results. The upstream information includes the characteristics of traffic participants, their historical movement information, and high-precision maps near the vehicle's location.

9. The device according to claim 8, characterized in that In the model training module, when training the model, the loss function uses position deviation and classification loss. The position deviation is the distance between the actual end position of the obstacle and the predicted intention point position. The classification loss uses a winner-takes-all strategy, that is, the difference between the probability of the predicted intention point with the highest probability and the probability of the intention point closest to the actual end point among all predicted intention points.

10. The device according to claim 9, characterized in that In the model deployment module, all obstacles are inferred and issued simultaneously according to the converted model, and the inference result is the short-term prediction point of the current obstacle.