A prediction method, device, intelligent driving system and vehicle

By combining multi-frame 3D point cloud images and high-precision map features, the prediction method solves the problem of insufficient accuracy in pedestrian trajectory prediction in existing technologies, and achieves higher accuracy in pedestrian trajectory prediction.

CN115273015BActive Publication Date: 2026-03-20HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-30
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing target trajectory and map-based prediction algorithms have insufficient prediction accuracy in new environments, while visual image-based prediction algorithms are susceptible to noise interference and have low accuracy in predicting pedestrian movement trajectories.

Method used

By acquiring multiple frames of 3D point cloud images, target detection and feature extraction are performed. A neural network model is used to predict the movement trajectory of pedestrians. The trajectory prediction is combined with high-precision map features to enhance environmental features and interactive information.

Benefits of technology

It improves the accuracy of pedestrian trajectory prediction, with an average distance error of 0.25m within three seconds, and enhances ranging accuracy and light robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115273015B_ABST
    Figure CN115273015B_ABST
Patent Text Reader

Abstract

The application provides a prediction method and device and a vehicle, and relates to the technical field of intelligent driving. The method comprises the following steps: obtaining multiple images and a high-precision map, processing the multiple images to obtain the features of each image, extracting the space-time features and interaction features of pedestrians according to the features of each image, obtaining more environment features and increasing the interaction information between the pedestrians and the surrounding environment, so that the subsequent prediction of the motion trajectory of the pedestrians is more accurate, and then extracting the map features by using the high-precision map, so that the pedestrian trajectory predicted by combining the space-time features of the pedestrians and the interaction features of the pedestrians is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent driving, and in particular to a prediction method and device, an intelligent driving system and a vehicle. BACKGROUND

[0002] With the development and popularization of intelligence, intelligent driving of vehicles has become a popular research direction. Intelligent driving systems can be divided into four key functional modules, namely positioning, environment perception, path planning and decision control, according to functional requirements. Among them, prediction functions such as predicting the road to be driven by the vehicle and the motion trajectory of pedestrians are mainly concentrated in the environment perception module. The prediction algorithms currently used to realize the prediction function include prediction based on target trajectory and map, prediction based on visual images, and so on.

[0003] For the existing algorithms for predicting based on the historical trajectory of the target and the map, the spatial coordinate points of the historical trajectory are used to predict the future trajectory. A large amount of historical data is needed to support the implementation of the prediction function of this algorithm, and if the vehicle is used for the first time, the location of the vehicle is a completely new environment, and so on, the prediction result of this algorithm will be greatly discounted. For the algorithm for predicting based on visual images, due to the lack of depth information in the captured images, the strong maneuverability, low speed and small target of pedestrians, and other defects, the tracking information of pedestrians generated according to the images is easily disturbed by noise, so the accuracy of the predicted motion trajectory of pedestrians is relatively low. Therefore, how to improve the accuracy of predicting the trajectory of vehicles or pedestrians is a problem that needs to be solved at present. SUMMARY

[0004] In order to solve the above problems, the embodiments of the present application provide a prediction method, device, intelligent driving system and vehicle.

[0005] In a first aspect, the application provides a trajectory prediction method, characterized in that: obtaining at least two frames of three-dimensional point cloud images, the at least two frames of three-dimensional point cloud images each comprising a first target and a second target, the at least two frames of three-dimensional point cloud images being three-dimensional point cloud images obtained after coordinate unification; performing target detection on the at least two frames of three-dimensional point cloud images to obtain respective feature maps corresponding to the at least two frames of three-dimensional point cloud images; extracting a position feature and a dynamic feature of the first target according to the feature maps, the position feature comprising position information of the first target in the feature maps, and the dynamic feature comprising a corresponding first region feature in the feature maps, the first region feature being determined according to the position information of the first target in the feature maps; determining an interaction feature of the first target by inputting the position feature and the dynamic feature of the first target and the position feature and the dynamic feature of the second target into a neural network model; and predicting a motion trajectory of the first target according to the position feature and the dynamic feature of the first target, the interaction feature of the first target, and a map feature of the first target, the map feature of the first target being obtained by encoding a stored map within a set range of a current position of the first target.

[0006] In this embodiment, the first target is taken as an example of a pedestrian. By obtaining multiple frames of images and a high-precision map, the multiple frames of images are processed to obtain features of each frame of image, and then the spatiotemporal feature and the interaction feature of the pedestrian are extracted according to the features of each frame of image, so as to obtain more environmental features and increase the interaction information between the pedestrian and the surrounding environment, thereby making the subsequent prediction of the motion trajectory of the pedestrian more accurate; and the map feature is extracted from the high-precision map, so that the pedestrian trajectory predicted by combining the spatiotemporal feature of the pedestrian and the interaction feature of the pedestrian is more accurate.

[0007] In one embodiment, the target detection on the at least two frames of three-dimensional point cloud images to obtain respective feature maps corresponding to the at least two frames of three-dimensional point cloud images comprises: encoding the at least two frames of three-dimensional point cloud images to extract shape features of the first target and the second target in each frame of three-dimensional point cloud image; and constructing a feature map corresponding to the at least two frames of three-dimensional point cloud images, the feature map comprising the shape features of the first target and the second target.

[0008] In this embodiment, the three-dimensional point cloud image is a relatively large memory-occupying image, and the three-dimensional point cloud image is converted into a feature map which is relatively small in memory occupation, so that the processing speed can be improved when the information in the three-dimensional point cloud image is used subsequently.

[0009] In one embodiment, the extraction of the position feature of the first target according to the feature maps comprises: inputting the respective feature maps corresponding to the at least two frames of three-dimensional point cloud images into a region extraction network model to obtain position information of the first target in each frame of feature map.

[0010] In an embodiment, the position feature and the dynamic feature of the first target are extracted according to the feature map, including: determining position information of the first target on the spliced feature map according to position information of the first target in each feature map of the at least two three-dimensional point cloud maps and the spliced feature, the spliced feature map being obtained by splicing the feature maps corresponding to the at least two three-dimensional point cloud maps in the feature dimension; determining a historical motion trajectory of the first target on the spliced feature map according to the position information of the first target on the spliced feature map; inputting the historical motion trajectory into a uniform speed model to obtain a first region of the first target on the spliced feature map; and extracting features in the first region of the spliced feature map.

[0011] In the embodiment, by adding environmental features in a set range around the target pedestrian, the prediction of the future trajectory of the target is more accurate.

[0012] In an embodiment, the interaction feature of the first target is determined, including: determining a first type target, the first type target being a target meeting a set rule, the at least two three-dimensional point cloud maps each including the first type target, and the first type target including the second target; and inputting the position feature and the dynamic feature of the first target and the first type target into the neural network model to obtain the interaction feature of the first target.

[0013] In an embodiment, the interaction feature of the first target is obtained by inputting the position feature and the dynamic feature of the first target and the first type target into the neural network model, including: inputting the position feature and the dynamic feature of the first target and the first type target into the neural network model to obtain interaction features between targets and targets; selecting the interaction feature of the first target and the first type target; and inputting the interaction feature of the first target and the first type target into the neural network model to obtain the interaction feature of the first target.

[0014] In an embodiment, the at least two three-dimensional point cloud maps are three-dimensional point cloud maps obtained after coordinate unification, including: performing coordinate conversion on three-dimensional point cloud maps other than a target three-dimensional point cloud map in the at least two three-dimensional point cloud maps with the coordinate system of the target three-dimensional point cloud map as a reference.

[0015] In an embodiment, the motion trajectory of the first target is predicted according to the position feature and the dynamic feature of the first target, the interaction feature of the first target, and the map feature of the first target, including: splicing the spatial feature of the first target, the interaction feature of the first target, and the map feature of the first target in the feature dimension to obtain a predicted trajectory feature of the first target; and inputting the predicted trajectory feature of the first target into a multilayer perception machine to obtain the motion trajectory of the first target.

[0016] In this embodiment, the feature dimension of the motion trajectory feature is reduced by inputting the obtained motion trajectory feature into the MLP, so as to shorten the prediction time, reduce the redundant features, reduce the noise, and obtain more accurate results.

[0017] In a second aspect, the present application provides a prediction device, comprising: a transceiver unit configured to obtain at least two frames of three-dimensional point cloud images, the at least two frames of three-dimensional point cloud images each comprising a first target and a second target, the at least two frames of three-dimensional point cloud images being three-dimensional point cloud images obtained after coordinate unification; a processing unit configured to perform target detection on the at least two frames of three-dimensional point cloud images, and obtain a feature map corresponding to each of the at least two frames of three-dimensional point cloud images; extract a position feature and a dynamic feature of the first target according to the feature map, the position feature comprising position information of the first target in the feature map, and the dynamic feature comprising a first region feature in the feature map, the first region feature being determined according to the position information of the first target in the feature map; determine an interaction feature of the first target, the interaction feature being obtained by inputting the position feature and the dynamic feature of the first target and a position feature and a dynamic feature of the second target into a neural network model; and predict a motion trajectory of the first target according to the position feature and the dynamic feature of the first target, the interaction feature of the first target, and a map feature of the first target, the map feature of the first target being obtained by encoding a stored map within a set range of a current position of the first target.

[0018] In an embodiment, the processing unit is specifically configured to encode the at least two frames of three-dimensional point cloud images, and extract shape features of the first target and the second target in each frame of three-dimensional point cloud image; and construct a feature map corresponding to the at least two frames of three-dimensional point cloud images, the feature map comprising the shape features of the first target and the second target.

[0019] In an embodiment, the processing unit is specifically configured to input the feature map corresponding to each of the at least two frames of three-dimensional point cloud images into a region extraction network model, and obtain position information of the first target in each frame of feature map.

[0020] In an implementation, the processing unit is specifically configured to determine, according to the feature maps corresponding to the at least two frames of three-dimensional point cloud maps respectively, the position information of the first target in each frame of feature map, and the spliced feature, determine the position information of the first target on the spliced feature map, the spliced feature map being obtained by splicing the feature maps corresponding to the at least two frames of three-dimensional point cloud maps in the feature dimension; determine the historical motion trajectory of the first target on the spliced feature map according to the position information of the first target on the spliced feature map; input the historical motion trajectory into a uniform speed model to obtain a first region of the first target on the spliced feature map; and extract the feature in the first region of the spliced feature map.

[0021] In an implementation, the processing unit is specifically configured to determine a first type of target, the first type of target being a target meeting a set rule, the at least two frames of three-dimensional point cloud maps each including the first type of target, and the first type of target including the second target; input the position feature and the dynamic feature of the first target and the first type of target into the neural network model to obtain the interaction feature of the first target.

[0022] In an implementation, the processing unit is specifically configured to input the position feature and the dynamic feature of the first target and the first type of target into the neural network model to obtain the interaction feature between targets; select the interaction feature of the first target and the first type of target; and input the interaction feature of the first target and the first type of target into the neural network model to obtain the interaction feature of the first target.

[0023] In an implementation, the processing unit is specifically configured to perform coordinate conversion on the other three-dimensional point cloud maps except the target three-dimensional point cloud map in the at least two frames of three-dimensional point cloud maps, with the coordinate system of the target three-dimensional point cloud map as a reference.

[0024] In an implementation, the processing unit is specifically configured to splice the spatial feature of the first target, the interaction feature of the first target, and the map feature of the first target in the feature dimension to obtain the predicted trajectory feature of the first target; and input the predicted trajectory feature of the first target into a multi-layer perception to obtain the motion trajectory of the first target.

[0025] In a third aspect, an intelligent driving system is provided, including at least one processor configured to execute instructions stored in a memory to perform the embodiments of the possible implementations of the first aspect.

[0026] In a fourth aspect, a vehicle is provided, including at least one memory configured to store a program, and at least one processor configured to execute the program stored in the memory, when the program stored in the memory is executed, the processor is configured to perform the embodiments of the possible implementations of the first aspect.

[0027] In a fifth aspect, the present application provides a computer storage medium, having stored therein instructions which, when executed on a computer, cause the computer to perform the embodiments of the first aspect.

[0028] In a sixth aspect, the present application provides a computer program product comprising instructions which, when executed on a computer, cause the computer to perform the embodiments of the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0029] The drawings needed to be used in the following embodiments or prior art description are briefly introduced.

[0030] Figure 1 An architecture schematic diagram of an intelligent driving system provided for the embodiments of the present application;

[0031] Figure 2 An architecture schematic diagram of trajectory prediction of an environment perception module provided for the embodiments of the present application;

[0032] Figure 3 An architecture schematic diagram of an image feature extraction unit provided for the embodiments of the present application;

[0033] Figure 4 A multi-frame feature map splicing schematic diagram provided for the embodiments of the present application;

[0034] Figure 5 An architecture schematic diagram of a space-time feature extraction unit provided for the embodiments of the present application;

[0035] Figure 6 A flow schematic diagram of a prediction method provided for the embodiments of the present application;

[0036] Figure 7 An architecture schematic diagram of a prediction device provided for the embodiments of the present application;

[0037] Figure 8 An architecture schematic diagram of a prediction device provided for the embodiments of the present application. DETAILED DESCRIPTION

[0038] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application.

[0039] An intelligent driving system is to detect the surrounding environment and the self state by using sensors, such as navigation positioning information, road information, obstacle information of other vehicles and pedestrians, self pose information and motion state information, etc., and to accurately control the vehicle driving speed and steering after certain decision planning algorithm, so as to realize automatic driving without the monitoring of the driver. Figure 1As shown, according to the functional requirements of the intelligent driving system 100, the system 100 can be divided into a positioning module 10, an environment perception module 20, a path planning module 30, and a decision control module 40.

[0040] The positioning module 10 is configured to obtain the position and navigation information of the vehicle through data collected by sensors in the sensor system, such as a global positioning system (GPS) unit, an inertial navigation system (INS) unit, an odometer, a camera, a radar, and the like.

[0041] Among them, the positioning technology can be divided into absolute positioning, relative positioning and combined positioning according to the positioning method. The absolute positioning refers to the realization through the GPS, that is, obtaining the absolute position and heading information of the vehicle on the earth through the satellite; the relative positioning refers to obtaining the acceleration and angular acceleration information through the INS, the odometer and other sensors according to the initial pose of the vehicle, and then integrating the information with respect to time, so as to obtain the current pose information relative to the initial pose; the combined positioning refers to combining the absolute positioning and the relative positioning to make up for the shortcomings of the single positioning method.

[0042] The environment perception module 20 is configured to perceive the environmental information and the vehicle state information around the vehicle through data collected by sensors in the sensor system, such as the GPS unit, the INS unit, the odometer, the camera, the radar (laser radar, millimeter wave radar, ultrasonic radar, and the like), and the like, and the position and navigation information of the vehicle obtained by the positioning module 10.

[0043] Among them, the environmental information can include the shape, direction, curvature, slope, lane of the road, the position, size, forward direction and speed of other vehicles or pedestrians, and the like; the vehicle state information can include the forward speed, acceleration, steering angle, body position and attitude of the vehicle, and the like.

[0044] The path planning module 30 is configured to plan a reasonable driving route for the vehicle through the position and navigation information of the vehicle obtained by the positioning module 10 and the environmental information and the vehicle state information around the vehicle perceived by the environment perception module 20. Among them, according to the range of path planning, it can be divided into global path planning and local path planning. The global path planning refers to planning a global path from the current position of the vehicle to the destination in the case of knowing the global map; the local path planning refers to planning a safe and smooth driving path in real time in the case of lane changing, turning, avoiding obstacles and the like according to the environmental perception information.

[0045] The decision control module 40 includes decision-making and control functions. The decision-making function is used to determine which lane the vehicle selects, whether to change lanes, whether to follow other vehicles, whether to detour, and whether to stop, based on data obtained from the positioning module 10, the environmental perception module 20, and the path planning module 30. The control function is used to execute the decision-making instructions issued by the decision-making function, control the vehicle to achieve the desired speed and steering angle, and control components such as turn signals, horn, doors, and windows.

[0046] In this embodiment, the process of predicting the trajectories of other vehicles and pedestrians around the vehicle is generally implemented in the environment perception module 20, but can also be implemented in the path planning module 30, depending on the application scenario of the prediction results, and is not limited here. The technical solution of this application will be described below using the prediction of pedestrian movement trajectories by the environment perception module 20 as an example.

[0047] like Figure 2 As shown, the environment perception module 20 can be divided into an image feature extraction unit 201, a target feature extraction unit 202, and a map feature extraction unit 203 according to the functions it performs.

[0048] Take a lidar sensor in a sensor system as an example. When the vehicle-controlled lidar sensor scans the environment around the vehicle in real time, the laser sensor emits a laser beam and scans the environment around the vehicle according to a certain trajectory, recording the reflected laser point information as it scans, obtaining a large number of laser points. Then, with one scanning cycle as one frame, multiple frames of laser point clouds are obtained. Since the reflected laser light carries information such as orientation and distance when it shines on the surface of an object, each frame of laser point cloud can be used to construct a three-dimensional point cloud map of the environment around the vehicle.

[0049] Since the vehicle is in motion, each frame of the 3D point cloud image is obtained from different positions of the vehicle. After obtaining multiple frames of 3D point cloud images, each frame can be unified under the same coordinate system. For example, after obtaining multiple frames of 3D point cloud images, the vehicle calculates the position information of each frame of the 3D point cloud image based on the acceleration, distance traveled, and positioning information collected by sensors such as accelerometers, odometers, and positioning devices. Then, using the coordinates of the last frame of the 3D point cloud image (i.e., the 3D point cloud image obtained at the current moment) as a reference, the other 3D point clouds obtained before that are transformed to the 3D point cloud coordinate system of the current moment, so that the 3D point cloud images of all frames are unified under the same coordinate system.

[0050] Among them, the multi-frame 3D point cloud map is a 3D point cloud map continuously acquired by the lidar in terms of time sequence. Alternatively, by setting an interval, a 3D point cloud map acquired by the lidar can be acquired every N frames (or every T seconds) to obtain a multi-frame 3D point cloud map.

[0051] Image feature extraction unit 201 is used to perform target detection on each frame of the three-dimensional point cloud image after receiving multiple frames of three-dimensional point cloud images, extract the features of pedestrians, vehicles and other objects in each frame of the three-dimensional point cloud image, and construct feature map B from the extracted features. j Then, based on feature map B j The stitched feature map B and the corresponding feature map B of the target pedestrian in each frame of the 3D point cloud map are obtained. j Location information on the screen.

[0052] like Figure 3 As shown, the image feature extraction unit 201, based on its function, can be divided into a point cloud feature encoder 2011, a point cloud stitcher 2012, and a region proposal network (RPN) unit 2013. The point cloud feature encoder 2011 encodes the 3D point cloud image of each frame to extract point, line, surface, and cylinder features from the 3D point cloud image data in each frame. Given that the application scenario of this application is to predict the movement trajectory of pedestrians, as well as other pedestrians, vehicles, and other objects that influence the future movement trajectory of the target pedestrian, the extracted features also include pedestrian features and vehicle features. The pedestrian and vehicle features extracted from each frame of the 3D point cloud image are used to construct a feature map, denoted as B. j , where j represents the frame number.

[0053] Point cloud stitcher 2012 obtains the feature map B corresponding to each frame of the 3D point cloud image. j Then, the feature map B corresponding to each frame of the 3D point cloud map is... j Following the chronological order, the features are concatenated along feature dimension C to obtain the concatenated feature map B, as shown below. Figure 4 As shown, this is to facilitate the subsequent extraction of the target's historical trajectory.

[0054] The RPN unit 2013 receives the feature map B corresponding to each frame of 3D point cloud image extracted by the point cloud feature encoder 2011. j Then, the feature maps (B1, B2, ..., B) are... j The input to the RPN model is used to extract the features of the target pedestrian, resulting in the target pedestrian features in each feature map B. j The position on, and should be in each feature map B j The position on the map is used as the detection bounding box (x, y, w, h, θ) to detect the location information of the target pedestrian on the stitched feature map B. Here, (x, y), w, h, and θ represent the center coordinates, width, height, and angle of the detection bounding box on the feature map, respectively.

[0055] The target feature extraction unit 202 is used to obtain the feature map B corresponding to each frame of the 3D point cloud image extracted by the image feature extraction unit 201. jThe target pedestrian is in feature map B of each frame. j The system extracts the location information of the target pedestrian in the stitched feature map B, the environmental features within the target pedestrian's activity range in the stitched feature map B, and the interaction features between the target pedestrian and other pedestrians, vehicles, and other objects.

[0056] like Figure 5 As shown, the target feature extraction unit 202 includes a spatial feature unit 2021 and an interaction feature unit 2022. The spatial feature unit 2021 includes a position feature unit 20211 and a dynamic feature unit 20212. The position features acquired by the position feature unit 20211 refer to the feature map B corresponding to each frame of the 3D point cloud. j The feature vector of the detection bounding box (x, y, w, h, θ) on the concatenated feature map B; the dynamic features acquired by the dynamic feature unit 20212, that is, the historical trajectory (x, y, w, h, θ) of the target pedestrian on the concatenated feature map B. j y j θ j Estimate the activity range of the target pedestrian, and then determine the feature vector of the pedestrian's activity range on feature map B based on the activity range of the target pedestrian.

[0057] The location feature unit 20211 receives the feature map B corresponding to each frame of the 3D point cloud map. j The target pedestrian is in feature map B of each frame. j The detection bounding box (x, y, w, h, θ) and the stitched feature map B are used to determine the location of the target pedestrian in each frame of feature map B. j The detection bounding box (x, y, w, h, θ) and feature map B on the map. j The positions of each detection box are extracted from the stitched feature map B, and then the positions of each extracted detection box are stitched together on the feature dimension C to obtain the position features of the target pedestrian.

[0058] After obtaining the positions of each detection box on the stitched feature map B, the dynamic feature unit 20212 connects the positions of each detection box to obtain the historical motion trajectory (x) of the target on the stitched feature map B. j y j θ j Then, the target's historical trajectory (x) is calculated. j y j θ j The input is fed into a uniform velocity model, which is then used to estimate the possible range of the target's movement (x). max x min y max y min), and finally taking the features on the spliced feature map B in the active range as dynamic features. The application makes the subsequent prediction of the future trajectory of the target more accurate by increasing the environmental features around the target pedestrian.

[0059] Optionally, after obtaining the position features and the dynamic features, the spatial feature unit 2021 inputs the position features and the dynamic features into a region of interest align (ROI Align) network respectively, performs clustering processing by using bilinear interpolation, and then splices in the feature dimension C to obtain the spliced spatial features of the target pedestrian, and then processes the spliced spatial features of the target pedestrian through a residual network (ResNet) model to obtain more accurate spatial features of the target pedestrian.

[0060] The interaction feature unit 2022 is configured to group pedestrians and vehicles according to distance information collected by a sensor system. The grouping manner can be distance-based grouping, for example, dividing the area collected by the distance sensor into M n x m size sub-areas, and other groups are divided in the same manner, or taking the target pedestrian as a reference, screening pedestrians or vehicles within a set distance d1 from the target pedestrian as a group, screening pedestrians or vehicles within a set distance d2 from the target pedestrian as a group, and other groups are divided in the same manner. The grouping manner is not limited herein.

[0061] After the grouping is completed, the interaction feature unit 2022 determines the group in which the target pedestrian is located, or determines the group corresponding to the screening rule, and then calculates the interaction features between the pedestrians and the pedestrians, and between the pedestrians and the vehicles in the group. The spatio-temporal features of each pedestrian and the spatio-temporal features of each vehicle (although the spatio-temporal features of the vehicle are not mentioned above, the implementation manner can be the same as the spatio-temporal features of the pedestrian, or the navigation route can be directly obtained, which is not limited herein) in the group are input into a graph neural network (GNN) model to obtain the interaction features between each pedestrian and pedestrian, pedestrian and vehicle, and vehicle and vehicle in the group. For example, the interaction feature unit 2022 takes the spatio-temporal features of each pedestrian and the spatio-temporal features of each vehicle in a group including the target pedestrian as nodes f of the GNN, and then calculates the interaction features v between each node f by formula (1) ij . Wherein, formula (1) is:

[0062]

[0063] Wherein, a and ψ are linear mapping functions, and i and j represent the serial numbers of pedestrians and vehicles.

[0064] The interaction feature unit 2022 obtains the interaction features between each pedestrian and pedestrian, pedestrian and vehicle, and vehicle and vehicle in the group, and then selects the interaction features belonging to the target pedestrian and other pedestrians and vehicles, and then inputs the selected interaction features into the GNN model again to obtain the interaction features of the target pedestrian. For example, the interaction feature unit 2022 inputs the interaction features of the target pedestrian and one pedestrian or one vehicle into the GNN model to obtain the interaction features of the target pedestrian and the one pedestrian or the one vehicle. ij As a node, the interaction features GNN(F) between the target pedestrian and the pedestrians and vehicles in the group are calculated by formula (2). Wherein, formula (2) is:

[0065] GNN(F) = softmax(V) · F; (2)

[0066] Wherein, F represents a set of nodes f, and V represents a set of interaction features of the node f

[0067] The map feature extraction unit 203 first vectorizes the high-precision map, processes the vectorized high-precision map by using a self-attention mechanism, selects the features of elements within a certain range of the current position, such as the features of pedestrian crosswalks, non-motor vehicle lanes, road surfaces, traffic lights and other elements, and then encodes the selected element features by using the self-attention mechanism to obtain global map features. For example, the map feature extraction unit 203 processes the high-precision map within a certain range of the current position by formula (3) to select the features of each element, and formula (3) is:

[0068] GNN(P) = softmax(P K P Q )P V ; (3)

[0069] Wherein, P is a GNN node feature matrix, P K , P Q , and P V are linear mappings thereof.

[0070] The map feature extraction unit 203 further encodes each node by formula (1) to obtain global map features, wherein each node is the feature of each element selected by the map feature extraction unit 203.

[0071] Finally, the environmental perception module 20 splices the spatial features of the target pedestrian, the interaction features of the target pedestrian and the global map features in the feature dimension C to obtain a pedestrian prediction trajectory feature with a relatively high feature dimension, and then inputs the prediction trajectory feature into a multi-layer perception (MLP) to extract through multiple calculation layers inside, map the input prediction trajectory feature with a relatively high feature dimension to a data set, and output a feature vector with a reasonable feature dimension, i.e., a predicted target pedestrian motion trajectory.

[0072] In the embodiments of the present application, after obtaining the multi-frame three-dimensional point cloud map and the high-precision map, the spatial features and the interaction features of the target pedestrian are extracted by processing the multi-frame three-dimensional point cloud map, so as to obtain more environmental features and increase the interaction information between the target pedestrian and the surrounding environment, so that the subsequent prediction of the motion trajectory of the pedestrian is more accurate; the map features are extracted by using the high-precision map, and finally the trajectory of the target pedestrian is predicted by splicing the space-time features of the target pedestrian, the interaction features of the target pedestrian and the map features. The scheme has a great improvement in accuracy, and the average distance error of the trajectory prediction within three seconds is 0.25 m, which is better than the existing method. Moreover, based on the laser point cloud information, the ranging accuracy and the light robustness are greatly improved compared with the algorithm using images.

[0073] Figure 6 A flowchart of a prediction method provided by the embodiments of the present application is shown in FIG. 1. As shown in FIG. 1, the embodiments of the present application provide a prediction method, and the specific implementation process is as follows: Figure 6

[0074] In step S601, at least two frames of three-dimensional point cloud maps are obtained. The at least two frames of three-dimensional point cloud maps each include a first target and a second target, and even more other targets, which can be pedestrians, vehicles and other objects. Here, "first" and "second" are only numbers for the targets and do not contain any specific meaning. The embodiments of the present application take the first target as an example of a target pedestrian for trajectory prediction, and the second target as a representative target that is other targets in the three-dimensional point cloud map that can affect the subsequent motion trajectory of the target pedestrian.

[0075] The embodiments of the present application take the collection of three-dimensional point cloud maps by a laser radar sensor as an example. When a vehicle controls the laser radar sensor to scan the environment around the vehicle in real time, the laser sensor emits a laser beam and scans the environment around the vehicle according to a certain trajectory, records the reflected laser point information while scanning, obtains a large number of laser points, and then obtains multiple frames of laser point clouds with one scanning period as one frame. Since the reflected laser carries information such as direction and distance when the laser is irradiated to the surface of an object, each frame of laser point cloud obtained can construct a three-dimensional point cloud map of the environment around the vehicle. ​

[0076] Since the vehicle is in a moving state, each frame of the three-dimensional point cloud image is obtained when the vehicle is in a different position. After a plurality of frames of three-dimensional point cloud images are obtained, each frame of the three-dimensional point cloud image can be unified in the same coordinate system. For example, after a plurality of frames of three-dimensional point cloud images are obtained, the vehicle calculates the position information of each frame of the three-dimensional point cloud image according to the acceleration, distance of movement, positioning and other information collected by the accelerometer, odometer, positioning device and other sensors, and then takes the coordinates of the last frame of the three-dimensional point cloud image (i.e. the three-dimensional point cloud image obtained at the current time) as the reference to convert the other three-dimensional point clouds obtained before the current time to the three-dimensional point cloud coordinate system at the current time, so that all frames of the three-dimensional point cloud image are unified in the same coordinate system.

[0077] In the formula, the plurality of frames of three-dimensional point cloud images are three-dimensional point cloud images continuously obtained by the lidar, or one frame of three-dimensional point cloud image collected by the lidar is obtained every N frames (or every T seconds) by setting an interval, and a plurality of frames of three-dimensional point cloud images are obtained.

[0078] In step S603, target detection is performed on the at least two frames of three-dimensional point cloud images to obtain the feature maps corresponding to the at least two frames of three-dimensional point cloud images respectively.

[0079] Specifically, after a plurality of frames of three-dimensional point cloud images are obtained, each frame of the three-dimensional point cloud image is encoded to extract the point, line, surface and cylindrical features of the three-dimensional point cloud image data in each frame. According to the application scenario of the present application, the motion trajectory of a pedestrian is predicted, and other pedestrians, vehicles and objects that affect the future motion trajectory of the target pedestrian, so the features extracted are also pedestrian features and vehicle features, and the pedestrian features and vehicle features extracted from each frame of the three-dimensional point cloud image form a feature map.

[0080] Optionally, the feature map B corresponding to each frame of the three-dimensional point cloud image is obtained j Then, the feature map B corresponding to each frame of the three-dimensional point cloud image is obtained j According to the time sequence, the feature maps are spliced in the feature dimension C to obtain the spliced feature map B, as shown in Figure 4 for subsequent extraction of the historical trajectory of the target.

[0081] In step S605, the position feature and dynamic feature of the first target are extracted according to the feature map. The first target is the object of the trajectory prediction of the present application, which can be a pedestrian or a vehicle, and the pedestrian is taken as an example here.

[0082] Specifically, the feature map B corresponding to each frame of the three-dimensional point cloud image is obtained j In the RPN model, the position of the target pedestrian feature on each feature map B is obtained by extracting the features of the target pedestrian, and j j ​The position of the target pedestrian on the feature map B is taken as a detection frame (x, y, w, h, θ) so as to detect the position information of the target pedestrian on the spliced feature map B subsequently. Wherein, (x, y), w, h and θ respectively represent the center coordinates, width, height and angle of the detection frame on the feature map.

[0083] In the process of extracting the position feature of the target pedestrian, according to the detection frame (x, y, w, h, θ) of the target pedestrian on each frame feature map B j j , the position of each detection frame is extracted on the spliced feature map B, and then the positions of the extracted detection frames are spliced in the feature dimension C to obtain the position feature of the target pedestrian.

[0084] In the process of extracting the dynamic feature of the target pedestrian, after obtaining the positions of the detection frames on the spliced feature map B, the positions of the detection frames are connected to obtain the historical motion trajectory (x j , y j , θ j ) of the target on the spliced feature map B, and then the historical motion trajectory (x j , y j , θ j ) of the target is input into an input constant speed model to estimate the possible activity range (x max , x min , y max , y min ) of the activity of the target, and finally the features on the spliced feature map B within the activity range are taken as the dynamic feature. The present application increases the environmental features around the target pedestrian, so that the subsequent prediction of the future trajectory of the target is more accurate.

[0085] Optionally, after obtaining the position feature and the dynamic feature, the position feature and the dynamic feature are respectively input into an ROIAlign network model, and after clustering processing by using bilinear interpolation, the spliced target pedestrian spatial feature is obtained by splicing in the feature dimension C, and the more accurate target pedestrian spatial feature is obtained by processing through a ResNet model.

[0086] Step S607, determine the interaction feature of the first target.

[0087] ​Optionally, pedestrians and vehicles can be grouped based on the distance information collected by the sensor system. Grouping can be done by distance, such as dividing the area collected by the distance sensor into M n×m sub-regions, and so on. Alternatively, based on the target pedestrian, pedestrians or vehicles within a set distance d1 from the target pedestrian can be grouped into one group, pedestrians or vehicles within a set distance d2 from the target pedestrian can be grouped into another group, and so on. This application does not limit the grouping method.

[0088] After grouping, the group containing the target pedestrian is determined, or the group corresponding to the filtering rule is determined, where the second target is included. In calculating the interaction features between pedestrians and between pedestrians and vehicles within the group, the spatiotemporal features of each pedestrian and each vehicle in the group are input into the GNN model to obtain the interaction features between each pedestrian and pedestrian, pedestrian and vehicle, and vehicle and vehicle within the group. After obtaining the interaction features between each pedestrian and pedestrian, pedestrian and vehicle, and vehicle and vehicle within the group, the interaction features between the target pedestrian and other pedestrians and vehicles are selected. These selected interaction features are then input into the GNN model again to obtain the interaction features of the target pedestrian.

[0089] Step S609: Based on the positional and dynamic characteristics of the first target, the interaction characteristics of the first target, and the map characteristics of the first target, predict the trajectory of the first target.

[0090] Before proceeding, the high-precision map is first processed. The high-precision map is first vectorized, and then the vectorized high-precision map is processed using a self-attention mechanism to select the features of elements within a certain range of the current location, such as pedestrian crossings, non-motorized vehicle lanes, road surfaces, traffic lights, etc. Then, the self-attention mechanism is used to encode the features of the selected elements to obtain the global map features.

[0091] After obtaining the spatial features, interaction features, and global map features of the target pedestrian, these features are concatenated along feature dimension C to obtain a pedestrian prediction trajectory feature with a relatively high feature dimension. This prediction trajectory feature is then input into the MLP, where it is extracted through multiple internal computation layers. The input prediction trajectory feature with a relatively high feature dimension is mapped onto a dataset, thereby outputting a feature vector with a reasonable feature dimension, which is the predicted target pedestrian movement trajectory.

[0092] Figure 7 This is a schematic diagram of the structure of a trajectory prediction device provided in an embodiment of this application. The trajectory prediction device 700 can be a computing device or apparatus (e.g., a vehicle, terminal, etc.), or a device within an apparatus (e.g., an ISP or SoC). It can achieve, for example...Figures 1 to 6 The trajectory prediction method and the above-mentioned optional embodiments are shown. As shown in the figure Figure 7 The trajectory prediction device 700 includes a transceiver unit 701 and a processing unit 702.

[0093] In this application, the trajectory prediction device 700 specifically realizes the process as follows: the transceiver unit 701 is configured to acquire at least two frames of three-dimensional point cloud images, the at least two frames of three-dimensional point cloud images each including a first target and a second target, and the at least two frames of three-dimensional point cloud images being three-dimensional point cloud images acquired after coordinate unification; the processing unit 702 is configured to perform target detection on the at least two frames of three-dimensional point cloud images to acquire respective feature maps of the at least two frames of three-dimensional point cloud images; according to the feature maps, a position feature and a dynamic feature of the first target are extracted, the position feature including position information of the first target in the feature map, and the dynamic feature including a corresponding first region feature in the feature map, the first region feature being determined according to the position information of the first target in the feature map; an interaction feature of the first target is determined, the interaction feature being obtained by inputting the position feature and the dynamic feature of the first target and the position feature and the dynamic feature of the second target into a neural network model; and a motion trajectory of the first target is predicted according to the position feature and the dynamic feature of the first target, the interaction feature of the first target, and a map feature of the first target, the map feature of the first target being obtained by encoding a stored map within a set range of a current position of the first target.

[0094] The transceiver unit 701 is configured to perform S601 in the trajectory prediction method and any optional example thereof. The processing unit 702 is configured to perform S603, S605, S607, and S609 in the trajectory prediction method and any optional example thereof. For details, refer to the detailed description in the method examples, which will not be repeated here.

[0095] It should be understood that the trajectory prediction device in the embodiments of the present application can be realized by software, for example, a computer program or instructions with the above-mentioned functions, and the corresponding computer program or instructions can be stored in a memory inside a terminal, and the above-mentioned functions are realized by a processor reading the corresponding computer program or instructions inside the memory. Alternatively, the trajectory prediction device in the embodiments of the present application can also be realized by hardware. The processing unit 702 is a processor (such as NPU, GPU, processor in a system chip), and the transceiver unit 701 is a transceiver circuit or an interface circuit. Alternatively, the trajectory prediction device in the embodiments of the present application can also be realized by a combination of a processor and a software module.

[0096] It should be understood that the details of the device in the embodiments of the present application can refer to the related content shown in the Figures 1 to 6 The embodiments of the present application will not be repeated here.

[0097] Figure 8 This is a schematic diagram of another trajectory prediction device provided in an embodiment of this application. The trajectory prediction device 800 can be a computing device or apparatus (e.g., a vehicle, terminal, etc.), or it can be a device within an apparatus (e.g., an ISP or SoC). It can achieve, for example... Figures 1 to 6 The trajectory prediction method and the above-described optional embodiments are shown. Figure 8 As shown, the trajectory prediction device 800 includes: a processor 801 and an interface circuit 802 coupled to the processor 801. It should be understood that, although... Figure 8 Only one processor and one interface circuit are shown in the diagram. The trajectory prediction device 800 may include a number of other processors and interface circuits.

[0098] In this application, the trajectory prediction device 800 is specifically implemented as follows: the interface circuit 802 is used to acquire at least two frames of three-dimensional point cloud images, each of which includes a first target and a second target. The at least two frames of three-dimensional point cloud images are obtained after coordinate unification. The processor 801 is used to perform target detection on the at least two frames of three-dimensional point cloud images and acquire feature maps corresponding to each of the at least two frames of three-dimensional point cloud images. Based on the feature maps, the position features and dynamic features of the first target are extracted. The position features include the position information of the first target in the feature map, and the dynamic features include the position information of the first target in the feature map. The corresponding first region feature is determined based on the position information of the first target in the feature map; the interaction feature of the first target is determined by inputting the position feature and dynamic feature of the first target and the position feature and dynamic feature of the second target into a neural network model; based on the position feature and dynamic feature of the first target, the interaction feature of the first target and the map feature of the first target, the trajectory of the first target is predicted, and the map feature of the first target is obtained by encoding the map within a set range of the current position of the first target.

[0099] The interface circuit 802 is used to communicate with other components of the terminal, such as memory or other processors. The processor 801 is used to interact with other components via the interface circuit 802. The interface circuit 802 can be an input / output interface for the processor 801.

[0100] For example, processor 801 reads computer programs or instructions from a coupled memory via interface circuit 802, and decodes and executes these computer programs or instructions. It should be understood that these computer programs or instructions may include the aforementioned terminal function programs, or the function programs of the trajectory prediction device applied within the terminal. When the corresponding function program is decoded and executed by processor 801, the terminal or the trajectory prediction device within the terminal can implement the scheme in the trajectory prediction method provided in the embodiments of this application.

[0101] Optionally, the terminal function programs are stored in a memory outside the trajectory prediction device 800. When the terminal function programs are decoded and executed by the processor 801, part or all of the terminal function programs are temporarily stored in the memory.

[0102] Optionally, the terminal function programs are stored in a memory inside the trajectory prediction device 800. When the terminal function programs are stored in the memory inside the trajectory prediction device 800, the trajectory prediction device 800 can be arranged in the terminal of the embodiments of the present application.

[0103] Optionally, part of the terminal function programs are stored in a memory outside the trajectory prediction device 800, and the other part of the terminal function programs are stored in a memory inside the trajectory prediction device 800.

[0104] It should be understood that, Figures 7 to 8 Any of the trajectory prediction devices shown can be combined with each other, Figures 7 to 8 Any of the trajectory prediction devices shown and the design details of the optional embodiments can be combined with each other, and can also be combined with Figure 6 Any of the trajectory prediction methods shown and the design details of the optional embodiments. Here, it is not repeated.

[0105] The present application provides a computer readable storage medium, which stores a computer program, when the computer program is executed in a computer, the computer executes any of the above methods.

[0106] The present application provides a computing device, comprising a memory and a processor, the memory stores executable code, and the processor executes the executable code to implement any of the above methods.

[0107] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of the present application.

[0108] Moreover, various aspects or features of the embodiments disclosed herein can be implemented as a method, apparatus, or article of manufacture using standard programming and / or engineering techniques. The term "article of manufacture" as used herein is intended to encompass a computer program accessible from any computer- readable device, carrier, or media. For example, computer-readable media can include but are not limited to magnetic storage devices (e.g., hard disk, floppy disk, magnetic strips, etc.), optical disks (e.g., compact disk (CD), digital versatile disk (DVD), etc.), smart cards, and flash memory devices (e.g., EPROM, card, stick, or key drive, etc.). Additionally, various storage media described herein can represent one or more devices and / or other machine-readable media for storing information. The term "machine- readable medium" can include, without being limited to, wireless channels and various other media capable of storing, containing, and / or carrying instruction(s) and / or data.

[0109] In the above-described embodiments, Figure 7 and Figure 8 The trajectory prediction apparatus can be implemented totally or partially by software, hardware, firmware, or any combination thereof. When implemented by software, it can be implemented totally or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When loaded and executed by a computer, the computer program instructions totally or partially generate the flow or function according to the embodiments of the present application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, server, or data center to another via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. that includes one or more available media sets. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., Solid State Disk (SSD)), etc.

[0110] It should be understood that the size of the serial number of the processes described above does not mean the execution order in various embodiments of the embodiments of the present application, and the execution order of the processes should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0111] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the system, device and unit described above can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.

[0112] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the unit is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0113] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0114] If the function is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application essentially or the parts that contribute to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or an access network device, etc.) execute all or part of the steps of the method of the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0115] The above is only a specific implementation of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the embodiments of the present application, which should be covered within the protection scope of the embodiments of the present application.

Claims

1. A trajectory prediction method, characterized in that, include: At least two frames of 3D point cloud images are acquired, each of which includes a first target and a second target. The at least two frames of 3D point cloud images are 3D point cloud images obtained after coordinate unification. Target detection is performed on the at least two frames of 3D point cloud images to obtain the feature maps corresponding to each of the at least two frames of 3D point cloud images; Based on the feature map, the position features and dynamic features of the first target are extracted. The position features include the position information of the first target in the feature map, and the dynamic features include the corresponding first region features in the feature map. The first region features are determined based on the position information of the first target in the feature map. The interaction features of the first target are determined by inputting the positional features and dynamic features of the first target with the positional features and dynamic features of the second target into a neural network model; Based on the location features and dynamic features of the first target, the interaction features of the first target, and the map features of the first target, the movement trajectory of the first target is predicted. The map features of the first target are obtained by encoding a map within a specified range of the current location of the first target.

2. The method according to claim 1, characterized in that, Target detection is performed on the at least two frames of 3D point cloud images to obtain the feature maps corresponding to each of the at least two frames of 3D point cloud images, including: The at least two frames of 3D point cloud images are encoded, and the shape features of the first target and the second target in each frame of 3D point cloud image are extracted; Construct feature maps corresponding to the at least two frames of 3D point cloud images, wherein the feature maps include the shape features of the first target and the second target.

3. The method according to claim 1, characterized in that, The step of extracting the position features of the first target based on the feature map includes: The feature maps corresponding to the at least two frames of 3D point cloud images are input into the region extraction network model to obtain the position information of the first target in the feature map of each frame.

4. The method according to claim 1, characterized in that, The step of extracting the positional and dynamic features of the first target based on the feature map includes: Based on the feature maps corresponding to the at least two frames of 3D point cloud images, the position information of the first target in the feature map of each frame, and the stitched features, the position information of the first target on the stitched feature map is determined. The stitched feature map is obtained by stitching the feature maps corresponding to the at least two frames of 3D point cloud images in the feature dimension. Based on the position information of the first target on the stitched feature map, determine the historical movement trajectory of the first target on the stitched feature map; The historical motion trajectory is input into the uniform velocity model to obtain the first region of the first target on the stitched feature map; Extract the features located in the first region from the stitched feature map.

5. The method according to any one of claims 1-4, characterized in that, Determining the interaction characteristics of the first target includes: A first type of target is determined, which is a target that conforms to a set rule, and the at least two frames of 3D point cloud maps both include the first type of target, which includes the second target; The positional and dynamic features of the first target and the first type of target are input into the neural network model to obtain the interaction features of the first target.

6. The method according to claim 5, characterized in that, The step of inputting the positional and dynamic features of the first target and the first type of target into the neural network model to obtain the interaction features of the first target includes: The positional and dynamic features of the first target and the first type of target are input into the neural network model to obtain the interaction features between the targets. Select the interaction features between the first target and the first type of target; The interaction features of the first target and the first type of target are input into the neural network model to obtain the interaction features of the first target.

7. The method according to claim 1, characterized in that, The at least two frames of 3D point cloud images are 3D point cloud images obtained after coordinate unification, including: Using the coordinate system of the target 3D point cloud image as a reference, coordinate transformation is performed on the other 3D point cloud images in the at least two frames of 3D point cloud images, excluding the target 3D point cloud image.

8. The method according to claim 1, characterized in that, The step of predicting the trajectory of the first target based on its positional features, dynamic features, interaction features, and map features includes: The spatial features, interaction features, and map features of the first target are concatenated along the feature dimension to obtain the predicted trajectory features of the first target. The predicted trajectory features of the first target are input into a multilayer perceptron to obtain the motion trajectory of the first target.

9. A trajectory prediction device, characterized in that, include: The transceiver unit is used to acquire at least two frames of three-dimensional point cloud images, each of which includes a first target and a second target. The at least two frames of three-dimensional point cloud images are three-dimensional point cloud images obtained after coordinate unification. The processing unit is used to perform target detection on the at least two frames of three-dimensional point cloud images and obtain the feature maps corresponding to each of the at least two frames of three-dimensional point cloud images. Based on the feature map, the position features and dynamic features of the first target are extracted. The position features include the position information of the first target in the feature map, and the dynamic features include the corresponding first region features in the feature map. The first region features are determined based on the position information of the first target in the feature map. The interaction features of the first target are determined by inputting the positional features and dynamic features of the first target with the positional features and dynamic features of the second target into a neural network model; Based on the location features and dynamic features of the first target, the interaction features of the first target, and the map features of the first target, the movement trajectory of the first target is predicted. The map features of the first target are obtained by encoding a map within a specified range of the current location of the first target.

10. The apparatus according to claim 9, characterized in that, The processing unit is specifically used for The at least two frames of 3D point cloud images are encoded, and the shape features of the first target and the second target in each frame of 3D point cloud image are extracted; Construct feature maps corresponding to the at least two frames of 3D point cloud images, wherein the feature maps include the shape features of the first target and the second target.

11. The apparatus according to claim 9, characterized in that, The processing unit is specifically used for The feature maps corresponding to the at least two frames of 3D point cloud images are input into the region extraction network model to obtain the position information of the first target in the feature map of each frame.

12. The apparatus according to claim 9, characterized in that, The processing unit is specifically used for Based on the feature maps corresponding to the at least two frames of 3D point cloud images, the position information of the first target in the feature map of each frame, and the stitched features, the position information of the first target on the stitched feature map is determined. The stitched feature map is obtained by stitching the feature maps corresponding to the at least two frames of 3D point cloud images in the feature dimension. Based on the position information of the first target on the stitched feature map, determine the historical movement trajectory of the first target on the stitched feature map; The historical motion trajectory is input into the uniform velocity model to obtain the first region of the first target on the stitched feature map; Extract the features located in the first region from the stitched feature map.

13. The apparatus according to any one of claims 9-12, characterized in that, The processing unit is specifically used for A first type of target is determined, which is a target that conforms to a set rule, and the at least two frames of 3D point cloud maps both include the first type of target, which includes the second target; The positional and dynamic features of the first target and the first type of target are input into the neural network model to obtain the interaction features of the first target.

14. The apparatus according to claim 13, characterized in that, The processing unit is specifically used for The positional and dynamic features of the first target and the first type of target are input into the neural network model to obtain the interaction features between the targets. Select the interaction features between the first target and the first type of target; The interaction features of the first target and the first type of target are input into the neural network model to obtain the interaction features of the first target.

15. The apparatus according to claim 9, characterized in that, The processing unit is specifically used for Using the coordinate system of the target 3D point cloud image as a reference, coordinate transformation is performed on the other 3D point cloud images in the at least two frames of 3D point cloud images, excluding the target 3D point cloud image.

16. The apparatus according to claim 9, characterized in that, The processing unit is specifically used for The spatial features, interaction features, and map features of the first target are concatenated along the feature dimension to obtain the predicted trajectory features of the first target. The predicted trajectory features of the first target are input into a multilayer perceptron to obtain the motion trajectory of the first target.

17. An intelligent driving system, comprising at least one processor, the processor being configured to execute instructions stored in a memory to perform the method as claimed in any one of claims 1-8.

18. A vehicle, characterized in that, include: At least one memory for storing programs; At least one processor is configured to execute a program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to perform the method as described in any one of claims 1-8.

19. A computer storage medium storing instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1-8.

20. A computer program product comprising instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Sparse point cloud multi-target tracking method fusing spatio-temporal information

    CN112561966A

  • Vehicle speed determination method, and vehicle

    WO2020233436A1