Vehicle Trajectory Prediction Method and Device, Trajectory Prediction Model Training Method and Device
By using a convolutional neural network layer to replace the fully connected network layer in vehicle trajectory prediction, the problem of large amount of calculations of the LSTM model is solved, and more efficient vehicle trajectory prediction is achieved.
Patent Information
- Application Number
- CN202211537834.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-01
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-12-01
AI Technical Summary
The existing vehicle trajectory prediction methods use LSTM models to calculate a lot and have low trajectory prediction efficiency.
The convolutional neural network layer is used to replace the fully connected network layer, and vehicle trajectory prediction is carried out by resetting information and updating information, and the convolutional neural network layer has a small computing capacity to improve prediction speed and efficiency.
It improves the speed and efficiency of vehicle trajectory prediction, reduces the calculation amount, and improves the prediction accuracy.
Smart Images

Figure CN115761429B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine learning, and particularly to a vehicle trajectory prediction method and apparatus, and a trajectory prediction model training method and apparatus. Background Art
[0002] In order to better implement autonomous driving, the autonomous driving control device of a vehicle (such as an on-vehicle computer) needs to reasonably predict the trajectory of the vehicle. When predicting the trajectory of the vehicle, it is necessary to fully understand the historical trajectory information of the vehicle to ensure the accuracy of the predicted vehicle trajectory. Based on this, currently, most use the long short-term memory (LSTM) network as the model basis to train a trajectory prediction model capable of predicting the vehicle trajectory. However, due to the design of the long short-term memory network, its computational complexity is large and the trajectory prediction efficiency is low. Summary of the Invention
[0003] The existing vehicle trajectory prediction uses a trajectory prediction model trained based on the LSTM model, with a large computational complexity and low trajectory prediction efficiency.
[0004] To solve the above technical problems, this application is proposed. Embodiments of this application provide a vehicle trajectory prediction method and apparatus, and a trajectory prediction model training method and apparatus. In the trajectory prediction solution provided by this application, when predicting the vehicle trajectory based on the hidden state information of the previous frame image and the current frame image, a reset information and an update information are obtained through a convolutional neural network layer with a smaller amount of computation, and a predicted image frame is obtained based on the reset information and the update information. Therefore, the prediction speed and prediction efficiency of vehicle trajectory prediction can be improved.
[0005] According to one aspect of this application, a vehicle trajectory prediction method is provided, including: first, obtaining the hidden state information of the previous frame image and the current frame image; then, based on the first feature fusion layer in the trajectory prediction model, fusing the current frame image and the hidden state information of the previous frame image to obtain a first feature map; then, based on the first convolutional neural network layer in the trajectory prediction model, processing the first feature map to obtain reset information; the first convolutional neural network layer includes a first activation function; at the same time, based on the second convolutional neural network layer in the trajectory prediction model, processing the first feature map to obtain update information, and the second convolutional neural network layer includes a second activation function; finally, based on the prediction layer in the trajectory prediction model, processing the reset information, the update information, the current image frame, and the hidden state information of the previous frame image to obtain a predicted frame image.
[0006] Based on the above solution, when predicting the vehicle trajectory based on the hidden state information of the previous frame image and the current frame image, both the reset information and the update information are obtained through the convolutional neural network layer. Compared with the reset information and update information of GRU, which are based on the fully connected network layer, since the computational amount of the convolutional neural network layer is smaller than that of the fully connected network layer. Therefore, when predicting the vehicle trajectory, determining the reset information and update information through the convolutional neural network layer can improve the prediction speed and prediction efficiency of the vehicle trajectory. GRU is obtained by simplifying the structure based on LSTM. Compared with LSTM, GRU has a smaller computational amount. Therefore, compared with LSTM, the vehicle trajectory prediction method provided by this application can further improve the prediction speed and prediction efficiency of the vehicle trajectory when predicting the vehicle trajectory.
[0007] According to one aspect of the present application, a method for training a trajectory prediction model is provided, including: First, obtain multiple groups of sample images and first trajectory images corresponding to the sample images one by one. Among them, the sample image includes the hidden state information of the second trajectory image and the third trajectory image; the second trajectory image is the previous frame trajectory image of the third trajectory image in the vehicle trajectory video sequence, and the first trajectory image is the next frame trajectory image of the third trajectory image in the vehicle trajectory video sequence. Then, based on the first feature fusion layer in the initial trajectory prediction model, fuse the third trajectory image and the hidden state information of the second trajectory image to obtain a first training feature map. Then, based on the first convolutional neural network layer in the initial trajectory prediction model, process the first training feature map to obtain training reset information. The first convolutional neural network layer includes a first activation function. At the same time, based on the second convolutional neural network layer in the initial trajectory prediction model, process the first training feature map to obtain training update information. The second convolutional neural network layer includes a second activation function. Then, based on the prediction layer in the initial trajectory prediction model, process the training reset information, training update information, the hidden state information of the second trajectory image, and the third trajectory image to obtain an initial predicted frame image. Finally, use the initial predicted frame image as the initial training output of the initial trajectory prediction model, and the first trajectory image as the supervision information, and iteratively train the initial trajectory prediction model to obtain the trained trajectory prediction model.
[0008] Based on the above technical solution, when training the trajectory prediction model, since both the training reset information and the training update information are obtained through the convolutional neural network layer, compared with the reset information and the update information of the GRU which are obtained based on the fully connected network layer, because the computational amount of the convolutional neural network layer is smaller than that of the fully connected network layer, determining the training reset information and the training update information through the convolutional neural network layer can improve the training speed and training efficiency of the trajectory prediction model. The GRU is obtained by simplifying the structure on the basis of the LSTM. Compared with the LSTM, the computational amount of the GRU is smaller. Therefore, compared with the LSTM, when training the trajectory prediction model in this application, the training speed and training efficiency of the trajectory prediction model can be further improved.
[0009] According to one aspect of the present application, a vehicle trajectory prediction device is provided, including: an acquisition module, configured to acquire the hidden state information of the previous frame image and the current frame image; the previous frame image is the previous frame image of the current frame image in the vehicle trajectory video sequence; a processing module, configured to fuse the current frame image and the hidden state information of the previous frame image acquired by the acquisition module based on the first feature fusion layer in the trajectory prediction model to obtain a first feature map; the processing module is further configured to process the first feature map based on the first convolutional neural network layer in the trajectory prediction model to obtain reset information; the first convolutional neural network layer includes a first activation function; the processing module is further configured to process the first feature map based on the second convolutional neural network layer in the trajectory prediction model to obtain update information, and the second convolutional neural network layer includes a second activation function; the processing module is further configured to process the reset information, the update information, the current image frame, and the hidden state information of the previous frame image based on the prediction module in the trajectory prediction model to obtain a predicted frame image.
[0010] According to one aspect of the present application, there is provided a device for training a trajectory prediction model, including: an acquisition module, configured to acquire multiple sets of sample images and first trajectory images corresponding to the sample images one by one; the sample images include the hidden state information of a second trajectory image and a third trajectory image; wherein, the second trajectory image is the previous frame trajectory image of the third trajectory image in a vehicle trajectory video sequence, and the first trajectory image is the next frame trajectory image of the third trajectory image; a training module, configured to fuse the third trajectory image and the hidden state information of the second trajectory image based on a first feature fusion layer in an initial trajectory prediction model to obtain a first training feature map; the training module is further configured to process the first training feature map based on a first convolutional neural network layer in the initial trajectory prediction model to obtain training reset information; the first convolutional neural network layer includes a first activation function; the training module is further configured to process the first training feature map based on a second convolutional neural network layer in the initial trajectory prediction model to obtain training update information, and the second convolutional neural network layer includes a second activation function; the training module is further configured to process the training reset information, the training update information, the hidden state information of the second trajectory image, and the third trajectory image based on a prediction layer in the initial trajectory prediction model to obtain an initial prediction frame image; the training module is further configured to use the initial prediction frame image as the initial training output of the initial trajectory prediction model, and the first trajectory image acquired by the acquisition module as supervision information, and iteratively train the initial trajectory prediction model to obtain a trained trajectory prediction model.
[0011] According to one aspect of the present application, there is provided a computer-readable storage medium storing a computer program for executing the method provided in any of the above aspects.
[0012] According to one aspect of the present application, there is provided an electronic device, which includes: a processor; a memory for storing executable instructions of the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the method provided in any of the above aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0014] Figure 1 FIG. 1 is a schematic structural diagram of an LSTM provided by the prior art.
[0015] Figure 2It is a schematic structural diagram of an electronic device provided by this application.
[0016] Figure 3 It is a schematic flowchart of a vehicle trajectory prediction method provided by this application Figure 1 。
[0017] Figure 4 It is a schematic structural diagram of a trajectory prediction model provided by this application Figure 1 。
[0018] Figure 5 It is a schematic flowchart of a vehicle trajectory prediction method provided by this application Figure 2 。
[0019] Figure 6 It is a schematic structural diagram of a trajectory prediction model provided by this application Figure 2 。
[0020] Figure 7 It is a schematic flowchart of a vehicle trajectory prediction method provided by this application Figure 3 。
[0021] Figure 8 It is a schematic structural diagram of a trajectory prediction model provided by this application Figure 3 。
[0022] Figure 8A It is a schematic structural diagram of a trajectory prediction model provided by this application Figure 4 。
[0023] Figure 9 It is a schematic flowchart of a trajectory prediction model training method provided by this application Figure 1 。
[0024] Figure 10 It is a schematic flowchart of a trajectory prediction model training method provided by this application Figure 2
[0025] Figure 11 It is a schematic flowchart of a trajectory prediction model training method provided by this application Figure 3 。
[0026] Figure 12 It is a schematic flowchart of a trajectory prediction model training method provided by this application Figure 4 。
[0027] Figure 13 It is a schematic structural diagram of a vehicle trajectory prediction device provided by this application Figure 1 。
[0028] Figure 14 It is a schematic structural diagram of a vehicle trajectory prediction device provided by this application Figure 2 。
[0029] Figure 15 is a schematic structure diagram of a vehicle trajectory prediction device provided by the present application Figure 3 。
[0030] Figure 16 is a schematic structure diagram of a trajectory prediction model training device provided by the present application Figure 1 。
[0031] Figure 17 is a schematic structure diagram of a trajectory prediction model training device provided by the present application Figure 2 。
[0032] Figure 18 is a schematic structure diagram of a trajectory prediction model training device provided by the present application Figure 3 。
[0033] Figure 19 is a schematic structure diagram of a trajectory prediction model training device provided by the present application Figure 4 。
[0034] Figure 20 is a schematic structure diagram of another electronic device provided by the present application. Detailed implementation manners
[0035] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.
[0036] Hereinafter, the terms "first" and "second" are only used for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise stated, the meaning of "a plurality" is two or more. "A and / or B" includes the following three combinations: only A, only B, and the combination of A and B.
[0037] Generally, the trajectory prediction of a vehicle is performed based on the vehicle trajectory image of the vehicle in the past period of time. Based on this, in the related art, most of the trajectory predictions use a long short-term memory (LSTM) network as a model to train a trajectory prediction model capable of predicting the vehicle trajectory.
[0038] The long short-term memory network is composed of at least one memory model as shown in Figure 1 shown. Referring to Figure 1 shown, this memory model includes three gating modules, namely the forget gate ft , the input gate i t and the output gate o t . The input of this memory model includes: the cell state information c at the previous moment t-1 , the hidden state information h at the previous moment t-1 and the input information x at the current moment t . The input of this memory model needs to be processed by each computing module inside this memory model to output the cell state information c t and the hidden state information h t . Among them, the forget gate f t , the input gate i t and the output gate o t all need to be obtained through the processing of the fully connected network layer, with a large amount of computation and low computational efficiency.
[0039] In order to reduce the number of model parameters, the vehicle trajectory can be predicted based on the gated recurrent unit (GRU). However, when predicting the vehicle trajectory based on GRU, the reset information and the update information are also obtained based on the fully connected network layer, so the computational amount is still large and the computational efficiency is low.
[0040] In view of the above problems, the embodiments of the present application provide a vehicle trajectory prediction method. In the trajectory prediction scheme provided by the present application, when predicting the vehicle trajectory based on the hidden state information of the previous frame image and the current frame image, both the reset information and the update information are obtained through the convolutional neural network layer. Compared with the reset information and the update information of GRU obtained based on the fully connected network layer, the computational amount of the convolutional neural network layer is much smaller than that of the fully connected network layer. Therefore, when predicting the vehicle trajectory, determining the reset information and the update information through the convolutional neural network layer can improve the prediction speed and prediction efficiency of the vehicle trajectory prediction. And GRU is obtained by simplifying the structure on the basis of LSTM. Compared with LSTM, GRU has a smaller computational amount. Therefore, compared with LSTM, the vehicle trajectory prediction method provided by the present application can further improve the prediction speed and prediction efficiency of the vehicle trajectory prediction when predicting the vehicle trajectory.
[0041] Figure 2 As shown in the structural diagram of an electronic device according to an exemplary embodiment, the vehicle trajectory prediction method provided by the embodiments of the present application can be applied to Figure 2 the electronic device 200 shown, as Figure 2 shown, the electronic device 200 includes one or more processors 201 and a memory 202.
[0042] The processor 201 can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 200 to perform desired functions.
[0043] The memory 202 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 201 may run the program instructions to implement the vehicle trajectory prediction method or trajectory prediction model training method of each embodiment of the present application described above and / or other desired functions.
[0044] In one example, the electronic device 200 may further include: an input device 203 and an output device 1604, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0045] Of course, for simplicity, Figure 2 only some of the components related to the present application in the electronic device 200 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 200 may further include any other appropriate components.
[0046] The vehicle trajectory prediction method provided by the embodiments of the present application will be described below with reference to the accompanying drawings.
[0047] Figure 3 A vehicle trajectory prediction method provided by an embodiment of the present application. Refer to Figure 3 As shown, the method may include S301 - S305:
[0048] S301. Obtain the hidden state information of the previous frame image and the current frame image.
[0049] Among them, the previous frame image is the previous frame image of the current frame image in the vehicle trajectory video sequence. In the embodiments of the present application, the vehicle trajectory video sequence here may be multiple trajectory images of the vehicle obtained by the electronic device in real time.
[0050] In the embodiments of the present application, since the prediction of the vehicle trajectory needs to consider the historical trajectory during the vehicle driving process, the hidden state information of the previous frame image here not only includes the relevant feature information of the previous frame image, but also can include the relevant feature information of all the frame images in the vehicle trajectory video sequence before the current frame image. Therefore, the image predicted based on the relevant feature information of all the frame images before the current frame image greatly improves the prediction accuracy of the vehicle driving trajectory.
[0051] Exemplarily, the electronic device can obtain the current frame image by controlling various sensors or image acquisition devices on the vehicle, or can analyze and obtain the current frame image according to the data obtained by various sensors or image acquisition devices on the vehicle.
[0052] S302. Based on the first feature fusion layer in the trajectory prediction model, fuse the current frame image and the hidden state information of the previous frame image to obtain a first feature map.
[0053] In the embodiments of the present application, the trajectory prediction model is used to predict the next frame trajectory image based on the historical trajectory image information and the current frame image. Refer to Figure 4 As shown, the obtained current frame image x t and the hidden state information h t of the previous frame image are input into the trajectory prediction model. First, the first feature fusion layer 31 in the trajectory prediction model will fuse the current frame image x t and the hidden state information h t of the previous frame image. Specifically, the first feature fusion layer 31 can fuse the current frame image x t and the hidden state information h t of the previous frame image through a fusion concat operation, so as to obtain a first feature map.
[0054] Exemplarily, the concat operation in the present application can specifically refer to stacking x t and h t on the channel layer. Taking (width, height, number of channels channels) to represent x t and h t as an example, the first feature map can be (width, height, 2 * channels).
[0055] S303. Based on the first convolutional neural network layer in the trajectory prediction model, process the first feature map to obtain reset information.
[0056] Among them, the reset information r t is used to determine the degree of forgetting. For example, the reset information r tFor determining how much information of the previous state information (i.e., the hidden state information h of the previous frame image) t ) is written into the subsequent candidate hidden state information h t+1' . The reset information r t is obtained by the electronic device based on processing the first feature map by the first convolutional neural network layer. In some embodiments, each matrix element value in the reset information r t is a value between 0 and 1.
[0057] Exemplarily, the first convolutional neural network layer includes a first activation function. Exemplarily, the first activation function can be a sigmoid function, and its expression can be the following formula (1):
[0058]
[0059] Referring to Figure 4 as shown, after the first feature fusion layer 31 obtains or outputs the first feature map, the first convolutional neural network layer 32 (which can also be called a reset gate) can receive the first feature map and process the first feature map to obtain the reset information r t .
[0060] Exemplarily, the operation process of the first convolutional neural network layer can be represented by the following formula (2):
[0061] r t = sigmoid(conv(concat(h t , x t ))) (2)
[0062] Wherein, conv() represents a convolution operation, and concat() represents a fusion operation.
[0063] It can be understood that since the reset information of the existing GRU is obtained by processing the input information based on a fully connected network layer, and when the fully connected network layer processes the input information, each node has to be fully connected to each node in the next layer, and each connection has parameters participating in the operation, so the calculation amount of the fully connected network layer is large and the operation efficiency is low. However, the reset information in the present application is obtained by performing a convolution operation on the first feature map and the convolution kernel of the first convolutional neural network layer 32. Since the operation of the convolutional neural network layer is only related to the convolution kernel size and the number of channels of the output feature map, the operation amount of the convolutional neural network layer can be further reduced compared with that of the fully connected network layer, so the operation speed and operation efficiency of the convolutional neural network layer are greatly improved.
[0064] S304. Process the first feature map based on the second convolutional neural network layer in the trajectory prediction model to obtain updated information.
[0065] Among them, the update information z t is used to determine the retention degree. For example, the update information z t is used to determine how much information of the previous state information (i.e., the hidden state information h of the previous frame image t ) should be combined into the current state information, and how much information in the candidate hidden state information h t+1' needs to be retained. The update information z t is obtained by the electronic device based on the second convolutional neural network layer processing the first feature map. In some embodiments, the update information z t the numerical value of each matrix element in is a numerical value between 0 and 1.
[0066] Exemplarily, the second convolutional neural network layer includes a second activation function. Exemplarily, the second activation function can be the same as the first activation function, that is, the second activation function can also be the sigmoid function, and its expression can be the above formula (1).
[0067] Referring to Figure 4 as shown, after the first feature fusion layer 31 obtains or outputs the first feature map, the second convolutional neural network layer 33 (which can also be called the update gate) can receive the first feature map and process the operation of the first feature map to obtain the update information z t .
[0068] Exemplarily, the operation process of the second convolutional neural network layer can be expressed by the following formula (3):
[0069] z t = sigmoid(conv(concat(h t , x t ))) (3)
[0070] Among them, conv() represents the convolution operation, and concat() represents the fusion operation.
[0071] It can be understood that since the update information of the existing GRU is obtained by the fully connected network layer processing the input information, and when the fully connected network layer processes the input information, each node has to be fully connected to each node in the next layer, and each connection has parameters participating in the operation, so the calculation amount of the fully connected network layer is large and the operation efficiency is low. And the update information in this application is obtained by performing a convolution operation on the first feature map and the convolution kernel of the second convolutional neural network layer 33. Since the operation of the convolutional neural network layer is only related to the convolution kernel size and the number of channels of the output feature map, the operation amount of the convolutional neural network layer can be further reduced compared with that of the fully connected network layer, so the operation speed and operation efficiency of the convolutional neural network layer are greatly improved.
[0072] S305. Process the reset information, update information, current image frame, and hidden state information of the previous frame image based on the prediction layer in the trajectory prediction model to obtain a predicted frame image.
[0073] After obtaining the reset information r in S303 t , and obtaining the update information z in S304 t , the electronic device can Figure 4 process the reset information r t , update information z t , current image frame x t , and the hidden state information h of the previous frame image t through the prediction layer 34 in the trajectory prediction model shown in t to complete the prediction of the current frame image x t and obtain the predicted frame image y.
[0074] When the vehicle trajectory prediction method provided by the embodiment of the present application performs vehicle trajectory prediction based on the hidden state information of the previous frame image and the current frame image, both the reset information and the update information are obtained through the convolutional neural network layer. Compared with the reset information and update information of the existing GRU obtained based on the fully connected network layer, since the computational amount of the convolutional neural network layer is smaller than that of the fully connected network layer. Therefore, when predicting the vehicle trajectory, determining the reset information and update information through the convolutional neural network layer can improve the prediction speed and prediction efficiency of vehicle trajectory prediction. And GRU is obtained by simplifying the structure on the basis of LSTM. Compared with LSTM, GRU has a smaller computational amount. Therefore, compared with LSTM, when the vehicle trajectory prediction method provided by the present application predicts the vehicle trajectory, it can not only further improve the prediction speed and prediction efficiency of vehicle trajectory prediction, but also has a higher prediction accuracy for the vehicle driving trajectory.
[0075] In some embodiments, in combination with Figure 2 , referring to Figure 5 shown, the above S305 may specifically include S3051 - S3054:
[0076] S3051. Multiply the reset information element by element with the hidden state information of the previous frame image based on the selection layer in the prediction layer to obtain a second feature map.
[0077] Referring to Figure 6 shown, after the first convolutional neural network layer 32 in the trajectory prediction model obtains the reset information r t , the selection layer 341 in the prediction layer 34 of the trajectory prediction model can select the hidden state information h of the previous frame image according to the reset information r t and then determine the hidden state information h of the previous frame image t t The content to be written into the candidate hidden state information h t+1' or to determine the hidden state information h of the previous frame image t needs to participate in generating the candidate hidden state information h t+1' content.
[0078] In some embodiments, the selection layer 341 in the prediction layer 34 can multiply the reset information r t element-wise with the hidden state information h of the previous frame image t to select the hidden state information h of the previous frame image t and obtain a second feature map.
[0079] S3052. Based on the second feature fusion layer in the prediction layer, fuse the second feature map with the current image frame to obtain a third feature map.
[0080] Refer to Figure 6 As shown, after the selection layer 341 in the prediction layer 34 of the trajectory prediction model obtains the second feature map, the second feature fusion layer 342 in the prediction layer 34 of the trajectory prediction model can fuse the second feature map with the current frame image x t to obtain a third feature map. In this way, when generating the candidate hidden state information h based on the third feature map t+1' it can make the subsequently generated candidate hidden state information h t+1' be able to combine the information of the current image frame x t information.
[0081] Exemplarily, the second feature fusion layer 342 can fuse the second feature map with the current frame image through a fusion concat operation to obtain a third feature map. The concat operation in this application can refer to stacking the two images (or feature maps) to be fused on the channel layer.
[0082] S3053. Based on the third convolutional neural network layer and the activation function layer in the prediction layer, process the third feature map in sequence to obtain the candidate hidden state information.
[0083] Refer to Figure 6 As shown, after the second feature fusion layer 342 in the prediction layer 34 of the trajectory prediction model obtains the third feature map, the third convolutional neural network layer 343 in the prediction layer 34 can perform a convolution operation on the third feature map and the convolution kernel in the third convolutional neural network layer 343, and then input the processed third feature map into the activation function layer 344. The activation function layer 344 processes the processed third feature map again to obtain the candidate hidden state information h t+1' .
[0084] Exemplarily, the activation function layer includes a third activation function, which can be tanh, and its expression can be the following formula (4):
[0085]
[0086] Exemplarily, the candidate hidden state information h t+1' can be calculated through the following formula (5).
[0087] h_t+1′ = tanh(conv(concat(r t *h t , x t ))) (5)
[0088] Among them, conv() represents the convolution operation, concat() represents the fusion operation, and * represents element-wise multiplication.
[0089] It can be understood that since the computational amount of the convolutional neural network layer is small, before inputting the third feature map into the activation function layer, the third convolutional neural network layer 343 performs a convolution operation on the third feature map and the convolution kernel, which can further reduce the computational amount in the vehicle trajectory prediction process, thereby further improving the prediction speed and prediction efficiency of the vehicle trajectory prediction.
[0090] S3054. Process the candidate hidden state information, the update information, and the hidden state information of the previous frame image based on the output layer in the prediction layer to obtain the predicted frame image.
[0091] Refer to Figure 6 As shown, after obtaining the candidate hidden state information h t+1' , the output layer 345 in the prediction layer 34 of the trajectory prediction model can combine the update information z t and the hidden state information h t of the previous frame image to predict the current frame image x t and obtain the predicted frame image y t .
[0092] When the vehicle trajectory prediction method provided by the implementation of this application performs vehicle trajectory prediction based on the hidden state information of the previous frame image and the current frame image, since both the reset information and the update information are obtained through the convolutional neural network layer, and the computational complexity of the convolutional neural network layer is smaller than that of the fully connected network layer. Therefore, compared with the prior art, when the vehicle trajectory prediction method provided by this application predicts the vehicle trajectory, by determining the reset information and the update information through the convolutional neural network layer, the prediction speed and prediction efficiency of vehicle trajectory prediction can be improved. Moreover, when obtaining the candidate hidden state information, the third feature map will be first subjected to convolutional processing, so that the computational complexity in the vehicle trajectory prediction process can be further reduced, and the prediction speed and prediction efficiency of vehicle trajectory prediction can be improved.
[0093] In some embodiments, in combination with Figure 5 , with reference to Figure 7 shown, the above S3054 may include S1 - S4:
[0094] S1. Based on the selective memory sub - layer in the output layer, multiply the candidate hidden state information and the update information element - by - element to obtain the fourth feature map.
[0095] The update information z t is used to determine how much information in the candidate hidden state information h t+1' needs to be retained, that is, the update information z t can selectively retain the candidate hidden state information h t+1' . Based on this, in combination with Figure 6 , with reference to Figure 8 shown, after the activation function layer 344 in the prediction layer 34 outputs the candidate hidden state information h t+1' , the selective memory sub - layer 3451 in the output layer 345 can multiply the update information z t and the candidate hidden state information h t+1' element - by - element to selectively retain the candidate hidden state information h t+1' and obtain the fourth feature map.
[0096] S2. Based on the selective forgetting sub - layer in the output layer, perform subtraction processing on each element in the update information to obtain the weight of the forgetting information, and multiply the weight of the forgetting information and the hidden state information of the previous frame image element - by - element to obtain the fifth feature map.
[0097] The update information z t is also used to determine how much information in the hidden state information h t of the previous frame image needs to be retained. Since when predicting the vehicle trajectory, a part of the hidden state information h t of the previous frame image needs to be removed / forgotten, so the update information z tPerform subtraction on each element in [[]] to obtain the weight of the forgotten information, and this weight of the forgotten information can determine the degree of forgetting.
[0098] Exemplarily, performing subtraction on each element in the updated information based on the selective forgetting sublayer in the output layer, the weight of the forgotten information obtained includes: subtracting each element in the updated information from 1 to obtain the weight of the forgotten information.
[0099] Referring to Figure 8 As shown, after the second convolutional neural network layer 33 of the trajectory prediction model outputs the updated information h t , the selective forgetting sublayer 3452 in the output layer 345 can subtract each element in the updated information z t from 1, thereby obtaining the weight of the forgotten information. Then, the selective forgetting sublayer 3452 can multiply the weight of the forgotten information by the hidden state information of the previous frame image, thereby obtaining the fifth feature map. This fifth feature map is the feature map after forgetting a part of the information of the hidden state information h t of the previous frame image.
[0100] S3. Based on the synthesis sublayer in the output layer, add the fourth feature map and the fifth feature map element by element to obtain the hidden state information of the current frame image.
[0101] Referring to Figure 8 As shown, after the selective forgetting sublayer 3452 in the output layer 345 of the trajectory prediction model outputs the fifth feature map and the selective memory sublayer 3451 in the output layer 345 outputs the fourth feature map, the synthesis sublayer 3453 in the output layer 345 can add the fourth feature map and the fifth feature map element by element. By adding the fourth feature map and the fifth feature map element by element, the unimportant information in the hidden state information h t of the previous frame image can be forgotten and the important information in the current frame image can be added to obtain the hidden state information h t+1 of the current frame image.
[0102] Exemplarily, the hidden state information h t+1 of the current frame image can be calculated by the following formula (6):
[0103] h t+1 =(1 - z t ) * h t + z t * h′ (6)
[0104] where, * represents element-wise multiplication.
[0105] It should be noted that in the foregoing embodiments Figure 4 , Figure 6 and Figure 8An example is given with the trajectory prediction model including a convolutional gated recurrent unit. The trajectory prediction model may also include multiple convolutional gated recurrent units. Refer to Figure 8A As shown, the hidden state information of the current frame image output by the convolutional gated recurrent unit 1 will be used as the hidden state information of the previous frame image input to the next convolutional gated recurrent unit (such as Figure 8A the convolutional gated recurrent unit 2 shown). It can be understood that when the trajectory prediction model includes multiple convolutional gated recurrent units, the model parameters of these multiple convolutional gated recurrent units are the same. In practical applications, the trajectory prediction model usually includes a single convolutional gated recurrent unit.
[0106] S4. Process the hidden state information of the current frame image based on the fourth convolutional neural network layer in the output layer to obtain the predicted frame image.
[0107] Refer to Figure 8 As shown, since there is a certain difference between the hidden state information h t+1 of the current frame image and the real image (for example, the hidden state information h t+1 of the current frame image is some feature maps that users cannot understand), after the output layer 345 outputs the hidden state information h t+1 of the current frame image, the fourth convolutional neural network layer 3454 in the output layer 345 of the trajectory prediction model can process the hidden state information h t+1 of the current frame image to obtain the predicted frame image y t . It can be understood that by processing the hidden state information h t+1 of the current frame image through the fourth convolutional neural network layer 3454, image information that users can understand can be obtained, and this image information is the information of the predicted frame image y t .
[0108] Exemplarily, the fourth convolutional neural network layer includes a fourth activation function, and this fourth activation function can be relu, and its expression can be the following formula (7):
[0109] y = max(0, x) (7)
[0110] In the embodiments of the present application, the predicted frame image is the prediction situation of the next frame image predicted from the current frame image.
[0111] Exemplarily, the predicted frame image y t can be calculated through the following formula (8).
[0112] y t = max(0, conv(h t+1 )) (8)
[0113] Among them, conv() represents the convolution operation.
[0114] Based on the above technical solution, the output layer in the prediction layer of the trajectory prediction model can process the candidate hidden state information, update information, and the hidden state information of the previous frame image to obtain the predicted frame image, thereby realizing the prediction of the vehicle trajectory.
[0115] In the embodiments of the present application, in order to improve the accuracy of vehicle trajectory prediction, model training can be performed in advance (at least before S302) to obtain the trajectory prediction model used in the foregoing embodiments. Based on this, the embodiments of the present application further provide a training method for a trajectory prediction model. Refer to Figure 9 As shown, a training method for a trajectory prediction model provided by the present application may include S901 - S906:
[0116] S901. Obtain multiple groups of sample images and the first trajectory images corresponding to the sample images one by one.
[0117] Among them, the sample images include the hidden state information of the second trajectory image and the third trajectory image; where the second trajectory image is the previous frame trajectory image of the third trajectory image in the vehicle trajectory video sequence, and the first trajectory image is the next frame trajectory image of the third trajectory image in the vehicle trajectory video sequence.
[0118] In some embodiments, the electronic device can obtain at least one vehicle trajectory video sequence, and then extract image frames from the at least one vehicle trajectory video sequence as samples, and select multiple consecutive frames of images from the samples as the initial samples. The initial samples include multiple groups of sample images, and each group of sample images includes a sample pair of the second trajectory image, the third trajectory image, and the first trajectory image.
[0119] The hidden state information of the above - mentioned second trajectory image can be the initial hidden state information, or can be obtained by inputting the hidden state information of the previous frame image of the second trajectory image and the second trajectory image into the initial trajectory prediction model.
[0120] After obtaining multiple groups of sample images and the first trajectory images corresponding to each group of sample images one by one, the hidden state information of the second trajectory image and the third trajectory image in each group of sample images can be input into the initial trajectory prediction model to obtain the predicted frame image; then, based on the predicted frame image, using the first trajectory image corresponding to this group of sample images as the supervision information, the initial trajectory prediction model is iteratively trained to obtain the trained trajectory prediction model. This trained trajectory prediction model can more accurately predict the vehicle trajectory. The training process of the trajectory prediction model is introduced below through S902 - S906.
[0121] S902. Based on the first feature fusion layer in the initial trajectory prediction model, fuse the hidden state information of the third trajectory image and the second trajectory image to obtain the first training feature map.
[0122] In some embodiments, during initial training, the parameters in the initial trajectory prediction model can be set to 0.
[0123] It can be understood that the specific implementation of S902 can refer to the specific implementation of S302 in the foregoing embodiments, which will not be elaborated here.
[0124] S903. Based on the first convolutional neural network layer in the initial trajectory prediction model, process the first training feature map to obtain training reset information.
[0125] Among them, the first convolutional neural network layer includes a first activation function. Exemplarily, the first activation function is sigmoid.
[0126] S904. Based on the second convolutional neural network layer in the initial trajectory prediction model, process the first training feature map to obtain training update information.
[0127] Among them, the second convolutional neural network layer includes a second activation function. Exemplarily, the second activation function is sigmoid.
[0128] It can be understood that the specific implementations of S903 and S904 can refer to the specific implementations of S302 and S303 in the foregoing embodiments, which will not be elaborated here.
[0129] S905. Based on the prediction layer in the initial trajectory prediction model, process the training reset information, training update information, the hidden state information of the second trajectory image, and the third trajectory image to obtain the initial prediction frame image.
[0130] In some embodiments, in combination with Figure 9 , referring to Figure 10 shown, S905 may specifically include S9051 - S9054:
[0131] S9051. Based on the selection layer in the prediction layer, multiply the training reset information element - by - element with the hidden state information of the second trajectory image to obtain the second training feature map.
[0132] The specific implementation of S9051 can refer to S3051 in the foregoing embodiments, which will not be elaborated here.
[0133] S9052. Based on the second feature fusion layer in the prediction layer, fuse the second training feature map and the third trajectory image to obtain the third training feature map.
[0134] S9053. Process the third training feature map successively based on the third convolutional neural network layer and the activation function layer in the prediction layer to obtain training candidate hidden state information.
[0135] Among them, the activation function layer includes a third activation function. Exemplarily, the third activation function can be tanh.
[0136] For the specific implementation of S9052 and S9053, reference can be made to S3052 and S3053 in the foregoing embodiments, which will not be elaborated here.
[0137] S9054. Process the training candidate hidden state information, the training update information, and the hidden state information of the second trajectory image based on the output layer in the prediction layer to obtain an initial predicted frame image.
[0138] Based on the above implementation solution, when obtaining candidate hidden state information, the third feature map will be first subjected to convolutional processing. Since the amount of computation of the convolutional neural network layer is small, the number of parameters in the training process of the trajectory prediction model can be reduced, and the training speed and training efficiency of the trajectory prediction model can be improved.
[0139] In some embodiments, in combination with Figure 10 , referring to Figure 11 shown, S9054 may specifically include X1 - X4:
[0140] X1. Multiply the training candidate hidden state information and the training update information element by element based on the selective memory sublayer in the output layer to obtain a fourth training feature map.
[0141] X2. Perform subtraction processing on each element in the training update information based on the selective forgetting sublayer in the output layer to obtain the weight of the training forgetting information, and multiply the weight of the training forgetting information and the hidden state information of the second trajectory image element by element to obtain a fifth training feature map.
[0142] When the implementation manner of S901 is a possible implementation manner as described after the foregoing S901, X2 may specifically be: perform subtraction processing on each element in the training update information based on the selective forgetting sublayer in the output layer to obtain the weight of the training forgetting information, and multiply the weight of the training forgetting information and the initial hidden state information element by element to obtain a fifth training feature map.
[0143] X3. Add the fourth training feature map and the fifth training feature map element by element based on the comprehensive sublayer in the output layer to obtain the hidden state information of the third trajectory image.
[0144] X4. Process the hidden state information of the third trajectory image based on the fourth convolutional neural network layer in the output layer to obtain an initial predicted frame image.
[0145] Among them, the fourth convolutional neural network layer includes a fourth activation function. Exemplarily, the fourth activation function is relu.
[0146] The specific implementation of X1 - X4 can refer to S1 - S4 in the foregoing embodiments, and will not be elaborated here.
[0147] S906. Use the initial predicted frame image as the initial training output of the initial trajectory prediction model, and the first trajectory image as the supervision information, and iteratively train the initial trajectory prediction model to obtain the trained trajectory prediction model.
[0148] In some embodiments, in combination Figure 9 , with reference to Figure 12 as shown, S906 may specifically include: S9061 and S9062:
[0149] S9061. Determine the loss value according to the initial predicted frame image and the first trajectory image.
[0150] Exemplarily, S9061 may specifically use a loss function to determine the loss value. Exemplarily, the loss function may be specifically the following formula (9):
[0151] loss = ∑(x t+1 - y t ) 2 (9)
[0152] Among them, loss is the loss value, x t+1 is the next - frame trajectory image corresponding to the input of the initial trajectory prediction model, and y t is the predicted frame image output by the initial trajectory prediction model.
[0153] S9062. Iteratively update the initial trajectory prediction model according to the loss value to obtain the trained trajectory prediction model.
[0154] It can be understood that through continuous iterative optimization of the loss value, a trajectory prediction model that can accurately achieve vehicle trajectory prediction can be obtained.
[0155] Based on the technical solutions corresponding to the foregoing S901 - S906, when training the trajectory prediction model, since both the training reset information and the training update information are obtained through the convolutional neural network layer, compared with the existing GRU where the reset information and the update information are based on the fully - connected network layer, because the computational amount of the convolutional neural network layer is smaller than that of the fully - connected network layer, determining the training reset information and the training update information through the convolutional neural network layer can improve the training speed and training efficiency of the trajectory prediction model. And GRU is obtained by simplifying the structure based on LSTM. Compared with LSTM, GRU has a smaller computational amount. Therefore, compared with LSTM, when training the trajectory prediction model in this application, the training speed and training efficiency of the trajectory prediction model can be further improved.
[0156] It can be understood that in order to implement the above functions, the above - mentioned electronic device includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, the embodiments of this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but this implementation should not be considered to exceed the scope of the embodiments of this application.
[0157] In the case of dividing each functional module corresponding to each function, the embodiments of this application also provide a vehicle trajectory prediction device. As Figure 13 shown, it is a schematic structural diagram of a vehicle trajectory prediction device provided by the embodiments of this application. The device may include: an acquisition module 1301 and a processing module 1302.
[0158] Among them, the acquisition module 1301 is used to acquire the hidden state information of the previous frame image and the current frame image; the previous frame image is the previous frame image of the current frame image in the vehicle trajectory video sequence; the processing module 1302 is used to fuse the hidden state information of the current frame image and the previous frame image acquired by the acquisition module 1301 based on the first feature fusion layer in the trajectory prediction model to obtain a first feature map; the processing module 1302 is further used to process the first feature map based on the first convolutional neural network layer in the trajectory prediction model to obtain reset information; the first convolutional neural network layer includes a first activation function; the processing module 1302 is further used to process the first feature map based on the second convolutional neural network layer in the trajectory prediction model to obtain update information, and the second convolutional neural network layer includes a second activation function; the processing module 1302 is further used to process the reset information, the update information, the current image frame and the hidden state information of the previous frame image based on the prediction module in the trajectory prediction model to obtain a predicted frame image.
[0159] In some embodiments, in combination with Figure 13 , with reference to Figure 14 shown, the processing module 1302 may include a selection unit 13021, a fusion unit 13022, a candidate unit 13023 and a processing unit 13024. Among them, the selection unit 13021 is used to multiply the reset information and the hidden state information of the previous frame image element by element based on the selection layer in the prediction layer to obtain a second feature map; the fusion unit 13022 is used to fuse the second feature map obtained by the selection unit 13021 and the current image frame based on the second feature fusion layer in the prediction layer to obtain a third feature map; the candidate unit 13023 is used to process the third feature map obtained by the fusion unit 13022 in sequence based on the third convolutional neural network layer and the activation function layer in the prediction layer to obtain candidate hidden state information, and the activation function layer includes a third activation function; the processing unit 13024 is used to process the update information, the hidden state information of the previous frame image and the candidate hidden state information obtained by the candidate unit 13023 based on the output layer in the prediction layer to obtain a predicted frame image.
[0160] In some embodiments, in combination with Figure 14 , with reference to Figure 15As shown, the processing unit 13024 may specifically include a first sub-unit 1501, a second sub-unit 1502, a third sub-unit 1503, and a fourth sub-unit 1504. Among them, the first sub-unit 1501 is configured to multiply the update information and the candidate hidden state information obtained by the candidate unit 13023 element by element based on the selective memory sub-layer in the output layer to obtain a fourth feature map; the second sub-unit 1502 is configured to perform a subtraction process on each element in the update information based on the selective forgetting sub-layer in the output layer to obtain the weight of the forgetting information, and multiply the weight of the forgetting information and the hidden state information of the previous frame image element by element to obtain a fifth feature map; the third sub-unit 1503 is configured to add the fourth feature map obtained by the first sub-unit 1501 and the fifth feature map obtained by the second sub-unit 1502 element by element based on the integration sub-layer in the output layer to obtain the hidden state information of the current frame image; the fourth sub-unit 1504 is configured to process the hidden state information of the current frame image obtained by the third sub-unit 1503 based on the fourth convolutional neural network layer in the output layer to obtain a predicted frame image, and the fourth convolutional neural network layer includes a fourth activation function.
[0161] Regarding the vehicle trajectory prediction device in the above embodiments, the specific manners in which each module performs operations and the corresponding beneficial effects have been described in detail in the embodiments of the vehicle trajectory prediction method described above, and will not be elaborated here.
[0162] In the case of dividing each function into corresponding function modules, an embodiment of the present application further provides a vehicle trajectory prediction device. As Figure 16 shown, it is a schematic structural diagram of a trajectory prediction model training device provided by an embodiment of the present application. The device may include: an acquisition module 1601 and a training module 1602.
[0163] Specifically, an acquisition module 1601 is configured to acquire multiple sets of sample images and first trajectory images corresponding to the sample images one by one; the sample images include the hidden state information of the second trajectory image and the third trajectory image; the second trajectory image is the previous frame trajectory image of the third trajectory image in the vehicle trajectory video sequence, and the first trajectory image is the next frame trajectory image of the third trajectory image in the vehicle trajectory video sequence; a training module 1602 is configured to fuse the third trajectory image and the hidden state information of the second trajectory image based on the first feature fusion layer in the initial trajectory prediction model to obtain a first training feature map; the training module 1602 is further configured to process the first training feature map based on the first convolutional neural network layer in the initial trajectory prediction model to obtain training reset information; the first convolutional neural network layer includes a first activation function; the training module 1602 is further configured to process the first training feature map based on the second convolutional neural network layer in the initial trajectory prediction model to obtain training update information, and the second convolutional neural network layer includes a second activation function; the training module 1602 is further configured to process the training reset information, the training update information, the hidden state information of the second trajectory image, and the third trajectory image based on the prediction layer in the initial trajectory prediction model to obtain an initial prediction frame image; the training module 1602 is further configured to use the initial prediction frame image as the initial training output of the initial trajectory prediction model, and use the first trajectory image acquired by the acquisition module 1601 as supervision information to iteratively train the initial trajectory prediction model to obtain a trained trajectory prediction model.
[0164] In some embodiments, in combination with Figure 16 , with reference to Figure 17 as shown, the training module 1602 includes a training selection unit 16021, a training fusion unit 16022, a training candidate unit 16023, and a training processing unit 16024. Among them, the training selection unit 16021 is configured to multiply the training reset information and the hidden state information of the second trajectory image element by element based on the selection layer in the prediction layer to obtain a second training feature map; the training fusion unit 16022 is configured to fuse the third trajectory image and the second training feature map obtained by the training selection unit 16021 based on the second feature fusion layer in the prediction layer to obtain a third training feature map; the training candidate unit 16023 is configured to sequentially process the third training feature map obtained by the training fusion unit 16022 based on the third convolutional neural network layer and the activation function layer in the prediction layer to obtain training candidate hidden state information, and the activation function layer includes a third activation function; the training processing unit 16024 is configured to process the training update information, the hidden state information of the second trajectory image, and the training candidate hidden state information obtained by the training candidate unit 16023 based on the output layer in the prediction layer to obtain an initial prediction frame image.
[0165] In some embodiments, in combination with Figure 17 , with reference toFigure 18 As shown, the training processing unit 16024 may specifically include a first training subunit 1801, a second training subunit 1802, a third training subunit 1803, and a fourth training subunit 1804.
[0166] Among them, the first training subunit 1801 is used to multiply the training update information and the training candidate hidden state information obtained by the training candidate unit 16023 element by element based on the selective memory sublayer in the output layer to obtain a fourth training feature map; the second training subunit 1802 is used to perform subtraction processing on each element in the training update information based on the selective forgetting sublayer in the output layer to obtain the weight of the training forgetting information, and multiply the weight of the training forgetting information by the hidden state information of the second trajectory image element by element to obtain a fifth training feature map; the third training subunit 1803 is used to add the fourth training feature map obtained by the first training subunit 1801 and the fifth training feature map obtained by the second training subunit 1802 element by element based on the synthesis sublayer in the output layer to obtain the hidden state information of the third trajectory image; the fourth training subunit 1804 is used to process the hidden state information of the third trajectory image obtained by the third training subunit 1803 based on the fourth convolutional neural network layer in the output layer to obtain an initial prediction frame image, and the fourth convolutional neural network layer includes a fourth activation function.
[0167] In some embodiments, in combination with Figure 16 , with reference to Figure 19 As shown, the training module 1602 further includes a loss unit 1901 and an iteration unit 1902. Among them, the loss unit 1901 is used to determine a loss value according to the initial prediction frame image obtained by the training processing unit 16024 and the first trajectory image obtained by the acquisition module 1601; the iteration unit 1902 is used to iteratively update the initial trajectory prediction model according to the loss value determined by the loss unit 1901 to obtain a trained trajectory prediction model.
[0168] The trained trajectory prediction model here is the trajectory prediction model used in the vehicle trajectory prediction method in the foregoing embodiments.
[0169] Regarding the control device of the vehicle cockpit in the above embodiments, the specific manners in which each module performs operations and the corresponding beneficial effects have been described in detail in the embodiments of the foregoing trajectory prediction model training method, and will not be elaborated here.
[0170] Figure 20 is a possible structural schematic diagram of an electronic device shown according to an exemplary embodiment. The electronic device may be the above-mentioned trajectory prediction model training device or vehicle trajectory prediction device, or may be a terminal or server including the trajectory prediction model training device and / or vehicle trajectory prediction device. As Figure 20As shown, the electronic device includes a processor 91 and a memory 92. Among them, the memory 92 is used to store instructions executable by the processor 91, and the processor 91 can implement the functions of each module in the trajectory prediction model training device and / or the vehicle trajectory prediction device in the above embodiments. Among them, at least one instruction is stored in the memory 92, and the at least one instruction is loaded and executed by the processor 91 to implement the trajectory prediction model training method and / or the vehicle trajectory prediction method provided by the above various method embodiments.
[0171] Among them, in a specific implementation, as an embodiment, the processor 91 (91-1 and 91-2) may include one or more CPUs, such as Figure 20 the CPU0 and CPU1 shown in. And as an embodiment, the electronic device may include multiple processors 91, such as Figure 20 the processor 91-1 and the processor 91-2 shown in. Each CPU in these processors 91 can be a single-core processor (Single-CPU) or a multi-core processor (Multi-CPU). The processor 91 here may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0172] The memory 92 can be a read-only memory 92 (read-only memory, ROM) or other types of static storage devices that can store static information and instructions, a random access memory (random access memory, RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (electrically erasable programmable read-only memory, EEPROM), a compact disc read-only memory (compactdisc read-only memory, CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk computer storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 92 can exist independently and be connected to the processor 91 through a communication bus 93. The memory 92 can also be integrated with the processor 91.
[0173] The communication bus 93 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The communication bus 93 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 20 it is only represented by a thick line in Figure 20 , but this does not mean that there is only one bus or one type of bus.
[0174] In addition, to facilitate information interaction between the electronic device and other devices (for example, when the electronic device is a terminal, it interacts with a server, or when the electronic device is a server, it interacts with a terminal), the electronic device includes a communication interface 94. The communication interface 94 uses any device such as a transceiver to communicate with other devices or a communication network, such as a control system, a Radio Access Network (RAN), a Wireless Local Area Networks (WLAN), etc. The communication interface 94 can include a receiving unit to implement the receiving function and a transmitting unit to implement the transmitting function. The communication interface 94, the processor 91, and the memory 92 are connected through the communication bus 93 to complete mutual communication.
[0175] The embodiment of the present application also provides a computer-readable storage medium storing computer instructions. When the computer instructions run on the electronic device, the electronic device is caused to execute the trajectory prediction model training method and / or the vehicle trajectory prediction method in the above method embodiments.
[0176] For example, the computer-readable storage medium can be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), magnetic tape, a floppy disk, an optical data storage device, etc.
[0177] The embodiment of the present application also provides a computer program product containing computer instructions. When the computer instructions run on the electronic device, the electronic device is caused to execute the trajectory prediction model training method and / or the vehicle trajectory prediction method in the above method embodiments.
[0178] Among them, the electronic device, computer-readable storage medium, or computer program product provided in the embodiments of the present application are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be elaborated here.
[0179] From the descriptions of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device (such as an electronic device) is divided into different functional modules to complete all or part of the functions described above. The specific working processes of the systems, devices (such as electronic devices), and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0180] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices (such as electronic devices), and methods can be implemented in other ways. For example, the device (such as an electronic device) embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0181] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0182] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0183] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs.
[0184] As described above, the foregoing is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A vehicle trajectory prediction method, comprising: Obtaining the hidden state information of the previous frame image and the current frame image; The previous frame image is the previous frame image of the current frame image in the vehicle trajectory video sequence; Fusing the hidden state information of the current frame image and the previous frame image based on the first connection layer feature fusion layer in the trajectory prediction model to obtain a first feature map; Processing the first feature map based on the first convolutional neural network layer in the trajectory prediction model to obtain reset information; the first convolutional neural network layer includes a first activation function; Processing the first feature map based on the second convolutional neural network layer in the trajectory prediction model to obtain updated information, the second convolutional neural network layer includes a second activation function; Processing the reset information, the updated information, the hidden state information of the current frame image and the previous frame image based on the prediction layer in the trajectory prediction model to obtain a predicted frame image.
2. The method according to claim 1, wherein The processing the reset information, the updated information, the hidden state information of the current frame image and the previous frame image based on the prediction layer in the trajectory prediction model to obtain a predicted frame image includes: Multiplying the reset information and the hidden state information of the previous frame image element by element based on the selection layer in the prediction layer to obtain a second feature map; Fusing the second feature map and the current frame image based on the second feature fusion layer in the prediction layer to obtain a third feature map; Processing the third feature map sequentially based on the third convolutional neural network layer and the activation function layer in the prediction layer to obtain candidate hidden state information, the activation function layer includes a third activation function; Processing the candidate hidden state information, the updated information and the hidden state information of the previous frame image based on the output layer in the prediction layer to obtain the predicted frame image.
3. The method according to claim 2, wherein The processing the candidate hidden state information, the updated information and the hidden state information of the previous frame image based on the output layer in the prediction layer to obtain the predicted frame image includes: Multiplying the candidate hidden state information and the updated information element by element based on the selective memory sublayer in the output layer to obtain a fourth feature map; Performing subtraction processing on each element in the updated information based on the selective forgetting sublayer in the output layer to obtain the weight of the forgetting information, and multiplying the weight of the forgetting information and the hidden state information of the previous frame image element by element to obtain a fifth feature map; Adding the fourth feature map and the fifth feature map element by element based on the synthesis sublayer in the output layer to obtain the hidden state information of the current frame image; Processing the hidden state information of the current frame image based on the fourth convolutional neural network layer in the output layer to obtain the predicted frame image, the fourth convolutional neural network layer includes a fourth activation function.
4. A trajectory prediction model training method, comprising: Obtaining multiple groups of sample images and first trajectory images corresponding to the sample images one by one; The sample image includes the hidden state information of the second trajectory image and the third trajectory image; wherein, the second trajectory image is the previous frame trajectory image of the third trajectory image in the vehicle trajectory video sequence, and the first trajectory image is the next frame trajectory image of the third trajectory image in the vehicle trajectory video sequence; Based on the first feature fusion layer in the initial trajectory prediction model, the third trajectory image and the hidden state information of the second trajectory image are fused to obtain a first training feature map; Based on the first convolutional neural network layer in the initial trajectory prediction model, the first training feature map is processed to obtain training reset information; the first convolutional neural network layer includes a first activation function; Based on the second convolutional neural network layer in the initial trajectory prediction model, the first training feature map is processed to obtain training update information, and the second convolutional neural network layer includes a second activation function; Based on the prediction layer in the initial trajectory prediction model, the training reset information, the training update information, the hidden state information of the second trajectory image, and the third trajectory image are processed to obtain an initial prediction frame image; Using the initial prediction frame image as the initial training output of the initial trajectory prediction model and the first trajectory image as supervision information, the initial trajectory prediction model is iteratively trained to obtain a trained trajectory prediction model.
5. The method according to claim 4, wherein, The processing of the training reset information, the training update information, the hidden state information of the second trajectory image, and the third trajectory image based on the prediction layer in the initial trajectory prediction model to obtain an initial prediction frame image includes: Based on the selection layer in the prediction layer, the training reset information is multiplied element-wise with the hidden state information of the second trajectory image to obtain a second training feature map; Based on the second feature fusion layer in the prediction layer, the second training feature map and the third trajectory image are fused to obtain a third training feature map; Based on the third convolutional neural network layer and the activation function layer in the prediction layer, the third training feature map is sequentially processed to obtain training candidate hidden state information, and the activation function layer includes a third activation function; Based on the output layer in the prediction layer, the training candidate hidden state information, the training update information, and the hidden state information of the second trajectory image are processed to obtain the initial prediction frame image.
6. The method according to claim 5, wherein The processing of the training candidate hidden state information, the training update information, and the hidden state information of the second trajectory image based on the output layer in the prediction layer to obtain the initial prediction frame image includes: Based on the selective memory sublayer in the output layer, the training candidate hidden state information and the training update information are multiplied element-wise to obtain a fourth training feature map; Based on the selective forgetting sublayer in the output layer, each element in the training update information is subtracted to obtain the weight of the training forgetting information, and the weight of the training forgetting information is multiplied element-wise with the hidden state information of the second trajectory image to obtain a fifth training feature map; Based on the synthesis sublayer in the output layer, the fourth training feature map and the fifth training feature map are added element by element to obtain the hidden state information of the third trajectory image; Based on the fourth convolutional neural network layer in the output layer, the hidden state information of the third trajectory image is processed to obtain the initial predicted frame image, and the fourth convolutional neural network layer includes a fourth activation function.
7. The method according to any one of claims 4-6, wherein Using the initial predicted frame image as the initial training output of the initial trajectory prediction model and the first trajectory image as the supervision information, iteratively training the initial trajectory prediction model to obtain the trained trajectory prediction model includes: Determining a loss value according to the initial predicted frame image and the first trajectory image; Iteratively updating the initial trajectory prediction model according to the loss value to obtain the trained trajectory prediction model.
8. A vehicle trajectory prediction device, comprising: An acquisition module, configured to acquire the hidden state information of the previous frame image and the current frame image; The previous frame image is the previous frame image of the current frame image in the vehicle trajectory video sequence; A processing module, configured to fuse the current frame image and the hidden state information of the previous frame image acquired by the acquisition module based on the first feature fusion layer in the trajectory prediction model to obtain a first feature map; The processing module is further configured to process the first feature map based on the first convolutional neural network layer in the trajectory prediction model to obtain reset information; the first convolutional neural network layer includes a first activation function; The processing module is further configured to process the first feature map based on the second convolutional neural network layer in the trajectory prediction model to obtain update information, and the second convolutional neural network layer includes a second activation function; The processing module is further configured to process the reset information, the update information, the current frame image, and the hidden state information of the previous frame image based on the prediction module in the trajectory prediction model to obtain a predicted frame image.
9. A trajectory prediction model training device, comprising: An acquisition module, configured to acquire multiple groups of sample images and the first trajectory images corresponding to the sample images one by one; The sample images include the hidden state information of the second trajectory image and the third trajectory image; wherein, the second trajectory image is the previous frame trajectory image of the third trajectory image in the vehicle trajectory video sequence, and the first trajectory image is the next frame trajectory image of the third trajectory image; A training module, configured to fuse the hidden state information of the third trajectory image and the second trajectory image based on the first feature fusion layer in the initial trajectory prediction model to obtain a first training feature map; The training module is further configured to process the first training feature map based on the first convolutional neural network layer in the initial trajectory prediction model to obtain training reset information; the first convolutional neural network layer includes a first activation function; The training module is further configured to process the first training feature map based on the second convolutional neural network layer in the initial trajectory prediction model to obtain training update information, and the second convolutional neural network layer includes a second activation function; The training module is further configured to process the training reset information, the training update information, the hidden state information of the second trajectory image, and the third trajectory image based on the prediction layer in the initial trajectory prediction model to obtain an initial predicted frame image; The training module is further configured to use the initial predicted frame image as the initial training output of the initial trajectory prediction model, and use the first trajectory image obtained by the acquisition module as supervision information to iteratively train the initial trajectory prediction model to obtain the trained trajectory prediction model.
10. A computer-readable storage medium storing a computer program for executing the vehicle trajectory prediction method according to any one of claims 1-3 above, or for executing the trajectory prediction model training method according to any one of claims 4-7 above.
11. An electronic device, comprising: a processor; a memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the vehicle trajectory prediction method according to any one of claims 1-3 above, or to execute the trajectory prediction model training method according to any one of claims 4-7 above.
Citation Information
Patent Citations
Method for training detection model, determining image updating information and updating high-precision map
CN113505834A
Audio generation method and device based on artificial intelligence, equipment and storage medium
CN113822017A