Image Prediction and Vehicle Behavior Planning Method, Device, System and Storage Medium

The method predicts future vehicle behavior using image processing and transformation techniques to enhance the ability to handle unexpected events and improve the explainability of decision-making in autonomous driving systems.

CN111414852BActive Publication Date: 2025-07-15UISEE TECH (ZHEJIANG) LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010196263.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-03-19
Publication Date
2025-07-15
Estimated Expiration
2040-03-19

AI Technical Summary

Technical Problem

Existing automatic driving technologies rely solely on current sensory information for vehicle behavior planning, which lacks the ability to handle unexpected events and lacks explainability in decision-making.

Method used

A method and system for image prediction and vehicle behavior planning that involves extracting features from current and predicted images using encoding and decoding networks, followed by convolutional transformations to generate expected acceleration and steering angles, enabling proactive response to potential events and enhancing decision-making transparency.

Benefits of technology

Enhances the ability to anticipate and respond to unexpected events by providing explainable vehicle behavior planning, improving safety and reliability in autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111414852B_ABST
    Figure CN111414852B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides an image prediction method, apparatus and system, a vehicle behavior planning method, apparatus and system, and a storage medium. The image prediction method includes: obtaining a current image I10 collected by a target vehicle at a current moment T1; extracting a feature F10 of the current image I10 through a first encoder EN0; for the moment T1 + i*Δt, in a prediction network N i to predict a feature F1 i‑1 based on one or more of the features F10 to F1 i , and reconstructing the feature F1 i to obtain a predicted image I1 i ' at the moment T1 + i*Δt, where i = 1, 2... m, m is an integer greater than or equal to 2, and Δt is a preset time period. According to the embodiment of the present invention, subsequent images can be predicted based on the current image collected by the vehicle, so that the change of the environment during the subsequent driving process of the vehicle can be predicted. These predicted images can be applied to vehicle behavior planning, thereby helping to improve the interpretability of behavior planning and helping to cope with emergencies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving technology, and more specifically to an image prediction method, device and system, a vehicle behavior planning method, device and system, and a storage medium. Background Art

[0002] In the field of autonomous driving, existing technologies mainly rely on perception information in the current state to complete the planning of vehicle behavior. This has two problems: one is that it cannot cope with emergencies, and the other is that the behavior made using this solution is not explainable. Summary of the invention

[0003] The present invention is proposed in view of the above problems. The present invention provides an image prediction method, device and system, a vehicle behavior planning method, device and system and a storage medium.

[0004] In one aspect, the present invention provides an image prediction method. The image prediction method comprises: obtaining a current image I10 captured by a target vehicle at a current time T1; extracting a feature F10 of the current image I10 through a first encoder EN0; and for the time T1+i*Δt, performing a prediction on the prediction network N i Based on feature F10 to feature F1 i-1 One or more of them to predict feature F1 i , and for feature F1 i Reconstruct to obtain the predicted image I1 at time T1+i*Δt i ', i = 1, 2...m, m is an integer greater than or equal to 2, and Δt is a preset period of time.

[0005] Another aspect of the present invention provides a vehicle behavior planning method, comprising: obtaining a current image I10 and predicted images I11', I12', ... I10 involved in the above-mentioned image prediction method; m '; Based on i=1, extract image I1 through the second encoder EN0' i-1 Features F1 i-1 '; Based on i=2,3...m, extract the predicted image I1 through the second encoder EN0' i-1 'Feature F1 i-1 '; Based on i=1,2...m, feature F1 i-1 'Enter the transformation convolution network CT with the first initial parameters for convolution to obtain the transformation matrix M1 i-1 ; Using transformation matrix M1 i-1 For feature F1 i-1 'Perform matrix transformation to obtain the transformation feature F1 i "; Through the second decoder DE0 'to the feature F1 i”Perform reconstruction to obtain the reconstructed image I1 i ”; Through the predicted image I1 i ' and the reconstructed image I1 i ”Calculate the first image loss function and train the transformation convolutional network CT based on the first image loss function to obtain the trained transformation convolutional network CT i-1 ; Based on the transformation convolutional network CT i-1 Output transformation matrix M1 i-1 Determine the expected acceleration and expected steering angle of the target vehicle at the moment of T1+(i-1)*Δt.

[0006] On the other hand, the present invention provides an image prediction device, including: an acquisition module for acquiring the current image I10 collected by the target vehicle at the current moment T1; an extraction module for extracting the feature F10 of the current image I10 through the first encoder EN0; a prediction module for predicting the feature F1 i in the prediction network N i-1 based on one or more of the features F10 to F1 i , and reconstruct the feature F1 i to obtain the predicted image I1 i ' at the moment of T1+i*Δt, where i = 1, 2... m, m is an integer greater than or equal to 2, and Δt is a preset time period.

[0007] On the other hand, the present invention provides a vehicle behavior planning device, including: an acquisition module for acquiring the current image I10 and the predicted images I11', I12'... I1 m ' involved in the above image prediction method; a first extraction module for extracting the feature F1 i-1 of the image I1 i-1 ' through the second encoder EN0' based on i = 1; a second extraction module for extracting the feature F1 i-1 of the predicted image I1 i-1 ' through the second encoder EN0' based on i = 2, 3... m; an input module for inputting the feature F1 i-1 ' into the transformation convolutional network CT with the first initial parameters for convolution based on i = 1, 2... m to obtain the transformation matrix M1 i-1 ; a transformation module for performing matrix transformation on the feature F1 i-1 using the transformation matrix M1 i-1 to obtain the transformed feature F1 i ”; a reconstruction module for reconstructing the feature F1 i ”through the second decoder DE0' based on i = 1, 2... m to obtain the reconstructed image I1i ”; a training module, for, based on i = 1, 2... m, through the prediction image I1 i ' and the reconstructed image I1 i ” calculate the first image loss function, and based on the first image loss function, train the transformation convolutional network CT to obtain the trained transformation convolutional network CT i-1 ; a determination module, for, based on i = 1, 2... m, based on the transformation convolutional network CT i-1 output the transformation matrix M1 i-1 to determine the expected acceleration and expected steering angle of the target vehicle at the moment of T1+(i - 1)*Δt.

[0008] On the other hand, the present invention provides an image prediction system, including a processor and a memory, wherein computer program instructions are stored in the memory, and when the computer program instructions are run by the processor, they are used to execute the above image prediction method.

[0009] On the other hand, the present invention provides a vehicle behavior planning system, including a processor and a memory, wherein computer program instructions are stored in the memory, and when the computer program instructions are run by the processor, they are used to execute the above vehicle behavior planning method.

[0010] On the other hand, the present invention provides a storage medium, on which program instructions are stored, and when the program instructions are run, they are used to execute the above image prediction method.

[0011] On the other hand, the present invention provides a storage medium, on which program instructions are stored, and when the program instructions are run, they are used to execute the above vehicle behavior planning method.

[0012] The image prediction method, device and system and vehicle behavior planning method, device and system and storage medium of the embodiments of the present invention can predict subsequent images based on the current images collected by the vehicle, so that the changes in the environment during the subsequent driving process of the vehicle can be predicted. These predicted images can be applied to vehicle behavior planning, which helps to improve the interpretability of behavior planning and helps to cope with emergencies. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] By describing the embodiments of the present invention in more detail in combination with the drawings, the above and other objects, features and advantages of the present invention will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification, and are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the drawings, the same reference numerals generally represent the same components or steps.

[0014] Figure 1Schematic flowchart showing an image prediction method according to an embodiment of the present invention;

[0015] Figure 2 Schematic diagram showing an image prediction model involved in the image prediction method according to an embodiment of the present invention;

[0016] Figure 3 Schematic diagram showing a prediction network according to an embodiment of the present invention;

[0017] Figure 4 Schematic flowchart showing a vehicle behavior planning method according to an embodiment of the present invention;

[0018] Figure 5 Schematic diagram showing a behavior planning model involved in the vehicle behavior planning method according to an embodiment of the present invention;

[0019] Figure 6 Schematic block diagram showing an image prediction device according to an embodiment of the present invention;

[0020] Figure 7 Schematic block diagram showing a vehicle behavior planning device according to an embodiment of the present invention;

[0021] Figure 8 Schematic block diagram showing an image prediction system according to an embodiment of the present invention; and

[0022] Figure 9 Schematic block diagram showing a vehicle behavior planning system according to an embodiment of the present invention. Detailed implementation manners

[0023] In order to make the objectives, technical solutions and advantages of the present invention more obvious, exemplary embodiments according to the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments of the present invention. It should be understood that the present invention is not limited by the exemplary embodiments described herein.

[0024] To solve the above problems, the present invention proposes an image prediction method and a vehicle behavior planning method. According to an embodiment of the present invention, future environmental information (i.e., predicted image) can be predicted based on the environmental information currently perceived by the vehicle (i.e., current image), and the predicted information can be used to generate control signals for the current state of the vehicle, such as desired acceleration and desired steering angle. This behavior planning method is a prediction-based planning method.

[0025] Predictive behavior planning can obtain information about the occurrence of an event through prediction before the event occurs and use the prediction information to guide behavior planning. Therefore, it can react in advance before an emergency occurs. In addition, through the prediction of the future, it is possible to know what results the current behavior will produce or what expectations the behavior is based on. Therefore, this solution can improve the interpretability of the planning system, which is an important criterion for the safe implementation of an autonomous driving system. It should be noted that the image prediction method provided in the embodiments of the present invention can be applied to various scenarios that require predicting the future state of a vehicle, including but not limited to the above-mentioned behavior planning. For example, the image prediction method can also be applied to trajectory planning, vehicle tracking, etc.

[0026] The driving state of a vehicle can be reflected by the images collected by the on-vehicle camera of the vehicle. The images can contain information about the vehicle's surrounding environment, such as information about other vehicles, pedestrians, roads, buildings, etc.

[0027] During the driving process of the vehicle, it is possible to predict the images within a certain period after based on the images collected in real time. If the final period to be predicted is relatively long, directly predicting may have a large error. In this case, it is possible to divide this period into several small periods and predict them segment by segment through a progressive prediction method until the image at the final moment is predicted. For example, assuming that the image 2 seconds later needs to be predicted finally, then 2 seconds can be divided into 10 parts, and the future image 0.2 seconds later is predicted each time. The next image prediction can be realized based on the information predicted previously. In this way, the accuracy of image prediction can be effectively improved. Based on such a prediction logic, the image prediction method 100 described in this article is proposed.

[0028] Figure 1 The schematic flowchart of the image prediction method 100 according to an embodiment of the present invention is shown. As Figure 1 shown, the image prediction method 100 includes steps S110 - S130.

[0029] In step S110, obtain the current image I10 collected by the target vehicle at the current moment T1.

[0030] The image prediction method 100 can run in the control device of any vehicle (referred to as the target vehicle). The vehicle can be equipped with an on-vehicle camera, and the on-vehicle camera can collect images around the vehicle in real time.

[0031] Assume that the current moment is represented by T1, and the on-vehicle camera collects an image at moment T1 to obtain the current image I10.

[0032] In step S120, extract the feature F10 of the current image I10 through the first encoder EN0. Optionally, the feature F10 can also be reconstructed through the first decoder DE0 to obtain the reconstructed image I10'.

[0033] The algorithm model involved in the image prediction method 100 (referred to as the image prediction model in this article) can be trained during the training phase and then used to perform actual predictions using the trained image prediction model during the application phase. The image prediction model may include a first encoder EN0 and a first decoder DE0, as well as prediction networks N1, N2... N m During the training phase, the first encoder EN0 and the first decoder DE0 can be trained as a whole. During the application phase, the trained first encoder can be used to extract the features F10 of the current image I10.

[0034] Both the first encoder EN0 and the first decoder DE0 can be implemented using any suitable network structure, such as a convolutional network structure. For example, the first encoder EN0 and the first decoder DE0 can each include one or more convolutional layers. Additionally, exemplarily, the first encoder EN0 may further include a downsampling layer, and the first decoder DE0 may further include an upsampling layer. In one example, the first encoder EN0 and the first decoder DE0 can be implemented using an auto-encoder (AE) or a variational auto-encoder (VAE), etc.

[0035] The first encoder EN0 and the first decoder DE0 can form a reconstruction network. The first encoder EN0 is used to extract features from the input image, and the first decoder DE0 is used to reconstruct the features extracted by the first encoder to restore them into an image. The features described in this article can be feature maps output by the network structure.

[0036] Figure 2 A schematic diagram showing the image prediction model involved in the image prediction method 100 according to an embodiment of the present invention. Refer to Figure 2 , which shows the first encoder EN0 and the first decoder DE0. The current image I10 can be input into the first encoder EN0 to extract the features F10 by the first encoder EN0. Optionally, the features F10 output by the first encoder EN0 can be input into the first decoder DE0. The first decoder DE0 can reconstruct the features F10 to obtain a reconstructed image I10'. The reconstructed image I10' output by the first decoder DE0 is the same size as the original image I10, equivalent to restoring the original image based on the features F10.

[0037] In step S130, for the time T1 + i*Δt, in the prediction network N i among them, based on one or more of the features F10 to F1 i-1 to predict the features F1 i , and for the features F1i Reconstructed to obtain the predicted image I1 at the moment of T1 + i*Δt i ', where i = 1, 2... m, m is an integer greater than or equal to 2, and Δt is a preset time period.

[0038] One or more of the features F10 to F1 i-1 Can be input into the prediction network N i To obtain the predicted image I1 i Output by the prediction network N i '. See Figure 2 Which shows the prediction network N i And shows the predicted image I1 i Output by each prediction network N i '.

[0039] For each moment after the current moment T1 that passes through the preset time period Δt, the image at that moment can be predicted. Δt can be of any suitable size, and the present invention does not limit this. For example, Δt can be 0.2 seconds.

[0040] For the first Δt moment after the current moment T1 (e.g., 0.2 seconds after T1), the feature F11 at that moment can be predicted by the prediction network N1 based on the feature F10, and the image I11' at that moment can be predicted based on the feature F11;

[0041] For the second Δt moment after the current moment T1 (e.g., 0.4 seconds after T1), the feature F12 at that moment can be predicted by the prediction network N2 based on the feature F10 and / or F11, and the image I12' at that moment can be predicted based on the feature F12;

[0042] For the third Δt moment after the current moment T1 (e.g., 0.6 seconds after T1), the feature F13 at that moment can be predicted by the prediction network N3 based on one or more of the features F10, F11, and F12, and the image I13' at that moment can be predicted based on the feature F13;

[0043] ……

[0044] For the mth Δt moment after the current moment T1 (e.g., 2 seconds after T1), the prediction network N m Based on one or more of the features F10, F11, F12... F1 m-1 Predicts the feature F1 at that moment m And based on the feature F1 m Predicts the image I1 at that moment m '.

[0045] When predicting the features at subsequent moments based on the features at the previous moment, an appropriate number of features can be selected for prediction as needed. Although Figure 2 shows each prediction network N i receives the feature F10 and the feature F1 i-1 (N1 only receives the feature F10) as input, but the features received by each prediction network N i can have other combinations.

[0046] In one example, when predicting each moment T1 + i*Δt, the feature F1 at that moment can be predicted based only on a single feature i . For example, regardless of the value of i, the feature F1 i-1 is predicted based only on the feature F1 i .

[0047] In another example, when predicting each moment T1 + i*Δt, the feature F1 at that moment can be predicted based on multiple features i . For example, based on i = 1, the feature F1 is predicted based on the feature F10 i (i.e., F11); based on i ≥ 2, the feature F1 is predicted based on the earliest feature F10 and the feature F1 i-1 closest in distance i . Optionally, in addition to the feature F10 and the feature F1 i-1 , some features at intermediate moments can be added to predict F1 i . For example, based on i = 1, the feature F1 is predicted based on the feature F10 i (i.e., F11); based on i = 2, the feature F1 is predicted based on the feature F10 and the feature F1 i-1 (i.e., F11) to predict the feature F1 i (i.e., F12); based on i ≥ 3, the feature F1 is predicted based on the feature F10, the feature F1 i-2 , the feature F1 i-1 to predict the feature F1 i . Optionally, regardless of the value of i, the feature F1 can also be predicted based on all the features from the feature F10 to the feature F1 i-1 i .

[0048] Exemplarily, each prediction network N i can include a decoder DE i . In the prediction network N i , after predicting and obtaining the feature F1 i , the feature F1 i can be input into the subsequent decoder DE i for reconstruction to obtain the predicted image I1 i '. Optionally, the decoder DE i ​It may share parameters with the above-mentioned first decoder DE0 (i.e., their parameters are the same). Of course, the parameters of the two may also be set independently.

[0049] According to the image prediction method of the embodiment of the present invention, subsequent images can be predicted based on the current image collected by the vehicle, so as to predict the changes in the environment during the subsequent driving process of the vehicle. These predicted images can be applied to vehicle behavior planning, thereby helping to improve the interpretability of behavior planning and helping to cope with emergencies.

[0050] According to the embodiment of the present invention, based on one or more of Feature F10 to Feature F1 i-1 to predict Feature F1 i may include: for each Feature F1 participating in the prediction among Feature F10 to Feature F1 i-1 , based on Feature F1 j calculate an attention mask S1 j ; perform a matrix inner product calculation on Feature F1 ij and the attention mask S1 j to obtain an attention feature FS1 ij ; input the attention feature FS1 ij into a fully connected layer or a convolutional layer for feature weighted summation to obtain a weighted feature FA1 ij ; fuse all the weighted features obtained in the prediction network N ij to obtain Feature F1 i ; where j ∈ {0, 1... i - 1}. i ; where j ∈ {0, 1... i - 1}.

[0051] The attention mask can reflect the position where the vehicle or the driver (agent) is looking at in the current state, that is, the position and state where it expects to be in the future. Therefore, the future position and state of the vehicle can be predicted through the attention mask.

[0052] Exemplarily, the attention mask can be obtained through a convolutional neural network (CNN). For example, calculating the attention mask S1 j based on Feature F1 ij may include: inputting Feature F1 j into the mask convolutional network CS i in the prediction network N ij for convolution to obtain the attention mask S1 ij , where the attention mask S1 ij is consistent with the height and width of Feature F1 j and the number of channels is 1, and each element in the attention mask S1 ij represents the response value of the position where the vehicle is going to drive to. Exemplarily, the attention mask S1 ijThe value of each element in [ ] can be any value within the range of [0, 1], and this value is a probability value. The larger the value, the greater the probability that it represents the position the vehicle is about to drive to.

[0053] For example, the original feature F1 j is a feature map containing 1024 channels. After convolution by the mask convolutional network CS ij , the 1024 channels can be compressed into 1 channel, while the height and width of the feature map remain unchanged, thereby obtaining the attention mask S1 ij .

[0054] Figure 3 Shows a schematic diagram of a prediction network according to an embodiment of the present invention. Refer to Figure 3 , which shows prediction networks N1 and N2. The prediction network N1 can include the mask convolutional network CS 10 . Inputting the feature F10 into the mask convolutional network CS 10 can obtain the attention mask S1 output by this network 10 . The prediction network N2 can include the mask convolutional networks CS 20 and CS 21 . Inputting the features F10 and F11 into the mask convolutional networks CS 20 and CS 21 respectively can obtain the attention masks S1 20 and S1 21 output by these networks respectively. In any prediction network N i , for each feature participating in the prediction, it is respectively input into its corresponding mask convolutional network to calculate the corresponding attention mask. The specific implementation method can be understood with reference to Figure 3 and related descriptions, and will not be listed one by one here.

[0055] After obtaining the attention mask S1 ij , the feature F1 j can be subjected to matrix inner product calculation with the attention mask S1 ij to thereby segment out the feature part under the viewed perspective and obtain the attention feature FS1 ij . Refer to Figure 3 , in the prediction network N1, based on the feature F10 and the attention mask S1 10 , the attention feature FS1 10 is calculated. In the prediction network N2, based on the feature F10 and the attention mask S1 20 , the attention feature FS1 20 is calculated, and based on the feature F11 and the attention mask S1 21 , the attention feature FS1 21 is calculated.

[0056] Prediction network N i may include a fully connected layer FC ij or a convolutional layer C ij Attention feature FS1 ij can be input into the fully connected layer FC ij or the convolutional layer C ij for weighted sum of features. Figure 3 The example shown is a fully connected layer. Those skilled in the art can understand the implementation method of replacing the fully connected layer with a convolutional layer (the convolutional layer implements the same function as the replaced fully connected layer), which will not be elaborated in this article. In addition, although Figure 3 not shown, those skilled in the art can understand that before inputting into the fully connected layer FC ij or the convolutional layer C ij the attention feature FS1 ij can be converted in form, and can be stretched into a one-dimensional vector, for example, the expression form is (C*H*W,1,1). In addition, after the output of the fully connected layer FC ij or the convolutional layer C ij the obtained weighted feature FA1 ij can be converted in form and reshaped into the same size as the original feature F1 j .

[0057] See Figure 3 , in the prediction network N1, the attention feature FS1 10 is input into the fully connected layer FC 10 to obtain the weighted feature FA1 10 . In the prediction network N2, the attention feature FS1 20 is input into the fully connected layer FC 20 to obtain the weighted feature FA1 20 , and the attention feature FS1 21 is input into the fully connected layer FC 21 to obtain the weighted feature FA1 21 .

[0058] Subsequently, for each prediction network N i , all the weighted features in this prediction network N i can be fused. When the number of weighted features is multiple, the fusion can be feature concatenation or adding the corresponding elements of the features. The fused feature is the required predicted feature F1 i . In the prediction network N1, only one weighted feature FA1 10 is obtained. Therefore, the fusion result of this feature is itself, that is, the weighted feature FA1 10 which is also the required predicted feature F11.

[0059] According to the above embodiments, the feature part that was attended to at a previous moment can be extracted by means of an attention mask, and then the position that the vehicle will drive towards at the next moment can be predicted.

[0060] According to an embodiment of the present invention, predicting feature F1 based on one or more of features F10 to F1 i-1 may include: based on i = 1, predicting feature F1 based on feature F10 i ; based on i ≥ 2, predicting feature F1 based on feature F10 and feature F1 i ; i-1 predicting feature F1. i .

[0061] For the next moment T1 + Δt of the current moment T1, the only previously obtained feature is F10. At this time, feature F11 can be predicted only based on this feature F10. For the subsequent remaining moments T1 + 2Δt, T1 + 3Δt, etc., the number of previously obtained (including extracted and predicted) features is increasing. At this time, it is possible to consider predicting feature F1 i-1 each time based on the earliest extracted feature F10 and the most recently predicted feature F1 i . The earliest extracted feature F10 is extracted from the initially acquired image I10, rather than indirectly predicted. Therefore, the earliest extracted feature F10 has relatively high reliability. And the most recently predicted feature F1 i-1 is the feature F1 closest to the current prediction i . Therefore, combining the earliest extracted feature F10 and the most recently predicted feature F1 i-1 to predict feature F1 i can better balance processing efficiency and prediction effect.

[0062] According to an embodiment of the present invention, predicting feature F1 based on one or more of features F10 to F1 i-1 includes: at least predicting feature F1 based on feature F10 i ; wherein, for different i, the parameters of the mask convolutional network CS i are independent of each other. i0

[0063] Referring to Figure 3 , the parameters of CS 10 and CS 20 can be independent of each other. For features F11, F12, F13, etc., the time gap from the feature F10 at moment T1 is gradually increasing. Looking at the states of subsequent moments T1 + Δt, T1 + 2Δt, T1 + 3Δt, etc. from moment T1, the attention situation will change. Therefore, different mask convolutional networks can be used to generate different attention masks, so that the predicted feature F1 iwill be more accurate.

[0064] According to an embodiment of the present invention, based on one or more of feature F10 to feature F1 i-1 to predict feature F1 i includes: at least based on feature F1 i-1 to predict feature F1 i ; wherein, for different i, the parameters of the masked convolutional network CS i(i-1) are shared.

[0065] For example, the parameters of CS 21 and CS 32 ( Figure 3 (not shown) can be independent of each other. The time gap between feature F11 and feature F10, the time gap between feature F12 and feature F11, the time gap between feature F13 and feature F12, etc. are the same. Therefore, each time the situation at the next moment T1 + i * Δt is viewed from the moment T1+(i - 1)*Δt, it is similar. Therefore, a masked convolutional network CS i(i-1) with the same parameters can be selected to calculate the attention mask of feature F i-1 . This scheme can reduce the data processing volume of the image prediction model during training and application, and can improve the processing efficiency. Of course, for different i, the parameters of the masked convolutional network CS i-1 input by the feature F1 i(i-1) can also be independent of each other.

[0066] According to an embodiment of the present invention, reconstructing feature F1 i to obtain the predicted image I1 i ' at the moment T1 + i * Δt may include: inputting feature F1 i into the decoder DE i in the prediction network N i , to obtain the predicted image I1 i '.

[0067] The decoder DE i can be implemented using any suitable network structure, for example, implemented using a convolutional network structure. Exemplarily, the decoder DE i may include an upsampling layer. Referring to Figure 3 , the decoder DE1 of the prediction network N1 and the decoder DE2 of the prediction network N2 are shown.

[0068] According to an embodiment of the present invention, method 100 may further include: reconstructing feature F10 through the first decoder DE0 to obtain the reconstructed image I10', wherein the decoder DE i shares parameters with the first decoder DE0. The decoder DE iSharing parameters with the first decoder DE0 can reduce the data volume of the image prediction model and speed up data processing. Of course, the parameters of the decoder DE i and the first decoder DE0 can also be independent of each other, which helps to improve the prediction accuracy of the image prediction model.

[0069] According to an embodiment of the present invention, method 100 may further include: obtaining (m + 1) sample images I20, I21,..., I2 collected by the first sample vehicle at times T2, T2 + Δt,..., T2 + mΔt m ; extracting the feature F20 of the sample image I20 through the first encoder EN0, and reconstructing the feature F20 through the first decoder DE0 to obtain the reconstructed image I20'; training the first encoder EN0 and the first decoder DE0 based on the sample image I20 and the reconstructed image I20'; for the time T2 + i*Δt, in the prediction network N i predict the feature F2 i-1 based on one or more of the features F20 to F2 i in, and reconstruct the feature F2 i to obtain the predicted image I2 i ' at the time T2 + i*Δt; training the prediction network N i based on the sample image I2 i and the predicted image I2 i '.

[0070] As described above, before applying the image prediction model, the model can be trained first. During training, the sample image I20 can be input into the reconstruction network composed of the first encoder EN0 and the first decoder DE0, and finally the reconstructed image I20' output by the first decoder DE0 can be obtained. The sample image I20 can be used as the annotation data (groundtruth), and the loss function (which can be called the first reconstruction loss function) can be calculated based on the sample image I20 and the reconstructed image I20', and the first encoder EN0 and the first decoder DE0 can be trained based on this loss function. Optionally, the first reconstruction loss function can be the mean square loss function (L2 loss function). Those skilled in the art can understand the training method based on the loss function, which will not be elaborated herein.

[0071] In addition, one or more of the features F20 to F2 i-1 can be input into the prediction network N i . The combination of features F20 to F2 i input into the prediction network N i-1 is the same as the combination of features F10 to F1 i input into the prediction network N i-1is consistent with the feature combination. For example, during the application stage, input the prediction network N i has features F10 and F1 i-1 In the case of, during the training stage, input the prediction network N i has features F20 and F2 i-1 .

[0072] In the prediction network N i any feature F2 participating in the prediction j undergoes the same processing as the above-mentioned feature F1 j which can be understood by referring to the above description and will not be elaborated here. Finally, a prediction image I2 i ' can be obtained at the output end of each prediction network N i '. Subsequently, the sample image I2 i can be used as the groundtruth, and based on the sample image I2 i and the prediction image I2 i ' calculate the loss function (which can be called the first prediction loss function), and train the prediction network N i based on this loss function. Optionally, the first prediction loss function can be the L2 loss function.

[0073] According to the above embodiment, the sample image I20 can be input into the image prediction model, that is, the reconstructed image I20' output by the first decoder DE0 and the prediction images I2 i output by each prediction network N i ' can be obtained. Subsequently, the images I20', I21'... I2 m ' and their corresponding sample images I20, I21... I2 m are used to calculate the loss function, and further train the parameters of the first encoder EN0, the first decoder DE0, and each prediction network N i . The above training method is simple to implement and has a small computational amount.

[0074] According to an embodiment of the present invention, the method 100 may further include: obtaining (m + 1) sample images I30, I31... I3 collected by the second sample vehicle at times T3, T3 + Δt... T3 + mΔt respectively m; Extract the feature F30 of the sample image I30 through the first encoder EN0, add a random Gaussian variable to the feature F30 to obtain a new feature F30', and reconstruct the new feature F30' through the first decoder DE0 to obtain a reconstructed image I30'; Use the first encoder EN0 and the first decoder DE0 as a generator and perform adversarial training with the first discriminator. Among them, in the adversarial training, use the sample image I30 as the positive sample and the reconstructed image I30' as the negative sample, and input them into the first discriminator for discrimination respectively; For the time at T3 + i*Δt, in the prediction network N i Based on one or more of the features F30 to F3 i-1 predict the feature F3 i , add a random Gaussian variable to the feature F3 i to obtain a new feature F3 i ', and reconstruct the new feature F3 i ' to obtain the predicted image I3 i ' at the time of T3 + i*Δt; Use the prediction network N i as a generator and perform adversarial training with the first discriminator. Among them, in the adversarial training, use the sample image I3 i as the positive sample and the predicted image I3 i ' as the negative sample, and input them into the first discriminator for discrimination respectively.

[0075] The second sample vehicle can be the same as or different from the first sample vehicle. The sample images I30, I31... I3 m can be the same as or different from the sample images I20, I21... I2 m .

[0076] Optionally, the image prediction model can be trained in an adversarial training manner. For the reconstruction network (including the first encoder EN0 and the first decoder DE0) and the prediction network in the image prediction model, a discriminator can be added to enhance the quality of image generation.

[0077] For example, after the first encoder EN0 outputs the feature F30 of the sample image I30, a random Gaussian variable z of the same size as F30 can be concatenated on F30. Subsequently, input the new feature F30' into the first decoder DE0 to obtain the reconstructed image I30'. Use the sample image I30 as the positive sample and the reconstructed image I30' as the negative sample, and input them into the first discriminator for discrimination respectively. Use the first encoder EN0 and the first decoder DE0 as a generator and perform adversarial training with the first discriminator. Optionally, during training, the parameters of the first discriminator can be updated first, and then the parameters of the generator can be updated using the updated first discriminator, and so on in a cycle.

[0078] In addition, one or more of Feature F30 to Feature F3 i-1 can be input into the prediction network N i . The combination of features of Feature F30 to Feature F3 i of the input prediction network N i-1 is consistent with the combination of features of Feature F10 to Feature F1 i of the input prediction network N i-1 . For example, during the application stage, the features of the input prediction network N i are F10 and F1 i-1 . In this case, during the training stage, the features of the input prediction network N i are F30 and F3 i-1 .

[0079] In the prediction network N i , any feature F3 j participating in the prediction undergoes the same processing as the above-mentioned Feature F1 j . It can be understood by referring to the above description and will not be elaborated here. Finally, a predicted image I3 i ' can be obtained at the output end of each prediction network N i . Subsequently, the sample image I3 i can be used as a positive sample, and the predicted image I3 i ' can be used as a negative sample and input into the first discriminator for discrimination respectively. The prediction network N i is used as a generator and undergoes adversarial training together with the above-mentioned first discriminator.

[0080] During adversarial training, the adversarial loss function of the adversarial network composed of the generator and the discriminator can be calculated. Those skilled in the art can understand the calculation method of the adversarial loss function, which will not be elaborated in this article. The above-mentioned adversarial network can be trained based on the adversarial loss function. In addition, the loss function (which can be called the second reconstruction loss function) can be calculated based on the sample image I30 and the reconstructed image I30', and the loss function (which can be called the second prediction loss function) can be calculated based on the sample image I3 i and the predicted image I3 i '. Optionally, the above-mentioned adversarial network can be trained based on the adversarial loss function, the second reconstruction loss function, and the second prediction loss function. Optionally, the second reconstruction loss function and the second prediction loss function can be L2 loss functions. Exemplarily, the adversarial loss function for training the first discriminator can be the Markov discriminator loss function (Patch GAN loss).

[0081] When predicting the subsequent predicted image I1 based on the current image I10 i'After (i = 1, 2... m), the behavior planning of the target vehicle can be performed based on the predicted image, such as calculating the expected acceleration and expected steering angle of the target vehicle at time T1 and subsequent moments. The basic idea of this behavior planning is based on the image at time T1+(i - 1)*Δt (the current image I10 or the predicted image I1 i-1 ') and the image at time T1+i*Δt (the predicted image I1 i ') to calculate the transformation matrix M1 i-1 , which enables the features obtained after the transformation of the features of the image at time T1+(i - 1)*Δt (the current image I10 or the predicted image I1 i-1 ') to approach as closely as possible the features of the image at time T1+i*Δt (the predicted image I1 i '). This transformation matrix represents the transformation of the current state, and the expected acceleration and expected steering angle at time T1+(i - 1)*Δt can be obtained from this transformation matrix. The vehicle behavior planning method based on this idea is described below.

[0082] According to another aspect of the present invention, a vehicle behavior planning method is provided. Figure 4 A schematic flowchart showing a vehicle behavior planning method 400 according to an embodiment of the present invention is shown. As Figure 4 shown, the vehicle behavior planning method 400 further includes steps S410 - S480.

[0083] In step S410, the current image I10 and the predicted images I11', I12'... I1 involved in the above image prediction method 100 are obtained m '.

[0084] Before behavior planning, the above image prediction method 100 can be run first to obtain the above current image I10 and subsequent predicted images I11', I12'... I1 m '.

[0085] In step S420, based on i = 1, the feature F1 of the image I1 is extracted through the second encoder EN0' i-1 '. i-1 '

[0086] In the case of i = 1, the image I1 i-1 is the current image I10. Since there is an actually acquired current image I10 at the current moment T1, the feature F10' of this current image I10 can be extracted, which may be the same as or different from the above feature F10.

[0087] Similarly to the above image prediction model, the algorithm model involved in the vehicle behavior planning method 400 (referred to as the behavior planning model in this article) can be trained during the training phase and then used to perform actual behavior planning using the trained behavior planning model during the application phase. The behavior planning model may include a second encoder EN0' and a second decoder DE0'.

[0088] Both the second encoder EN0' and the second decoder DE0' can be implemented using any suitable network structure, such as a convolutional network structure. For example, the second encoder EN0' and the second decoder DE0' may each include one or more convolutional layers. Additionally, exemplarily, the second encoder EN0' may further include a downsampling layer, and the second decoder DE0' may further include an upsampling layer. In one example, the second encoder EN0' and the second decoder DE0' can be implemented using an auto-encoder (AE) or a variational auto-encoder (VAE), etc.

[0089] The second encoder EN0' and the second decoder DE0' can form a reconstruction network. The second encoder EN0' is used to extract features from the input image, and the second decoder DE0' is used to reconstruct the features extracted by the second encoder EN0' to restore them into an image. Optionally, the parameters of the second encoder EN0' and the above first encoder EN0 can be shared or can be independent of each other. Optionally, the parameters of the second decoder DE0' and the above first decoder DE0 can be shared or can be independent of each other.

[0090] Figure 5 Schematic diagram showing the behavior planning model involved in the vehicle behavior planning method 400 according to an embodiment of the present invention. Refer to Figure 5 , showing the second encoder EN0' and the second decoder DE0'. The current image I10 (in the case of i = 1) or the predicted image I1 i-1 '(in the case of i = 2, 3... m) can be input into the second encoder EN0' so that the second encoder EN0' extracts the feature F1 i-1 '. Optionally, the feature F1 i-1 ' output by the second encoder EN0' can be input into the second decoder DE0'. The second decoder DE0' can reconstruct the feature F1 i-1 ' to obtain the reconstructed image I1 i-1 ”. The reconstructed image I1 i-1 ” output by the second decoder DE0' is the same size as the original image I10 or I1 i-1 ', equivalent to restoring the original image based on the feature F1 i-1 '.

[0091] In step S430, based on i = 2, 3... m, the prediction image I1 i-1 's feature F1 i-1 ' is extracted by the second encoder EN0'.

[0092] Based on i = 2, 3... m, there is a prediction image I1 i-1 ', and the feature F1 of this prediction image is extracted i-1 '.

[0093] In step S440, based on i = 1, 2... m, the feature F1 i-1 ' is input into the transformation convolutional network CT with the first initial parameters for convolution to obtain the transformation matrix M1 i-1 .

[0094] In steps S420 and S430, the feature extraction is processed in different cases. And in the subsequent steps after feature extraction (i.e., steps S440 - S480), regardless of the value of i, the same processing method is uniformly adopted.

[0095] See Figure 5 , the behavior planning model may also include the transformation convolutional network CT. The transformation convolutional network CT can be implemented using any suitable convolutional network structure. The transformation matrix M1 i-1 can be any suitable type of matrix. For example, it can be an affine transformation matrix (affine matrix), or it can be the transformation matrix corresponding to the affine matrix, etc.

[0096] In step S450, based on i = 1, 2... m, the transformation matrix M1 i-1 is used to perform a matrix transformation on the feature F1 i-1 ' to obtain the transformed feature F1 i ”.

[0097] See Figure 5 , the transformation matrix M1 output by the transformation convolutional network CT i-1 can be matrix-transformed (e.g., warped) with the feature F1 i-1 ' to obtain the transformed feature F1 i ”.

[0098] In step S460, based on i = 1, 2... m, the second decoder DE0' reconstructs the feature F1 i ” to obtain the reconstructed image I1 i ”.

[0099] As described above, the second encoder EN0' and the second decoder DE0' can form a reconstruction network and be trained together. In the application stage, the feature F1 i”Input the second decoder DE0' to reconstruct the feature by the second decoder DE0'. Refer to Figure 5 , the second decoder DE0' can output the reconstructed image I1 i ”.

[0100] In step S470, based on i = 1, 2... m, calculate the first image loss function through the predicted image I1 i ' and the reconstructed image I1 i ”, and train the transformation convolutional network CT based on the first image loss function to obtain the trained transformation convolutional network CT i-1 .

[0101] During the behavior planning process, except for the transformation convolutional network CT, the parameters of other network parts of the behavior planning model remain unchanged. After the training of the behavior planning model in the training stage, the transformation convolutional network CT has initial parameters (i.e., the first initial parameters). Subsequently, during the actual behavior planning, the parameters of the transformation convolutional network CT can be further adjusted (i.e., trained) to obtain a more accurate transformation matrix at each moment. For different i, the parameters of the transformation convolutional network CT are trained separately, and thus the corresponding transformation convolutional network CT at each moment can be obtained i-1 . Exemplarily, during training, the transformation convolutional network CT can be trained in a forward propagation manner based on the first image loss function.

[0102] When training the transformation convolutional network CT, it can be determined whether the Euclidean distance of the action is less than a preset threshold before and after updating the parameters of the transformation convolutional network CT. If not, continue the next round of training. If so, the training can be stopped and the trained transformation convolutional network CT can be obtained i-1 , where the action includes the expected acceleration and expected steering angle of the target vehicle at the moment of T1+(i - 1)*Δt. The preset threshold can be set as needed, such as 0.002, etc.

[0103] In step S480, based on i = 1, 2... m, based on the transformation matrix M1 i-1 output by the transformation convolutional network CT i-1 determine the expected acceleration and expected steering angle of the target vehicle at the moment of T1+(i - 1)*Δt.

[0104] Exemplarily, based on the transformation matrix M1 i-1 output by the transformation convolutional network CT i-1 determining the expected acceleration and expected steering angle of the target vehicle at the moment of T1+(i - 1)*Δt (step S480) can include: inputting the transformation matrix M i-1 into the behavior convolutional network CA for convolution to transform the matrix Mi-1 It is transformed from a size of 2*H*W to a size of 2*1*1. The two values in the transformed matrix respectively represent the desired acceleration and the desired steering angle of the target vehicle.

[0105] See Figure 5 , the behavior planning model may further include a behavior convolutional network CA. The parameters of the behavior convolutional network CA can also be trained well in the training stage, and the parameters of the behavior convolutional network CA are fixed in the application stage. The transformation matrix M i-1 is of size 2*H*W and can be convolved to transform it into a size of 2*1*1. The two values therein are respectively the desired acceleration and the desired steering angle of the target vehicle at the moment of T1+(i - 1)*Δt. At the moment of T1+(i - 1)*Δt, the control device of the target vehicle can control the movement of the target vehicle according to the desired acceleration and the desired steering angle.

[0106] Through the above method, the behavior of the vehicle can be planned based on the prediction of the future state of the vehicle (i.e., a series of predicted images). As described above, this can improve the interpretability of behavior planning and help cope with emergencies.

[0107] According to an embodiment of the present invention, the vehicle behavior planning method 400 may further include: obtaining a sample image I4; extracting the feature F4 of the sample image I4 through a second encoder EN0', and reconstructing the feature F4 through a second decoder DE0' to obtain a reconstructed image I4'; training the second encoder EN0' and the second decoder DE0' based on the sample image I4 and the reconstructed image I4'.

[0108] The sample image I4 can be any image. Exemplarily, a third reconstruction loss function can be calculated based on the sample image I4 and the reconstructed image I4', and the second encoder EN0' and the second decoder DE0' can be trained based on the third reconstruction loss function. Optionally, the third reconstruction loss function can be an L2 loss function.

[0109] The manner of training the second encoder EN0' and the second decoder DE0' based on the third reconstruction loss function is similar to the manner of training the first encoder EN0 and the first decoder DE0 based on the first reconstruction loss function above. This embodiment can be understood with reference to the corresponding description above and will not be elaborated here. This training method is simple to implement and has a small amount of calculation.

[0110] According to an embodiment of the present invention, the vehicle behavior planning method 400 may further include: obtaining sample images I40 and I41 collected by a third sample vehicle at times T4 and T4+Δt respectively, and the actual acceleration and actual steering angle of the third sample vehicle at time T4; extracting the feature F40 of the sample image I40 through a second encoder EN0'; inputting the feature F40 into a transformation convolutional network CT with second initial parameters for convolution to obtain a transformation matrix M40; performing matrix transformation on the feature F40 using the transformation matrix M40 to obtain a transformed feature F41'; reconstructing the feature F41' through a second decoder DE0' to obtain a reconstructed image I41'; calculating a second image loss function based on the sample image I41 and the reconstructed image I41', and training the transformation convolutional network CT based on the second image loss function to obtain a transformation convolutional network CT with first initial parameters; inputting the transformation matrix M40 output by the transformation convolutional network CT into a behavior convolutional network CA for convolution to determine the expected acceleration and expected steering angle of the third sample vehicle at time T4; calculating a behavior loss function based on the expected acceleration and expected steering angle of the third sample vehicle and the actual acceleration and actual steering angle, and training the behavior convolutional network CA based on the behavior loss function.

[0111] Any two of the third sample vehicle, the second sample vehicle, and the first sample vehicle described above may be the same or different, and the sample images I40 and I41 may be the same as or different from the sample images I30, I31... I3 m or the sample images I20, I21... I2 m in any pair of adjacent images.

[0112] The second initial parameters of the transformation convolutional network CT may be preset and will be transformed into the above-mentioned first initial parameters after training.

[0113] When training the transformation convolutional network CT and the behavior convolutional network CA, the parameters of the second encoder EN0' and the second decoder DE0' are fixed. The second encoder EN0' and the second decoder DE0' can be trained first, and then the transformation convolutional network CT and the behavior convolutional network CA can be trained after they are trained. Optionally, the parameters of the transformation convolutional network CT can be trained first, and then the parameters of the behavior convolutional network CA can be trained after it is trained.

[0114] During training, the second encoder EN0' and the second decoder DE0' can be used to process the sample image I40 to obtain the corresponding reconstructed image I41'. The sample image I41 can be used as the ground truth, and the loss (the second image loss function) between it and the reconstructed image I41' can be calculated. Then, based on this loss function, the parameters of the transformation convolutional network CT can be trained. Similar to the training in the application phase, it can be determined whether the Euclidean distance of the action is less than a preset threshold before and after updating the parameters of the transformation convolutional network CT. If not, the next round of training continues. If so, the training can be stopped and the trained transformation convolutional network CT with the first initial parameters can be obtained.

[0115] Subsequently, the expected acceleration and expected steering angle of the third sample vehicle at the T4 moment can be obtained through the trained transformation convolutional network CT with the first initial parameters. The actual acceleration and actual steering angle of the third sample vehicle can be used as the ground truth, and the loss (the behavior loss function) between them and the expected acceleration and expected steering angle can be calculated. Then, based on the behavior loss function, the behavior convolutional network CA can be trained, and finally the trained behavior convolutional network CA can be obtained.

[0116] The above training method is simple to implement and has a small computational amount.

[0117] According to an embodiment of the present invention, the vehicle behavior planning method 400 may further include: obtaining the sample image I5; extracting the feature F5 of the sample image I5 through the second encoder EN0', adding a random Gaussian variable to the feature F5 to obtain a new feature F5', and reconstructing the new feature F5' through the second decoder DE0' to obtain the reconstructed image I5'; using the second encoder EN0' and the second decoder DE0' as a generator to perform adversarial training with the second discriminator. In the adversarial training, the sample image I5 is used as the positive sample, and the reconstructed image I5' is used as the negative sample, and they are respectively input into the second discriminator for discrimination.

[0118] Optionally, the parameters of the second discriminator and the above first discriminator can be shared or independent of each other. Sharing parameters can reduce the number of parameters and improve the training speed of the model. Independent parameters can improve the processing accuracy of the model.

[0119] Similar to the above image prediction model, the behavior planning model can also be trained in an adversarial manner. The implementation manner of adversarial training of the second encoder EN0' and the second decoder DE0' based on the sample image I5 is similar to the implementation manner of adversarial training of the first encoder EN0 and the first decoder DE0 based on the sample image I30. You can refer to the corresponding description above to understand this embodiment, and details will not be repeated here.

[0120] As described above, during adversarial training, the adversarial loss function of the adversarial network composed of the generator and the discriminator can be calculated. The adversarial network composed of the second encoder EN0', the second decoder DE0', and the second discriminator can be trained based on the adversarial loss function. In addition, the loss function (which can be referred to as the fourth reconstruction loss function) can be calculated based on the sample image I5 and the reconstructed image I5'. Optionally, the above-mentioned adversarial network can be trained based on the adversarial loss function and the fourth reconstruction loss function. Optionally, the fourth reconstruction loss function can be the L2 loss function. Exemplarily, the adversarial loss function for training the second discriminator can be the Patch GAN loss.

[0121] According to an embodiment of the present invention, the vehicle behavior planning method 400 may further include: acquiring the sample images I50 and I51 collected by the fourth sample vehicle at times T5 and T5+Δt respectively, and the actual acceleration and actual steering angle of the fourth sample vehicle at time T5; extracting the feature F50 of the sample image I50 through the second encoder EN0'; inputting the feature F50 into the transformation convolutional network CT with the third initial parameters for convolution to obtain the transformation matrix M50; performing matrix transformation on the feature F50 using the transformation matrix M50 to obtain the transformed feature F51'; reconstructing the feature F51' through the second decoder DE0' to obtain the reconstructed image I51'; calculating the third image loss function based on the sample image I51 and the reconstructed image I51', and training the transformation convolutional network CT based on the third image loss function to obtain the transformation convolutional network CT with the first initial parameters; adding a random Gaussian variable to the transformation matrix M50 output by the transformation convolutional network CT to obtain a new transformation matrix M50', inputting the new transformation matrix M50' into the behavior convolutional network CA for convolution to determine the expected acceleration and expected steering angle of the fourth sample vehicle at time T5; using the behavior convolutional network CA as the generator to perform adversarial training with the third discriminator, where, during the adversarial training, the actual acceleration and actual steering angle are used as positive samples, and the expected acceleration and expected steering angle of the fourth sample vehicle are used as negative samples and input into the third discriminator for discrimination.

[0122] Any two of the fourth sample vehicle, the above-mentioned third sample vehicle, the above-mentioned second sample vehicle, and the above-mentioned first sample vehicle may be the same or different, the sample images I50 and I51 may be the same or different from the above-mentioned sample images I40 and I41, and the sample images I50 and I51 may be the same or different from the sample images I30, I31... I3 m or the sample images I20, I21... I2 m among any pair of adjacent images.

[0123] The parameters of the third discriminator and the second discriminator described above are independent of each other. The second discriminator is used to discriminate the authenticity of the input image, and the third discriminator is used to discriminate the authenticity of the input acceleration and rotation angle. Since their discrimination objects are different, the independent parameters are beneficial to improving the accuracy of the behavior planning model.

[0124] The third initial parameter can be arbitrary and can be the same as or different from the second initial parameter. The random Gaussian variable added to the transformation matrix M50 is consistent with the size of the transformation matrix M50.

[0125] Those skilled in the art can understand the implementation manner of adversarial training, which will not be elaborated here. By adopting the above solution, the quality of the acceleration and rotation angle generated by the behavior convolutional network CA can be improved through adversarial training.

[0126] It can be understood that the above second discriminator and third discriminator are only used in the training stage of the behavior planning model and are not used in the application stage of the behavior planning model (i.e., when actual behavior planning is performed).

[0127] According to another aspect of the present invention, an image prediction device is provided. Figure 6 A schematic block diagram of an image prediction device 600 according to an embodiment of the present invention is shown.

[0128] As Figure 6 shown, the image prediction device 600 according to an embodiment of the present invention includes an acquisition module 610, an extraction module 620, and a prediction module 630. Each of the modules can respectively execute the respective steps / functions of the image prediction method described above in combination with Figures 1-3 description. Only the main functions of the components of the image prediction device 600 will be described below, and the details already described above will be omitted.

[0129] The acquisition module 610 is used to acquire the current image I10 collected by the target vehicle at the current moment T1.

[0130] The extraction module 620 is used to extract the feature F10 of the current image I10 through the first encoder EN0.

[0131] The prediction module 630 is used to, for the (T1 + i*Δt)-th moment, in the prediction network N i based on one or more of the features F10 to F1 i-1 predict the feature F1 i , and reconstruct the feature F1 i to obtain the predicted image I1 i ' at the (T1 + i*Δt)-th moment, where i = 1, 2... m, m is an integer greater than or equal to 2, and Δt is a preset time period.

[0132] According to another aspect of the present invention, there is provided a vehicle behavior planning device. Figure 7 FIG. shows a schematic block diagram of a vehicle behavior planning device 700 according to an embodiment of the present invention.

[0133] As Figure 7 shown, the vehicle behavior planning device 700 according to an embodiment of the present invention includes an acquisition module 710, a first extraction module 720, a second extraction module 730, an input module 740, a transformation module 750, a reconstruction module 760, a training module 770, and a determination module 780. Each of the modules can respectively execute the respective steps / functions of the vehicle behavior planning method described above in conjunction with Figures 4-5 description. Only the main functions of the components of the vehicle behavior planning device 700 are described below, and the details already described above are omitted.

[0134] The acquisition module 710 is configured to acquire the current image I10 and the predicted images I11', I12'... I1 involved in the above image prediction method 100 m '.

[0135] The first extraction module 720 is configured to extract the feature F1 of the image I1 based on i = 1 through the second encoder EN0' i-1 '. i-1 '

[0136] The second extraction module 730 is configured to extract the feature F1 of the predicted image I1 based on i = 2, 3... m through the second encoder EN0' i-1 '. i-1 '

[0137] The input module 740 is configured to input the feature F1 based on i = 1, 2... m into the transformation convolutional network CT with the first initial parameters for convolution to obtain the transformation matrix M1 i-1 '. i-1

[0138] The transformation module 750 is configured to perform matrix transformation on the feature F1 based on i = 1, 2... m using the transformation matrix M1 i-1 to obtain the transformed feature F1 i-1 ''. i ''

[0139] The reconstruction module 760 is configured to reconstruct the feature F1 based on i = 1, 2... m through the second decoder DE0' i '' to obtain the reconstructed image I1 i ''.

[0140] The training module 770 is configured to, based on i = 1, 2... m, use the predicted image I1 i'and the reconstructed image I1 i ”Calculate the first image loss function, and train the transformation convolutional network CT based on the first image loss function to obtain the trained transformation convolutional network CT i-1 .

[0141] The determination module 780 is used to determine the expected acceleration and expected steering angle of the target vehicle at the moment T1+(i-1)*Δt based on i = 1, 2... m and based on the transformation matrix M1 i-1 output by the transformation convolutional network CT i-1 in the transformation convolutional network CT

[0142] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0143] Figure 8 Fig. 800 shows a schematic block diagram of an image prediction system 800 according to an embodiment of the present invention. The image prediction system 800 includes a memory 810 and a processor 820.

[0144] The memory 810 stores computer program instructions for implementing the corresponding steps in the image prediction method according to the embodiments of the present invention.

[0145] The processor 820 is used to run the computer program instructions stored in the memory 810 to execute the corresponding steps of the image prediction method according to the embodiments of the present invention.

[0146] In one embodiment, when the computer program instructions are run by the processor 820, they are used to execute the following steps: obtain the current image I10 collected by the target vehicle at the current moment T1; extract the feature F10 of the current image I10 through the first encoder EN0; for the moment T1+i*Δt, in the prediction network N i predict the feature F1 based on one or more of the features F10 to F1 i-1 in the prediction network N i , and reconstruct the feature F1 i to obtain the predicted image I1 at the moment T1+i*Δt i ', where i = 1, 2... m, m is an integer greater than or equal to 2, and Δt is a preset time period.

[0147] Figure 9FIG. 0 shows a schematic block diagram of a vehicle behavior planning system 900 according to an embodiment of the present invention. The vehicle behavior planning system 900 includes a memory 910 and a processor 920.

[0148] The memory 910 stores computer program instructions for implementing corresponding steps in the vehicle behavior planning method according to an embodiment of the present invention.

[0149] The processor 920 is configured to run the computer program instructions stored in the memory 910 to execute corresponding steps of the vehicle behavior planning method according to an embodiment of the present invention.

[0150] In one embodiment, when the computer program instructions are run by the processor 920, they are used to perform the following steps: obtaining the current image I10 and the predicted images I11', I12'... I1 m ' involved in the above image prediction method; based on i = 1, extracting the feature F1 i-1 of the image I1 i-1 ' through the second encoder EN0'; based on i = 2, 3... m, extracting the feature F1 i-1 of the predicted image I1 i-1 ' through the second encoder EN0'; based on i = 1, 2... m, inputting the feature F1 i-1 ' into the transformation convolutional network CT with the first initial parameters for convolution to obtain the transformation matrix M1 i-1 ; using the transformation matrix M1 i-1 to perform matrix transformation on the feature F1 i-1 ' to obtain the transformed feature F1 i ”; reconstructing the feature F1 i ” through the second decoder DE0' to obtain the reconstructed image I1 i ”; calculating the first image loss function based on the predicted image I1 i ' and the reconstructed image I1 i ”, and training the transformation convolutional network CT based on the first image loss function to obtain the trained transformation convolutional network CT i-1 ; determining the desired acceleration and desired steering angle of the target vehicle at the moment of T1+(i-1)*Δt based on the transformation matrix M1 i-1 output by the transformation convolutional network CT i-1

[0151] In addition, according to an embodiment of the present invention, a storage medium is further provided. Program instructions are stored on the storage medium, and when the program instructions are run by a computer or a processor, they are used to execute the corresponding steps of the image prediction method according to the embodiment of the present invention, and are used to implement the corresponding modules in the image prediction device according to the embodiment of the present invention. The storage medium may, for example, include a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above storage media.

[0152] In one embodiment, when the program instructions are run by a computer or a processor, they can cause the computer or the processor to implement each functional module of the image prediction device according to the embodiment of the present invention, and / or can execute the image prediction method according to the embodiment of the present invention.

[0153] In one embodiment, the program instructions are used to execute the following steps when running: obtaining a current image I10 collected by a target vehicle at a current moment T1; extracting a feature F10 of the current image I10 through a first encoder EN0; for the moment T1 + i*Δt, in a prediction network N i in which, based on one or more of the features F10 to F1 i-1 to predict a feature F1 i and reconstruct the feature F1 i to obtain a predicted image I1 i ' at the moment T1 + i*Δt, where i = 1, 2... m, m is an integer greater than or equal to 2, and Δt is a preset time period.

[0154] In addition, according to an embodiment of the present invention, a storage medium is further provided. Program instructions are stored on the storage medium, and when the program instructions are run by a computer or a processor, they are used to execute the corresponding steps of the vehicle behavior planning method according to the embodiment of the present invention, and are used to implement the corresponding modules in the vehicle behavior planning device according to the embodiment of the present invention. The storage medium may, for example, include a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above storage media.

[0155] In one embodiment, when the program instructions are run by a computer or a processor, they can cause the computer or the processor to implement each functional module of the vehicle behavior planning device according to the embodiment of the present invention, and / or can execute the vehicle behavior planning method according to the embodiment of the present invention.

[0156] In one embodiment, when the program instructions are running, they are used to perform the following steps: obtaining the current image I10 and the predicted images I11', I12'... I1 involved in the above image prediction method m '; based on i = 1, extracting the feature F1 of the image I1 i-1 through the second encoder EN0' i-1 '; based on i = 2, 3... m, extracting the feature F1 of the predicted image I1 i-1 ' through the second encoder EN0' i-1 '; based on i = 1, 2... m, inputting the feature F1 i-1 ' into the transform convolutional network CT with the first initial parameters for convolution to obtain the transform matrix M1 i-1 ; using the transform matrix M1 i-1 to perform matrix transformation on the feature F1 i-1 ' to obtain the transformed feature F1 i ”; reconstructing the feature F1 i ” through the second decoder DE0' to obtain the reconstructed image I1 i ”; calculating the first image loss function based on the predicted image I1 i ' and the reconstructed image I1 i ”, and training the transform convolutional network CT based on the first image loss function to obtain the trained transform convolutional network CT i-1 ; determining the expected acceleration and expected steering angle of the target vehicle at the moment of T1 + (i - 1)*Δt based on the transform matrix M1 i-1 output by the transform convolutional network CT i-1 .

[0157] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different systems for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0158] Similarly, it should be understood that, for the sake of streamlining the present invention and facilitating the understanding of one or more of the various aspects of the invention, in the description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, the system of the present invention should not be construed as reflecting the intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the corresponding claims, the inventive point lies in that the corresponding technical problem can be solved with features less than all the features of a single disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, where each claim itself serves as a separate embodiment of the present invention.

[0159] It should be noted that the above embodiments are illustrative of the present invention rather than restrictive thereof, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words may be construed as names.

[0160] As described above, it is only the specific implementation manner or the description of the specific implementation manner of the present invention, and the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, and all of them should be covered by the protection scope of the present invention. The protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. An image prediction method, comprising: Obtaining a current image I10 collected by a target vehicle at a current moment T1; Extracting a feature F10 of the current image I10 through a first encoder EN0; For the time T1 + i*Δt, in the prediction network N i one or more of the features F10 to F1 are used to predict the feature F1 i-1 in which one or more of the features F10 to F1 i are used to predict the feature F1 i-1 including: i ​ Based on i = 1, predicting the feature F1 based on the feature F10 i ; Based on i≥2, based on the feature F10 and the feature F1 i-1 Predict the feature F1 i ; And reconstruct the feature F1 i to obtain the predicted image I1 at the moment of T1 + i*Δt i ', where i = 1, 2... m, m is an integer greater than or equal to 2, and Δt is a preset time period. Among them, the time gap between adjacent features from the feature F10 to the feature F1 i is the same; Among them, predicting feature F1 based on one or more of the features F10 to feature F1 i-1 includes: i including: For the feature F10 to the feature F1 i-1 For each feature F1 participating in prediction j among them, based on the feature F1 j calculate the attention mask S1 ij ; For the feature F1 j and the attention mask S1 ij perform matrix inner product calculation to obtain the attention feature FS1 ij ; Input the attention feature FS1 ij into a fully connected layer or a convolutional layer for weighted sum of features to obtain a weighted feature FA1 ij ; All the weighted features obtained in the prediction network N i are fused to obtain the feature F1 i ; Where j ∈ {0, 1... i - 1}.

2. The method according to claim 1, wherein Based on the feature F1 j Calculate the attention mask S1 ij including: Input the feature F1 j into the prediction network N i wherein the mask convolutional network CS ij in it performs convolution to obtain the attention mask S1 ij , wherein the attention mask S1 ij has the same height and width as the feature F1 j and the number of channels is 1, and each element in the attention mask S1 ij represents the response value of the position where the vehicle will drive to.

3. The method according to claim 2, wherein predicting the feature F1 based on one or more of the features F10 to F1 i-1 includes: i ​ Predict the feature F1 based at least on the feature F10 i ; Among them, for different i, the parameters of the mask convolutional network CS i0 are independent of each other.

4. The method according to claim 2, wherein Predicting feature F1 based on one or more of the features F10 to F1 i-1 includes: i ​ Based at least on the feature F1 i-1 Predict the feature F1 i ; Among them, for different i, the parameters of the masked convolutional network CS i(i-1) are shared.

5. The method according to any one of claims 1 to 2, wherein The reconstruction of the feature F1 i is performed to obtain a predicted image I1 at the (T1 + i*Δt)-th moment i ' includes: Input the feature F1 i into the decoder DE i in the prediction network N i to obtain the predicted image I1 i '.

6. The method according to claim 5, further comprising: The feature F10 is reconstructed by the first decoder DE0 to obtain a reconstructed image I10', where the decoder DE i shares parameters with the first decoder DE0.

7. The method according to any one of claims 1 to 2, wherein The method further comprises: Obtain (m + 1) sample images I20, I21... I2 collected by the first sample vehicle at times T2, T2 + Δt... T2 + mΔt respectively m ; Extracting a feature F20 of the sample image I20 through the first encoder EN0, and reconstructing the feature F20 through a first decoder DE0 to obtain a reconstructed image I20'; Training the first encoder EN0 and the first decoder DE0 based on the sample image I20 and the reconstructed image I20'; For the time T2 + i*Δt, in the prediction network N i one or more of the features F20 to F2 i-1 are used to predict the feature F2 i , and the feature F2 i is reconstructed to obtain the predicted image I2 i ' at time T2 + i*Δt; Based on the sample image I2 i and the predicted image I2 i 'train the prediction network N i thereby.

8. The method according to any one of claims 1 to 2, further comprising: Obtain (m + 1) sample images I30, I31... I3 collected by the second sample vehicle at times T3, T3 + Δt... T3 + mΔt respectively m ; Extracting a feature F30 of the sample image I30 through the first encoder EN0, adding a random Gaussian variable to the feature F30 to obtain a new feature F30', and reconstructing the new feature F30' through a first decoder DE0 to obtain a reconstructed image I30'; Using the first encoder EN0 and the first decoder DE0 as a generator to perform adversarial training with a first discriminator. Wherein, in the adversarial training, the sample image I30 is used as a positive sample, and the reconstructed image I30' is used as a negative sample, and they are respectively input into the first discriminator for discrimination; For the time T3 + i*Δt, in the prediction network N i one or more of the features F30 to F3 i-1 are used to predict the feature F3 i , and a random Gaussian variable is added to the feature F3 i to obtain a new feature F3 i ', and the new feature F3 i ' is reconstructed to obtain the predicted image I3 i ' at time T3 + i*Δt; Use the prediction network N i as a generator and perform adversarial training with the first discriminator. Among them, in the adversarial training, use the sample image I3 i as a positive sample, and use the predicted image I3 i ' as a negative sample, and input them into the first discriminator for discrimination respectively.

9. A vehicle behavior planning method, comprising: Obtain the current image I10 and the predicted images I11', I12'... I1 involved in the image prediction method according to any one of claims 1 to 8 m '; Based on i = 1, extract the features F1 of the image I1 through the second encoder EN0' i-1 ; i-1 ' Based on i = 2, 3,..., m, extract the feature F1 i-1 ' of the predicted image I1 i-1 ' by means of the second encoder EN0'; Based on i = 1, 2,..., m, the feature F1 i-1 is input into the transform convolutional network CT with the first initial parameter for convolution to obtain the transform matrix M1 i-1 ; Adopt the transformation matrix M1 i-1 Perform matrix transformation on the feature F1 i-1 ' to obtain the transformed feature F1 i ”; The feature F1 is reconstructed by the second decoder DE0' i to obtain a reconstructed image I1 i "; By predicting the image I1 i ' and the reconstructed image I1 i ” calculate the first image loss function, and train the transform convolutional network CT based on the first image loss function to obtain the trained transform convolutional network CT i-1 ; Based on the transformed convolutional network CT i-1 The output transformation matrix M1 i-1 Determine the expected acceleration and expected steering angle of the target vehicle at the moment of T1+(i-1)*Δt.

10. The method according to claim 9, wherein, The transformation matrix M1 output based on the transformation convolutional network CT i-1 Determining the expected acceleration and expected steering angle of the target vehicle at the moment of T1+(i-1)*Δt includes: i-1 Determining the expected acceleration and expected steering angle of the target vehicle at the moment of T1+(i-1)*Δt includes: Input the transformation matrix M i-1 into the convolutional network CA for convolution, so as to transform the transformation matrix M i-1 from the size of 2*H*W to the size of 2*1*1, where the two values in the transformed matrix respectively represent the desired acceleration and the desired steering angle of the target vehicle.

11. The method according to claim 9, wherein, The second encoder EN0' shares parameters with the first encoder EN0.

12. The method according to any one of claims 9 to 11, wherein, The vehicle behavior planning method further comprises: Obtaining a sample image I4; Extracting a feature F4 of the sample image I4 through the second encoder EN0', and reconstructing the feature F4 through the second decoder DE0' to obtain a reconstructed image I4'; Training the second encoder EN0' and the second decoder DE0' based on the sample image I4 and the reconstructed image I4'.

13. The method according to claim 10, wherein The vehicle behavior planning method further comprises: Obtaining sample images I40 and I41 collected by a third sample vehicle at moments T4 and T4 + Δt respectively, and the actual acceleration and actual steering angle of the third sample vehicle at the moment T4; Extracting a feature F40 of the sample image I40 through the second encoder EN0'; Inputting the feature F40 into a transformation convolutional network CT with second initial parameters for convolution to obtain a transformation matrix M40; Performing matrix transformation on the feature F40 using the transformation matrix M40 to obtain a transformed feature F41'; Reconstructing the feature F41' through the second decoder DE0' to obtain a reconstructed image I41'; Calculating a second image loss function based on the sample image I41 and the reconstructed image I41', and training the transformation convolutional network CT based on the second image loss function to obtain a transformation convolutional network CT with the first initial parameters; Inputting the transformation matrix M40 output by the transformation convolutional network CT into a behavior convolutional network CA for convolution to determine the expected acceleration and expected steering angle of the third sample vehicle at the moment T4; Calculate a behavior loss function based on the expected acceleration and expected steering angle of the third sample vehicle and the actual acceleration and the actual steering angle, and train the behavior convolutional network CA based on the behavior loss function.

14. The method according to any one of claims 9 to 11, wherein The vehicle behavior planning method further includes: Obtain a sample image I5; Extract the feature F5 of the sample image I5 through the second encoder EN0', add a random Gaussian variable to the feature F5 to obtain a new feature F5', and reconstruct the new feature F5' through the second decoder DE0' to obtain a reconstructed image I5'; Use the second encoder EN0' and the second decoder DE0' as a generator, and perform adversarial training with a second discriminator. Wherein, in the adversarial training, use the sample image I5 as a positive sample and the reconstructed image I5' as a negative sample, and input them into the second discriminator for discrimination respectively.

15. The method according to claim 10, wherein The vehicle behavior planning method further includes: Obtain sample images I50 and I51 collected by a fourth sample vehicle at times T5 and T5+Δt respectively, and the actual acceleration and actual steering angle of the fourth sample vehicle at time T5; Extract the feature F50 of the sample image I50 through the second encoder EN0'; Input the feature F50 into a transformation convolutional network CT with third initial parameters for convolution to obtain a transformation matrix M50; Perform matrix transformation on the feature F50 using the transformation matrix M50 to obtain a transformed feature F51'; Reconstruct the feature F51' through the second decoder DE0' to obtain a reconstructed image I51'; Calculate a third image loss function based on the sample image I51 and the reconstructed image I51', and train the transformation convolutional network CT based on the third image loss function to obtain a transformation convolutional network CT with the first initial parameters; Add a random Gaussian variable to the transformation matrix M50 output by the transformation convolutional network CT to obtain a new transformation matrix M50', and input the new transformation matrix M50' into the behavior convolutional network CA for convolution to determine the expected acceleration and expected steering angle of the fourth sample vehicle at time T5; Use the behavior convolutional network CA as a generator, and perform adversarial training with a third discriminator. Wherein, in the adversarial training, use the actual acceleration and the actual steering angle as positive samples, and use the expected acceleration and expected steering angle of the fourth sample vehicle as negative samples, and input them into the third discriminator for discrimination respectively.

16. An image prediction device, comprising: An acquisition module, configured to acquire a current image I10 collected by a target vehicle at a current time T1; An extraction module, configured to extract a feature F10 of the current image I10 through a first encoder EN0; The prediction module is used to predict the network N at the time T1+i*Δt. i Based on the feature F10 to feature F1 i-1 One or more of them to predict feature F1 i , and for the feature F1 i Reconstruct to obtain the predicted image I1 at time T1+i*Δt i ', i = 1, 2 ... m, m is an integer greater than or equal to 2, Δt is a preset period, wherein the feature F1i is predicted based on one or more of the features F10 to F1i-1, including: Based on i = 1, predict the feature F1i based on the feature F10; Based on i≥2, predict the feature F1i based on the feature F10 and the feature F1i-1; and the time gaps between adjacent features from the feature F10 to the feature F1 i are the same; Among them, predicting Feature F1 based on one or more of the said Feature F10 to Feature F1 i-1 includes: i including: For the feature F10 to the feature F1 i-1 For each feature F1 participating in the prediction j among them, based on the feature F1 j calculate the attention mask S1 ij ; For the feature F1 j and the attention mask S1 ij perform an inner matrix product calculation to obtain the attention feature FS1 ij ; Input the attention feature FS1 ij into a fully connected layer or a convolutional layer for weighted sum of features to obtain a weighted feature FA1 ij ; All the weighted features obtained in the prediction network N i are fused to obtain the feature F1 i ; Wherein, j∈{0,1……i-1}.

17. A vehicle behavior planning device, comprising: An acquisition module, configured to acquire the current image I10 and the predicted images I11', I12'... I1 involved in the image prediction method according to any one of claims 1 to 8 m '; The first extraction module is used to extract the feature F1 of the image I1 through the second encoder EN0' based on i = 1 i-1 ; i-1 ' The second extraction module is configured to extract the predicted image I1 based on i = 2, 3,..., m through the second encoder EN0'. i-1 The feature F1 of i-1 '; An input module, for, based on i = 1, 2,..., m, convolving the feature F1 i-1 in a transform convolutional network CT with a first initial parameter to obtain a transform matrix M1 i-1 ; A transformation module, which is configured to perform matrix transformation on the feature F1 based on i = 1, 2,..., m using the transformation matrix M1 i-1 to obtain a transformed feature F1 i-1 ' i "; A reconstruction module, which is configured to reconstruct the feature F1 through a second decoder DE0' based on i = 1, 2... m i to obtain a reconstructed image I1 i "; A training module, which is used to calculate a first image loss function based on i = 1, 2,..., m, by predicting an image I1 i ' and the reconstructed image I1 i ” and train the transform convolutional network CT based on the first image loss function to obtain a trained transform convolutional network CT i-1 ; A determination module, configured to determine, based on the transformation convolutional network CT where i = 1, 2,..., m i-1 the output transformation matrix M1 i-1 and determine the desired acceleration and the desired steering angle of the target vehicle at the moment of T1 + (i - 1) * Δt.

18. An image prediction system, comprising a processor and a memory, wherein, Computer program instructions are stored in the memory, and when the computer program instructions are run by the processor, they are used to execute the image prediction method according to any one of claims 1 to 8.

19. A vehicle behavior planning system, comprising a processor and a memory, wherein, Computer program instructions are stored in the memory, and when the computer program instructions are run by the processor, they are used to execute the vehicle behavior planning method according to any one of claims 9 to 15.

20. A storage medium, on which program instructions are stored, and the program instructions are used to execute the image prediction method according to any one of claims 1 to 8 when running.

21. A storage medium, on which program instructions are stored, and the program instructions are used to execute the vehicle behavior planning method according to any one of claims 9 to 15 when running.

Citation Information

Patent Citations

  • Image processing methods and devices, equipment and memory medium

    CN109361934A

  • Method and device for determining movement strategy of unmanned vehicle

    CN110488821A