Information prediction method, method of training autonomous driving model, and device
By encoding perception data to generate prediction token sequences and control information within autonomous driving systems, the method addresses the challenge of relying on expensive annotations, enhancing the scalability and practicality of autonomous driving technology.
Patent Information
- Application Number
- JP2025040819
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-19
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-14
AI Technical Summary
Existing autonomous driving systems rely on expensive and labor-intensive accurate perception annotations for training, limiting their scalability and practicality.
The method involves obtaining perception data from vehicles, encoding image and driving data to generate prediction token sequences and control information using a generation model, thereby predicting control information without relying on large-scale accurate annotations.
This approach enables robust environmental perception and representation in autonomous driving systems, reducing the need for expensive annotations and improving scalability and practicality.
Smart Images

Figure 2025090781000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to technical fields such as computer vision and deep learning, and can be applied to scenarios such as autonomous driving.
Background Art
[0002] With the rapid development of artificial intelligence and autonomous driving technologies, end-to-end autonomous driving systems have attracted great attention due to the simplification of system architecture, reduction of error accumulation, global optimization functions, etc.
[0003] For example, an autonomous driving system can analyze perception data to predict control signals for a vehicle. The vehicle can control the braking system according to the predicted control signals to realize the autonomous driving of the vehicle. The introduction of this autonomous driving system can meet the need for convenient travel to a certain extent.
Summary of the Invention
[0004] The present disclosure provides an information prediction method, a training method for an autonomous driving model, an apparatus, a device, a medium, a program product, and an autonomous driving vehicle that can realize the prediction of control information without relying on large-scale and accurate annotation data.
[0005] According to a first aspect of the present disclosure, there is provided an information prediction method including: obtaining perception data including image data collected by sensors in a vehicle and driving data of the vehicle; encoding the image data to obtain an image token sequence corresponding to the image data; encoding the driving data to obtain driving features corresponding to the driving data; and generating, based on the driving features and the image token sequence, a prediction token sequence corresponding to the image token sequence and control information for the vehicle by using a generation model.
[0006] According to a second aspect of the present disclosure, a method for training an autonomous driving model is provided. The autonomous driving model includes an encoding layer and a generation model. The encoding layer includes a sequence encoding network and a driving data encoding network. The training method includes encoding image data among sample perception data using the sequence encoding network to obtain an image token sequence corresponding to the image data; encoding driving data among the sample perception data using the driving data encoding network to obtain driving features corresponding to the driving data; generating, based on the driving features and the image token sequence, a predicted token sequence corresponding to the image token sequence and predicted control information for the vehicle using the generation model; and training the autonomous driving model according to the predicted token sequence and the image token sequence.
[0007] According to a third aspect of the present disclosure, an information prediction device is provided, including a data acquisition module configured to acquire perception data including image data collected by sensors in a vehicle and driving data of the vehicle, a first encoding module configured to encode the image data to obtain an image token sequence corresponding to the image data, a second encoding module configured to encode the driving data to obtain driving features corresponding to the driving data, and a generation module configured to generate, based on the driving features and the image token sequence, a predicted token sequence corresponding to the image token sequence and control information for the vehicle using a generation model.
[0008] According to a fourth aspect of the present disclosure, a training apparatus for an autonomous driving model is provided. The autonomous driving model includes an encoding layer and a generation model, and the encoding layer includes a sequence encoding network and a driving data encoding network. The training apparatus includes: a first encoding module configured to encode image data among sample perception data using the sequence encoding network to obtain an image token sequence corresponding to the image data; a second encoding module configured to encode driving data among sample perception data using the driving data encoding network to obtain driving features corresponding to the driving data; a generation module configured to generate a predicted token sequence corresponding to the image token sequence and predicted control information for a vehicle using the generation model based on the driving features and the image token sequence; and a training module configured to train the autonomous driving model according to the predicted token sequence and the image token sequence.
[0009] According to a fifth aspect of the present disclosure, an electronic device is provided, which includes at least one processor and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor, and the instructions cause the at least one processor to execute an information prediction method provided by the present disclosure.
[0010] According to a sixth aspect of the present disclosure, an electronic device is provided, which includes at least one processor and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor, and the instructions cause the at least one processor to execute a training method for an autonomous driving model provided by the present disclosure.
[0011] According to a seventh aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, and the computer instructions cause a computer to execute the information prediction method or the training method for an autonomous driving model provided by the present disclosure.
[0012] According to an eighth aspect of the present disclosure, there is provided a computer program product, which is stored in at least one of a readable storage medium and an electronic device, and includes computer program instructions for implementing the information prediction method or the training method of the autonomous driving model provided by the present disclosure when executed by a processor.
[0013] According to a ninth aspect of the present disclosure, there is provided an autonomous driving vehicle equipped with the electronic device provided by the fifth aspect of the present disclosure.
[0014] It should be understood that the content described in this part is not intended to indicate the key points or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will be easily understood from the following description.
Brief Description of the Drawings
[0015] The drawings are for better understanding of the present technical solution and do not limit the present application.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Mode for Carrying Out the Invention
[0016] Hereinafter, exemplary embodiments of the present disclosure will be described with reference to the drawings. Here, various details of the embodiments of the present disclosure are included for easier understanding, and they should be considered exemplary. Therefore, those skilled in the art should understand that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clear and concise description, descriptions of well-known functions and configurations are omitted in the following description.
[0017] A major problem faced by end-to-end autonomous driving systems is how to establish robust environmental perception and representation. For example, the ability to establish robust environmental perception and representation can be improved by introducing auxiliary perception tasks. Specifically, existing autonomous driving systems usually rely on multi-task support plans and improve environmental representation by introducing auxiliary monitoring signals. However, these methods often require accurate perception annotations that are expensive and difficult to obtain. That is, accurately annotated perception data is required to use the annotation data as an auxiliary monitoring signal. Therefore, there are problems of being time-consuming and labor-intensive, and the introduction cost, scalability, and practicality of the autonomous driving model are limited to a certain extent.
[0018] The present disclosure provides an information prediction method, a method for training an autonomous driving model, an apparatus, a device, a medium, a program product, and an autonomous vehicle to solve the problems existing in the prior art. First, the application scenarios of the method and apparatus provided by the present disclosure will be described in combination with FIG. 1 below.
[0019] FIG. 1 is a schematic diagram of an application scenario of an information prediction method, a method for training an autonomous driving model, and an apparatus according to an embodiment of the present disclosure.
[0020] As shown in FIG. 1, the application scenario 100 of this embodiment may include a vehicle 110 and a server 120. The vehicle 110 is integrated with an autonomous driving system, and the server 120 may be, for example, a background management server that supports the operation of an intelligent driving system.
[0021] In one embodiment, the vehicle 110 may also be integrated with various types of sensors communicatively connected to the autonomous driving system, such as a vision camera and a radar-type ranging sensor. Among them, examples of the vision camera include a monocular camera, a binocular stereo camera, a panoramic vision camera, an infrared camera, etc. Examples of the radar-type ranging sensor include a lidar, a millimeter-wave radar, an ultrasonic radar, etc. The autonomous driving system may, for example, process the data collected by various sensors to determine the environmental information of the vehicle 110 and the control information for the vehicle 110, and cause the vehicle to travel according to the control information.
[0022] In one embodiment, the server 120 can provide, for example, an autonomous driving model 130 to an autonomous driving system. The autonomous driving system processes the image data collected by sensors according to the autonomous driving model 130, combines the current driving data of the vehicle, predicts control information for the vehicle, and enables the vehicle to automatically drive according to the control information. For example, the autonomous driving model 130 has predicting control information as the main task, generating images as the auxiliary task, synchronously outputs an image token sequence and the predicted control information, and can decode the output image token sequence to generate an image.
[0023] In one embodiment, the server 120 can train the autonomous driving model 130 in combination with annotation data so that the autonomous driving model has the ability to predict control information. The server 120 uses the image data to self-supervised train the autonomous driving model 130, enables the autonomous driving model to have the ability to generate images to execute the auxiliary task, and depending on the auxiliary task, can improve the accuracy of the autonomous driving model 130 when executing the main task.
[0024] In one embodiment, the autonomous driving system in the vehicle 110 can also send the data collected by the sensors to the server 120, and the server 120 can execute image generation and prediction of control information. Then, the server 120 sends the predicted control information to the autonomous driving system, and the autonomous driving system controls the movement of the vehicle according to the control information.
[0025] Note that the information prediction method provided by the present disclosure may be executed by a vehicle or an autonomous driving system in the vehicle, or may be executed by the server 120. Therefore, the information prediction device provided by the present disclosure may be provided in a vehicle or an autonomous driving system included in the vehicle, or may be provided in the server 120. The training method of the autonomous driving model provided by the present disclosure can be executed by the server 120. Therefore, the training device of the autonomous driving model provided by the present disclosure may be installed in the server 120.
[0026] Note that the number and type of the vehicle 110 and the server 120 in FIG. 1 are merely illustrative. Depending on the implementation requirements, there may be vehicles 110 and servers 120 with arbitrary numbers and types.
[0027] Hereinafter, the information prediction method provided by the present disclosure will be described in detail with reference to FIGS. 2 to 6.
[0028] FIG. 2 is a schematic diagram of a flowchart of the information prediction method according to an embodiment of the present disclosure.
[0029] As shown in FIG. 2, the information prediction method 200 of this embodiment can include operations S210 to S240.
[0030] In operation S210, perception data is acquired. The perception data includes at least image data collected by sensors in the vehicle and driving data of the vehicle.
[0031] According to an embodiment of the present disclosure, perception data can be acquired in real time during the execution of the information prediction method. The driving data of the vehicle may include one or more of navigation data at the current time, speed at the current time, control information for the vehicle at the current time, and the like. The control information may include one or more of an accelerator pedal angle, a brake pedal angle, a rotation direction of the steering wheel, and a rotation angle of the steering wheel. The driving data can be acquired, for example, by a vehicle control system in the vehicle or by an autonomous driving system in the vehicle, and the present disclosure is not limited thereto. The image data collected by the sensors may include, for example, an image of the environment at the current time collected by a camera in the vehicle.
[0032] In one embodiment, the acquired perceptual data may include, in addition to the image data and driving data collected at the current time, the image data and driving data collected at past times (e.g., one or more past times). The image data at past times and the image data at the current time may constitute an image data sequence, and the driving data at past times and the driving data at the current time may constitute a driving data sequence.
[0033] In operation S220, the image data is encoded and an image token sequence corresponding to the image data is obtained.
[0034] According to an embodiment of the present disclosure, for example, an environmental image can be first processed in blocks, and then each blocked image block can be encoded to obtain one token. By encoding a plurality of blocked image blocks, one token sequence, that is, an image token sequence, can be obtained. For example, a two-dimensional convolutional layer can be used to encode each image block to obtain one token.
[0035] When an image data sequence is obtained, each image data in the image data sequence is encoded to obtain one token sequence, and a plurality of token sequences can be obtained in total. Then, according to the collection order of the image data, the plurality of token sequences are stitched to obtain an image token sequence corresponding to the image data sequence.
[0036] In step S230, the driving data is encoded to obtain driving features corresponding to the driving data.
[0037] For example, the driving data can be represented by a series of numerical values including the speed of the vehicle, the two-dimensional coordinates of the target point in the navigation path, the pedal angle of the vehicle accelerator, and the like. For example, an embedding layer can be used to encode the driving data and map the driving data to a vector to obtain driving features.
[0038] When a driving data sequence is acquired, the driving data at each time in the driving data sequence may be encoded to obtain one feature. A plurality of features are summed for the driving data sequence. Then, according to the collection order of the driving data, the plurality of features are stitched to obtain a driving feature corresponding to the driving data sequence.
[0039] In operation S240, based on the driving feature and the image token sequence, a generation model is used to generate a predicted token sequence corresponding to the image token sequence and control information for the vehicle.
[0040] For example, the driving feature can be stitched with the image token sequence to obtain an input feature, which is input into the generation model to obtain a predicted token sequence and control information. The control information can be understood as a control signal of the vehicle and can include at least one of information such as the accelerator pedal angle, the brake pedal angle, the rotation direction of the steering wheel, and the rotation angle of the steering wheel.
[0041] In this embodiment, the generation model is a model that executes a plurality of tasks. The plurality of tasks to be executed include an image generation task and a control information prediction task. The image generation task is an image reconstruction task or an image prediction task, and the image generation task is an auxiliary task for the control information prediction task.
[0042] For example, the generation model can be a variational autoencoder (VAEs), a diffusion model, or an autoregressive model. The autoregressive model is, for example, a model based on a transformer architecture, but is not limited to this in the present disclosure.
[0043] In one embodiment, the image token sequence is a sequence X corresponding to the image data at time t t =(x t,1 ,x t,2 ,…,x t,nand the predicted token sequence generated by the generation model is, for example, the predicted sequence X corresponding to the image data at time t t ' =(x t,1 ' ,x t,2 ' ,…,x t,n ' ) and may also be the predicted sequence X corresponding to the image data at time t + 1 t+1 =(x t+1,1 ,x t+1,2 ,…,x t+1,n ).
[0044] After acquiring the control information, the automatic driving system in the vehicle can, for example, determine the driving parameters according to this control information and control the driving of the vehicle according to these driving parameters.
[0045] In the embodiments of the present disclosure, by introducing an auxiliary task of predicting the token sequence of the image, it is possible to predict the control information for the vehicle, thereby establishing robust environmental perception and representation in the process of predicting the control information. In addition, since the collected image can be used as a monitoring signal to implement the auxiliary task, there is no need to rely on accurate perception annotations, the difficulty of implementing the auxiliary task is effectively reduced, the wide deployment of the control information prediction task becomes easy, and the scalability and practicality of the control information prediction task are improved.
[0046] FIG. 3 is a schematic diagram of the implementation principle of the information prediction method according to the embodiments of the present disclosure.
[0047] In one embodiment, in the process of encoding image data into an image token sequence, fine-grained information in the image, such as the correlation information between image blocks corresponding to two tokens and the impact of small objects such as traffic lights on the prediction accuracy of control information, may be lost. In order to further ensure the prediction accuracy of the control information and enable the auxiliary task to better support the completion of the main task (predicting the control information), in this embodiment, the image data is encoded to obtain an image token sequence, and at the same time, the image features of the image data are extracted, that is, feature extraction is performed with the image data as a unit, and the extracted features can be used as part of the input data of the generation model.
[0048] As shown in FIG. 3, in Embodiment 300, the sequence encoding network 311 can be used to encode the image data 301 to obtain an image token sequence 302. In this Embodiment 300, the first convolutional network 312 can also be used to extract the image features of the image data 301 and obtain a first feature vector 303 representing the image features. This Embodiment 300 also encodes the driving data 304 using the driving data encoding network 313 to obtain driving features 305.
[0049] After obtaining the driving features 305, the first feature vector 303, and the image token sequence 302, the generation model 320 is used to perform the auxiliary task and the main task based on these features, that is, the predicted token sequence 306 and the control information 307 for the vehicle can be generated. Specifically, the driving features 305, the first feature vector 303, and the image token sequence 302 are stitched as input features and input into the generation model 320, and the generation model 320 outputs the predicted token sequence 306 and the control information 307.
[0050] In one embodiment, the generation model 320 may be an autoregressive model. When executing a task, the driving feature 305 and the first feature vector 303 are stitched and input into the generation model 320, and the generation model predicts to obtain the first token of the token sequence 306. Next, the driving feature 305, the first feature vector 303, and the first image token in this image token sequence 302 are stitched and input into the generation model 320, and the generation model predicts to obtain the second token of the token sequence 306. Similarly hereinafter, the token sequence 306 is generated. Finally, the driving feature 305, the first feature vector 303, and this image token sequence 302 are stitched and input into the generation model 320, and the generation model 320 predicts to obtain the control information 307.
[0051] In one embodiment, when the generation model 320 predicts a single token, the generated data is a probability vector including the probabilities of each token within a plurality of tokens. In this embodiment, the token corresponding to the highest probability within the probability vector can be used as the predicted token. The token corresponding to the image block can be understood as a visual vocabulary, and the number of probability values included in the probability vector can be understood as the number of a predetermined visual vocabulary.
[0052] In one embodiment, the driving data encoding network 313 can adopt, for example, an embedding layer structure. In this embodiment, instead of the first convolutional network 312, an arbitrary non-sequential network can also be used. Among them, the first convolutional network 312 can adopt network architectures such as VGG, ResNet, MobileNet, etc., and can adopt, for example, a lightweight convolutional network architecture, etc., and the present disclosure is not limited thereto.
[0053] Among these, the function of this first convolutional network 312 is to map image data to image features. For example, the image features of the original image with a spatial resolution of 1 / 2 m can be flattened to obtain a first feature vector.
[0054] In one embodiment, the driving data includes at least two data parts, such as the driving parameters of the vehicle and the navigation data of the vehicle. In this embodiment, different networks can be used to encode different parts of the driving data, and the representation ability of the encoded driving features can be improved. This is because using different encoding principles for different types of data may improve the accuracy of encoding.
[0055] For example, the navigation data may be a navigation map, that is, the navigation data may be data of the image modality. In this embodiment, a convolutional network can be used to encode the navigation data to obtain a second feature vector representing the navigation data. For example, the driving parameters may include at least one of parameters such as the speed of the vehicle, scalar acceleration, and rotation angle. In this embodiment, a fully connected layer or a multi-layer perceptron can be used to encode the driving parameters, thereby obtaining a third feature vector representing the driving parameters. Correspondingly, the above driving features include the second feature vector and the third feature vector of this embodiment.
[0056] FIG. 4 is a schematic diagram of the principle of encoding navigation data according to an embodiment of the present disclosure.
[0057] As shown in FIG. 4, in one embodiment 400, the navigation data can include the positions of at least two target points on the navigation path. This position may be represented, for example, by two-dimensional coordinates representing the position of the target point on the road plane where the vehicle is located. For example, these at least two target points may be two target points corresponding to the current position of the vehicle, and the other one may be the next passing point of the vehicle on the navigation path. These at least two target points may be plural to represent the entire navigation path or a sub-path on the navigation path close to the vehicle.
[0058] That is, in embodiment 400, only the positions of at least two target points on the navigation path can be extracted from the navigation map 401 as the navigation data 402. Thereby, interference with the prediction task of other information other than the navigation path in the navigation map 401 can be reduced, which is useful for improving the prediction accuracy of the control information.
[0059] For example, after obtaining the navigation data 402 in embodiment 400, a mask image 403 representing a path composed of at least two target points can be generated based on the positions of the at least two target points. By converting to a mask image, it is convenient to extract the features of the navigation path using a convolutional network. After the mask image 403 is obtained, the convolutional network 410 can be used to encode this mask image 403 to obtain a second feature vector 404 representing the navigation data.
[0060] Among them, examples of the convolutional network 410 for encoding the mask image 403 include, but are not limited to, ResNet-based networks, VGG-based networks, etc. In this embodiment, the mask image 403 is input into the convolutional network 410, and the convolutional network outputs image features in the form of a tensor, and the second feature vector 404 can be obtained by expanding the image features.
[0061] FIG. 5 is a schematic diagram of the principle of encoding image data according to an embodiment of the present disclosure to obtain an image token sequence.
[0062] In one embodiment, in the process of encoding image data to obtain an image token sequence, for example, quantization processing for discretizing the obtained image token sequence is performed, thereby facilitating the calculation of probabilities in the prediction task.
[0063] For example, in Embodiment 500, first, the encoder Encoder510 is used to encode the image data 501, and by mapping the pixel values of the image data to a more efficient representation space, the purpose of compressing information can be achieved, and the calculation efficiency of the generation model can be improved. By encoding with the encoder 510, an encoded feature sequence 502 is obtained. Subsequently, the quantizer 520 is used to perform quantization processing on the encoded feature sequence 502, map the continuous representation of the encoded feature sequence to integer values, and make a discrete representation. Through the quantization processing, an image token sequence 503 is obtained.
[0064]
Number
[0065]
Number
[0066] In one embodiment, the step of encoding image data to obtain an image token sequence can be realized by using vector quantization (VQ), which is a data compression technique. For example, an encoder in a pre-trained vector quantized generative adversarial network (VQGAN) or a vector quantized variational autoencoder (VQVAE) can be used to encode the image data and map the image data to a discretized vector representation in the VQ space. However, the present disclosure is not limited thereto.
[0067]
Number
[0068]
Number
[0069] FIG. 6 is a schematic diagram of the implementation principle of an information prediction method according to another embodiment of the present disclosure.
[0070] According to an embodiment of the present disclosure, when using a generative model to execute a prediction task, marks can be added to the sequence input to the generative model to prompt the generative model to execute an auxiliary task and a main task.
[0071] For example, when the auxiliary task is a task of regenerating the input image token sequence, that is, when the input image token sequence represents the image data at time t, the auxiliary task may be a task of predicting a token sequence representing the image data at time t. Then, a start marker is added to the head position of the image token sequence obtained by the above embodiment, whereby the generation model predicts the first token of the token sequence representing the image data at time t based on this start marker. The start marker is, for example, <c>Such as predefined markers, but the present disclosure is not limited thereto.
[0072] For example, a query marker can be added to the end position of the image token sequence to notify the generation model that the image token sequence has been fully input and start the main task of predicting control information for the vehicle. The query marker can be, for example, Although they are predefined markers such as, etc., the present disclosure is not limited thereto.
[0073] In one embodiment, markers can be added to both the start position and the end position of the image token sequence to obtain a marked token sequence. Then, for example, the driving characteristics obtained in the above embodiment can be stitched with the marked token sequence to obtain the input sequence of the generation model. Then, the input sequence is input into the generation model, and the generation model can generate a predicted token sequence (i.e., a token sequence representing the image data at time t) and control information.
[0074] In one embodiment, when performing the auxiliary task and the main task, for example, the first feature vector representing the image features obtained in the above embodiment may also be considered simultaneously.
[0075] In one embodiment, the driving data may include the aforementioned driving parameters and navigation data. The second convolutional network can be used to encode the navigation data, and the MLP can be used to encode the driving parameters.
[0076] For example, as shown in FIG. 6, in Embodiment 600, the acquired perception data includes image data 601 and driving data. The perception data is processed by the autonomous driving model, and a predicted token sequence and control information for the vehicle are generated.
[0077] Among these, the autonomous driving model may include an encoding layer and a generation model 620. The encoding layer may include networks for encoding the image data 601, the driving parameters 602, and the navigation data 603, respectively.
[0078] In one embodiment, the encoding layer includes a sequence encoding network 611 for encoding the image data 601 to obtain an image token sequence 604 corresponding to the image data 601.
[0079] In one embodiment, the encoding layer further includes a driving data encoding network, which is used to encode driving data to obtain driving features. In one embodiment, the driving data may include driving parameters 602 and navigation data 603. The driving data encoding network may include a multi-layer perceptron MLP612 and a second convolutional network 613. The multi-layer perceptron MLP612 is used to encode the driving parameters 602 to obtain a third feature vector 605. The second convolutional network 613 is used to encode the navigation data 603 to obtain a second feature vector 606.
[0080]
Number
[0081]
Number
[0082]
Number
[0083] In one embodiment, during the process of encoding the perception data, learnable position information can also be combined. This learnable position information is similar to the position markers added to the tokens when inputting into the Transformer architecture.
[0084] In one embodiment, after the encoding of the perception data is completed, for example, based on each vector and sequence obtained by encoding, an input sequence of the generation model 620 can be obtained.
[0085]
Number
[0086]
Number
[0087] According to the embodiments of the present disclosure, with the help of an autonomous driving model, end-to-end prediction of control information can be realized. In addition, since the image generation task is an auxiliary task that guides the execution of the main task, it is not necessary to rely on complex perception tasks and expensive perception annotations, and a low-cost and extensible solution can be provided for end-to-end autonomous driving technology.
[0088] To facilitate the implementation of the information prediction method of the embodiments of the present disclosure, the present disclosure also provides a training method for an autonomous driving model, which will be described in detail below with reference to FIG. 7.
[0089] FIG. 7 is a schematic diagram of a flowchart of a method for training an autonomous driving model according to an embodiment of the present disclosure.
[0090] As shown in FIG. 7, the training method 700 of the autonomous driving model of this embodiment can include operations S710 to S740. Here, the autonomous driving model can include an encoding layer and a generation model. Here, the encoding layer can include a sequence encoding network and a driving data encoding network.
[0091] In operation S710, a sequence encoding network is used to encode the image data in the sample perception data to obtain an image token sequence corresponding to the image data.
[0092] According to the embodiments of the present disclosure, the sample perception data includes image data collected at past times and the driving data of the vehicle at the time when the image data was collected.
[0093] The sequence encoding network includes, for example, a structure that performs block processing on an image and a two-dimensional convolutional layer, and encodes the image data in the sample perceptual data using the same principle as the principle described in the above operation S220 to obtain an image token sequence. In one embodiment, the sequence encoding network can encode the image data using vector quantization compression technology.
[0094] In operation S720, the driving data encoding network is used to encode the driving data in the sample perceptual data to obtain driving features corresponding to the driving data.
[0095] According to an embodiment of the present disclosure, the implementation principle of this operation S720 is the same as the implementation principle of the above operation S220, and will not be repeated here. The driving data encoding network may adopt the structure of an embedding layer, or may include an MLP and a convolutional network as shown in the above embodiment 600.
[0096] In operation S730, based on the driving features and the image token sequence, a generation model is used to generate a predicted token sequence corresponding to the image token sequence and predicted control information for the vehicle.
[0097] According to an embodiment of the present disclosure, the implementation principle of operation S730 is the same as the implementation principle of the above operation S230. When the generation model is an autoregressive model, the generation model generates a series of probability vectors that correspond one-to-one to the tokens in the predicted token sequence. In this embodiment, the token corresponding to the maximum probability value in the probability vector is used as the token in the predicted token sequence corresponding to this probability vector.
[0098] In operation S740, an automatic driving model is trained based on the predicted token sequence and the image token sequence.
[0099] For example, in this embodiment, based on the difference between the predicted token sequence and the image token sequence, the loss value of the automatic driving model can be determined. Next, the automatic driving model is trained with the aim of minimizing the loss value.
[0100] In one embodiment, when the generation model generates a series of probability vectors, this embodiment can determine the loss value of the automatic driving model based on the probability value corresponding to the i-th image token in the image token sequence within the i-th probability vector. For example, this loss value can be calculated using a classification loss function.
[0101]
Number
[0102]
Number
[0103] According to this embodiment, since the monitoring signal is the image data itself in the sample perception data, self-supervised training of the automatic driving model can be realized without requiring accurate perception annotation data.
[0104] In one embodiment, when training the automatic driving model, for example, by not adjusting the network parameters included in the sequence encoding network in the automatic driving model, the stability of the image token sequence functioning as the monitoring signal can be easily improved, and the training accuracy and training efficiency of the automatic driving model can be improved.
[0105] In one embodiment, the sample perception data may include perception data at several times before time t in addition to the perception data at time t. For the perception data at each time, an image token sequence and driving features can be obtained in the same way. When the generation model predicts the predicted token sequence and the control information at time t+1, the prediction accuracy can be improved by considering the perception data at several times before time t at the same time.
[0106] For example, when the sample perception data includes the perception data from the first time to the T-th time at past times, first, in the process of training the automatic driving model, based on the perception data at the first time, a predicted token sequence corresponding to the image data at the first time is predicted, and one loss value can be obtained based on this predicted token sequence. Subsequently, based on the perception data at the first time and the perception data at the second time, a predicted token sequence corresponding to the image data at the second time is predicted, and one loss value can be obtained based on the predicted token sequence. Similarly, T loss values can also be obtained. In this embodiment, the sum of the T loss values is used as the loss value for the automatic driving model to generate the token sequence. The automatic driving model is trained with the goal of minimizing the loss value.
[0107]
Number
[0108] For example, when the sample perception data includes the perception data from the first time to the T-th time at past times, in the process of training the automatic driving model, first, the control information at the second time can be predicted based on the perception data at the first time. Based on the predicted control information and the actual control information at the second time, a loss value for generating the control information can be obtained. Subsequently, based on the perception data at the first time and the perception data at the second time, the control information at the third time can be predicted, and based on the predicted control information and the actual control information at the third time, a loss value for generating the control information can be obtained. Similarly, T loss values can also be obtained. In this embodiment, the sum of the T loss values is used as the loss value for the automatic driving model to generate the control information. The automatic driving model is trained with the goal of minimizing the loss value.
[0109]
Number
[0110] In one embodiment, the aforementioned encoding layer may further include a first convolutional network. This first convolutional network is similar to the first convolutional network of Embodiment 600 described above. In this embodiment, the first convolutional network can be used to extract the image features of the image data and obtain a first feature vector representing the image features. Next, using the generation model, a predicted token sequence and predicted control information are generated based on the driving features, the first feature vector, and the image token sequence.
[0111] In one embodiment, the driving data includes the historical driving parameters of the vehicle and the historical navigation data of the vehicle. The driving data encoding network may include the second convolutional network and the multi-layer perceptron in the aforementioned embodiment 600. When using the second convolutional network to encode the historical navigation data, a second feature vector representing the navigation data can be obtained. When using the multi-layer perceptron to encode the historical driving parameters, a third feature vector representing the driving parameters can be obtained. The driving features include this second feature vector and the third feature vector.
[0112] In one embodiment, the sequence encoding network includes an encoder and a quantizer. The sequence encoding network can encode the image data using the same principle as described in the aforementioned embodiment 500, thereby obtaining an image token sequence.
[0113] In one embodiment, since the principle of generating the predicted token sequence and the predicted control information by the generation model is the same as the principle described in the aforementioned embodiment 600, it will not be repeated here.
[0114] In one embodiment, the method of training the autonomous driving model can train the autonomous driving model through two tasks. One of them is the autoregressive generation task of the next token (i.e., the task of generating the predicted token sequence), and the other is the action prediction of the planner (e.g., brakes, accelerator, steering wheel, etc.) (i.e., the task of predicting the control information). Except for the different output heads, these two tasks share most of the network parameters. In this way, when optimizing the shared parameters of the generation task, the model can appropriately learn the dependencies between the input tokens, thereby establishing an appropriate environmental representation and facilitating the implementation of the action prediction task.
[0115] Based on the information prediction method provided by the present disclosure, the present disclosure also provides an information prediction device. This device will be described in detail below with reference to FIG. 8.
[0116] FIG. 8 is a structural block diagram of an information prediction device according to an embodiment of the present disclosure.
[0117] As shown in FIG. 8, the information prediction device 800 of this embodiment can include a data acquisition module 810, a first encoding module 820, a second encoding module 830, and a first generation module 840.
[0118] The data acquisition module 810 is used to acquire perception data including image data collected by sensors in the vehicle and driving data of the vehicle. In one embodiment, the data acquisition module 810 can be used to execute the above-mentioned operation S210, but will not be described in detail here.
[0119] The first encoding module 820 is used to encode the image data to obtain an image token sequence corresponding to the image data. In one embodiment, the first encoding module 820 may be used to execute the above-mentioned operation S220, but will not be described in detail here.
[0120] The second encoding module 830 is used to encode the driving data to obtain driving features corresponding to the driving data. In one embodiment, the second encoding module 830 may be used to execute the above-mentioned operation S230, but will not be described in detail here.
[0121] The first generation module 840 is used to use a generation model based on the driving features and the image token sequence to generate a prediction token sequence corresponding to the image token sequence and control information for the vehicle. In one embodiment, the first generation module 840 may be used to execute the above-mentioned operation S240, but will not be described in detail here.
[0122] According to an embodiment of the present disclosure, an image token sequence is a discrete feature of image data. The above-described information prediction device 800 may also include a first feature extraction module used to extract image features of the image data using a first convolutional network and obtain a first feature vector representing the image features. Specifically, the first generation module 840 may be used to generate a prediction token sequence and control information using a generation model based on driving features, the first feature vector, and the image token sequence.
[0123] According to an embodiment of the present disclosure, the above driving data includes vehicle driving parameters and vehicle navigation data. The above second encoding module 830 may include a first encoding sub-module and a second encoding sub-module. The first encoding sub-module is used to encode navigation data using a second convolutional network and obtain a second feature vector representing the navigation data. The second encoding sub-module is used to encode driving parameters using a multi-layer perceptron and obtain a third feature vector representing the driving parameters. Here, the driving features include the second feature vector and the third feature vector.
[0124] According to an embodiment of the present disclosure, the above navigation data may include the positions of at least two target points on the navigation path. The above first encoding sub-module may include an image generation unit and an image encoding unit. The image generation unit is used to generate a mask image representing a path composed of at least two target points based on the positions of the at least two target points. The image encoding unit is used to encode the mask image using a second convolutional network and obtain a second feature vector.
[0125] According to an embodiment of the present disclosure, the first encoding module 820 may include a third encoding sub-module and a first quantization sub-module. The third encoding sub-module is used to encode image data using an encoder to obtain an encoded feature sequence. The first quantization sub-module is used to perform quantization processing on the encoded feature sequence using a quantizer to obtain an image token sequence.
[0126] According to an embodiment of the present disclosure, the first generation module may include a first marker addition sub-module, a first sequence acquisition sub-module, and a first generation sub-module. The first marker addition sub-module is used to add a start marker to the start position of the image token sequence and add a query marker to the end position of the image token sequence to obtain a marked token sequence. The first sequence acquisition sub-module is used to obtain an input sequence of the generation model based on the driving feature and the marked token sequence. The first generation sub-module is used to input the input sequence into the generation model to obtain a predicted token sequence and control information generated by the generation model.
[0127] The present disclosure also provides a training apparatus for an autonomous driving model based on the method for training an autonomous driving model provided in the present disclosure. This apparatus will be described in detail below with reference to FIG. 9.
[0128] FIG. 9 is a structural block diagram of a training apparatus for an autonomous driving model according to an embodiment of the present disclosure.
[0129] As shown in FIG. 9, the training apparatus 900 for the autonomous driving model of this embodiment may include a third encoding module 910, a fourth encoding module 920, a second generation module 930, and a training module 940. Among them, the autonomous driving model includes an encoding layer and a generation model, and the encoding layer includes a sequence encoding network and a driving data encoding network.
[0130] The third encoding module 910 is used to encode the image data in the sample perception data using a sequence encoding network and obtain an image token sequence corresponding to the image data. In one embodiment, the third encoding module 910 may be used to execute the above-described operation S710, which will not be described in detail here.
[0131] The fourth encoding module 920 is used to encode the driving data in the sample perception data using a driving data encoding network and obtain driving features corresponding to the driving data. In one embodiment, the fourth encoding module 920 may be used to execute the above-described operation S720, which will not be described in detail here.
[0132] The second generation module 930 is used to generate a predicted token sequence corresponding to the image token sequence and predicted control information for the vehicle using a generation model based on the driving features and the image token sequence. In one embodiment, the second generation module 930 may be used to execute the above-described operation S730, which will not be described in detail here.
[0133] The training module 940 is used to train an autonomous driving model according to the predicted token sequence and the image token sequence. In one embodiment, the training module 940 may be used to execute the above-described operation S740, which will not be described in detail here.
[0134] According to an embodiment of the present disclosure, the sample perception data also includes actual control information. The training module 940 can also be used to train the autonomous driving model according to the difference between the actual control information and the predicted control information.
[0135] According to an embodiment of the present disclosure, the training module 940 is specifically used to train other model structures in the autonomous driving model except for the sequence encoding network.
[0136] According to an embodiment of the present disclosure, the encoding layer also includes a first convolutional network. The training device 900 of the autonomous driving model may also include a second feature extraction module used to extract image features of image data using the first convolutional network and obtain a first feature vector representing the image features. Specifically, the second generation module 930 is used to generate a predicted token sequence and predicted control information using a generation model based on driving features, the first feature vector, and an image token sequence.
[0137] According to an embodiment of the present disclosure, the driving data includes vehicle historical driving parameters and vehicle historical navigation data, and the driving data encoding network includes a second convolutional network and a multi-layer perceptron. The fourth encoding module 920 may include a fourth encoding sub-module and a fifth encoding sub-module. The fourth encoding sub-module is used to encode historical navigation data using the second convolutional network and obtain a second feature vector representing the navigation data. The fifth encoding sub-module is used to encode historical driving parameters using a multi-layer perceptron and obtain a third feature vector representing the driving parameters. The driving features include the second feature vector and the third feature vector.
[0138] According to an embodiment of the present disclosure, the sequence encoding network includes an encoder and a quantizer. The third encoding module 910 may include a sixth encoding sub-module and a second quantization sub-module. The sixth encoding sub-module is used to encode image data using an encoder to obtain an encoded feature sequence. The second quantization sub-module is used to perform quantization processing on the encoded feature sequence using a quantizer to obtain an image token sequence. Here, the sequence encoding network processes the image data based on the vector quantization compression technology.
[0139] According to an embodiment of the present disclosure, the second generation module 930 may include a second marker addition sub-module, a second sequence acquisition sub-module, and a second generation sub-module. The second marker addition sub-module is used to add a start marker to the start position of the image token sequence and add a query marker to the end position of the image token sequence to obtain a marked token sequence. The second sequence acquisition sub-module is used to obtain an input sequence of the generation model based on the driving feature and the marked token sequence. The second generation sub-module is used to input the input sequence into the generation model to obtain a predicted token sequence and predicted control information generated by the generation model. Here, the generation model includes an autoregressive model.
[0140] It should be noted that in the technical solution of the present disclosure, all processes such as the collection, storage, use, processing, transmission, provision, disclosure, and application of the user's personal information comply with the provisions of relevant laws and regulations, and necessary confidentiality measures are taken, and it does not violate public order and good customs. In the technical solution of the present disclosure, the user's permission or consent is obtained before obtaining or collecting the user's personal information.
[0141] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0142] FIG. 10 shows a schematic block diagram of an exemplary electronic device 1000 capable of implementing the information prediction method or the training method of the autonomous driving model according to an embodiment of the present disclosure. The electronic device represents various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can further represent various forms of mobile devices, such as personal digital processes, mobile phones, smartphones, wearable devices, and other similar computing devices. The members, their connections and relationships, and their functions shown in this specification are merely exemplary and do not limit the implementation of the present disclosure described and / or claimed in this specification.
[0143] As shown in FIG. 10, the device 1000 includes a computing unit 1001 and can execute various appropriate operations and processes based on a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. The RAM 1003 can further store various programs and data necessary for the operation of the device 1000. The computing unit 1001, the ROM 1002, and the RAM 1003 are interconnected by a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0144] A plurality of components in the device 1000 are connected to the I / O interface 1005 and include an input unit 1006 such as a keyboard and a mouse, an output unit 1007 such as various types of displays and speakers, a storage unit 1008 such as a magnetic disk and an optical disk, and a communication unit 1009 such as a network card, a modem, and a wireless communication transceiver. The communication unit 1009 allows the device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0145] The computing unit 1001 can be various general-purpose and / or dedicated processing modules with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, computing units for various running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes each of the methods and processes described above, such as an information prediction method or a training method for an autonomous driving model. For example, in some embodiments, the information prediction method or the training method for the autonomous driving model is implemented as a computer software program and is tangibly included in a machine-readable medium such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed into the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the information prediction method or the training method for the autonomous driving model described above can be executed. Alternatively, in other embodiments, the computing unit 1001 may be arranged to execute the information prediction method or the training method for the autonomous driving model in any other suitable manner (e.g., firmware).
[0146] The various embodiments of the systems and techniques described in this specification can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can be implemented in one or more computer programs executed and / or interpreted in a programmable system including at least one programmable processor, the programmable processor may be a dedicated or general-purpose programmable processor, and receives data and instructions from a memory system, at least one input device, and at least one output device, and can transmit the data and instructions to the memory system, the at least one input device, and the at least one output device, which can be included.
[0147] The program code for implementing the methods of the present disclosure can be created in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a dedicated computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations defined in the flowchart and / or block diagram are implemented. The program code may be executed entirely by a machine, partially by a machine, partially by a device as an independent software package and partially by a remote machine, or entirely by a remote machine or server.
[0148] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be either a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium includes, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium include, but are not limited to, one or more electrical connections based on one or more lines, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0149] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well. For example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input received from the user can be in any form (including voice input, speech input, or tactile input).
[0150] The systems and techniques described herein can be implemented in a computing system that includes background components (e.g., a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser, through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such background components, middleware components, or front-end components. Components of the system can be connected to each other by digital data communication in any form or medium (e.g., a communication network). Exemplifications of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0151] A computer system can include clients and servers. Clients and servers are generally separated from each other and usually interact via a communication network. A client-server relationship is created by computer programs that run on corresponding computers and have a client-server relationship with each other. A server can be a cloud server, also called a cloud computing server or cloud host, which is a host product within a cloud computing service system, so as to solve the drawbacks such as the difficulty in management and the weak business scalability in a conventional physical host and a VPS service (Virtual Private Server, abbreviated as "VPS"). The server can be a server of a distributed system or a server combined with a blockchain.
[0152] Based on the electronic device according to the information prediction method provided in the present disclosure, the present disclosure also provides an autonomous vehicle including an electronic device for implementing the information prediction method.
[0153] In one embodiment, the autonomous vehicle also includes a sensor for collecting image data, and the electronic device can predict the control information for the next moment based on the image data and the driving data of the autonomous vehicle, and control the driving of the autonomous vehicle based on the control information.
[0154] It should be understood that various forms of flows shown above may be used, and each step may be sorted again, added, or deleted. For example, each step described in the present disclosure may be executed in parallel, sequentially, or in a different order, and the present specification is not limited here as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved.
[0155] The above specific embodiments do not limit the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure should all be included within the protection scope of the present disclosure. < / c>
Claims
1. acquiring sensory data including image data collected by sensors within a vehicle and driving data of the vehicle; encoding the image data to obtain a sequence of image tokens corresponding to the image data; encoding the driving data to obtain driving characteristics corresponding to the driving data; and generating, based on the driving characteristics and the image token sequence, a generative model for generating a predicted token sequence corresponding to the image token sequence and control information for the vehicle. Information prediction methods.
2. the sequence of image tokens are discrete features of the image data; The method comprises: extracting image features of the image data using a first convolutional network to obtain a first feature vector representing the image features; generating a predicted token sequence corresponding to the image token sequence and control information for the vehicle using a generative model based on the driving features and the image token sequence; generating the predicted token sequence and the control information using the generative model based on the driving characteristics, the first feature vector, and the image token sequence. The method of claim 1.
3. The driving data includes driving parameters of the vehicle and navigation data of the vehicle; Encoding the driving data to obtain driving characteristics corresponding to the driving data includes: encoding the navigation data using a second convolutional network to obtain a second feature vector representative of the navigation data; encoding the driving parameters using a multi-layer perceptron to obtain a third feature vector representative of the driving parameters; Here, the driving features include the second feature vector and a third feature vector. The method of claim 1.
4. the navigation data includes positions of at least two landmarks on a navigation path; Encoding the navigation data using the second convolutional network to obtain a second feature vector representative of the navigation data includes: generating a mask image representing a path defined by the at least two target points based on the positions of the at least two target points; and encoding the mask image using the second convolutional network to obtain the second feature vector. The method according to claim 3.
5. Encoding the image data to obtain a sequence of image tokens corresponding to the image data includes: encoding the image data using an encoder to obtain a sequence of encoded features; quantizing the encoded feature sequence using a quantizer to obtain the image token sequence. The method of claim 1.
6. generating a predicted token sequence corresponding to the image token sequence and control information for the vehicle using a generative model based on the driving features and the image token sequence; adding a start marker to the beginning of the image token sequence and adding a query marker to the end of the image token sequence to obtain a marked token sequence; obtaining an input sequence for the generative model based on the driving features and the marked token sequence; inputting the input sequence into the generative model to obtain the predicted token sequence and the control information generated by the generative model. The method of claim 1.
7. A method for training an autonomous driving model, comprising: The autonomous driving model includes an encoding layer and a generative model; The coding layer includes a sequence coding network and a driving data coding network; The method comprises: encoding image data of the sample sensory data using the sequence encoding network to obtain a sequence of image tokens corresponding to the image data; encoding driving data of the sample sensory data using the driving data encoding network to obtain driving characteristics corresponding to the driving data; generating a predicted token sequence corresponding to the image token sequence and predicted control information for a vehicle using the generative model based on the driving characteristics and the image token sequence; training the autonomous driving model in response to the predicted token sequence and the image token sequence. How to train an autonomous driving model.
8. The sample sensory data further includes actual control information; The method and training the autonomous driving model according to a difference between the actual control information and the predicted control information. The method according to claim 7.
9. Training the autonomous driving model includes: Training a model structure other than the sequence coding network among the autonomous driving models. The method according to claim 7.
10. the encoding layer further includes a first convolutional network; The method comprises: extracting image features of the image data using the first convolutional network to obtain a first feature vector representing the image features; generating a predicted token sequence corresponding to the image token sequence and predicted control information for a vehicle using the generative model based on the driving features and the image token sequence, generating the predicted token sequence and the predicted control information using the generative model based on the driving characteristics, the first feature vector, and the image token sequence. The method according to claim 7.
11. the driving data includes historical driving parameters of a vehicle and historical navigation data of the vehicle, and the driving data encoding network includes a second convolutional network and a multi-layer perceptron; encoding driving data of the sample sensory data using the driving data encoding network to obtain driving characteristics corresponding to the driving data, encoding the historical navigation data using the second convolutional network to obtain a second feature vector representative of the navigation data; and encoding the historical driving parameters using the multi-layer perceptron to obtain a third feature vector representative of the driving parameters; Here, the driving features include the second feature vector and the third feature vector. The method according to claim 7.
12. the sequence coding network includes an encoder and a quantizer; encoding image data of sample sensory data using the sequence encoding network to obtain a sequence of image tokens corresponding to the image data; encoding the image data using the encoder to obtain a sequence of encoded features; quantizing the encoded feature sequence using the quantizer to obtain the image token sequence; Wherein the sequence coding network processes the image data based on a vector quantization compression technique. The method according to claim 7.
13. generating a predicted token sequence corresponding to the image token sequence and predicted control information for a vehicle using the generative model based on the driving features and the image token sequence, adding a start marker to the beginning of the image token sequence and adding a query marker to the end of the image token sequence to obtain a marked token sequence; obtaining an input sequence for the generative model based on the driving features and the marked token sequence; inputting the input sequence into the generative model to obtain the predicted token sequence and the predicted control information generated by the generative model; Here, the generative model includes an autoregressive model. The method according to claim 7.
14. a data acquisition module that acquires image data collected by sensors in a vehicle and sensory data including driving data of the vehicle; a first encoding module for encoding the image data to obtain a sequence of image tokens corresponding to the image data; a second encoding module that encodes the driving data to obtain driving characteristics corresponding to the driving data; a generation module that uses a generative model to generate a predicted token sequence corresponding to the image token sequence and control information for the vehicle based on the driving characteristics and the image token sequence. Information prediction device.
15. the sequence of image tokens are discrete features of the image data; The apparatus comprises: a feature extraction module for extracting image features of the image data using a first convolutional network to obtain a first feature vector representing the image features; The generation module is adapted to generate the predicted token sequence and the control information using the generative model based on the driving characteristics, the first feature vector, and the image token sequence.
15. The apparatus of claim 14.
16. The driving data includes driving parameters of the vehicle and navigation data of the vehicle; The second encoding module comprises: a first encoding sub-module for encoding the navigation data using a second convolutional network to obtain a second feature vector representative of the navigation data; a second encoding sub-module for encoding the driving parameters using a multi-layer perceptron to obtain a third feature vector representative of the driving parameters; Here, the driving features include the second feature vector and the third feature vector.
16. Apparatus according to claim 14 or 15.
17. the navigation data includes positions of at least two landmarks on a navigation path; The first encoding sub-module: an image generation unit for generating a mask image representing a path formed by the at least two target points based on the positions of the at least two target points; an image encoding unit for encoding the mask image using the second convolutional network to obtain the second feature vector.
17. The apparatus of claim 16.
18. The first encoding module comprises: a third encoding sub-module for encoding the image data using an encoder to obtain an encoded feature sequence; a quantization submodule for quantizing the encoded feature sequence using a quantizer to obtain the image token sequence.
15. The apparatus of claim 14.
19. The generation module includes: a marker addition submodule for adding a start marker to the beginning of the image token sequence and adding a query marker to the end of the image token sequence to obtain a marked token sequence; a sequence acquisition submodule for acquiring an input sequence for the generative model based on the driving features and the marked token sequence; a generation submodule for inputting the input sequence into the generative model to obtain the predicted token sequence and the control information generated by the generative model.
15. The apparatus of claim 14.
20. A training device for an autonomous driving model, The autonomous driving model includes an encoding layer and a generative model; The coding layer includes a sequence coding network and a driving data coding network; The apparatus comprises: a first encoding module for encoding image data of the sample sensory data using the sequence encoding network to obtain an image token sequence corresponding to the image data; a second encoding module for encoding driving data of the sample sensory data using the driving data encoding network to obtain driving characteristics corresponding to the driving data; a generation module for generating, based on the driving features and the image token sequence, a predicted token sequence corresponding to the image token sequence and predicted control information for a vehicle using the generative model; a training module for training the autonomous driving model in response to the predicted token sequence and the image token sequence. Training equipment for autonomous driving models.
21. The sample sensory data further includes actual control information; The training module The method is further used for training the autonomous driving model according to a difference between the actual control information and the predicted control information.
21. The apparatus of claim 20.
22. The training module includes: The autonomous driving model is used to train a model structure other than the sequence coding network.
22. Apparatus according to claim 20 or 21.
23. the encoding layer further includes a first convolutional network; The apparatus comprises: a feature extraction module that extracts image features of the image data using the first convolutional network to obtain a first feature vector representing the image features; The generation module includes: and generating the predicted token sequence and the predicted control information using the generative model based on the driving features, the first feature vector, and the image token sequence.
21. The apparatus of claim 20.
24. the driving data includes historical driving parameters of a vehicle and historical navigation data of the vehicle, and the driving data encoding network includes a second convolutional network and a multi-layer perceptron; The second encoding module comprises: a first encoding sub-module for encoding the historical navigation data using the second convolutional network to obtain a second feature vector representative of the navigation data; a second encoding sub-module for encoding the historical driving parameters using the multi-layer perceptron to obtain a third feature vector representative of the driving parameters; Here, the driving features include the second feature vector and the third feature vector.
22. Apparatus according to claim 20 or 21.
25. the sequence coding network includes an encoder and a quantizer; The first encoding module comprises: a third encoding sub-module for encoding the image data using the encoder to obtain a sequence of encoded features; a quantization sub-module for quantizing the encoded feature sequence using the quantizer to obtain the image token sequence; Wherein the sequence coding network processes the image data based on a vector quantization compression technique.
21. The apparatus of claim 20.
26. The generation module includes: a marker addition submodule for adding a start marker to the beginning of the image token sequence and adding a query marker to the end of the image token sequence to obtain a marked token sequence; a sequence acquisition submodule that acquires an input sequence for the generative model based on the driving features and the marked token sequence; a generation submodule for inputting the input sequence into the generative model to obtain the predicted token sequence and the predicted control information generated by the generative model; Here, the generative model includes an autoregressive model.
21. The apparatus of claim 20.
27. At least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor performs the method of any one of claims 1 to 6. electronic equipment.
28. At least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor performs a method according to any one of claims 7 to 13. electronic equipment.
29. A non-transitory computer-readable storage medium having computer instructions stored thereon, The computer instructions cause a computer to carry out a method according to any one of claims 1 to 13. A non-transitory computer-readable storage medium.
30. A computer program product comprising computer program instructions which, when stored on a readable storage medium and / or an electronic device, implements the method according to any one of claims 1 to 13 when executed by a processor.
31. An autonomous vehicle comprising the electronic device according to claim 27.
Citation Information
Patent Citations
Systems and methods for providing future object localization
US20210082283A1
Training a codebook for trajectory determination
US20240104934A1