Automatic driving track decoding method based on single token large model and related device
By using a large language model based on single token and multi-layer perception model for trajectory decoding in autonomous driving technology, the problem of too long decoding time in the existing technology is solved, and more efficient real-time and safety of autonomous driving is achieved.
Patent Information
- Application Number
- CN202411980374.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-16
AI Technical Summary
In existing autonomous driving technology, the decoding time is too long, which affects real-time and safety.
The large language model based on a single token is used to decode the autonomous driving trajectory. Multiple coded tokens encoded by multiple different encoders are input into the large language model, and a decoded token is output, and the trained multi-layer perception model is decoded to output the planned trajectory.
The model structure is simplified, the decoding time is shortened, and the real-time requirements and safety of autonomous driving are ensured.
Smart Images

Figure CN120012917A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving technology, specifically to technical fields such as end-to-end and multimodal large models, and in particular to an autonomous driving trajectory decoding method and related devices based on a single-token large model. Background Art
[0002] With the rapid development of autonomous driving technology, driving decisions are crucial, and the key to driving decisions is accurate prediction of trajectories. Since the output of large language models is generally represented by language tokens, researchers prefer to map the output of planned trajectories to language. For example, if the trajectory for the next 5 seconds is<X_t1,Y_t1,X_t2,Y_t2,...,X_t5,Y_t5> This 10-dimensional vector expression will be decoded into the tokens corresponding to these 10 vectors. These tokens will also be further optimized by a trajectory decoder to obtain the final trajectory output.
[0003] Since the output of the large language model is autoregressive, that is, each token must wait for the previous token to be output before it can be inferred by the entire network, and the more tokens are output, the longer the inference time is, which affects the real-time performance of autonomous driving, especially for time-sensitive tasks. Summary of the invention
[0004] The present application provides an autonomous driving trajectory decoding method and related devices based on a single-token large model to solve the problem in the prior art that the decoding time is too long, affecting the real-time requirements and safety of autonomous driving.
[0005] The technical solution is as follows:
[0006] In the first aspect, a method for decoding an autonomous driving trajectory based on a single token large model is provided, comprising:
[0007] Obtain multiple encoded tokens obtained by encoding through multiple different encoders;
[0008] Inputting the plurality of encoded tokens into a trained large language model, and outputting a decoded token; wherein the large language model is obtained based on repeated training of a single decoded token at the output end;
[0009] The one decoded token is input into a trajectory decoder for decoding, and a planning trajectory is output; wherein the trajectory decoder is a trained multi-layer perception model, and the trajectory decoder and the large language model are jointly trained.
[0010] In a possible implementation, the large language model is trained in the following manner:
[0011] Acquire a historical vehicle data set as a training sample, wherein the historical vehicle data set includes vehicle data of multiple modes;
[0012] Encode the corresponding vehicle data based on encoders of different modes to obtain multiple encoding tokens;
[0013] The multiple encoding tokens are input into a preset large language model for autoregressive training, and the required single decoding token is learned at the output end based on the classification loss function to obtain a trained large language model.
[0014] In a possible implementation, during the training of the large language model, the method further includes:
[0015] Receives the single decoded token required for learning output when training a large language model;
[0016] The single decoded token is used as a training sample, input into the decoder model and trained based on the regression loss function to obtain a trained trajectory decoder.
[0017] In a possible implementation, the decoding token is an output identifier of the end text, which is used to use its own feature expression to assist the large language model in understanding the current semantic environment.
[0018] In a possible implementation, the trajectory decoder is a multi-layer perceptron (MLP) model.
[0019] In the second aspect, an autonomous driving trajectory decoding device based on a single token large model is provided, comprising:
[0020] An acquisition module, used to acquire multiple encoding tokens obtained by encoding with multiple different encoders;
[0021] A prediction module, configured to input the plurality of encoded tokens into a trained large language model and output a decoded token; wherein the large language model is obtained based on repeated training of a single decoded token at the output end;
[0022] A decoding module is used to input the one decoded token into a trajectory decoder for decoding, and output a planned trajectory; wherein the trajectory decoder is a trained multi-layer perception model, and the trajectory decoder and the large language model are jointly trained.
[0023] In a third aspect, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the method of the above-mentioned aspect and any possible implementation manner.
[0024] In a fourth aspect, an electronic device is provided, including:
[0025] at least one processor; and
[0026] a memory communicatively connected to the at least one processor; wherein,
[0027] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any possible implementation manner and the aspects described above.
[0028] According to a fifth aspect, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the above-mentioned aspects and any possible implementation method.
[0029] In a sixth aspect, an autonomous driving vehicle is provided, comprising the electronic device as described above.
[0030] The beneficial effects of the technical solution provided by this application include at least:
[0031] It can be seen from the above technical solution that the embodiment of the present application can obtain multiple encoding tokens obtained by encoding through multiple different encoders; input the multiple encoding tokens into the trained large language model, and output a decoding token; wherein the large language model is obtained based on the repeated training of a single decoding token at the output end; input the one decoding token into the trajectory decoder for decoding, and output a planned trajectory; wherein the trajectory decoder is a trained multi-layer perception model, and the trajectory decoder and the large language model are jointly trained. In this way, during training, the supervised large language model only outputs one decoding token, and the MLP network is used to decode and restore the trajectory, thereby simplifying the model structure, shortening the decoding time, and ensuring the real-time requirements and safety of autonomous driving.
[0032] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0034] Figure 1 This is a schematic diagram of the steps of an autonomous driving trajectory decoding method based on a single token large model provided in an embodiment of the present application.
[0035] Figure 2a It is a schematic diagram of the steps of a large language model training method provided in another embodiment of the present application.
[0036] Figure 2b This is a schematic diagram of the steps of a training method for a large language model and a trajectory decoder provided in another embodiment of the present application.
[0037] Figure 3a This is a schematic diagram of the autonomous driving trajectory decoding architecture based on a single token large model provided in an embodiment of the present application.
[0038] Figure 3b It is a schematic diagram of the training and prediction of the autonomous driving trajectory decoding architecture based on the single token large model provided in an embodiment of the present application.
[0039] Figure 4 This is a structural block diagram of an autonomous driving trajectory decoding device based on a single token large model provided in yet another embodiment of the present application.
[0040] Figure 5 It is a block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0041] The following is a description of exemplary embodiments of the present application in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0042] Obviously, the described embodiments are only part of the embodiments of the present application, but not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.
[0043] It should be noted that the terminal devices involved in the embodiments of the present application may include but are not limited to mobile phones, personal digital assistants (PDAs), wireless handheld devices, tablet computers and other smart devices; display devices may include but are not limited to personal computers, televisions and other devices with display functions.
[0044] In addition, the term "and / or" in this article is only a description of the association relationship between the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0045] In view of the fact that the output of the existing large language model takes a long time, it affects the real-time performance of autonomous driving, and is particularly unfriendly to time-sensitive tasks. To this end, the embodiment of the present application proposes an autonomous driving trajectory decoding solution based on a single-token large model. The main inventive concept is: obtaining multiple encoding tokens obtained by encoding through multiple different encoders; inputting the multiple encoding tokens into the trained large language model, and outputting a decoding token; wherein the large language model is obtained based on repeated training of a single decoding token at the output end; inputting the one decoding token into the trajectory decoder for decoding, and outputting a planned trajectory; wherein the trajectory decoder is a trained multi-layer perception model, and the trajectory decoder and the large language model are jointly trained. In this way, during training, the supervised large language model only outputs one decoding token, and uses the MLP network to decode and restore the trajectory, thereby adopting an end-to-end simplified model structure, shortening the decoding time, and ensuring the real-time requirements and safety of autonomous driving.
[0046] Reference Figure 1 As shown, it is a schematic diagram of the steps of an autonomous driving trajectory decoding method based on a single token large model provided by an embodiment of the present application. The execution subject of the autonomous driving trajectory decoding method can be an autonomous driving trajectory decoding device based on a single token large model, wherein the device can be a hardware device or software module with computing, storage, data processing and other functions, such as a terminal device such as a computer, a tablet computer, a smart phone, a smart wearable device, or a functional module or component integrated or installed in the terminal device, and the present application does not limit this.
[0047] like Figure 1 As shown, the autonomous driving trajectory decoding method based on the single token large model may include the following steps:
[0048] Step 102: Obtain multiple encoded tokens obtained by encoding through multiple different encoders.
[0049] In the present application, the multiple different encoders may be encoding modules for multiple different input information, for example, there may be a vehicle information encoder, an obstacle information encoder, a CNN network encoder, etc. The specific encoding object type is not limited here, for example, it may be input information such as images, videos, laser radars, historical trajectories, etc. The encoding results of each different encoder, i.e., multiple encoding tokens, are obtained respectively.
[0050] Step 104: input the plurality of encoded tokens into a trained large language model, and output a decoded token; wherein the large language model is obtained based on repeated training of a single decoded token at the output end.
[0051] Optionally, refer to Figure 2a As shown, the large language model can be trained in the following ways:
[0052] Step 202: Acquire a historical vehicle data set as a training sample, wherein the historical vehicle data set includes vehicle data of multiple modalities.
[0053] The vehicle data of multiple modalities may include input information such as images, videos, lidar, historical trajectories, etc.
[0054] Step 204: Encode the corresponding vehicle data based on encoders of different modes to obtain multiple encoded tokens.
[0055] Step 206: Input the plurality of encoded tokens into a preset large language model for autoregressive training, and learn a required single decoded token based on a classification loss function at the output end.
[0056] Thus, by repeatedly executing steps 202 to 206 for iterative training, a trained large language model is obtained.
[0057] When step 206 is implemented specifically, the large language model can be subjected to autoregressive training, including pre-training, instruction tuning, or low-rank tuning, and the required single token, such as or |end| token, can be learned at the output end through the classification loss function and back propagation.
[0058] In the present application, the large language model uses any available open source model, such as InternLM, LLaMa, Qwen, etc., with a parameter size ranging from 4 billion to 13 billion. The training time is generally selected until the loss function converges.
[0059] Step 106: Input the decoded token into a trajectory decoder for decoding, and output a planned trajectory; wherein the trajectory decoder is a trained multi-layer perception model, and the trajectory decoder and the large language model are jointly trained.
[0060] Further, refer to Figure 2b As shown, based on Figure 2a The training method shown in the figure, in the process of training the large language model, the method also includes the following steps:
[0061] Step 208: Receive the required single decoded token outputted by learning when training the large language model.
[0062] Step 210: Use the single decoded token as a training sample, input it into the decoder model and perform training based on the regression loss function.
[0063] After the single token is output successfully, the MLP network decoding trajectory is added after the single token, and the planning trajectory is obtained through training and learning through the L2 loss function.
[0064] As above Figure 2a and 2b As shown, steps 202 to 210 are repeatedly iterated and jointly multi-task trained through a large amount of data, so that a large language model and a trajectory decoder are trained simultaneously.
[0065] Optionally, the trajectory decoder is a multi-layer perceptron MLP model.
[0066] Optionally, the one decoding token is an output identifier of the end text, which is used to use its own feature expression to assist the large language model in understanding the current semantic environment.
[0067] Reference Figure 3a As shown, it is a schematic diagram of the autonomous driving trajectory decoding architecture based on a single token large model provided in an embodiment of the present application.
[0068] Get the encoded tokens received from different encoders, Figure 3a In the example, 10 encoded tokens are input into the LLM model for decoding, as shown in the 10 boxes on the left of the LLM model. A decoded token can be obtained. As shown in the box on the right of the LLM model, a token is output. Then the token is input into the MLP model, and the desired planning trajectory is obtained.
[0069] Reference Figure 3b As shown, it is a schematic diagram of the training and prediction of the autonomous driving trajectory decoding architecture based on the single token large model provided in an embodiment of the present application.
[0070] First, a large number of encoding tokens can be collected as training samples, input into the initialization large language model for training, and the output decoding token is input into the initialization MLP model for training; by repeatedly iteratively training these two models, a trained LLM model and trajectory decoder are obtained. After that, the encoding token to be processed is sent to the LLM model, and a required decoding token is output. The decoding token is sent to the trajectory decoder, and the required planning trajectory is output.
[0071] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0072] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0073] Figure 4 The structure block diagram of an automatic driving trajectory decoding device based on a single token large model provided by an embodiment of the present application is shown as follows: Figure 4 As shown. The autonomous driving trajectory decoding device 400 based on a single-token large model of this embodiment may include an acquisition module 401, a prediction module 402 and a decoding module 403. Among them, the acquisition module 401 is used to obtain multiple encoded tokens obtained by encoding processing through multiple different encoders. The prediction module 402 is used to input the multiple encoded tokens into the trained large language model, and output a decoded token; wherein the large language model is obtained based on repeated training of a single decoded token at the output end. The decoding module 403 is used to input the one decoded token into a trajectory decoder for decoding, and output a planned trajectory; wherein the trajectory decoder is a trained multi-layer perception model, and the trajectory decoder and the large language model are jointly trained.
[0074] It should be noted that part or all of the autonomous driving trajectory decoding device based on the single-token large model of this embodiment may be an application located in the local terminal, or may also be a functional unit such as a plug-in or software development kit (SDK) set in the application located in the local terminal, or may also be a processing engine located in a network-side server, or may also be a distributed system located on the network side, for example, a processing engine or distributed system in an autonomous driving platform on the network side, etc. This embodiment does not specifically limit this.
[0075] It is understandable that the application may be a local program (nativeApp) installed on the local terminal, or may be a webpage program (webApp) of a browser on the local terminal, which is not limited in this embodiment.
[0076] Optionally, in a possible implementation of this embodiment, the autonomous driving trajectory decoding device based on a single-token large model also includes: a first training module; the large language model is trained based on the first training module in the following manner: obtaining a historical vehicle data set as a training sample, wherein the historical vehicle data set contains vehicle data of multiple modalities; encoding the corresponding vehicle data based on encoders of different modalities to obtain multiple encoded tokens; inputting the multiple encoded tokens into a preset large language model for autoregressive training, and learning the required single decoding token at the output end based on the classification loss function to obtain a trained large language model.
[0077] Optionally, in a possible implementation of this embodiment, the autonomous driving trajectory decoding device based on a single-token large model also includes: a second training module; during the training of the large language model, the second training module is also used to receive the required single decoding token output when training the large language model; the single decoding token is used as a training sample, input into the decoder model and train and learn based on the regression loss function to obtain a trained trajectory decoder.
[0078] Optionally, in a possible implementation of this embodiment, the dimension of the planned trajectory output by the trajectory decoder is the same as the number of the encoding tokens.
[0079] Optionally, in a possible implementation of this embodiment, the decoding token is an output identifier of the end text, which is used to use its own feature expression to assist the large language model in understanding the current semantic environment.
[0080] In this embodiment, multiple encoding tokens can be obtained by encoding through multiple different encoders; the multiple encoding tokens are input into the trained large language model to output a decoding token; wherein the large language model is obtained by repeated training of a single decoding token at the output end; the decoding token is input into the trajectory decoder for decoding, and the planned trajectory is output; wherein the trajectory decoder is a trained multi-layer perception model, and the trajectory decoder and the large language model are jointly trained. In this way, during training, the supervised large language model only outputs one decoding token, and the MLP network is used to decode and restore the trajectory, thereby simplifying the model structure, shortening the decoding time, and ensuring the real-time requirements and safety of autonomous driving.
[0081] An embodiment of the present application provides a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement the method for autonomous driving trajectory decoding based on a single-token large model as described above.
[0082] An embodiment of the present application provides an electronic device, comprising a processor and a memory, wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the method for autonomous driving trajectory decoding based on a single token large model as described above.
[0083] An embodiment of the present application provides an autonomous driving vehicle, comprising the electronic device as described above. Specifically, the autonomous driving vehicle may be a vehicle of level L2 or above.
[0084] In the technical solution of this application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the relevant laws and regulations and do not violate public order and good morals.
[0085] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement an embodiment of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0086] like Figure 5As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0087] Multiple components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0088] The computing unit 501 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above, such as the method for decoding the autonomous driving trajectory based on a single token large model. For example, in some embodiments, the method for decoding the autonomous driving trajectory based on a single token large model may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the method for decoding the autonomous driving trajectory based on the single token large model described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured in any other appropriate manner (e.g., by means of firmware) to perform the method for autonomous driving trajectory decoding based on a single token large model.
[0089] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0090] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, implements the functions / operations specified in the flow chart and / or block diagram. The program code can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0091] In the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0092] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0093] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0094] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0095] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps disclosed in this application can be performed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in this application can be achieved, and this document does not limit this.
[0096] The above specific implementations do not constitute a limitation on the protection scope of this application. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of this application should be included in the protection scope of this application.
Claims
1. A method for decoding autonomous driving trajectories based on a single-token large model, characterized in that: include: Obtain multiple encoded tokens obtained by encoding through multiple different encoders; Inputting the plurality of encoded tokens into a trained large language model, and outputting a decoded token; wherein the large language model is obtained based on repeated training of a single decoded token at the output end; The one decoded token is input into a trajectory decoder for decoding, and a planning trajectory is output; wherein the trajectory decoder is a trained multi-layer perception model, and the trajectory decoder and the large language model are jointly trained.
2. The method according to claim 1, characterized in that The large language model is trained in the following way: Acquire a historical vehicle data set as a training sample, wherein the historical vehicle data set includes vehicle data of multiple modes; Encode the corresponding vehicle data based on encoders of different modes to obtain multiple encoding tokens; The multiple encoding tokens are input into a preset large language model for autoregressive training, and the required single decoding token is learned at the output end based on the classification loss function to obtain a trained large language model.
3. The method according to claim 2, characterized in that In the process of training the large language model, the method further includes: Receives the single decoded token required for learning output when training a large language model; The single decoded token is used as a training sample, input into the decoder model and trained based on the regression loss function to obtain a trained trajectory decoder.
4. The method according to claim 2 or 3, characterized in that The one decoding token is an output mark of the end text, which is used to use its own feature expression to assist the large language model in understanding the current semantic environment.
5. The method according to claim 2 or 3, characterized in that: The trajectory decoder is a multi-layer perceptron MLP model.
6. An autonomous driving trajectory decoding device based on a single token large model, characterized in that: include: An acquisition module, used to acquire multiple encoding tokens obtained by encoding through multiple different encoders; A prediction module, configured to input the plurality of encoding tokens into a trained large language model and output a decoding token; wherein the large language model is obtained based on repeated training of a single decoding token at the output end; A decoding module is used to input the one decoded token into a trajectory decoder for decoding, and output a planned trajectory; wherein the trajectory decoder is a trained multi-layer perception model, and the trajectory decoder and the large language model are jointly trained.
7. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-5.
9. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 5.
10. An autonomous driving vehicle comprising the electronic device as claimed in claim 7.