Method and device for controlling vehicle, method and device for training multi-modal model and vehicle
By introducing prompt words to understand semantics in multimodal models, the problems of vehicle driving stability and deployment cost in multimodal models in autonomous driving are solved, and more stable, safe and efficient vehicle control is achieved.
Patent Information
- Application Number
- CN202510864572.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-25
AI Technical Summary
The existing multimodal large models have poor vehicle driving stability in autonomous driving, and the model deployment cost is high, making it difficult to meet the actual scenario needs.
By receiving multimodal perception data, the trained multimodal model uses the driving action plan based on the prompt word, generate driving action information for the specified period in the future, and control the vehicle's driving based on this information. The prompt word is used to control the multimodal model to semantically understand the perceptual data, improve the interpretability and comprehensibility of the model output, and reduce training difficulty and deployment cost.
It improves the stability and safety of vehicle driving control, reduces the training difficulty and deployment cost of multimodal models, and improves the generalization ability and model adaptability in diverse autonomous driving scenarios.
Smart Images

Figure CN120363948A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, specifically to technical fields such as vehicle control, deep learning, and autonomous driving. More specifically, it relates to a method for controlling a vehicle, a method for training a multimodal model, a device, a vehicle, a device, a medium, and a program product. Background Art
[0002] With the rapid development of artificial intelligence technology, the autonomous driving function of vehicles can be realized based on artificial intelligence algorithms, so as to improve the operation efficiency and travel efficiency in scenarios such as logistics transportation, mining operations, and public transportation travel through the autonomous driving function of vehicles.
[0003] In the process of implementing the inventive concept of the present invention, the inventors found that the driving stability of vehicles traveling based on the autonomous driving function is poor and it is difficult to meet the actual scenario requirements. Summary of the Invention
[0004] In view of the above problems, this application provides a method for controlling a vehicle, a method for training a multimodal model, a device, a vehicle, a device, a medium, and a program product.
[0005] According to a first aspect of this application, there is provided a method for controlling a vehicle, including: receiving multimodal perception data for vehicle driving; based on a prompt, using a trained multimodal model to perform a driving action plan for the vehicle based on the multimodal perception data to obtain driving action information for a future specified time period, where the driving action information represents the driving actions of the vehicle, and where the prompt is used to control the multimodal model to use the driving action information as a target to perform semantic understanding on the multimodal perception data; and controlling the vehicle to drive based on the driving action information.
[0006] Another aspect of this application also provides a method for training a multimodal model, including: receiving sample multimodal perception data for vehicle driving and labeled driving action information, where the labeled driving action information represents the driving actions of the vehicle in a sample specified time period; based on a sample prompt, using a multimodal model to perform a driving action plan for the vehicle based on the sample multimodal perception data to obtain sample driving action information for a sample specified time period, where the sample prompt is used to control the multimodal model to use the sample driving action information as a target to perform semantic understanding on the sample multimodal perception data; training the multimodal model based on the sample driving action information and the labeled driving action information to obtain a trained multimodal model.
[0007] Another aspect of the present application further provides a device for controlling a vehicle, including: a perception data receiving module, configured to receive multi-modal perception data for vehicle driving; a driving action information obtaining module, configured to, based on a prompt, use a trained multi-modal model to perform driving action planning on the vehicle based on the multi-modal perception data, so as to obtain driving action information for a future specified time period, where the driving action information represents the driving actions of the vehicle, and the prompt is used to control the multi-modal model to use the driving action information as a target to perform semantic understanding on the multi-modal perception data; and a control module, configured to control the vehicle to drive based on the driving action information.
[0008] Another aspect of the present application further provides a device for training a multi-modal model, including: a sample receiving module, configured to receive sample multi-modal perception data for vehicle driving and labeled driving action information, where the labeled driving action information represents the driving actions of the vehicle during a sample specified time period; a sample driving action information obtaining module, configured to, based on a sample prompt, use the multi-modal model to perform driving action planning on the vehicle based on the sample multi-modal perception data, so as to obtain sample driving action information for the sample specified time period, and the sample prompt is used to control the multi-modal model to use the sample driving action information as a target to perform semantic understanding on the sample multi-modal perception data; and a training module, configured to train the multi-modal model based on the sample driving action information and the labeled driving action information to obtain a trained multi-modal model.
[0009] Another aspect of the present application further provides a vehicle, characterized by including: a processor configured to execute the method for controlling a vehicle as described above.
[0010] Another aspect of the present application further provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, where the one or more processors execute the one or more computer programs to implement the steps of the method as described above.
[0011] Another aspect of the present application further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the method as described above are implemented.
[0012] Another aspect of the present application further provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the method as described above are implemented.
[0013] According to an embodiment of the present application, since the prompt can be used to control a trained multi-modal model to semantically understand multi-modal perception data with the driving action information as the target, the multi-modal model can, under the control of the prompt, generate driving action information representing the driving actions of the vehicle in a specified future period by sufficiently understanding the multi-modal perception data for vehicle driving. Thus, the vehicle driving can be controlled by the driving action information more accurately to improve the stability of vehicle driving control. At the same time, the output result of the multi-modal model can have a strong correlation with the driving actions for controlling vehicle driving by outputting the driving action information, which improves the interpretability and understandability of the output result of the multi-modal model. Furthermore, a multi-modal model with strong prediction accuracy can be trained based on relatively small training resources, reducing the training difficulty of the multi-modal model and further improving the generalization ability of the multi-modal model to control the vehicle in diverse autonomous driving scenarios, and reducing the model deployment cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Through the following description of the embodiments of the present application with reference to the drawings, the above content and other objects, features and advantages of the present application will become clearer. In the drawings:
[0015] Figure 1 The application scenario diagram of the method and device for controlling a vehicle according to an embodiment of the present application is shown;
[0016] Figure 2 The flowchart of the method for controlling a vehicle according to an embodiment of the present application is shown;
[0017] Figure 3A The schematic diagram of the principle of the multi-modal model according to an embodiment of the present application is shown;
[0018] Figure 3B The schematic diagram of the principle of the decoder according to an embodiment of the present application is shown;
[0019] Figure 4 The application scenario diagram of the method for controlling a vehicle according to an embodiment of the present application is shown;
[0020] Figure 5 The flowchart of the method for training a multi-modal model according to an embodiment of the present application is shown;
[0021] Figure 6 The structural block diagram of the device for controlling a vehicle according to an embodiment of the present application is shown;
[0022] Figure 7 The structural block diagram of the device for training a multi-modal model according to an embodiment of the present application is shown;
[0023] Figure 8A block diagram of an electronic device suitable for implementing a method of controlling a vehicle and a method of training a multimodal model according to an embodiment of the present application is shown. Detailed implementation manners
[0024] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present application. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a thorough understanding of the embodiments of the present application. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present application.
[0025] The terms used herein are merely for describing specific embodiments and are not intended to limit the present application. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0026] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0027] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).
[0028] In the technical solution of the present application, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0029] The inventors found that deep learning models with a large number of model parameters, such as Multimodal Large Models (MLLMs), are applied to scenarios such as autonomous driving to plan driving trajectories to control vehicle driving. For example, multimodal large models such as Vision-Language-Action (VLA) models can combine vehicle perception data by introducing natural language instructions expressed in language to perform driving path planning. However, controlling a vehicle based on multimodal large models is prone to problems such as poor vehicle driving stability and high model deployment costs.
[0030] Embodiments of the present application provide a method for controlling a vehicle, a method for training a multimodal model, an apparatus, a vehicle, a device, a medium, and a program product. The method for controlling a vehicle includes: receiving multimodal perception data for vehicle driving; based on a prompt, using a trained multimodal model to perform driving action planning on the vehicle based on the multimodal perception data to obtain driving action information for a future specified time period, where the driving action information represents the driving action of the vehicle, and the prompt is used to control the multimodal model to use the driving action information as a target to perform semantic understanding on the multimodal perception data; and controlling the vehicle to drive based on the driving action information.
[0031] According to the embodiments of the present application, since the prompt can be used to control the trained multimodal model to use the driving action information as a target to perform semantic understanding on the multimodal perception data, the multimodal model can, under the control of the prompt, generate driving action information representing the driving action of the vehicle in the future time period by relatively fully understanding the multimodal perception data for vehicle driving. Thus, the vehicle can be controlled to drive with relatively accurate driving action information to improve the stability of vehicle driving control. At the same time, the output result of the multimodal model can be made to have a strong correlation with the driving action for controlling vehicle driving by having the multimodal model output the driving action information, improving the interpretability and understandability of the output result of the multimodal model. Furthermore, a multimodal model with strong prediction accuracy can be trained based on relatively small training resources to reduce the training difficulty of the multimodal model and further improve the generalization ability of the multimodal model to control the vehicle in diverse autonomous driving scenarios, reducing the model deployment cost.
[0032] Figure 1 The application scenario diagram of the method and apparatus for controlling a vehicle according to the embodiments of the present application is shown.
[0033] As Figure 1As shown, the application scenario 100 according to this embodiment may include a first vehicle 101, a second vehicle 102, a third vehicle 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first vehicle 101, the second vehicle 102, the third vehicle 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0034] The first vehicle 101, the second vehicle 102, and the third vehicle 103 can each interact with the server 105 through the network 104 to receive or send messages, etc. The first vehicle 101, the second vehicle 102, and the third vehicle 103 can also interact with each other based on the network 104 to receive or send messages. The first vehicle 101, the second vehicle 102, and the third vehicle 103 can be driverless vehicles, or can also include vehicles driven by driver users.
[0035] The first vehicle 101, the second vehicle 102, and the third vehicle 103 can be any type of vehicle, such as a mining truck, a freight truck, a sedan, etc.
[0036] The server 105 can be a server that provides various services, such as a server that provides background management for the motion state and operation state of the user's use of the first vehicle 101, the second vehicle 102, and the third vehicle 103 (for example only). The background management server can analyze and process data such as broadcast messages sent by the vehicle received, and feedback the processing results to any one of the vehicles.
[0037] It should be noted that the method for controlling a vehicle provided in the embodiments of the present application can generally be executed by any one of the first vehicle 101, the second vehicle 102, and the third vehicle 103. The device for controlling a vehicle provided in the embodiments of the present application can also be set in any one of the first vehicle 101, the second vehicle 102, and the third vehicle 103.
[0038] Alternatively, the method for controlling a vehicle provided in the embodiments of the present application can generally also be executed by the server 105. Correspondingly, the device for controlling a vehicle provided in the embodiments of the present application can generally be set in the server 105. The method for controlling a vehicle provided in the embodiments of the present application can also be executed by a server or a server cluster that is different from the server 105 and can communicate with the first vehicle 101, the second vehicle 102, the third vehicle 103, and / or the server 105. Correspondingly, the device for controlling a vehicle provided in the embodiments of the present application can also be set in a server or a server cluster that is different from the server 105 and can communicate with the first vehicle 101, the second vehicle 102, the third vehicle 103, and / or the server 105.
[0039] It should be understood, Figure 1The numbers of vehicles, networks, and servers therein are merely illustrative. According to implementation requirements, there can be any number of vehicles, networks, and servers.
[0040] Figure 2 The flowchart of a method for controlling a vehicle according to an embodiment of the present application is shown.
[0041] As Figure 2 shown, the method for controlling a vehicle in this embodiment includes operation S210 to operation S230.
[0042] In operation S210, multi-modal perception data for vehicle driving is received.
[0043] In operation S220, based on the prompt words, using the trained multi-modal model, a driving action plan for the vehicle is made based on the multi-modal perception data, and driving action information for a future specified time period is obtained.
[0044] In operation S230, the vehicle is controlled to drive based on the driving action information.
[0045] According to an embodiment of the present application, the multi-modal perception data for vehicle driving may include perception data related to the vehicle driving environment. For example, image data representing obstacle information such as the position and speed of obstacles around the vehicle, but not limited to this. It may also include data related to road information such as lane lines, or may also include perception data of any other modality such as voice modality and text modality. The specific data information type of the multi-modal perception data in the embodiments of the present application is not limited.
[0046] In some embodiments, the multi-modal perception data includes at least two of image modality data, text modality data, and map data.
[0047] Exemplarily, the image modality data may include image modality data representing the driving environment of the vehicle. For example, roadside road environment images, in-vehicle perspective images, etc. collected for the vehicle's driving environment.
[0048] Exemplarily, the text modality data may include data related to text semantics such as voice data and text data representing the driving intention of controlling the vehicle. For example, the text modality data is the voice control information "Unload goods at unloading point A" of the driver or relevant personnel for the vehicle.
[0049] Exemplarily, the map data is related to the vehicle's driving environment. For example, the map data may be high-precision map data, standard map data, etc. related to the vehicle driving area.
[0050] In some embodiments, the multi-modal perception data may further include point cloud data representing the vehicle driving environment, such as lidar point cloud data, millimeter wave radar perception data, etc.
[0051] It should be noted that the embodiments of the present application do not limit the acquisition method of multimodal perception data. For example, it can be collected based on in-vehicle sensors such as in-vehicle cameras and in-vehicle lidar, or the vehicle can also obtain it from other terminals or servers through a communication module.
[0052] According to the embodiments of the present application, the driving action information represents the driving actions of the vehicle. The driving actions of the vehicle can be related to the driving motion state during the driving process. For example, the driving action information can represent any type of driving action such as the driving speed, driving direction, braking action, and accelerating action of the vehicle. The embodiments of the present application do not limit the specific types of driving actions represented by the driving action information, as long as they can represent the driving state of the vehicle.
[0053] According to the embodiments of the present application, the multimodal model can be a multimodal large model (Multimodal Large Models, MLLMs) for processing multimodal data. The multimodal model can be used to process cross-modal multimodal perception data to more fully understand the perception data related to the vehicle driving environment. The prompt is used to control the multimodal model to use the driving action information as the target for semantic understanding of the multimodal perception data. Thus, the prompt and the multimodal perception data can be processed by the multimodal model to control the multimodal model to use the driving action information as the regression target of the model algorithm, so as to improve the interpretability and reliability of the driving action information output by the multimodal model as the output result. Furthermore, the vehicle can be stably driven according to the driving action information to improve the driving stability and safety of the vehicle and enhance the driving efficiency of the vehicle. It should be noted that the multimodal model in the method for controlling a vehicle provided by the embodiments of the present disclosure can be a trained multimodal model, and the embodiments of the present disclosure will not elaborate herein.
[0054] In some embodiments, the driving action information includes at least one of the following: driving speed data for a specified moment; driving curvature data for a specified moment; driving heading angle data for a specified moment.
[0055] According to the embodiments of the present application, the future specified time period includes one or more specified moments. For example, the future specified time period can be a time period of 3 seconds after the current moment, and the multiple specified moments can be the 1st second, 2nd second, and 3rd second after the current moment.
[0056] According to an embodiment of the present application, the driving speed data may represent the speed value that the vehicle should have when driving to a specified moment. For example, the speed at the 1st second after the current moment should be 3 m / s. The driving heading angle data may represent the heading angle that the vehicle should have when driving to a specified moment. The driving curvature data may represent the degree of bending of the driving action of the vehicle at a specified moment. For example, the driving curvature data may be represented as a driving curvature value of 1 rad / m at the 1st second after the current moment.
[0057] It should be noted that the driving speed data, driving curvature data, driving heading angle data, etc. in the driving action information may be driving action data in a preset coordinate system. For example, they may be the driving speed value and driving heading angle value in the geodetic coordinate system. The embodiment of the present application does not limit the specific setting method of the preset coordinate system, and it can be selected based on actual needs.
[0058] In some embodiments, controlling the vehicle to drive based on the driving action information may include: determining a power control instruction and a steering control instruction based on the driving speed data and driving curvature data in the driving action information; and controlling the vehicle to drive in a future specified time period based on the power control instruction and the steering control instruction.
[0059] According to an embodiment of the present application, the power control instruction may be an instruction for controlling the degree of vehicle power output or braking control. For example, the power control instruction may include an accelerator opening instruction, a gear setting instruction, a braking degree control instruction, etc.
[0060] According to an embodiment of the present application, the steering control instruction may include an instruction for indicating the heading angle that the vehicle should have at a specified moment. For example, the steering control instruction may include a steering wheel rotation angle control instruction of the vehicle, etc.
[0061] In some embodiments, the future specified time period may include multiple specified moments, and each specified moment corresponds to the driving speed data and driving curvature data. The driving action information at multiple specified moments output by the multi-modal model may be represented as . Where M ego is a set of driving action information at multiple specified moments, m t is the driving action information at the specified moment t, and the driving action information m t includes the driving speed data and driving curvature data. For example, the driving action information may be m t =(k t , s t ), where s t represents the driving speed data at the specified moment t, and k t represents the driving curvature data at the specified moment t.
[0062] It should be noted that the driving speed data and the driving curvature data can be based on data representations of any data type, and the data type includes but is not limited to numerical values, vectors, strings, etc.
[0063] According to an embodiment of the present application, by performing driving action planning on multi-modal perception data under the control of a prompt word through a multi-modal model to output driving speed data and driving curvature data, the driving power control degree and the driving steering control degree in the vehicle driving action can be represented according to the driving speed data and the driving curvature data respectively. Thus, the power control instruction and the steering control instruction can be directly generated based on the driving speed data and the driving curvature data to control the vehicle driving, which can avoid generating the power control signal or the steering control signal by using the trajectory points output by the model to control the vehicle driving, resulting in difficulty in understanding the control method of the vehicle driving action, and avoid the problem of reduced stability and reliability caused by controlling the vehicle driving based on the trajectory points, thereby improving the driving efficiency and safety of the vehicle.
[0064] According to an embodiment of the present application, based on the prompt word, performing driving action planning on the vehicle based on the multi-modal perception data by using the trained multi-modal model may include: using the action embedding feature of the action type representing the driving action information as the prompt word, and using the multi-modal model to perform feature fusion on the multi-modal perception data to obtain an intermediate fusion feature related to the action type; and determining the driving action information for the action type according to the intermediate fusion feature.
[0065] According to an embodiment of the present application, the action embedding feature representing the action type as the prompt word can be used to control the multi-modal model to take the driving action information corresponding to the action type as the target and perform semantic understanding on the multi-modal perception data. For example, the action type can be the driving speed type and the driving curvature type, and multiple action embedding features can correspond to the driving speed type and the driving curvature type respectively.
[0066] Thus, the multi-modal model can fully fuse the multi-modal perception data representing the vehicle driving environment or the vehicle control intention under the control condition of taking the driving action information corresponding to the action type as the target, so that the intermediate fusion feature can represent the driving action semantics related to the action type on the basis of more accurately fusing the semantic attributes of the multi-modal perception data. Furthermore, the driving action information for the action type can be determined based on the intermediate fusion feature related to the action type, so that the action information corresponding to the action type can accurately adapt to the driving environment semantics and the text control semantics represented by the multi-modal perception data, thereby improving the stability of vehicle control according to the driving action information.
[0067] In some embodiments, the decoder of the multi-modal model can be used to process the intermediate fusion features corresponding to each of the multiple action types to obtain the driving action information corresponding to the multiple action types.
[0068] In some embodiments, based on a multimodal model, feature fusion is performed on multimodal perception data using action embedding features as prompt words, and the intermediate fusion features corresponding to the action type after fusion are used to determine the driving action information of the numerical type, which can avoid a large amount of computational overhead generated by outputting planned trajectory points through the multimodal model, so as to save the computational overhead and control delay generated by executing and deploying the multimodal model to control vehicle driving, realize the flexibility and adaptability of deploying the multimodal large model on the vehicle side to execute the method provided in the embodiments of the present application, and reduce problems such as poor driving stability and reduced driving reliability caused by control delay, and improve the efficiency and stability of vehicle control.
[0069] According to the embodiments of the present application, the prompt words include action embedding features corresponding to respective multiple action types, the multiple prompt words are arranged in sequence, and the arrangement positions of the multiple prompt words correspond to multiple specified moments in a future specified period.
[0070] In some embodiments, the action embedding feature can be represented as an empty placeholder (token) vector to prompt the multimodal model of the action type of the action embedding feature. The multiple action embedding features can be arranged based on the following manner to represent the driving action information related to multiple specified moments and action types. For example, the multiple action embedding features are based on the feature sequence E emb ={(e11,e21), (e12,e22),……(e1t,e2t)}, where e1t represents the action embedding feature corresponding to the action type of the driving speed type at the t-th specified moment, e2t represents the action embedding feature corresponding to the action type of the driving curvature type at the t-th specified moment, and (e1t,e2t) can represent a prompt word for prompting the multimodal model to predict the driving speed data and driving curvature data at the t-th specified moment.
[0071] In a comparative example of the present application, the input data of the multimodal model is multimodal perception data, and the output is the trajectory points at specified moments in a future specified period. The features corresponding to the multimodal perception data include visual perception features and vehicle control text features, and the visual perception features and vehicle control text features are represented as X input1 =(V env ,L query ): where X input is the input feature input to the multimodal model, V env is the visual perception feature (or image modality feature), and L query is the control text feature of the text modality.
[0072] The multimodal model passes the input X inputand the trajectory points at the (t-1)-th specified moment, and output the trajectory points at the t-th specified moment. The multi-modal model obtains the vehicle driving trajectory by smoothing the trajectory points at T specified moments in the future specified time period.
[0073] In one embodiment of the present application, the input data of the multi-modal model is an action embedding feature sequence E obtained by arranging multi-modal perception data and multiple action embedding features in sequence emb . The input data of the multi-modal model is X input2 =(V env , L query , E emb ). The multi-modal model can execute a parallel data processing process under the prompt control of the action embedding feature sequence and output the driving action information at multiple specified moments.
[0074] According to the embodiment of the present application, arranging multiple prompt words in sequence to prompt the multi-modal model to generate the driving action information at multiple specified moments can realize using the multi-modal model to parallelly output multiple sequentially continuous driving action instructions of the vehicle in the future specified time period, so as to realize using the multi-modal model to parallelly calculate multiple driving action instructions. This can avoid the redundant computational overhead generated by using the multi-modal model to iteratively generate the t-th trajectory point based on the (t-1)-th trajectory point generated in the current round and the multi-modal input data, and reduce the time delay of generating the control instructions for controlling the vehicle by parallelly generating the driving action information at multiple specified moments, thereby reducing the computational overhead of the computing device for controlling the vehicle while improving the vehicle control efficiency, avoiding affecting the driving stability and reliability of the vehicle due to a large time delay in the output of the control signal, and improving the driving safety of the vehicle in scenarios such as autonomous driving.
[0075] According to the embodiment of the present application, using the multi-modal model to perform feature fusion on the multi-modal perception data includes: determining the perception features representing each modal perception data in the multi-modal perception data; based on multiple action embedding features arranged in sequence as prompt words, fusing multiple perception features according to the attention mechanism to obtain intermediate fusion features respectively related to multiple specified moments in the future specified time period.
[0076] According to the embodiment of the present application, the perception features corresponding to the perception data of the modality can be obtained by performing feature extraction on the perception data of at least two modalities such as image modality data, text modality data, and map data. For example, the multi-modal perception data can be feature-embedded based on the feature embedding layer of the attention network algorithm to obtain the perception features of each modal perception data. However, it is not limited to this. Other types of neural network algorithms such as convolutional neural network algorithms can also be used to perform feature extraction on the perception data of at least two modalities to obtain the perception features corresponding to the perception data of the modality. The embodiment of the present application does not limit the specific algorithm type for determining the perception features.
[0077] According to an embodiment of the present application, the encoder of the multimodal model can be used to perform attention fusion on multiple perception features and action embedding features corresponding to multiple specified moments, so as to obtain intermediate fusion features related to each of the multiple specified moments. Multiple driving action information is determined based on the multiple intermediate fusion features.
[0078] According to an embodiment of the present application, determining the driving action information for the action type according to the intermediate fusion features may include: fusing the multiple intermediate fusion features based on the attention mechanism to obtain target fusion features related to each of the multiple specified moments; and performing driving action detection based on the multiple target fusion features to obtain the driving action information for the action type at the multiple specified moments.
[0079] In some embodiments, the decoder of the trained multimodal model can be used to fuse the multiple intermediate fusion features based on the attention mechanism, so that the intermediate fusion features corresponding to the multiple specified moments can fully learn the relevance and continuity between the driving action information at multiple future specified moments, and the target fusion features corresponding to the multiple specified moments can represent the driving action under the condition of fully learning the feature semantics contained in the multimodal data, so as to improve the adaptability of the driving action information represented by the target fusion features to the driving conditions represented by multimodal perception data such as the driving environment and control text, improve the adaptability between the driving action information corresponding to the multiple specified moments and the action type output subsequently, and further improve the control accuracy and reliability of the vehicle through the driving action information.
[0080] In some implementations, by fusing the multiple intermediate fusion features to determine the target fusion features at the multiple specified moments, and further determining the driving action information of each action type at the multiple specified moments according to the multiple target fusion features, it is possible to implement parallel decoding processing on the encoded features that fuse the feature semantics of the multimodal perception data. As a result, the multimodal model can directly output the driving action information at the multiple specified moments, avoiding the model inference method in which the multimodal model needs to rely on at least the prediction result of the previous round as input each time it outputs a current prediction result, and avoiding the model inference duration of repeated feature decoding. Therefore, the device computing duration and computing overhead of devices such as vehicles and servers deploying the multimodal model in the process of controlling the vehicle can be reduced, the real-time requirement of vehicle control can be met, and the method provided by the embodiment of the present application can perform real-time control on the vehicle in devices with limited computing power such as vehicles, improving the flexibility of large model deployment.
[0081] Figure 3A A schematic diagram of the principle of the multimodal model according to an embodiment of the present application is shown.
[0082] Figure 3BThe schematic diagram of the principle of the decoder according to the embodiment of the present application is shown.
[0083] As Figure 3A shown, the trained multi-modal model includes an encoder, a decoder, and a detection layer. Among them, the encoder and the decoder can be constructed based on the Transformer network algorithm, and the detection layer can be constructed based on the multi-layer perceptron algorithm. The multiple perception features corresponding to the multi-modal perception data are respectively the image-modal perception feature 310 and the text-modal perception feature 320. The image-modal perception feature 310 can be determined by performing feature extraction on the image-modal data representing the driving environment, and the text-modal perception feature 320 can be determined by performing feature extraction on the text-modal data for controlling the vehicle.
[0084] The prompt words can include multiple action embedding features 330 arranged in a preset order. Among them, each square in the multiple action embedding features 330 can be understood as an action embedding feature corresponding to an action type. Input the image-modal perception feature 310, the text-modal perception feature 320, and the multiple action embedding features 330 into the encoder of the multi-modal model. The multiple action embedding features 330 can be used as the embedding vectors corresponding to the placeholders, so as to utilize the encoder to perform feature fusion on the image-modal perception feature 310, the text-modal perception feature 320, and the multiple action embedding features 330 based on the attention network algorithm, and obtain multiple intermediate fusion features 340. Among them, each square in the multiple intermediate fusion features 340 represents an intermediate fusion feature corresponding to a specified moment, and the arrangement order of the multiple intermediate fusion features 340 corresponds to the arrangement order of the multiple action embedding features 330.
[0085] Combined Figure 3A and Figure 3B shown, the intermediate fusion features 340 corresponding to multiple specified moments are input into the decoder, and the decoder performs feature fusion on the multiple intermediate fusion features based on the attention mechanism to obtain multiple target fusion features 350 corresponding to multiple specified moments. The multiple target fusion features 350 are processed by the detection layer constructed based on the multi-layer perceptron algorithm to perform a one-time detection output on the driving action information at multiple specified moments, and obtain the driving action information 360 corresponding to multiple specified moments. The driving action information can include the driving action data {m1, m2, …, mt} at t specified moments, where mt represents the driving speed data and the driving curvature data at the t-th specified moment.
[0086] In one embodiment, the multiple target fusion features can be expressed as H = VLA(X input ), where VLA() represents the algorithm execution process of the encoder and the decoder, and X input represents the multi-modal perception data and the multiple action embedding features input into the multi-modal model, and H is the multiple target fusion features.
[0087] In some embodiments, the action embedding features are determined based on the following operations: determining a plurality of preset text tokens for the multimodal model; performing data fusion on the text embedding features corresponding to the plurality of preset text tokens respectively to obtain action embedding features for at least one action type.
[0088] According to an embodiment of the present application, the preset text token can be understood as a placeholder (or called token) for a preset text word, and the text embedding feature corresponding to the preset text token can be stored in a preset text word feature library, so as to perform feature embedding on the text words input to the multimodal model based on the preset text word feature library.
[0089] According to an embodiment of the present application, by performing data fusion on the text embedding features of the preset text tokens, the fused action embedding features can be made different from the text embedding features, thereby avoiding the confusion between the action embedding features and the perceptual features of the multimodal perception data input to the multimodal model and generating model understanding interference.
[0090] In some embodiments, by adding the action embedding features for the action type to the preset text word feature library to update the embedding feature set, the multimodal model can perform semantic understanding with the driving action information for a specific action type as the target, and improve the semantic understanding adaptation ability of the multimodal model for the driving action information.
[0091] In some embodiments, at least one of an average value algorithm, a standard deviation algorithm, and a convolution fusion algorithm can be used to perform data fusion on the plurality of text embedding features. The embodiment of the present application does not limit the specific algorithm type for performing data fusion, as long as it can satisfy that the action embedding features are different from the text embedding features in the preset text word feature library.
[0092] In one example, an average value vector of the text embedding features in the preset text word feature library is calculated based on the average value algorithm, and the average value vector is determined as the action embedding feature related to at least one action type, and a mapping relationship is established between the action embedding feature and the action type, so as to make the action embedding feature adapt to the input rule requirements of the multimodal model.
[0093] In some embodiments, the method for controlling a vehicle may further include: determining driving speed data and driving curvature data respectively related to a plurality of specified moments from the driving action information; performing an integration operation on the driving speed data and driving curvature data respectively related to the plurality of specified moments to obtain driving heading angles for the plurality of specified moments; performing trajectory point planning based on the driving heading angles and driving speed data for the plurality of specified moments to obtain a plurality of planned trajectory points; and determining a planned driving trajectory based on the plurality of planned trajectory points.
[0094] In one embodiment, after the multimodal model predicts the driving speed data and driving curvature data at t specified times, the heading angle θ at the specified time t is calculated by performing an integration operation on the driving speed data and driving curvature data based on the following formula (1). t .
[0095] ; Formula (1)
[0096] where k(t) represents the driving curvature value for the driving curvature data at the specified time t, s(t) represents the driving speed value for the driving speed data at the specified time t, and θ t-1 represents the heading angle at time t - 1.
[0097] It should be noted that the driving speed data and driving curvature data at the starting specified time in the future specified period of the vehicle can be collected based on sensors, so as to calculate the heading angles at multiple subsequent specified times using formula (1).
[0098] In one embodiment, trajectory point planning based on the driving heading angle and driving speed data for multiple specified times may include: for the driving heading angle and driving speed data at any specified time, calculating the speed components according to the driving heading angle and driving speed data related to the specified time to obtain the speed components at any specified time; performing an integration calculation operation on the speed components for multiple specified times to obtain the planned trajectory points for the specified time.
[0099] For example, the speed components of the vehicle in the x-axis direction and y-axis direction at the specified time t can be determined based on the following formulas (2) and (3).
[0100] ; Formula (2)
[0101] ; Formula (3)
[0102] where and respectively represent the speed components of the vehicle in the x-axis direction and y-axis direction at the specified time t, and s t is the driving speed value at the specified time t.
[0103] In one example, the integration calculation operation can be performed on the speed components for multiple specified times based on the following formula (4) to obtain the trajectory point (x t , y t ) at the specified time t
[0104] ; Formula (4)
[0105] where xt and y t respectively represent the x - coordinate and y - coordinate of the trajectory point at the specified time t in the preset coordinate system. τ is the integration variable. In formula (4), the preset trajectory point (x0, y0) can be used as the starting trajectory point to perform the integration operation shown in formula (4).
[0106] In one embodiment, the initial preset trajectory point (x0, y0) can also be used as the data input of the following formula (5), and combined with the driving speed data and driving curvature data at t specified times to calculate the velocity components v x and v y , and perform the cumulative trapezoidal algorithm operation of the following formula (5) to obtain the heading angle and trajectory points.
[0107] ; formula (5)
[0108] wherein, is the time step.
[0109] For example, the initial preset trajectory point (x0, y0)=(0, 0), the initial heading angle θ0 = 0, the time step s, t > i ≥ 1, the driving action information at multiple specified times output by the multi - modal model is shown in Table 1.
[0110] Table 1
[0111]
[0112] According to the prediction results of the vehicle driving speed and driving curvature in the vehicle driving action information, the corresponding vehicle trajectory point set within the final future specified period of 3s is shown in Table 2
[0113] Table 2
[0114]
[0115] According to the embodiments of the present application, by performing vehicle driving trajectory planning based on driving action information, a large amount of computing power consumption caused by iteratively predicting multiple trajectory points by setting a trajectory point detection head for the multi - modal model can be avoided, the computing power resource requirements of the multi - modal model deployed on the vehicle end can be reduced, and the control timeliness and accuracy for the vehicle can be improved.
[0116] According to the method provided by the embodiments of the present application, the driving trajectory can be generated based on driving action information with strong interpretability, such as driving curvature data and driving speed data, which can provide a stable and highly interpretable way for driving trajectory planning generation, so as to improve a new way for vehicle driving trajectory planning.
[0117] It should be noted that in the embodiments of the present application, the trajectory points, planned trajectory points, and vehicle trajectory points represent the same meaning, and the planned trajectory and driving trajectory represent the same or similar meanings. The embodiments of the present application will not be elaborated herein.
[0118] Figure 4 The application scenario diagram of the method for controlling a vehicle according to an embodiment of the present application is shown.
[0119] As Figure 4 shown, in this application scenario, the computing device of vehicle 401 is deployed with a trained multimodal model. During the driving process of vehicle 401, the input multimodal perception data 410 can be processed by the multimodal model of the computing device to output prediction results 420 for multiple future specified time periods.
[0120] Among them, the multimodal perception data 410 includes driving environment images, the control text of the driver, and a prompt word, and the prompt word may include "Please provide the driving speed and driving curvature of the vehicle within the next 3 seconds, and the output mode is {(k1, s1), (k2, s2), (k3, s3), (k4, s4), (k5, s5), (k6, s6)}".
[0121] The prediction results 420 output by the multimodal model may include the text reply content "The driving speed and driving curvature within the next 3 seconds are predicted, and the output result is {(0.02, 0.01), (0.03, 0.2), (0.00, 0.0), (0.1, 0.01), (0.02, 0.3), (0.4, 0.01)}", where "{(0.02, 0.01), (0.03, 0.2), (0.00, 0.0), (0.1, 0.01), (0.02, 0.3), (0.4, 0.01)}" are the driving speed values and driving curvature values at 6 specified moments every 0.5 seconds within the next 3 seconds. The driving of vehicle 401 can be controlled based on the driving action information in the prediction results 420 to implement the autonomous driving function of vehicle 401.
[0122] The method for controlling a vehicle provided by the embodiments of the present application improves the interpretability of the model output results by determining the output results of the multimodal model as driving action information, so as to facilitate the rapid docking of the vehicle control system with the model output results, and predicts the driving action information at multiple specified moments through parallel decoding to replace the iterative reasoning of trajectory points, realizing the one-time output of multiple driving actions, improving the reasoning efficiency of the multimodal model, meeting the actual operation requirements of the vehicle control system to control the vehicle driving in real time, and optimizing the regression mechanism of the multimodal model by using the action embedding feature representing the action type as a structured prompt word, enabling the vehicle to be widely used in complex autonomous driving scenarios with extremely high requirements for real-time performance and safety, and having high practical value and engineering promotion prospects.
[0123] Embodiments of the present application also provide a method for training a multimodal model. The trained multimodal model determined based on the method for training a multimodal model provided by the embodiments of the present application can be applied to the method for controlling a vehicle in the above embodiments.
[0124] Figure 5 The flowchart of the method for training a multimodal model according to an embodiment of the present application is shown.
[0125] As Figure 5 shown, the method for training a multimodal model includes operation S510 to operation S530.
[0126] In operation S510, receive sample multimodal perception data for vehicle driving and labeled driving action information.
[0127] In operation S520, based on the sample prompt words, use the multimodal model to plan the driving actions of the vehicle based on the sample multimodal perception data, and obtain sample driving action information for the sample specified time period.
[0128] In operation S530, train the multimodal model based on the sample driving action information and the labeled driving action information to obtain a trained multimodal model.
[0129] In one example, the sample multimodal perception data includes sample data in multiple modalities such as images and speech texts.
[0130] According to the embodiment of the present application, the labeled driving action information represents the driving actions of the vehicle during the sample specified time period. For example, it can be the actual driving actions of the vehicle at multiple sample specified moments during the sample specified time period.
[0131] According to the embodiment of the present application, the sample prompt words are used to control the multimodal model to take the sample driving action information as the target and perform semantic understanding on the sample multimodal perception data.
[0132] It should be noted that the technical terms involved in the method for training a multimodal model provided by the embodiments of the present application, including but not limited to sample driving action information, sample multimodal perception data, etc., have the same or similar technical attributes as the corresponding technical terms involved in the method for controlling a vehicle provided by the embodiments of the present application, including but not limited to driving action information, multimodal perception data, etc. The embodiments of the present application will not elaborate herein.
[0133] In some embodiments, the sample prompt words include sample action embedding features corresponding to respective multiple sample action types. The multiple prompt words are arranged in sequence, and the arrangement positions of the multiple prompt words correspond to multiple specified moments in the sample specified time period.
[0134] In one embodiment, the sample prompt can be expressed in the following form:
[0135] Please provide the driving speed and driving curvature of the vehicle within the next 3 seconds. The output mode is: [ ([MASK],[MASK]), ([MASK], [MASK]), ..., ([MASK], [MASK]) ], where [MASK] represents the sample action embedding feature. ([MASK], [MASK]) indicates that the prediction target is 6 sets of predicted sample driving speed data and sample driving curvature data within 3 seconds after the current moment for the sample-specified time period. A sample self-defined moment is determined every 0.5s. The labeled driving action information is 6 sets of driving speed values and driving curvature values actually traveled by the vehicle at multiple specified moments within the specified time period.
[0136] In some embodiments, the sample driving action information and the labeled driving action information are respectively the first parameter value and the second parameter value corresponding to the sample action type. For example, the first parameter value and the second parameter value are respectively the predicted curvature value and the labeled curvature value corresponding to the driving curvature type corresponding to the specified moment t. Another example is that the first parameter value and the second parameter value are respectively the first array and the second array. The first array includes the sample curvature value and the sample speed value corresponding to the specified moment t, and the second array includes the labeled curvature value and the labeled speed value corresponding to the specified moment t.
[0137] In some embodiments, training the multimodal model based on the sample driving action information and the labeled driving action information may include: processing the first parameter value and the second parameter value using a loss function to obtain a loss value; and training the multimodal model based on the loss value.
[0138] For example, the first parameter value and the second parameter value can be processed using a loss function based on the following formula (6).
[0139] ; Formula (6)
[0140] where L pos represents the loss value, is the second parameter value of the labeled driving action information, m tLet \(v\) be the first parameter value of the sample driving action information, \(t\) be the serial number of the sample specified time, and \(T\) be the number of sample specified times. Among them, the first parameter value and the second parameter value used to calculate the loss value are of numerical type, for example, they can be speed value and curvature value. By processing the difference between the predicted result of numerical type and the label through the loss function, the error caused by the loss calculation for other data types such as trajectory point coordinates can be avoided, so as to optimize the prediction accuracy of the multi-modal model during the training process, avoid the error caused by the multi-modal model processing the text-based embedded feature representation, and enable the multi-modal model to be deployed in computing devices such as vehicle terminals, realizing better vehicle control performance and engineering applicability.
[0141] In some embodiments, the multi-modal model inference process can be based on an in-vehicle computing platform. First, convert the trained multi-modal model on the training platform, and at the in-vehicle computing platform end, convert the multi-modal model into a format supported by the in-vehicle computing platform end. The model inference can convert the driving action information into a driving trajectory according to actual needs for display on the in-vehicle side for the driver to view, or directly use the driving speed value and driving curvature value in the driving action information to control the vehicle driving.
[0142] Based on the above method for controlling a vehicle and the method for training a multi-modal model, the present application also provides a device for controlling a vehicle and a device for training a multi-modal model. The following will be combined with Figure 6 and Figure 7 to describe the device in detail.
[0143] Figure 6 FIG. shows a structural block diagram of a device for controlling a vehicle according to an embodiment of the present application.
[0144] As Figure 6 shown, the device 600 for controlling a vehicle in this embodiment includes a perception data receiving module 610, a driving action information obtaining module 620, and a control module 630.
[0145] The perception data receiving module 610 is configured to receive multi-modal perception data for vehicle driving.
[0146] The driving action information obtaining module 620 is configured to, based on a prompt word, use the trained multi-modal model to perform driving action planning on the vehicle based on the multi-modal perception data, and obtain driving action information for a future specified time period. The driving action information represents the driving action of the vehicle. Among them, the prompt word is used to control the multi-modal model to take the driving action information as the target and perform semantic understanding on the multi-modal perception data.
[0147] The control module 630 is configured to control the vehicle driving based on the driving action information.
[0148] According to an embodiment of the present application, the driving action information acquisition module includes: an intermediate fusion feature acquisition sub-module and a driving action information determination sub-module.
[0149] The intermediate fusion feature acquisition sub-module is configured to use the action embedding feature representing the action type of the driving action information as a prompt word, and utilize a multi-modal model to perform feature fusion on multi-modal perception data to obtain an intermediate fusion feature related to the action type, wherein the prompt word is used to control the multi-modal model to take the driving action information corresponding to the action type as the target and perform semantic understanding on the multi-modal perception data.
[0150] The driving action information determination sub-module is configured to determine the driving action information for the action type according to the intermediate fusion feature.
[0151] According to an embodiment of the present application, the intermediate fusion feature acquisition sub-module includes: a first determination unit and a first acquisition unit.
[0152] The first determination unit is configured to determine the perception features representing each modality perception data in the multi-modal perception data.
[0153] The first acquisition unit is configured to, based on a plurality of action embedding features arranged in sequence as prompt words, fuse a plurality of perception features according to the attention mechanism to obtain intermediate fusion features respectively related to a plurality of specified moments in a future specified time period, wherein a plurality of driving action information is determined based on the plurality of intermediate fusion features.
[0154] According to an embodiment of the present application, the driving action information determination sub-module includes: a second acquisition unit and a third acquisition unit.
[0155] The second acquisition unit is configured to fuse a plurality of intermediate fusion features according to the attention mechanism to obtain target fusion features respectively related to a plurality of specified moments.
[0156] The third acquisition unit is configured to perform driving action detection based on the plurality of target fusion features to obtain the driving action information for the action type at a plurality of specified moments.
[0157] According to an embodiment of the present application, the action embedding feature is determined based on the following operations: determining a plurality of preset text tokens for the multi-modal model; performing data fusion on the text embedding features respectively corresponding to the plurality of preset text tokens to obtain the action embedding feature for at least one action type.
[0158] According to an embodiment of the present application, the prompt word includes the action embedding features respectively corresponding to a plurality of action types, the plurality of prompt words are arranged in sequence, and the arrangement positions of the plurality of prompt words correspond to a plurality of specified moments in a future specified time period.
[0159] According to an embodiment of the present application, the driving action information includes at least one of the following: driving speed data for a specified moment; driving curvature data for a specified moment; driving heading angle data for a specified moment.
[0160] According to an embodiment of the present application, the control module includes: a first determination sub-module and a control sub-module.
[0161] The first determination sub-module is configured to determine a power control instruction and a steering control instruction based on the driving speed data and the driving curvature data in the action information.
[0162] The control sub-module is configured to control the vehicle to drive in a future specified time period based on the power control instruction and the steering control instruction.
[0163] According to an embodiment of the present application, the device for controlling a vehicle further includes: a first determination module, a first acquisition module, a trajectory point acquisition module, and a second determination module.
[0164] The first determination module is configured to determine the driving speed data and the driving curvature data respectively related to a plurality of specified moments from the driving action information, where the future specified time period includes a plurality of specified moments.
[0165] The first acquisition module is configured to perform an integration operation on the driving speed data and the driving curvature data respectively related to a plurality of specified moments to obtain the driving heading angles for the plurality of specified moments.
[0166] The trajectory point acquisition module is configured to perform trajectory point planning based on the driving heading angles and the driving speed data for the plurality of specified moments to obtain a plurality of planned trajectory points.
[0167] The second determination module is configured to determine a planned driving trajectory based on the plurality of planned trajectory points.
[0168] According to an embodiment of the present application, the trajectory point acquisition module includes: a speed component acquisition sub-module and a planned trajectory point acquisition sub-module.
[0169] The speed component acquisition sub-module is configured to perform speed component calculation according to the driving heading angle and the driving speed data related to a specified moment for the driving heading angle and the driving speed data at any specified moment to obtain the speed component at any specified moment.
[0170] The planned trajectory point acquisition sub-module is configured to perform an integration calculation operation on the speed components for the plurality of specified moments to obtain the planned trajectory points for the specified moment.
[0171] According to an embodiment of the present application, the multi-modal perception data includes at least two of the following: image modal data characterizing the driving environment of the vehicle; text modal data characterizing the driving intention of controlling the vehicle; map data related to the driving environment.
[0172] Figure 7 The structural block diagram of an apparatus for training a multimodal model according to an embodiment of the present application is shown.
[0173] As Figure 7 shown, the apparatus 700 for training a multimodal model in this embodiment includes a sample receiving module 710, a sample driving action information obtaining module 720, and a training module 730.
[0174] The sample receiving module 710 is configured to receive sample multimodal perception data for vehicle driving and labeled driving action information, where the labeled driving action information represents the driving action of the vehicle during the sample specified period.
[0175] The sample driving action information obtaining module 720 is configured to, based on a sample prompt word, use a multimodal model to perform driving action planning on the vehicle based on the sample multimodal perception data, so as to obtain sample driving action information for the sample specified period, where the sample prompt word is used to control the multimodal model to use the sample driving action information as a target to perform semantic understanding on the sample multimodal perception data.
[0176] The training module 730 is configured to train the multimodal model based on the sample driving action information and the labeled driving action information to obtain a trained multimodal model.
[0177] According to an embodiment of the present application, the sample driving action information and the labeled driving action information are respectively a first parameter value and a second parameter value corresponding to a sample action type; wherein, the training module includes: a loss value determination sub-module and a training sub-module.
[0178] The loss value determination sub-module is configured to process the first parameter value and the second parameter value by using a loss function to obtain a loss value; and
[0179] The training sub-module is configured to train the multimodal model based on the loss value.
[0180] According to an embodiment of the present application, the sample prompt word includes sample action embedding features respectively corresponding to multiple sample action types, the multiple prompt words are arranged in sequence, and the arrangement positions of the multiple prompt words correspond to multiple specified moments in the future specified period.
[0181] An embodiment of the present application further provides a vehicle, including: a processor configured to execute the method for controlling a vehicle provided in the embodiment of the present application.
[0182] Figure 8 The block diagram of an electronic device suitable for implementing the method for controlling a vehicle and the method for training a multimodal model according to an embodiment of the present application is shown.
[0183] As Figure 8As shown, the electronic device 800 according to an embodiment of the present application includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. The processor 801 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 801 can also include on-board memory for caching purposes. The processor 801 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present application.
[0184] In the RAM 803, various programs and data required for the operation of the electronic device 800 are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. The processor 801 performs various operations of the method flow according to an embodiment of the present application by executing the program in the ROM 802 and / or the RAM 803. It should be noted that the program can also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 can also perform various operations of the method flow according to an embodiment of the present application by executing the program stored in the one or more memories.
[0185] According to an embodiment of the present application, the electronic device 800 may further include an input / output (I / O) interface 805, and the input / output (I / O) interface 805 is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the input / output (I / O) interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A driver 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the driver 810 as needed so that a computer program read from it can be installed into the storage section 808 as needed.
[0186] The present application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist alone without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiments of the present application is implemented.
[0187] According to an embodiment of the present application, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disk, hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, the computer-readable storage medium may include the above-described ROM 802 and / or RAM 803 and / or one or more memories other than ROM 802 and RAM 803.
[0188] An embodiment of the present application also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to cause the computer system to implement the method provided by the embodiments of the present application.
[0189] When the computer program is executed by the processor 801, the above functions defined in the system / apparatus of the embodiments of the present application are executed. According to an embodiment of the present application, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.
[0190] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 809, and / or be installed from the removable medium 811. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0191] In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, the above functions defined in the system of the embodiments of the present application are executed. According to the embodiments of the present application, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0192] According to the embodiments of the present application, the program code for executing the computer program provided by the embodiments of the present application can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by connecting through the Internet using an Internet service provider).
[0193] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0194] Those skilled in the art can understand that the features described in the various embodiments of the present application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present application. In particular, without departing from the spirit and teachings of the present application, the features described in the various embodiments of the present application can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present application.
[0195] The embodiments of the present application have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present application, those skilled in the art can make various substitutions and modifications, and all of these substitutions and modifications should fall within the scope of the present application.
Claims
1. A method for controlling a vehicle, characterized in that, The method includes: Receiving multimodal perception data for vehicle driving; Based on a prompt word, using a trained multimodal model to plan driving actions for the vehicle based on the multimodal perception data, obtaining driving action information for a future specified time period, where the driving action information represents the driving actions of the vehicle. Among them, the prompt word is used to control the multimodal model to use the driving action information as the target and perform semantic understanding on the multimodal perception data; and Controlling the vehicle to drive based on the driving action information.
2. The method according to claim 1, wherein The step of using the trained multimodal model to plan driving actions for the vehicle based on the multimodal perception data based on the prompt word includes: Using the action embedding feature representing the action type of the driving action information as the prompt word, and using the multimodal model to perform feature fusion on the multimodal perception data to obtain an intermediate fusion feature related to the action type. Among them, the prompt word is used to control the multimodal model to use the driving action information corresponding to the action type as the target and perform semantic understanding on the multimodal perception data; Determining the driving action information for the action type according to the intermediate fusion feature.
3. The method according to claim 2, wherein The step of using the multimodal model to perform feature fusion on the multimodal perception data includes: Determining perception features representing each modality perception data in the multimodal perception data; Based on a plurality of the action embedding features arranged in sequence as the prompt word, fusing the plurality of perception features according to the attention mechanism to obtain intermediate fusion features respectively related to a plurality of specified moments in the future specified time period, where a plurality of the driving action information is determined based on the plurality of intermediate fusion features.
4. The method according to claim 3, characterized in that, The step of determining the driving action information for the action type according to the intermediate fusion feature includes: Fusing a plurality of the intermediate fusion features according to the attention mechanism to obtain target fusion features respectively related to the plurality of specified moments; and Performing driving action detection based on the plurality of target fusion features to obtain driving action information for the action type at the plurality of specified moments.
5. The method according to claim 2, characterized in that, The action embedding feature is determined based on the following operations: Determining a plurality of preset text tokens for the multimodal model; Performing data fusion on the text embedding features corresponding to the plurality of preset text tokens respectively to obtain an action embedding feature for at least one of the action types.
6. The method according to claim 2, wherein, The prompt word includes action embedding features corresponding to the plurality of action types respectively, the plurality of prompt words are arranged in sequence, and the arrangement positions of the plurality of prompt words correspond to the plurality of specified moments in the future specified time period.
7. The method according to any one of claims 1 to 6, characterized in that The driving action information includes at least one of the following: Driving speed data for a specified moment; Driving curvature data for the specified moment; Driving heading angle data for the specified moment.
8. The method according to claim 1, characterized in that, The step of controlling the vehicle to drive based on the driving action information includes: Determining a power control instruction and a steering control instruction based on the driving speed data and the driving curvature data in the driving action information; and Based on the power control instruction and the steering control instruction, control the vehicle to travel in the future specified time period.
9. The method according to claim 1, characterized in that The method further includes: Determine driving speed data and driving curvature data respectively related to a plurality of specified moments from the driving action information, wherein the future specified time period includes a plurality of the specified moments; Perform an integration operation on the driving speed data and driving curvature data respectively related to a plurality of the specified moments to obtain driving heading angles for a plurality of the specified moments; Based on the driving heading angles and driving speed data for a plurality of the specified moments, perform trajectory point planning to obtain a plurality of planned trajectory points; and Determine the planned driving trajectory based on the plurality of planned trajectory points.
10. The method according to claim 9, wherein The performing trajectory point planning based on the driving heading angles and driving speed data for a plurality of the specified moments includes: For the driving heading angle and driving speed data of any specified moment, perform speed component calculation according to the driving heading angle and driving speed data related to the specified moment to obtain the speed component of the any specified moment; Perform an integration calculation operation on the speed components for a plurality of the specified moments to obtain the planned trajectory points for the specified moment.
11. The method according to claim 1, wherein The multi-modal perception data includes at least two of the following: Image modal data characterizing the driving environment of the vehicle; Text modal data characterizing the driving intention of controlling the vehicle; Map data related to the driving environment.
12. A method for training a multi-modal model, characterized in that, Includes: Receive sample multi-modal perception data for vehicle driving and labeled driving action information, where the labeled driving action information represents the driving action of the vehicle in the sample specified time period; Based on the sample prompt words, use a multi-modal model to perform driving action planning on the vehicle based on the sample multi-modal perception data to obtain sample driving action information for the sample specified time period, where the sample prompt words are used to control the multi-modal model to use the sample driving action information as the target and perform semantic understanding on the sample multi-modal perception data; Train the multi-modal model based on the sample driving action information and the labeled driving action information to obtain a trained multi-modal model.
13. The method according to claim 12, characterized in that, The sample driving action information and the labeled driving action information are respectively a first parameter value and a second parameter value corresponding to the sample action type; wherein, the training the multi-modal model based on the sample driving action information and the labeled driving action information includes: Process the first parameter value and the second parameter value using a loss function to obtain a loss value; and Train the multi-modal model based on the loss value.
14. The method according to claim 12, wherein The sample prompt words include sample action embedding features respectively corresponding to a plurality of sample action types, the plurality of prompt words are arranged in order, and the arrangement positions of the plurality of prompt words correspond to a plurality of specified moments in the sample specified time period.
15. A device for controlling a vehicle, characterized in that, The device includes: A perception data receiving module, configured to receive multi-modal perception data for vehicle driving; A driving action information acquisition module, configured to, based on a prompt, use a trained multimodal model to perform driving action planning on the vehicle based on the multimodal perception data, to obtain driving action information for a specified future time period, where the driving action information represents the driving action of the vehicle, and where the prompt is used to control the multimodal model to use the driving action information as a target to perform semantic understanding on the multimodal perception data; and A control module, configured to control the vehicle to drive based on the driving action information.
16. An apparatus for training a multi-modal model, characterized in that, The apparatus includes: A sample receiving module, configured to receive sample multimodal perception data for vehicle driving and labeled driving action information, where the labeled driving action information represents the driving action of the vehicle during a sample specified time period; A sample driving action information acquisition module, configured to, based on a sample prompt, use a multimodal model to perform driving action planning on the vehicle based on the sample multimodal perception data, to obtain sample driving action information for a sample specified time period, where the sample prompt is used to control the multimodal model to use the sample driving action information as a target to perform semantic understanding on the sample multimodal perception data; A training module, configured to train the multimodal model based on the sample driving action information and the labeled driving action information, to obtain a trained multimodal model.
17. A vehicle, characterized in that, Including: A processor, configured to execute the method according to any one of claims 1 to 11.
18. An electronic device, characterized in that, Including: One or more processors; A storage device, configured to store one or more programs, where, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 14.
19. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 14.
20. A computer program product, characterized in that, Including a computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Multi-source sensor data fusion system and multi-source sensor data fusion method
CN110794406A
Training method of automatic driving model and control information acquisition method
CN118393876A
Vehicle control method and device, vehicle and computer readable storage medium
CN118419067A
Automatic driving path optimization control method and device integrating environment perception and prediction
CN119568197A
Automatic driving method and system based on end-to-end and multi-modal large model
CN120003527A