Driving decision generation method and device, model training method and device, equipment and storage medium

Through data processing technology based on reinforcement learning and multimodal learning models, and combining voice data and environmental information to generate personalized driving decisions, the problem of insufficient user needs of intelligent driving systems is solved, and the flexibility and user experience of the system are improved.

CN120359161APending Publication Date: 2025-07-22SZ ZHUOYU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202580000464.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The driving decision interaction interface of existing intelligent driving systems is usually standardized and has a single type of adjustment, which makes it difficult to adapt to users' personalized needs and lack of flexibility.

Method used

Through a data processing model trained based on reinforcement learning models or multimodal models and deep learning models, combining user voice data and driving environment information, personalized driving decisions are generated and autonomous mobile devices are controlled.

Benefits of technology

It realizes a driving experience that is more in line with the personalized needs of users, improves the flexibility and adaptability of the intelligent driving system, and enhances the naturalness and ease of use of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120359161A_ABST
    Figure CN120359161A_ABST
Patent Text Reader

Abstract

The invention discloses a driving decision generation method and device, a model training method and device, equipment and a storage medium. The method comprises the steps of processing voice data input by a user and pre-acquired driving environment information based on a data processing model, and generating a driving decision; and controlling the autonomous mobile device to run according to the driving decision. Wherein the data processing model is obtained by training based on a reinforcement learning model, or is obtained by training based on a multi-modal model and a deep learning model. Through the method, personalized processing of the voice data is realized, so that the driving decision which better meets the personalized demand of the user is generated, and the control flexibility and the user experience of the intelligent driving system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent driving technology, and particularly to a method for generating driving decisions, a method for model training, a device, a device, and a storage medium. Background Art

[0002] With the rapid development of intelligent driving technology, autonomous mobile devices equipped with intelligent driving systems can comprehensively analyze data from various sensors, such as radar, lidar, and ultrasonic sensors, etc., to comprehensively perceive the surrounding environment, process complex traffic conditions in real time, identify road signs, pedestrians, and other vehicles, and make intelligent path planning and speed adjustment according to the current traffic rules and environmental conditions.

[0003] In related technologies, users can interact with the intelligent driving system through preset driving decision interaction interfaces (such as lever for lane change, lever for speed adjustment, lever for vehicle distance adjustment), allowing users to control the vehicle in different driving scenarios. However, the preset interaction interfaces are usually standardized and have a single adjustment type, resulting in insufficient flexibility of the intelligent driving system in adapting to users' personalized needs. Summary of the Invention

[0004] The present application provides a method for generating driving decisions, a method for model training, a device, a device, and a storage medium, so as to improve the flexibility of the intelligent driving system in adapting to users' personalized needs.

[0005] In a first aspect, the present application provides a method for generating driving decisions, the method comprising:

[0006] Based on a data processing model, process the voice data input by the user and the pre-acquired driving environment information to generate a driving decision; the data processing model is trained based on a reinforcement learning model, or is trained based on a multimodal model and a deep learning model;

[0007] Control the autonomous mobile device to travel according to the driving decision.

[0008] In a possible implementation manner, the processing the voice data input by the user and the pre-acquired driving environment information based on the data processing model to generate a driving decision includes:

[0009] Based on a vision-language large model, process the voice data input by the user and the pre-acquired driving environment information to obtain a target feature vector, the target feature vector being used to represent at least one of the following features: the driving intention feature of the user, the category feature of the driving scenario, the driving style feature, and the state feature of the autonomous mobile device;

[0010] Process the target feature vector through an assisted driving model to generate a driving decision;

[0011] Wherein, the data processing model is obtained by combining the vision - language large model and the assisted driving model.

[0012] In a possible implementation manner, the method further includes:

[0013] Push the execution information of the driving decision to the user, where the execution information includes: the current operation being executed by the autonomous mobile device, and / or the expected operation to be executed.

[0014] In a possible implementation manner, the method further includes:

[0015] Obtain the driving environment information collected during the movement of the autonomous mobile device, where the driving environment information includes at least one of the following: the motion state information of the autonomous mobile device, the motion state information of surrounding traffic participants, and road information;

[0016] Obtain the voice data input by the user through the audio collection device in the autonomous mobile device.

[0017] In a possible implementation manner, the voice data includes a control instruction, and / or a driving style instruction;

[0018] The control instruction includes at least one of the following: increasing the driving speed, decreasing the driving speed, overtaking the target ahead, bypassing an obstacle, selecting a target road, driving into a feeder road, and docking at a target area;

[0019] The driving style instruction includes at least one of the following: the user's driving preference, the state of the autonomous mobile device, the user state, and the driving task time limit requirement.

[0020] In a possible implementation manner, the driving decision includes at least one of the following: planning a driving route, speed adjustment, overtaking, obstacle avoidance, driving into a target lane, driving into a feeder road, lane keeping, parking, and emergency braking.

[0021] In a possible implementation manner, the method further includes:

[0022] Receive the data processing model sent by the server and deploy the data processing model.

[0023] In a second aspect, the present application provides a model training method for driving decisions. The data processing model is trained based on a multi - modal model and a deep learning model. The method includes:

[0024] Obtain multiple first training samples, where each first training sample includes data to be processed and a driving scenario label corresponding to the data to be processed; the data to be processed includes voice data input by a user and driving environment information, and the driving environment information includes at least one of the following: motion state information of an autonomous mobile device, motion state information of surrounding traffic participants, and road information;

[0025] Train the multi-modal model according to the multiple first training samples to obtain a vision-language large model;

[0026] Construct multiple second training samples according to the feature vector corresponding to each first training sample and the driving decision corresponding to each first training sample. Each second training sample includes the feature vector corresponding to a first training sample and the driving decision corresponding to the first training sample;

[0027] Train the deep learning model according to the multiple second training samples to obtain an assisted driving model;

[0028] Perform combined processing on the vision-language large model and the assisted driving model to obtain the data processing model.

[0029] In a possible implementation manner, the driving scenario label includes the category of the driving scenario and the state of the autonomous mobile device. The category of the driving scenario is any one of the following: overtaking, detouring, accelerating, decelerating, or lane changing; the state of the autonomous mobile device includes in preparation, in execution, or completed.

[0030] In a possible implementation manner, if the data processing model is trained based on a reinforcement learning model, the method further includes:

[0031] Obtain multiple training samples, where each training sample includes data to be processed and a driving decision corresponding to the data to be processed; the data to be processed includes voice data input by a user and driving environment information, and the driving environment information includes at least one of the following: motion state information of an autonomous mobile device, motion state information of surrounding traffic participants, and road information;

[0032] Train the reinforcement learning model according to the multiple training samples to obtain the data processing model.

[0033] In a possible implementation manner, the method further includes:

[0034] Send the data processing model to at least one autonomous mobile device.

[0035] In a third aspect, the present application provides a device for generating driving decisions, and the device includes:

[0036] A first processing module, configured to process the voice data input by the user and the pre-acquired driving environment information based on a data processing model to generate a driving decision; the data processing model is obtained by training based on a reinforcement learning model, or is obtained by training based on a multimodal model and a deep learning model;

[0037] A second processing module, configured to control the autonomous mobile device to travel according to the driving decision.

[0038] In a fourth aspect, the present application provides a model training device for driving decisions, and the device includes:

[0039] An acquisition module, configured to acquire a plurality of first training samples, each first training sample including data to be processed and a driving scenario label corresponding to the data to be processed; the data to be processed includes voice data input by the user and driving environment information, and the driving environment information includes at least one of the following: motion state information of the autonomous mobile device, motion state information of surrounding traffic participants, and road information;

[0040] A first training module, configured to train a multimodal model according to the plurality of first training samples to obtain a vision-language large model;

[0041] A construction module, configured to construct a plurality of second training samples according to the feature vector corresponding to each first training sample and the driving decision corresponding to each first training sample, each second training sample including the feature vector corresponding to a first training sample and the driving decision corresponding to the first training sample;

[0042] A second training module, configured to train a deep learning model according to the plurality of second training samples to obtain an assisted driving model;

[0043] A combination module, configured to perform a combination process on the vision-language large model and the assisted driving model to obtain the data processing model.

[0044] In a fifth aspect, the present application provides an electronic device, including: a processor, a memory, and a communication interface;

[0045] The memory stores computer execution instructions;

[0046] The processor executes the computer execution instructions stored in the memory to implement the method according to any one of the first aspects.

[0047] In a sixth aspect, the present application provides an autonomous mobile device, including: the electronic device according to the fifth aspect.

[0048] In a seventh aspect, the present application provides a server, including: a processor, a memory, and a communication interface;

[0049] The memory stores computer-executable instructions;

[0050] The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of the second aspect.

[0051] In an eighth aspect, the present application provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the method according to any one of the first aspect or the second aspect when executed by a processor.

[0052] In a ninth aspect, the present application provides a program product, including: a computer program, which causes the computer to execute the method according to any one of the first aspect or the second aspect when the program product runs on the computer.

[0053] In a tenth aspect, the present application provides a computer program, which is used to execute the method according to any one of the first aspect or the second aspect when executed by a processor.

[0054] The present application provides a method for generating a driving decision, a method for model training, a device, a device, and a storage medium. It includes processing the voice data input by the user and the pre-acquired driving environment information based on a data processing model to generate a driving decision, and controlling the autonomous mobile device to travel according to the driving decision. Among them, the data processing model is trained based on a reinforcement learning model, or is trained based on a multi-modal model and a deep learning model. In the above process, a driving decision that better meets the personalized needs of the user can be generated through natural language interaction, combined with the pre-acquired driving environment information of the autonomous mobile device, improving the flexibility of the intelligent driving system in adapting to the personalized needs of the user. Description of the Drawings

[0055] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0056] Figure 1 It is a schematic diagram of the application scenario provided by the embodiment of the present application;

[0057] Figure 2 It is a schematic flowchart of the first embodiment of the method for generating a driving decision provided by the present application;

[0058] Figure 3 It is a schematic flowchart of the second embodiment of the method for generating a driving decision provided by the present application;

[0059] Figure 4Schematic flowchart of Embodiment 1 of the model training method for driving decision-making provided by this application;

[0060] Figure 5 Schematic structural diagram of a multi-modal model provided by an embodiment of this application;

[0061] Figure 6 Schematic flowchart of Embodiment 2 of the model training method for driving decision-making provided by this application;

[0062] Figure 7 Schematic diagram of the data processing model training and application example provided by an embodiment of this application;

[0063] Figure 8 Schematic structural diagram of Embodiment 1 of the driving decision generation device provided by this application;

[0064] Figure 9 Schematic structural diagram of Embodiment 2 of the driving decision generation device provided by this application;

[0065] Figure 10 Schematic structural diagram of Embodiment 1 of the model training device for driving decision-making provided by this application;

[0066] Figure 11 Schematic structural diagram of Embodiment 2 of the model training device for driving decision-making provided by this application;

[0067] Figure 12 Schematic structural diagram of the electronic device provided by an embodiment of this application;

[0068] Figure 13 Schematic structural diagram of the server provided by an embodiment of this application. Detailed implementation manners

[0069] In the embodiments of this application, the term "and / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0070] In the embodiments of this application, the term "multiple" means two or more, and other quantifiers are similar.

[0071] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part rather than all of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.

[0072] Figure 1 It is a schematic diagram of the application scenario provided by the embodiment of this application. Please refer to Figure 1 , in an autonomous mobile device, a driving decision interaction interface and an intelligent driving system can be preset. When the intelligent driving system is working, the user can adjust the driving decisions generated by the intelligent driving system through the driving decision interaction interface. This enables the user to adjust the driving behavior of the device in a changing traffic environment to adapt to the current road conditions and personal preferences.

[0073] Optionally, the autonomous mobile device can be a flying car, an intelligent driving vehicle, a mobile robot, an intelligent ship, or other devices that can move and are configured with an intelligent driving system.

[0074] For example, in the application scenario of an intelligent driving vehicle, when the intelligent driving system is working, the user can adjust the following distance during the driving of the intelligent driving vehicle by adjusting the lever distance, improving the driving safety and comfort.

[0075] In the related art, the preset driving decision interaction interface is usually standardized and has a single adjustment type. When interacting with the intelligent driving system, only limited options can be used to adjust the driving behavior of the vehicle. This results in a lack of flexibility in the intelligent driving system to adapt to the personalized needs of users.

[0076] To address the above problems, the inventors considered generating driving decisions that better meet the personalized needs of users by using natural language interaction and combining the driving environment information obtained by the autonomous mobile device. Accordingly, the inventors found through multiple experiments that a data processing model trained based on a reinforcement learning model or based on a multi-modal model and a deep learning model can be used to process the voice data input by the user and the pre-obtained driving environment information, generate driving decisions, and control the autonomous mobile device to drive, thereby achieving a more efficient and personalized driving experience. Based on this, this application proposes a method for generating driving decisions, aiming to improve the flexibility of the intelligent driving system to better adapt to the personalized needs of users.

[0077] The following uses specific embodiments to elaborate in detail on the technical solution of this application and how the technical solution of this application solves the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the embodiments of this application in conjunction with the accompanying drawings.

[0078] Figure 2 This is a schematic flowchart of the first embodiment of the method for generating a driving decision provided by this application. Please refer to Figure 2 This method includes:

[0079] S201. Based on a data processing model, process the voice data input by the user and the pre-acquired driving environment information to generate a driving decision.

[0080] The execution subject of the embodiment of this application can be an electronic device or a device for generating a driving decision provided in the electronic device. The device for generating a driving decision can be implemented by software or by a combination of software and hardware. The device for generating a driving decision can be a processor in the electronic device. For the sake of easy understanding, in the following, the execution subject is taken as an electronic device as an example for description.

[0081] In this step, the electronic device can comprehensively analyze the voice data input by the user and the pre-acquired driving environment information based on the pre-acquired data processing model, and generate a driving decision that conforms to the current driving environment and user needs. Among them, the data processing model is trained based on a reinforcement learning model or trained based on a multi-modal model and a deep learning model.

[0082] For the sake of easy understanding, the following will use an intelligent driving vehicle as an example of an autonomous mobile device for description.

[0083] Regarding the voice data, it can refer to the audio signals provided to the electronic device by the user in the form of voice in the context of driving decision generation. These signals can include control instructions and / or driving style instructions.

[0084] Optionally, the control instructions can include at least one of the following: increasing the driving speed, decreasing the driving speed, overtaking the target ahead, bypassing an obstacle, selecting a target road, entering a feeder road, and docking at a target area;

[0085] Optionally, the driving style instructions can include at least one of the following: the user's driving preference, the state of the autonomous mobile device, the user state, and the driving task time limit requirement.

[0086] Regarding driving environment information, it refers to the detailed information about the current driving conditions and the surrounding environment of an autonomous mobile device obtained by an electronic device through various sensors and data sources, which may include at least one of the following: the motion state information of the autonomous mobile device, the motion state information of surrounding traffic participants, and road information.

[0087] Regarding driving decisions, they can be specific operation plans generated by a data processing model after comprehensively processing voice data and driving environment information, and are used to guide the actual driving behavior of an autonomous mobile device.

[0088] In an alternative embodiment, driving decisions may include at least one of the following: planning a driving route, speed adjustment, overtaking, obstacle avoidance, entering a target lane, entering a secondary road, lane keeping, parking, and emergency braking.

[0089] The reinforcement learning model is a learning framework trained through a trial-and-error and feedback mechanism. In this model, the autonomous mobile device acts as an intelligent agent, which can continuously interact with its surrounding environment, perceive the environmental state, and make a series of action decisions in combination with the user's voice data input. For example, the reinforcement learning model can be a learning model adopting a Deep Q-Network (DQN) structure.

[0090] The multi-modal model can enhance the environmental perception of an autonomous mobile device by integrating information of multiple modalities such as visual, sensor data, and the voice data input by the user.

[0091] The deep learning model aims to simulate the learning process of the human brain using a multi-layer neural network structure, and can extract features and patterns from a large amount of data, so as to output predictions or decisions for specific tasks. For example, the deep learning model can be a learning model adopting a convolutional neural network structure.

[0092] It should be noted that the reinforcement learning model can be a unified model obtained by integrating the functions of the multi-modal model and the deep learning model.

[0093] For example, a data processing model can be trained based on the reinforcement learning model. The data processing model can process the voice data input by the user and the pre-acquired driving environment information to generate driving decisions. Among them, the voice data can be "Overtake the vehicle A ahead"; the driving environment information can include: the driving speed of the intelligent driving vehicle is 60 kilometers per hour (km / h), the driving speed of the vehicle A ahead is 30 km / h, and there are no other vehicles in the left lane of the current lane; the driving decision can be overtaking, and for example, the planned overtaking trajectory can be output.

[0094] In an alternative embodiment, a vision-language large model can be trained based on a multi-modal model, and an assisted driving model can be trained based on a deep learning model. The data processing model is obtained by combining the vision-language large model and the assisted driving model. Then, based on the vision-language large model, the speech data input by the user and the pre-acquired driving environment information can be processed to obtain a target feature vector, and the target feature vector can be processed by the assisted driving model to generate a driving decision. Among them, the target feature vector can represent at least one of the following features: the driving intention feature of the user, the category feature of the driving scenario, the driving style feature, and the state feature of the autonomous mobile device.

[0095] For example, based on the vision-language large model, the speech data "overtake the vehicle A ahead" input by the user and the pre-acquired driving environment information including: "the driving speed of the intelligent driving vehicle is 60 kilometers per hour (km / h), the driving speed of the vehicle A ahead is 30 km / h, and there are no other vehicles in the left lane of the current lane" can be processed to obtain a target feature vector 1, and the target feature vector 1 can be processed by the assisted driving model to generate a driving decision. Among them, the driving decision can be overtaking.

[0096] S202. Control the autonomous mobile device to travel according to the driving decision.

[0097] In this step, the autonomous mobile device can be controlled to travel according to the driving decision generated by the data processing model.

[0098] For example, if the driving decision generated by the data processing model is overtaking, then based on the overtaking driving decision, the intelligent driving vehicle can be controlled to overtake the vehicle A ahead. Exemplarily, the intelligent driving vehicle can change lanes to the left lane at a speed of 60 km / h to overtake the vehicle A traveling at a speed of 30 km / h in the current lane.

[0099] In the embodiments of the present application, the electronic device can use the data processing model trained through the reinforcement learning model or trained through the multi-modal model and the deep learning model to process the speech data input by the user and the pre-acquired driving environment information, generate a driving decision, and control the autonomous mobile device to travel. In the above process, by analyzing the speech data input by the user, the data processing model can generate personalized driving decisions to better meet the preferences and needs of the user.

[0100] Furthermore, in the driving decision generation method provided by the embodiments of the present application, users can achieve fine control of autonomous mobile devices through simple voice data, without being limited to traditional standardized interfaces (such as lever for lane change, lever for speed adjustment, lever for vehicle distance adjustment) to control the devices, thereby improving the flexibility and adaptability of the intelligent driving system, better meeting the personalized needs of users, and making the user experience more natural and intuitive.

[0101] Based on Figure 2 the embodiments shown below, in combination with Figure 3 , the above driving decision generation method will be further described in detail.

[0102] Figure 3 is a schematic flowchart of the second embodiment of the driving decision generation method provided by the present application. Please refer to Figure 3 , the data processing model is obtained by combining a vision-language large model and an assisted driving model. The method includes:

[0103] S301. Obtain driving environment information collected during the movement of the autonomous mobile device.

[0104] In this step, the electronic device can obtain the driving environment information collected during the movement of the autonomous mobile device. The driving environment information includes at least one of the following: the motion state information of the autonomous mobile device, the motion state information of surrounding traffic participants, and road information.

[0105] The motion state information of the autonomous mobile device may include but is not limited to: current position, driving speed, acceleration magnitude, driving direction, vehicle attitude (pitch angle, roll angle, yaw angle), current lane position, and power system state (such as battery charge, fuel level). For example, it can be obtained through components such as sensors installed on the autonomous mobile device.

[0106] The motion state information of surrounding traffic participants may include but is not limited to: the distance and direction, relative speed, driving direction, and traffic participant type (such as vehicle, pedestrian, bicycle) of surrounding traffic participants relative to the autonomous mobile device. For example, it can be obtained by collecting data such as environmental images or point clouds around the autonomous mobile device.

[0107] The road information may include road type (such as highway, urban road, rural road), lane width, road surface condition (such as wet, potholed, icy), road signs and markings (such as speed limit signs, stop signs, lane lines), traffic signal status, intersection information (such as crossroads, roundabouts), and terrain information (such as slope, curve radius). For example, it can be obtained through navigation or map data.

[0108] For example, an electronic device can obtain driving environment information collected during the driving of a smart driving vehicle. The driving environment information includes: the driving speed of the smart driving vehicle is 60 km / h, the driving speed of vehicle A is 30 km / h, vehicle A and the smart driving vehicle are driving in the same lane (the current lane), and vehicle A is driving in front of the smart driving vehicle. There are no other traffic participants in the left lane of the current lane; both the current lane and the left lane are urban roads.

[0109] S302. Obtain voice data input by the user through an audio collection device in the autonomous mobile device.

[0110] In this step, the electronic device can obtain voice data input by the user through an audio collection device in the autonomous mobile device.

[0111] Optionally, the voice data input by the user can be obtained through a built-in microphone or an external microphone set in the autonomous mobile device. Among them, the voice data includes control instructions and / or driving style instructions;

[0112] The control instructions include at least one of the following: increasing the driving speed, decreasing the driving speed, overtaking the target ahead, bypassing obstacles, selecting a target road, entering the auxiliary lane, and docking at the target area;

[0113] For example, when the user is not satisfied with the driving speed of the smart driving vehicle, the user can issue an instruction to increase the driving speed or decrease the driving speed; when there is a vehicle with a slow driving speed in front of the smart driving vehicle, the user can issue an instruction to overtake the vehicle ahead; when there are temporary obstacles on the road (such as a construction area or a parked vehicle), the user can issue an instruction to bypass the obstacles; when the user hopes to change the current route to select a faster or more preferred road, the user can issue an instruction to select a target road; when the user hopes that the vehicle enters the auxiliary lane of the highway for the upcoming exit or for a slower driving speed, the user can issue an instruction to enter the auxiliary lane; when the user hopes to dock at a specific location (such as a store entrance or a parking lot), the user can issue an instruction to dock at the target area.

[0114] The driving style instructions include at least one of the following: the user's driving preference, the state of the autonomous mobile device, the user state, and the driving task time limit requirement.

[0115] Regarding the user's driving preference, it can refer to the personalized requirements and habits of the user for the driving experience. For example, the driving style instruction can be: Please drive smoothly to the destination.

[0116] Regarding the state of the autonomous mobile device, it can refer to the current operating state and health status of the autonomous mobile device. For example, the driving style instruction can be: The battery power is running low.

[0117] Regarding the user state, it may include the user's current physiological and psychological conditions, such as fatigue level, attention concentration, emotional state, etc. For example, the driving style instruction can be: I'm tired.

[0118] Regarding the driving task time limit requirement, it may refer to the time limit or deadline requirement related to the current driving task. For example, the driving style instruction can be: The time is a bit tight. Please drive safely and quickly.

[0119] For example, through the built-in microphone in the intelligent driving vehicle, the voice data 1 input by the user can be obtained as: Overtake the vehicle A ahead.

[0120] Another example, through the built-in microphone in the intelligent driving vehicle, the voice data 2 input by the user can be obtained as: Please drive smoothly to the destination.

[0121] S303. Based on the vision-language large model, process the voice data input by the user and the pre-obtained driving environment information to obtain the target feature vector.

[0122] In this step, based on the vision-language large model trained by the multi-modal model, the voice data input by the user and the pre-obtained driving environment information can be processed to obtain the target feature vector. Among them, the target feature vector is used to represent at least one of the following features: the driving intention feature of the user, the category feature of the driving scene, the driving style feature, and the state feature of the autonomous mobile device.

[0123] The driving intention feature of the user may refer to the driving goal or plan of the user at a specific moment. For example, the user may hope to overtake, change lanes, or exit the highway at a certain exit.

[0124] The category features of the driving scene may include but are not limited to: overtaking, detouring, accelerating, decelerating, or changing lanes.

[0125] The driving style features may include but are not limited to: aggressive style, sports style, cautious style, and smooth style.

[0126] Among them, each driving style feature can correspond to a high-dimensional style vector, and the style vector can be represented by the target feature vector. Each style vector represents a specific driving preference.

[0127] For example, style vector 1 is used to represent the smooth style. Even at the expense of efficiency, it still gives priority to pursuing a comfortable driving experience; style vector 2 is used to represent the sports style, striving to achieve efficient driving on the premise of ensuring safety; style vector 3 is used to represent the cautious style, which can be the behavior of choosing to yield to ensure safety when facing the situation of being cut in; style vector 4 is used to express the aggressive style, which can be still hoping to maintain its own road rights without actively yielding on the premise of ensuring safety.

[0128] It should be understood that the above style vectors are only examples and do not cover all possible driving style characteristics. In actual applications, driving style characteristics can be adjusted and extended according to different driving environments, vehicle types, and personal preferences. Each style vector quantifies the user's driving behavior tendency or driving preference through specific numerical values in a multi-dimensional feature space, thereby providing a basis for optimizing the personalized driving experience. This representation method of high-dimensional style vectors can help develop more intelligent driving assistance systems, enabling them to better adapt to the needs and expectations of different drivers.

[0129] The state characteristics of the autonomous mobile device can include being in preparation, in progress, and completed.

[0130] For example, based on a vision-language large model, processing the voice data "overtake the vehicle A ahead" input by the user and the pre-acquired driving environment information can obtain a target feature vector 1, which is used to represent that the driving intention feature of the user is to overtake the vehicle A ahead, the category feature of the driving scenario is overtaking, and the state feature of the autonomous mobile device is in preparation.

[0131] Another example is that based on a vision-language large model, processing the voice data "please drive smoothly to the destination" input by the user and the pre-acquired driving environment information can obtain a target feature vector 2, which is used to represent the driving style feature as a smooth style.

[0132] S304. Process the target feature vector through an assisted driving model to generate a driving decision.

[0133] In this step, the assisted driving model can be an end-to-end model, that is, it can directly output driving decisions (such as driving trajectories, etc.) based on the input perception data (such as image or point cloud data, etc.). An assisted driving model trained based on a deep learning model can process the target feature vector output by the vision-language large model to generate a driving decision. Among them, the driving decision includes at least one of the following: planning a driving path, speed adjustment, overtaking, obstacle avoidance, entering a target lane, entering a secondary road, lane keeping, parking, and emergency braking.

[0134] For example, based on the assisted driving model, processing the target feature vector 1 generates driving decisions including a speed adjustment decision and an overtaking decision.

[0135] In an alternative embodiment, the target feature vector obtained based on the vision-language large model can be further processed by an assisted driving model, and the generated driving decision can correspond to a driving style (or driving preference), that is, each driving style can have at least one driving decision. That is, through the assisted driving model, the target feature vector (also the style vector) can be recognized, and at least one driving decision of the driving style corresponding to the target feature vector can be determined.

[0136] For example, based on the assisted driving model, processing the target feature vector 2 can generate driving decisions corresponding to the smooth style, and the driving decisions corresponding to the smooth style can include speed adjustment decisions and lane keeping decisions.

[0137] It should be understood that in the driving decision system of an autonomous mobile device, safety is the primary influencing factor. Therefore, when the user's voice command does not match the operations that can be safely performed in the current scenario, the intelligent driving system will give priority to ensuring safe driving.

[0138] For example, an intelligent driving vehicle is driving on a two-lane urban road and is in the rightmost lane. The voice data input by the user is "overtake the vehicle A ahead", but when the driving environment information shows that there are obstacles in the left lane of the rightmost lane, or there is a solid yellow line between the rightmost lane and the left lane, indicating that overtaking is prohibited. Although the user hopes to overtake the vehicle A ahead, the intelligent driving system (such as the assisted driving model) can, according to the requirements of safety and traffic regulations, output a lane keeping decision and temporarily not perform the overtaking operation. Among them, the lane keeping decision can include continuing to drive in the rightmost lane and maintaining a safe distance from the vehicle A ahead. At this time, the intelligent driving system can also prompt the reason why overtaking is not possible currently through voice, so as to prompt the user to pay attention to the driving environment and give a safe command.

[0139] S305. Control the autonomous mobile device to drive according to the driving decision.

[0140] In this step, the autonomous mobile device can be controlled to drive according to the driving decision generated by the assisted driving model.

[0141] Optionally, if the driving decision includes multiple operations (such as speed adjustment and overtaking are required simultaneously), these operations can be executed in parallel, and the specific execution order can be dynamically adjusted according to the real-time traffic conditions and safety.

[0142] For example, according to the speed adjustment decision included in the driving decision, the driving speed of the intelligent driving vehicle can be increased from 60 km / h to 80 km / h. At the same time, according to the overtaking decision in the driving decision, the intelligent driving vehicle can be controlled to change lanes to the left lane. After changing lanes, according to the increased driving speed of 80 km / h, quickly overtake the vehicle A.

[0143] For another example, according to the speed adjustment decision included in the driving decision, the driving speed of the intelligent driving vehicle can be reduced from 60 km / h to 30 km / h. At the same time, according to the lane keeping decision included in the driving decision, the intelligent driving vehicle can be controlled to follow the vehicle A in front and drive safely and smoothly.

[0144] For yet another example, according to the driving style instruction "The battery power is running low", a driving decision corresponding to an aggressive style can be generated. The driving decision can include a speed adjustment decision, and the speed adjustment decision can be to increase the driving speed to reach the destination before the battery runs out. According to the driving style instruction "I'm tired", a driving decision corresponding to a cautious style can be generated. The driving decision can include a parking decision, and the parking decision can automatically find the nearest safe parking area for the user to rest. According to the driving style instruction "Time is a bit tight, please drive safely and quickly", a driving decision corresponding to a sporty style can be generated. The driving decision can include a decision to drive into the target lane, and the target lane decision can be to select a faster lane or a lane more suitable for the current speed and route.

[0145] In an alternative implementation, the purpose of voice controlling an autonomous mobile device can also be achieved based on a simple language model and rule-based decision planning.

[0146] For example, the voice data input by the user is "Drive into the nearest gas station". Through the language model, "Drive into" can be recognized as the action instruction, and "the nearest gas station" as the target location. The rule-based decision planning module can combine the current state of the intelligent driving vehicle (such as the remaining fuel quantity) and environmental information (such as the current location, the location and distance of nearby gas stations) to formulate an optimal path from the current location to the nearest gas station.

[0147] S306. Push the execution information of the driving decision to the user.

[0148] In this step, during the process of controlling the autonomous mobile device to drive according to the driving decision, the execution information of the driving decision can be pushed to the user. Among them, the execution information includes: the current operation being executed by the autonomous mobile device, and / or the expected operation to be executed.

[0149] Optionally, the execution information of the driving decision can be displayed to the user through a visual interface and / or reminded to the user through voice prompts.

[0150] For example, on the in-vehicle display screen of the intelligent driving vehicle, the adjustment process of the driving speed from 60 km / h to 80 km / h can be dynamically displayed; when the vehicle is performing an overtaking operation, the in-vehicle display screen can highlight the current location of the vehicle, the target lane, and the expected path to overtake vehicle A.

[0151] Meanwhile, voice prompts can provide users with real-time information. For example: "Please note that the driving speed of the intelligent driving vehicle is increasing to 80 km / h", or "Prepare to change lanes to the left lane to overtake the vehicle ahead."

[0152] Optionally, when the user's voice command does not match the operations that can be safely executed in the current scenario, the intelligent driving system will prioritize safe driving and push the execution information of the driving decision to the user through voice prompts. For example, the user is informed by voice that overtaking is not safe because there are obstacles or a solid yellow line in the left lane, and the vehicle is staying in the current lane and driving at a safe distance.

[0153] In the embodiments of the present application, the electronic device can obtain the driving environment information collected by the autonomous mobile device during movement, and obtain the voice data input by the user through the audio collection device. The vision-language large model processes the voice data and the driving environment information to generate a target feature vector. The assisted driving model can process the target feature vector to generate a driving decision. Furthermore, the driving of the autonomous mobile device can be controlled according to the driving decision, and information about the execution of the driving decision can be pushed to the user. In the above process, based on the general understanding ability of the vision-language large model, the input is no longer limited to a preset driving decision interaction interface, but can be input through natural language dialogue. The vision-language large model can combine this input with the actual driving scenario to generate a feature vector for the assisted driving model to use. The assisted driving model can generate a driving decision according to the feature vector and execute corresponding actions, thereby improving the flexibility of the intelligent driving system to meet the personalized needs of users.

[0154] Furthermore, in the method for generating a driving decision provided by the embodiments of the present application, the electronic device generates a feature vector through the user's voice input by using the understanding ability of the vision-language large model. Then, the assisted driving model processes the feature vector, achieving the purpose of the user controlling the behavior of the autonomous mobile device through voice. This method comprehensively improves the human-machine interaction method of intelligent driving, which is no longer limited to traditional joystick buttons or a specified number of interfaces, but supports natural language interaction in the whole scene, greatly improving the usability of intelligent driving products and the flexibility of user use.

[0155] In a possible design, based on any of the above embodiments, the method for generating a driving decision may further include: receiving a data processing model sent by the server and deploying it.

[0156] On the one hand, after receiving the data processing model sent by the server, the autonomous mobile device can deploy it in the intelligent driving system of the autonomous mobile device, where the data processing model can be trained based on a reinforcement learning model, or trained based on a multi-modal model and a deep learning model.

[0157] On the other hand, in an intelligent driving system, the update and optimization of the data processing model are important links to improve the system performance and adaptability. By receiving the new data processing model sent by the server, the autonomous mobile device can obtain the latest algorithms and model updates in a timely manner. These updates can include improved vision-language large models, assisted driving models, or other related algorithm modules.

[0158] In an alternative embodiment, the server can regularly push the newly trained and optimized data processing model to the autonomous mobile device. The new data processing model can be updated based on the latest driving data, user feedback, or technological progress. After receiving the new model, the autonomous mobile device can automatically deploy it to replace or enhance the existing model.

[0159] In another alternative embodiment, after receiving the new data processing model, the autonomous mobile device can display it on the visualization interface so that the user can understand and choose whether to apply these updates.

[0160] For example, after the electronic device of the autonomous mobile device receives the new data processing model, it can display the relevant information of the new model on the in-vehicle display screen, such as the version number of the model, the updated content, the improvement points, and the expected performance improvement. This helps the user understand the advantages and changes of the new model, and the user can deploy the new model through the installation option provided on the in-vehicle display screen.

[0161] In the embodiments of the present application, by deploying an advanced data processing model, the autonomous mobile device can efficiently process voice data and driving environment information, generate driving decisions, and control the autonomous mobile device to drive. In addition, by regularly receiving and deploying the latest data processing model sent by the server, the autonomous mobile device can continuously optimize its performance and adaptability, so as to maintain efficient operation under changing driving conditions and provide a smoother and more personalized driving experience for users.

[0162] In the above Figure 2 or Figure 3 Based on the above-described embodiments, below, in combination with Figure 4 , the training process of the data processing model will be described in detail. The training process of the data processing model can be performed on the server.

[0163] Figure 4 This is a schematic flowchart of the first embodiment of the model training method for driving decisions provided by the present application. Please refer to Figure 4 , the data processing model is trained based on a multi-modal model and a deep learning model. The method may include:

[0164] S401. Obtain multiple first training samples, where each first training sample includes data to be processed and a driving scenario label corresponding to the data to be processed.

[0165] In this step, the server can obtain multiple first training samples. Each training sample includes data to be processed and a driving scenario label corresponding to the data to be processed. The data to be processed includes voice data input by the user and driving environment information, and the driving environment information includes at least one of the following: motion state information of the autonomous mobile device, motion state information of surrounding traffic participants, and road information.

[0166] Optionally, the driving scenario label includes the category of the driving scenario and the state of the autonomous mobile device. The category of the driving scenario includes, but is not limited to: overtaking, bypassing, accelerating, decelerating, or lane changing; the state of the autonomous mobile device includes, but is not limited to: in preparation, in execution, or completed.

[0167] For example, scenario label 1 can be that the intelligent driving vehicle is preparing for overtaking, scenario label 2 can be that the intelligent driving vehicle is accelerating, and scenario label 2 can be that the intelligent driving vehicle has completed lane changing.

[0168] Optionally, the server can obtain the driving environment information of the autonomous mobile device through multiple methods such as mobile application upload, Application Programming Interface (API), cloud storage service, Internet of Things platform, data synchronization mechanism, or third-party service integration.

[0169] Through the scenario annotation system, corresponding voice data can be added to each driving environment information. Further, through the scenario annotation system, a common scenario label can also be added to each driving environment information and the voice data corresponding to each driving environment.

[0170] For example, the server can obtain multiple first training samples including first training sample 1. First training sample 1 can include driving environment information 1, driving environment information 1 corresponds to voice data 1, and driving environment information 1 and voice data 1 correspond to a common scenario label 1. Among them, driving environment information 1 can include: the driving speed of the intelligent driving vehicle is 60 kilometers per hour (km / h), the driving speed of the vehicle A ahead is 30 km / h, and there are no other vehicles in the left lane of the current lane; voice data 1 can include: overtaking the vehicle A ahead, and scenario label 1 can include: the intelligent driving vehicle is preparing for overtaking.

[0171] S402. Train a multimodal model based on the multiple first training samples to obtain a vision-language large model.

[0172] In this step, according to multiple first training samples, using the scene labels included in the samples (including the categories of driving scenarios and the states of autonomous mobile devices), the voice data input by the user, and the driving environment information, the multimodal model can be trained to obtain a vision-language large model.

[0173] Optionally, by training the multimodal model, the multimodal model can obtain the general behavior understanding ability for intelligent driving scenarios, and realize the output of the feature vector corresponding to the scene label according to the input of natural language, driving environment information (such as panoramic area images, driving speed, trajectory), etc.

[0174] Regarding the feature vector, it can be the intermediate result (implicit feature vector) output from the target layer of the multimodal model. The target layer can be the feature processing layer with the richest semantics in the multimodal model, or the feature processing layer at a preset number of layers away from the output layer.

[0175] Figure 5 It is a schematic structural diagram of a multimodal model provided by an embodiment of the present application. Please refer to Figure 5 , the multimodal model can include an input layer, a convolutional layer A, a pooling layer A, a convolutional layer B, a pooling layer B, and an output layer. Among them, the input layer is used to input voice data and driving environment information, and the output layer is used to output the scene label corresponding to the voice data and driving environment information. If the convolutional layer A is the feature processing layer with the richest semantics in the multimodal model, the feature vector corresponding to the scene label can be output through the convolutional layer A.

[0176] Optionally, if the feature processing layer at a preset number of layers of 2 away from the output layer is selected as the target layer, the target layer can be the convolutional layer B, and then the feature vector corresponding to the scene label can be output through the convolutional layer B.

[0177] S403. Construct multiple second training samples according to the feature vector corresponding to each first training sample and the driving decision corresponding to each first training sample. Each second training sample includes the feature vector corresponding to a first training sample and the driving decision corresponding to the first training sample.

[0178] In this step, after obtaining the feature vector corresponding to each first training sample through the multimodal model, multiple second training samples can be constructed according to the feature vector corresponding to each first training sample and the driving decision corresponding to each first training sample. Each second training sample includes the feature vector corresponding to a first training sample and the driving decision corresponding to the first training sample.

[0179] For example, the feature vector 1 is output by the most semantically rich feature processing layer of the multimodal model. The first training sample 1 used to obtain the feature vector 1 can have a corresponding driving decision 1. Based on the feature vector 1 and the driving decision 1, the second training sample 1 can be constructed. Among them, the driving decision 1 can be the ego vehicle trajectory when the intelligent driving vehicle executes an overtaking decision.

[0180] S404. Train the deep learning model according to multiple second training samples to obtain an assisted driving model.

[0181] In this step, the deep learning model can be trained according to the constructed multiple second training samples to obtain an assisted driving model. Optionally, the deep learning model can be a learning model adopting a convolutional neural network structure.

[0182] In another alternative implementation, the assisted driving model can also be obtained by training an Artificial Intelligence (AI) model. By training the AI model, the obtained assisted driving model can effectively adapt to the requirements of driving tasks and improve its decision-making ability and performance in the actual driving environment.

[0183] S405. Perform combined processing on the vision-language large model and the assisted driving model to obtain a data processing model.

[0184] In this step, by combining the trained vision-language large model with the assisted driving model, a comprehensive data processing model can be constructed. This model can generate corresponding driving decisions based on the voice data input by the user and the pre-acquired driving environment information.

[0185] In an embodiment of the present application, multiple first training samples can be obtained, and each sample includes data to be processed and its corresponding driving scenario label. By using the multiple first training samples to train a multimodal model, a vision-language large model can be obtained. Then, according to the feature vectors of each first training sample and the corresponding driving decisions, multiple second training samples are constructed, and each second training sample includes a feature vector and the corresponding driving decision. By using the multiple second training samples to train a deep learning model, an assisted driving model can be obtained. Finally, by combining the vision-language large model and the assisted driving model, a data processing model can be generated. In the above process, through the joint training of the multimodal model and the deep learning model, the feature vectors generated by the multimodal model are transmitted to the deep learning model. By learning the correlation between these features and the driving trajectory, the deep learning model can master the ability to output corresponding driving behaviors and trajectories according to the feature vectors provided by the multimodal model. This joint training can improve the model's understanding and decision-making ability for complex scenarios, enabling it to make more accurate and reliable judgments based on the user's voice data and the changing driving environment.

[0186] Figure 6 It is a schematic flowchart of the second embodiment of the model training method for driving decisions provided by the present application. Please refer to Figure 6 , the data processing model is trained based on a reinforcement learning model, and the method may include:

[0187] S601. Obtain multiple training samples, and each training sample includes data to be processed and the driving decision corresponding to the data to be processed.

[0188] In this step, the server can obtain multiple training samples, and each training sample includes data to be processed and the driving decision corresponding to the data to be processed. The data to be processed includes voice data input by the user and driving environment information, and the driving environment information includes at least one of the following: motion state information of the autonomous mobile device, motion state information of surrounding traffic participants, and road information.

[0189] For example, the server can obtain a plurality of training samples including training sample 1 and training sample 2. Training sample 1 can include driving environment information 1, voice data 1 corresponding to driving environment information 1, and driving decision 1 corresponding to driving environment information 1 and voice data 1. Among them, driving environment information 1 can include: the driving speed of the intelligent driving vehicle is 60 kilometers per hour (km / h), the driving speed of the vehicle A ahead is 30 km / h, and there are no other vehicles in the left lane of the current lane; voice data 1 can include: overtaking vehicle A ahead, and driving decision 1 can be an overtaking decision; training sample 2 can include driving environment information 2, voice data 2 corresponding to driving environment information 2, and driving decision 2 corresponding to driving environment information 2 and voice data 2; driving environment information 2 can include: the driving speed of the intelligent driving vehicle is 20 km / h, and there is a traffic cone in front of the intelligent driving vehicle. Voice data 2 can include: bypassing the obstacle traffic cone, and driving decision 2 can be a decision to bypass the obstacle.

[0190] S602. Train the reinforcement learning model according to the multiple training samples to obtain a data processing model.

[0191] In this step, the reinforcement learning model can be trained according to the multiple training samples to obtain a data processing model. Optionally, the reinforcement learning model can be a learning model adopting a deep Q-network structure.

[0192] In addition, the data processing model can also be obtained by training an AI model.

[0193] In the embodiment of the present application, by obtaining a plurality of training samples, each sample includes the data to be processed and its corresponding driving decision, and the data to be processed includes the voice data input by the user and the driving environment information. The server can train the reinforcement learning model to generate a data processing model. The generated data processing model can accurately capture the user's intention and generate a driving decision that meets the user's personalized needs according to the driving environment information, improving the user's driving experience.

[0194] In a possible design, based on the embodiment of any of the above model training methods, the model training method for driving decisions can further include: sending the data processing model to at least one autonomous mobile device.

[0195] On the one hand, the server can send the data processing model obtained by training based on the reinforcement learning model, or the data processing model obtained by training based on the multi-modal model and the deep learning model to the autonomous mobile device. After receiving the data processing model sent by the server, the autonomous mobile device can deploy it in the intelligent driving system of the autonomous mobile device.

[0196] On the other hand, the server can send the updated and optimized new data processing model to at least one autonomous mobile device through a preset cycle. For example, a fixed time interval (such as 1 month) can be set in the server to regularly send the updated and optimized data processing model to at least one autonomous mobile device.

[0197] In the embodiments of the present application, by deploying the trained data processing model in the autonomous mobile device, the autonomous mobile device can be enabled to have an understanding ability similar to that of a user, including recognizing the relationships between various elements in the road environment and understanding how to drive in these driving environments. In addition, the autonomous mobile device can also evaluate the impact of various sudden anomalies on its operation, so as to make more accurate and safe driving decisions in complex and dynamic traffic environments.

[0198] Figure 7 Schematic diagram of the training and application examples of the data processing model provided by the embodiments of the present application. Please refer to Figure 7 , the data processing model is trained based on a multimodal model and a deep learning model, and the autonomous mobile device is an intelligent driving vehicle, including two processes: training and inference.

[0199] During the training process, based on the scene annotation system, the categories of driving scenes, the states of the intelligent driving vehicle, and the behavior of the vehicle itself can be manually annotated to obtain driving scene labels. The driving scene labels can include the categories of driving scenes and the states of the intelligent driving vehicle. Among them, the categories of driving scenes can include overtaking the vehicle on the far right, bypassing an electric vehicle, accelerating, changing lanes to the left, driving on the ramp, driving through the Electronic Toll Collection (ETC) lane; the states of the intelligent driving vehicle can include being in preparation, being in execution, or having been completed; the behavior of the vehicle itself can be obtained through the actual driving trajectory generated when the user drives the vehicle.

[0200] By training the multimodal model with natural language input, driving environment information, and scene labels, a vision-language large model can be obtained. Among them, the natural language input can be the voice data input by the user, and the driving environment information can include the image data collected by the camera of the intelligent driving vehicle.

[0201] Furthermore, according to the feature output (feature vector) of the vision-language large model and the vehicle's own trajectory, the deep learning model can be trained to obtain an assisted driving model.

[0202] During the inference process, the trained vision-language large model and the assisted driving model can be deployed in the intelligent driving system of the autonomous mobile device. When the intelligent driving system is working, it can process the natural language input and the pre-acquired driving environment information to obtain driving decisions that meet the personalized needs of users, achieving the purpose of controlling the intelligent driving vehicle with natural language input.

[0203] It should be noted that all the training data in the training samples come from high-quality manual annotations and human driving trajectories. Therefore, there is a strong correlation among the voice data, the real vehicle scenario, and the execution trajectory. After the data processing model is deployed in the intelligent driving system of the autonomous mobile device, even if the user inputs voice data that conflicts with the scenario, the intelligent driving system can automatically ignore the voice data instead of executing against the physical laws. For example, on a section without a service road, when the user inputs the voice data "Drive into the service road", the intelligent driving system will not control the autonomous mobile device to rush towards the road edge; at a traffic light intersection or when the vehicle in front decelerates, when the user inputs the voice data "Increase the driving speed", the intelligent driving system will not control the autonomous mobile device to force an acceleration against safety and traffic regulations.

[0204] In the example of this application, the vision-language large model can be trained by using scene labels, the categories of driving scenarios, and the states of intelligent driving vehicles, so that the vision-language large model can obtain the general behavior understanding ability for autonomous driving scenarios; according to the input of natural language and driving environment information, the feature output corresponding to the scene label is output. The multi-modal model and the deep learning model are jointly trained. During the training process, the feature output of the multi-modal model can be given to the deep learning model, and the deep learning model learns the correlation between the feature output of the multi-modal model and the driving trajectory, so that the assisted driving model trained based on the deep learning model can output corresponding behaviors and trajectories according to the feature output of the vision-language large model.

[0205] Furthermore, by deploying the vision-language large model and the assisted driving model on the vehicle side, multi-modal interaction control based on natural voice commands can be realized. The user's needs are collected through the voice interaction entrance, the environment perception and semantic understanding are carried out with the help of the vision-language large model, and after feature fusion, the assisted driving model is driven to generate corresponding driving decisions, thus completing the full-link closed loop of "voice-vision-behavior". This innovative interaction paradigm breaks through the limitations of traditional lever interactions, realizes full-scenario natural language interaction, enhances the user's freedom of controlling the intelligent driving system through the intuitiveness of natural language interaction, and improves the usability and flexibility of intelligent driving products.

[0206] Figure 8 This is the structural schematic diagram of the first embodiment of the driving decision generation device provided by this application. Please refer to Figure 8 , the driving decision generation device 10 includes:

[0207] The first processing module 11 is configured to process the voice data input by the user and the pre-acquired driving environment information based on a data processing model to generate a driving decision; the data processing model is obtained by training based on a reinforcement learning model, or is obtained by training based on a multimodal model and a deep learning model;

[0208] The second processing module 12 is configured to control the autonomous mobile device to travel according to the driving decision.

[0209] The driving decision generation device provided by the embodiments of the present application can execute the technical solutions shown in the above method embodiments, and its implementation principle and beneficial effects are similar, and will not be elaborated here.

[0210] In a possible implementation manner, the first processing module 11 is specifically configured to:

[0211] Process the voice data input by the user and the pre-acquired driving environment information based on a vision-language large model to obtain a target feature vector, where the target feature vector is used to represent at least one of the following features: the driving intention feature of the user, the category feature of the driving scene, the driving style feature, and the state feature of the autonomous mobile device;

[0212] Process the target feature vector through an assisted driving model to generate a driving decision;

[0213] Wherein, the data processing model is obtained by combining the vision-language large model and the assisted driving model.

[0214] In a possible implementation manner, the second processing module 12 is further configured to:

[0215] Push the execution information of the driving decision to the user, where the execution information includes: the current operation being performed by the autonomous mobile device, and / or the expected operation to be performed.

[0216] In a possible implementation manner, the first processing module 11 is further configured to:

[0217] Obtain the driving environment information collected during the movement of the autonomous mobile device, where the driving environment information includes at least one of the following: the motion state information of the autonomous mobile device, the motion state information of surrounding traffic participants, and road information;

[0218] Obtain the voice data input by the user through the audio collection device in the autonomous mobile device.

[0219] In a possible implementation manner, the voice data includes a control instruction, and / or a driving style instruction;

[0220] The control instructions include at least one of the following: increasing the driving speed, decreasing the driving speed, overtaking the target ahead, bypassing an obstacle, selecting a target road, entering a secondary road, and docking at a target area;

[0221] The driving style instructions include at least one of the following: the driving preferences of the user, the status of the autonomous mobile device, the user status, and the driving task time limit requirements.

[0222] In a possible implementation manner, the driving decision includes at least one of the following: planning a driving route, speed adjustment, overtaking, obstacle avoidance, entering a target lane, entering a secondary road, lane keeping, parking, and emergency braking.

[0223] Figure 9 This is a schematic structural diagram of the second embodiment of the driving decision generation device provided by this application. Based on Figure 8 the shown embodiment, please refer to Figure 9 , the driving decision generation device 10 further includes:

[0224] A deployment module 13, configured to receive a data processing model sent by a server and deploy the data processing model.

[0225] The driving decision generation device provided by the embodiment of this application can execute the technical solutions shown in the above method embodiment, and its implementation principle and beneficial effects are similar, which will not be elaborated here.

[0226] Figure 10 This is a schematic structural diagram of the first embodiment of the model training device for driving decisions provided by this application. Please refer to Figure 10 , the model training device 20 for driving decisions includes:

[0227] An acquisition module 21, configured to acquire a plurality of first training samples, each of which includes data to be processed and a driving scenario label corresponding to the data to be processed; the data to be processed includes voice data input by the user and driving environment information, and the driving environment information includes at least one of the following: motion state information of the autonomous mobile device, motion state information of surrounding traffic participants, and road information;

[0228] A first training module 22, configured to train a multi-modal model according to the plurality of first training samples to obtain a vision-language large model;

[0229] A construction module 23, configured to construct a plurality of second training samples according to the feature vector corresponding to each first training sample and the driving decision corresponding to each first training sample, and each second training sample includes the feature vector corresponding to a first training sample and the driving decision corresponding to the first training sample;

[0230] The second training module 24 is configured to train a deep learning model based on the multiple second training samples to obtain an assisted driving model;

[0231] The combination module 25 is configured to perform a combination process on the vision-language large model and the assisted driving model to obtain the data processing model.

[0232] The model training device for driving decision-making provided in the embodiments of the present application can execute the technical solutions shown in the above method embodiments, and the implementation principles and beneficial effects are similar, and will not be elaborated here.

[0233] In a possible implementation manner, the driving scenario label includes the category of the driving scenario and the state of the autonomous mobile device. The category of the driving scenario is any one of the following: overtaking, detouring, accelerating, decelerating, or lane changing; the state of the autonomous mobile device includes in preparation, in execution, or completed.

[0234] In a possible implementation manner, the data processing model is trained based on a reinforcement learning model;

[0235] The obtaining module 21 is further configured to obtain a plurality of training samples, each training sample including data to be processed and the driving decision corresponding to the data to be processed; the data to be processed includes voice data input by a user and driving environment information, and the driving environment information includes at least one of the following: motion state information of the autonomous mobile device, motion state information of surrounding traffic participants, and road information;

[0236] The first training module 22 is further configured to train a reinforcement learning model based on the plurality of training samples to obtain the data processing model.

[0237] Figure 11 This is a schematic structural diagram of the second embodiment of the model training device for driving decision-making provided by the present application. On the basis of the embodiment shown in Figure 10 please refer to Figure 11 , the model training device 20 for driving decision-making further includes:

[0238] The sending module 26 is configured to send the data processing model to at least one autonomous mobile device.

[0239] The model training device for driving decision-making provided in the embodiments of the present application can execute the technical solutions shown in the above method embodiments, and the implementation principles and beneficial effects are similar, and will not be elaborated here.

[0240] Figure 12 This is a schematic structural diagram of the electronic device provided in the embodiments of the present application. Please refer to Figure 12, the electronic device 30 may include a processor 31, a memory 32, and a communication interface 34. Exemplarily, the processor 31, the memory 32, and the communication interface 34 are interconnected via a bus 33.

[0241] The memory 32 stores computer-executable instructions;

[0242] The processor 31 executes the computer-executable instructions stored in the memory 32, such that the processor 31 executes the method for generating a driving decision provided in the above method embodiments. The implementation principle and beneficial effects are similar, and will not be elaborated here.

[0243] An embodiment of the present application provides an autonomous mobile device, including Figure 12 the electronic device shown, to implement the method for generating a driving decision in the above embodiments. The implementation principle and beneficial effects are similar, and will not be elaborated here.

[0244] Figure 13 It is a schematic structural diagram of the server provided by an embodiment of the present application. Please refer to Figure 13 , the server 40 may include a processor 41, a memory 42, and a communication interface 44. Exemplarily, the processor 41, the memory 42, and the communication interface 44 are interconnected via a bus 43.

[0245] The memory 42 stores computer-executable instructions;

[0246] The processor 41 executes the computer-executable instructions stored in the memory 42, such that the processor 41 executes the method for training a model for driving decision provided in the above method embodiments. The implementation principle and beneficial effects are similar, and will not be elaborated here.

[0247] An embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method described in the above method embodiments.

[0248] An embodiment of the present application provides a program product, including: a computer program, when the program product runs on a computer, it causes the computer to execute the method described in the above method embodiments.

[0249] An embodiment of the present application provides a computer program, when the computer program is executed by a processor, it is used to execute the method described in the above method embodiments.

[0250] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0251] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0252] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0253] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0254] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0255] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0256] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0257] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0258] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A method for generating driving decisions, characterized in that, The method includes: Based on a data processing model, processing the voice data input by the user and the pre-acquired driving environment information to generate a driving decision; the data processing model is trained based on a reinforcement learning model or is trained based on a multi-modal model and a deep learning model; Controlling the autonomous mobile device to travel according to the driving decision.

2. The method according to claim 1, wherein The processing of the voice data input by the user and the pre-acquired driving environment information based on the data processing model to generate a driving decision includes: Based on a vision-language large model, processing the voice data input by the user and the pre-acquired driving environment information to obtain a target feature vector, where the target feature vector is used to represent at least one of the following features: the driving intention feature of the user, the category feature of the driving scenario, the driving style feature, and the state feature of the autonomous mobile device; Processing the target feature vector through an assisted driving model to generate a driving decision; Among them, the data processing model is obtained by combining the vision-language large model and the assisted driving model.

3. The method according to claim 1 or 2, characterized in that, The method further includes: Pushing the execution information of the driving decision to the user, where the execution information includes: the current operation being performed by the autonomous mobile device and / or the expected operation to be performed.

4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Obtaining the driving environment information collected during the movement of the autonomous mobile device, where the driving environment information includes at least one of the following: the motion state information of the autonomous mobile device, the motion state information of surrounding traffic participants, and road information; Obtaining the voice data input by the user through the audio collection device in the autonomous mobile device.

5. The method according to any one of claims 1 to 4, characterized in that The voice data includes control instructions and / or driving style instructions; The control instructions include at least one of the following: increasing the driving speed, decreasing the driving speed, overtaking the target ahead, bypassing an obstacle, selecting a target road, entering a secondary road, and docking at a target area; The driving style instructions include at least one of the following: the driving preference of the user, the state of the autonomous mobile device, the user state, and the driving task time limit requirement.

6. The method according to any one of claims 1 to 5, characterized in that The driving decision includes at least one of the following: planning a driving path, speed adjustment, overtaking, obstacle avoidance, entering a target lane, entering a secondary road, lane keeping, parking, and emergency braking.

7. The method according to any one of claims 1-6, characterized in that, The method further includes: Receiving the data processing model sent by the server and deploying the data processing model.

8. A method for training a model for driving decision-making, characterized in that, The data processing model is trained based on a multi-modal model and a deep learning model, and the method includes: Obtaining a plurality of first training samples, where each first training sample includes data to be processed and a driving scenario label corresponding to the data to be processed; the data to be processed includes the voice data input by the user and the driving environment information, and the driving environment information includes at least one of the following: the motion state information of the autonomous mobile device, the motion state information of surrounding traffic participants, and road information; Training the multi-modal model according to the plurality of first training samples to obtain a vision-language large model; Construct a plurality of second training samples according to the feature vector corresponding to each first training sample and the driving decision corresponding to each first training sample. Each second training sample includes the feature vector corresponding to a first training sample and the driving decision corresponding to the first training sample; Train the deep learning model according to the plurality of second training samples to obtain an assisted driving model; Perform a combination process on the vision-language large model and the assisted driving model to obtain the data processing model.

9. The method according to claim 8, characterized in that, The driving scenario label includes the category of the driving scenario and the state of the autonomous mobile device. The category of the driving scenario is any one of the following: overtaking, detouring, accelerating, decelerating, or lane changing; the state of the autonomous mobile device includes in preparation, in execution, or completed.

10. The method according to claim 8 or 9, characterized in that If the data processing model is trained based on a reinforcement learning model, the method further includes: Obtain a plurality of training samples. Each training sample includes data to be processed and the driving decision corresponding to the data to be processed; the data to be processed includes voice data input by the user and driving environment information, and the driving environment information includes at least one of the following: motion state information of the autonomous mobile device, motion state information of surrounding traffic participants, and road information; Train the reinforcement learning model according to the plurality of training samples to obtain the data processing model.

11. The method according to claim 8 or 10, characterized in that, The method further includes: Send the data processing model to at least one autonomous mobile device.

12. A driving decision generation device, characterized in that The device includes: A first processing module, configured to process the voice data input by the user and the pre-acquired driving environment information based on a data processing model to generate a driving decision; the data processing model is trained based on a reinforcement learning model or based on a multi-modal model and a deep learning model; A second processing module, configured to control the autonomous mobile device to drive according to the driving decision.

13. A model training device for driving decision-making, characterized in that, The device includes: An acquisition module, configured to acquire a plurality of first training samples. Each first training sample includes data to be processed and the driving scenario label corresponding to the data to be processed; the data to be processed includes voice data input by the user and driving environment information, and the driving environment information includes at least one of the following: motion state information of the autonomous mobile device, motion state information of surrounding traffic participants, and road information; A first training module, configured to train a multi-modal model according to the plurality of first training samples to obtain a vision-language large model; A construction module, configured to construct a plurality of second training samples according to the feature vector corresponding to each first training sample and the driving decision corresponding to each first training sample. Each second training sample includes the feature vector corresponding to a first training sample and the driving decision corresponding to the first training sample; A second training module, configured to train a deep learning model according to the plurality of second training samples to obtain an assisted driving model; A combination module, configured to perform a combination process on the vision-language large model and the assisted driving model to obtain the data processing model.

14. An electronic device, characterized in that, Includes: A processor, a memory, and a communication interface; The memory stores computer execution instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.

15. An autonomous mobile device, characterized in that, An electronic device according to claim 14.

16. A server, characterized in that, Comprising: a processor, a memory, and a communication interface; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 8 to 11.

17. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the method according to any one of claims 1 to 11.

18. A program product, characterized in that, Comprising: a computer program, which, when the program product runs on a computer, causes the computer to execute the method according to any one of claims 1 to 11 above.

19. A computer program, characterized in that, When the computer program is executed by a processor, it is used to execute the method according to any one of claims 1 to 11 above.