Vehicle control method, device, vehicle, medium and program product
By combining in-vehicle camera data and driver voice data to generate autonomous driving control commands, the problem of insufficient understanding of computer vision autonomous driving technology in complex scenarios is solved, thereby improving the accuracy and efficiency of autonomous driving.
Patent Information
- Application Number
- CN202510180125.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-02-18
AI Technical Summary
Existing computer vision-based autonomous driving technologies may not fully understand or understand slowly in some scenarios, leading to reduced safety and efficiency of autonomous driving.
By combining scene images captured by vehicle cameras and driver voice data, target driving control commands are generated through large language models and visual language models. The scene descriptions expressed by voice data and driving commands are integrated to generate accurate autonomous driving control commands.
It improves the efficiency and accuracy of driving control command generation, enhances the vehicle's autonomous driving capabilities, and improves safety and autonomous driving processing capabilities, especially in complex driving scenarios such as passing other vehicles on narrow roads.
Smart Images

Figure CN119636811B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of automatic driving, and particularly relates to a vehicle control method and device, a vehicle, a medium and a program product. BACKGROUND
[0002] Automatic driving requires that a vehicle has sufficient understanding of a real-world environment. Although computer vision-based automatic driving technology has been greatly developed, so that the automatic driving can adapt to most driving scenarios, there are still some scenarios in which the computer vision-based automatic driving solution has insufficient understanding or slow understanding, thereby reducing the safety and efficiency of automatic driving. Therefore, the automatic driving technology in the related art still needs to be improved. SUMMARY
[0003] To overcome the problems in the related art, the present disclosure provides a vehicle control method and device, a vehicle, a medium and a program product.
[0004] According to a first aspect of an embodiment of the present disclosure, a vehicle control method is provided, which comprises:
[0005] obtaining environment data collected by a vehicle-mounted sensor;
[0006] in a case where it is determined according to the environment data that the vehicle is in a target driving scenario, obtaining a scene image collected by a vehicle-mounted camera and voice data of a driver;
[0007] generating a target driving control instruction according to the scene image and the voice data;
[0008] controlling the vehicle according to the target driving control instruction.
[0009] Optionally, the generating a target driving control instruction according to the scene image and the voice data comprises:
[0010] performing voice recognition on the voice data to obtain text data corresponding to the voice data;
[0011] performing intent recognition on the text data to obtain an intent type corresponding to the text data;
[0012] generating a target driving control instruction according to the scene image, the text data and the intent type.
[0013] Optionally, the generating a target driving control instruction according to the scene image, the text data and the intent type comprises:
[0014] In a case where the intention type is a scene description intention, performing scene extraction processing on the text data by a large language model to obtain first scene description text;
[0015] Performing scene extraction processing on the scene image and the first scene description text by a visual language model to obtain second scene description text;
[0016] Generating a target driving control instruction according to the second scene description text.
[0017] Optionally, the generating of the target driving control instruction according to the scene image, the text data, and the intention type comprises:
[0018] In a case where the intention type is a driving instruction intention, performing control instruction extraction processing on the text data by a large language model to obtain a first sub-driving control instruction;
[0019] Performing scene extraction processing on the scene image by a visual language model to obtain third scene description text;
[0020] Generating a second sub-driving control instruction according to the third scene description text;
[0021] Fusing the first sub-driving control instruction and the second sub-driving control instruction to obtain the target driving control instruction.
[0022] Optionally, the method further comprises:
[0023] In a case where it is determined according to the environment data that the vehicle is in a target driving scene, outputting prompt information, the prompt information being used to prompt the driver to perform scene description and / or driving instruction description by voice.
[0024] Optionally, the voice data is voice data of the driver within a preset time period after the prompt information is output.
[0025] According to a second aspect of the embodiments of the present disclosure, a vehicle control device is provided, and the device comprises:
[0026] A first acquisition module configured to acquire environment data collected by an on-board sensor;
[0027] A second acquisition module configured to, in a case where it is determined according to the environment data that the vehicle is in a target driving scene, acquire a scene image collected by an on-board camera and voice data of a driver;
[0028] A generation module configured to generate a target driving control instruction according to the scene image and the voice data;
[0029] a control module configured to automatically control driving of the vehicle according to the target driving control instruction.
[0030] According to a third aspect of embodiments of the present disclosure, a computer-readable storage medium is provided, which stores computer program instructions, and the computer program instructions, when executed by a processor, implement the steps of the vehicle control method provided in the first aspect of the present disclosure.
[0031] According to a fourth aspect of embodiments of the present disclosure, a vehicle is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the steps of the vehicle control method provided in the first aspect of the present disclosure when invoking the executable instructions stored on the memory.
[0032] According to a fifth aspect of embodiments of the present disclosure, a computer program product is provided, comprising a computer program, and the computer program, when executed by a processor, implements the steps of the vehicle control method mentioned in the first aspect of the present disclosure.
[0033] The technical solutions provided by the embodiments of the present disclosure can include the following beneficial effects:
[0034] First, the environment data collected by the vehicle-mounted sensor is acquired for driving scene judgment, and in the case that it is determined according to the environment data that the vehicle is in a target driving scene, the scene image collected by the vehicle-mounted camera and the voice data of the driver are acquired, then the target driving control instruction can be generated according to the scene image and the voice data, and then the vehicle can be automatically controlled according to the target driving control instruction. Since the voice data of the driver is more accurate in the current scene, and is also easier to understand than image data, therefore, by combining the scene image collected by the vehicle-mounted camera and the voice data of the driver to generate the target driving control instruction, not only the efficiency of generating the driving control instruction can be improved, but also the accuracy of generating the driving control instruction can be improved, and thus the automatic driving capability of the vehicle can be enhanced. In addition, in the case that it is determined that the vehicle is in a target driving scene, the scene image collected by the vehicle-mounted camera and the voice data of the driver are acquired, which can also guarantee the necessity of voice data acquisition and save resources.
[0035] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0036] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0037] Figure 1is a flowchart of a vehicle control method according to an example embodiment of the present disclosure.
[0038] Figure 2 is a structural diagram of a vehicle control system according to an example embodiment of the present disclosure.
[0039] Figure 3 is a structural block diagram of a vehicle control device according to an example embodiment of the present disclosure.
[0040] Figure 4 is a block diagram of a vehicle according to an example embodiment of the present disclosure. DETAILED DESCRIPTION
[0041] The example embodiments will be described in detail herein with reference to the accompanying drawings. In the following description, the same numbers refer to the same or similar elements unless otherwise represented. The implementations described in the following example embodiments do not represent all implementations consistent with the present disclosure. Instead, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0042] It should be noted that all actions of obtaining signals, information or data in the present application are carried out in compliance with the corresponding data protection regulations and policies of the country where the device is located, and with the authorization of the owner of the corresponding device.
[0043] Figure 1 is a flowchart of a vehicle control method according to an example embodiment, which can be applied to a vehicle terminal device alone, a cloud device alone, or part of the steps applied to a vehicle terminal device and the other part of the steps applied to a cloud device, as shown in Figure 1 The vehicle control method includes the following steps:
[0044] S110, obtaining environment data collected by a vehicle-mounted sensor.
[0045] It can be understood that the vehicle-mounted sensor can include a vehicle-mounted camera, a vehicle-mounted radar, an ultrasonic sensor, etc.
[0046] In the embodiments of the present disclosure, the environment data collected by the vehicle-mounted sensor can be obtained according to actual needs, and the driving scene in which the vehicle is currently located can be determined through the collected environment data.
[0047] For example, video and picture data can be captured by a vehicle-mounted camera, and the driving scene in which the vehicle is currently located can be recognized in combination with the sensor results of the radar, etc.
[0048] In some embodiments, the target driving scene can be a scene that is relatively difficult to understand or understand slowly, for example, a narrow road meeting scene, a narrow road overtaking scene, etc.
[0049] In some embodiments, the target driving scene can be a scene that is relatively difficult to understand or understand slowly, for example, a narrow road meeting scene, a narrow road overtaking scene, etc.
[0050] Here, the scene image can be one or more images captured by the vehicle-mounted camera, which can be obtained as needed.
[0051] In the embodiments of the present disclosure, in the case where it is determined according to the environmental data that the vehicle is in the target driving scene, the scene image captured by the vehicle-mounted camera and the voice data of the driver are further obtained. In this way, the necessity of voice data acquisition can be ensured.
[0052] S130, generating a target driving control instruction according to the scene image and the voice data.
[0053] S140, automatically driving the vehicle according to the target driving control instruction.
[0054] With the above method, the environmental data collected by the vehicle-mounted sensor is first obtained for driving scene judgment, and in the case where it is determined according to the environmental data that the vehicle is in the target driving scene, the scene image captured by the vehicle-mounted camera and the voice data of the driver are further obtained. Then, the target driving control instruction can be generated according to the scene image and the voice data, and then the vehicle can be automatically driven according to the target driving control instruction. Since the voice data of the driver is more accurate in the current scene, and is also easier to understand than image data, the target driving control instruction is generated by combining the scene image captured by the vehicle-mounted camera and the voice data of the driver. Not only can the efficiency of driving control instruction generation be improved, but also the accuracy of driving control instruction generation can be improved, thereby the vehicle automatic driving capability can be enhanced. In addition, in the case where it is determined that the vehicle is in the target driving scene, the scene image captured by the vehicle-mounted camera and the voice data of the driver are further obtained, which can ensure the necessity of voice data acquisition and save resources.
[0055] It should be noted that the data used in the vehicle control method of the embodiments of the present disclosure can be obtained by the vehicle-side device or the cloud-side device through direct or indirect means. Direct acquisition includes but is not limited to processing or calculating data by itself, and indirect acquisition includes but is not limited to obtaining data through data transmission process between other devices.
[0056] In some embodiments, in step S130, generating a target driving control instruction according to the scene image and the voice data can include the following steps:
[0057] perform speech recognition on the voice data to obtain text data corresponding to the voice data;
[0058] perform intent recognition on the text data to obtain an intent type corresponding to the text data;
[0059] generate a target driving control instruction according to the scene image, the text data, and the intent type.
[0060] In the embodiments of the present disclosure, the voice data of the driver can be various, for example, can include scene description, or can include driving instructions issued by the vehicle, etc. Exemplarily, the scene description is, for example, a front mountain road meeting, and the right side is a mountain, and the driving instruction is, for example, a front meeting, and a right side road along deceleration driving.
[0061] In this case, in order to further improve the accuracy of generating the target driving control instruction, the voice data can be first subjected to speech recognition to obtain text data corresponding to the voice data, and then the text data can be subjected to intent recognition to obtain an intent type corresponding to the text data, so that subsequently a target driving control instruction can be generated according to the scene image, the text data, and the intent type.
[0062] In some embodiments, the voice data can be subjected to speech recognition by an automatic speech recognition (ASR) model to obtain text data corresponding to the voice data.
[0063] In some embodiments, the text data can be subjected to intent recognition by a large language model to obtain an intent type corresponding to the text data.
[0064] The large language model (LLM) refers to a deep learning model trained using a large amount of text data, and can generate natural language text or understand the meaning of language text. The large language model can process various natural language tasks based on a prompt, such as text classification, question answering, dialogue, etc. The prompt represents a command or instruction, and is used to indicate that the large language model needs to perform an action or generate an output, i.e., to indicate what action the model should take or what output the model should generate when performing a specific task.
[0065] By adopting the above technical solutions, for voice data of different intent types, different methods can be used to generate a target driving control instruction, so that the accuracy and efficiency of generating the target driving control instruction can be improved.
[0066] In some embodiments, the intent type can include a scene description intent, in which case, generating the target driving control instruction according to the scene image, the text data, and the intent type can include the following steps:
[0067] In the case where the intent type is a scene description intent, performing scene extraction processing on the text data by a large language model to obtain first scene description text;
[0068] Performing scene extraction processing on the scene image and the first scene description text by a vision language model to obtain second scene description text;
[0069] Generating the target driving control instruction according to the second scene description text.
[0070] Vision-Language Models (VLM) are a class of multi-modal artificial intelligence systems that can simultaneously process and understand image and text information. By combining visual data (such as images or video frames) with corresponding text descriptions, they can perform well in a variety of tasks, such as image caption generation, visual question answering, and image-text retrieval.
[0071] In the embodiments of the present disclosure, in the case where the identified intent type corresponding to the text data is a scene description intent, first, scene extraction processing is performed on the text data by a large language model to obtain first scene description text, then scene extraction processing is performed on the scene image and the first scene description text by a vision language model to obtain second scene description text, and finally, the target driving control instruction can be generated according to the second scene description text output by the vision language model.
[0072] By integrating the scene description expressed by the voice data, the above-mentioned solution enables a deep understanding of the real-world environment, improves the accuracy of scene description, and thus improves the accuracy of driving control instruction generation and enhances the vehicle automatic driving capability.
[0073] In some embodiments, the intent type can also include a driving instruction intent, in which case, generating the target driving control instruction according to the scene image, the text data, and the intent type can include the following steps:
[0074] In the case where the intent type is a driving instruction intent, performing control instruction extraction processing on the text data by a large language model to obtain a first sub-driving control instruction;
[0075] Performing scene extraction processing on the scene image by a vision language model to obtain third scene description text;
[0076] Generating a second sub-driving control instruction according to the third scene description text;
[0077] The first sub-driving control instruction and the second sub-driving control instruction are fused to obtain a target driving control instruction.
[0078] In the case where it is identified that the intention type corresponding to the text data is a driving instruction intention, in the embodiment of the present disclosure, on the one hand, the text data can be directly processed by the large language model to obtain the first sub-driving control instruction, and on the other hand, the scene image can be processed by the visual language model to obtain the third scene description text, and the second sub-driving control instruction can be generated according to the third scene description text. Finally, the first driving control instruction issued by the driver and the second driving control instruction automatically generated can be fused to obtain the target driving control instruction used for automatically driving the vehicle.
[0079] In some embodiments, when the first sub-driving control instruction and the second sub-driving control instruction are fused to obtain the target driving control instruction, the content of the first sub-driving control instruction can be more referenced.
[0080] By fusing the driving instruction expressed by the voice data when generating the target driving control instruction, the above scheme can realize that the driver can guide the automatic driving decision in special cases, thereby improving the efficiency of generating the target driving control instruction and the accuracy of generating the driving control instruction, and further enhancing the automatic driving capability of the vehicle.
[0081] In some embodiments, the vehicle control method of the embodiment of the present disclosure can further include the following steps:
[0082] In the case where it is determined according to the environment data that the vehicle is in the target driving scene, prompt information is output, and the prompt information is used to prompt the driver to describe the scene and / or the driving instruction through voice.
[0083] In the case where it is determined according to the environment data that the vehicle is in the target driving scene, in the embodiment of the present disclosure, prompt information can be output to timely prompt the driver to describe the scene and / or the driving instruction through voice. In this way, the timeliness of voice data acquisition can be improved.
[0084] In some embodiments, the output prompt information is also used to prompt the driver that the target driving scene will be encountered, so that the driver can prepare in advance.
[0085] In some embodiments, the voice data is the voice data of the driver within a preset time period after the prompt information is output.
[0086] In the embodiments of the present disclosure, if all voice data of the driver is collected and used as voice data for generating the target driving control instruction, on the one hand, the workload of collection is large, and on the other hand, the voice data at other times may disturb the normal automatic driving decision. Therefore, only the voice data of the driver in a preset time period after the output of the prompt information is collected as voice data for generating the target driving control instruction.
[0087] In some embodiments, a driving route trajectory corresponding to the target driving scene can be generated according to the target driving control instruction, and the driving route trajectory is displayed to the driver.
[0088] The execution process of the vehicle control method of the embodiments of the present disclosure will be described below in conjunction with the structure schematic diagram of the vehicle control system shown in Figure 2
[0089] As shown in Figure 2 The vehicle control system includes a vehicle-side device and a cloud-side device, wherein the vehicle-side device is deployed with a driving scene recognition model, a voice recognition model and a driving decision module, and the cloud-side device is deployed with a visual language model (VLM) and a large language model (LLM).
[0090] The vehicle-mounted sensor in the vehicle continuously collects environmental data and sends it to the vehicle-side device. The environmental data is recognized by the driving scene recognition model in the vehicle-side device to determine whether it is a narrow road meeting scene.
[0091] When it is determined to be a narrow road meeting scene, a prompt information can be output to the driver, for example, the prompt information is output through voice. At this time, the driver can speak the scene description and / or driving instruction according to the voice prompt. The voice collection device in the vehicle can collect the voice data of the driver in a preset time period after the output of the prompt information, and send the voice data to the vehicle-side device. The voice data is recognized by the voice recognition model in the vehicle-side device to obtain text data corresponding to the voice data.
[0092] In addition, the vehicle-side device can send the scene image collected by the vehicle-mounted camera and the text data to the cloud-side device. The cloud-side device can recognize the intention of the text data through the large language model to obtain the intention type corresponding to the text data.
[0093] If the intent type is a scene description intent, the cloud device can continue to perform scene extraction processing on the text data through the large language model to obtain first scene description text, and then perform scene extraction processing on the scene image and the first scene description text through the visual language model to obtain second scene description text. Subsequently, the cloud device can send the second scene description text to the vehicle-side device, and the vehicle-side device can subsequently process the second scene description text through the driving decision module to generate the target driving control instruction.
[0094] If the intent type is a driving instruction intent, the cloud device can continue to perform control instruction extraction processing on the text data through the large language model to obtain a first sub-driving control instruction. In addition, the cloud device can perform scene extraction processing on the scene image through the visual language model to obtain third scene description text. Subsequently, the cloud device can send the third scene description text and the first sub-driving control instruction to the vehicle-side device. The vehicle-side device can subsequently process the third scene description text through the driving decision module to generate a second sub-driving control instruction. The vehicle-side device can further fuse the first sub-driving control instruction and the second sub-driving control instruction through the driving decision module to obtain the target driving control instruction.
[0095] After obtaining the target driving control instruction, the vehicle-side device can automatically control the vehicle according to the target driving control instruction.
[0096] Through testing, the vehicle control method described above can enhance the automatic driving capability of complex driving scenes such as narrow road passing, reduce the driving difficulty, improve the safety, and effectively handle the passing of complex scenes such as mountain roads, rural roads, and communities.
[0097] Figure 3 is a structural block diagram of a vehicle control apparatus 300 according to an example embodiment, with reference to Figure 3 The vehicle control apparatus 300 includes:
[0098] A first acquisition module 310 configured to acquire environmental data collected by an on-board sensor.
[0099] A second acquisition module 320 configured to, in a case where it is determined according to the environmental data that the vehicle is in a target driving scene, acquire a scene image collected by an on-board camera and voice data of a driver.
[0100] A generation module 330 configured to generate a target driving control instruction according to the scene image and the voice data.
[0101] A control module 340 configured to automatically control the vehicle according to the target driving control instruction.
[0102] In some embodiments, the generating module 330 includes:
[0103] a speech recognition sub-module configured to perform speech recognition on the speech data to obtain text data corresponding to the speech data;
[0104] an intent recognition sub-module configured to perform intent recognition on the text data to obtain an intent type corresponding to the text data;
[0105] a generating sub-module configured to generate a target driving control instruction according to the scene image, the text data, and the intent type.
[0106] In some embodiments, the generating sub-module includes:
[0107] a first generating unit configured to, in a case where the intent type is a scene description intent, perform scene extraction processing on the text data by a large language model to obtain first scene description text, perform scene extraction processing on the scene image and the first scene description text by a visual language model to obtain second scene description text, and generate a target driving control instruction according to the second scene description text:
[0108] In some embodiments, the generating sub-module includes:
[0109] a second generating unit configured to, in a case where the intent type is a driving instruction intent, perform control instruction extraction processing on the text data by a large language model to obtain a first sub-driving control instruction, perform scene extraction processing on the scene image by a visual language model to obtain third scene description text, generate a second sub-driving control instruction according to the third scene description text, and fuse the first sub-driving control instruction and the second sub-driving control instruction to obtain the target driving control instruction.
[0110] In some embodiments, the vehicle control apparatus 300 further includes:
[0111] an output module configured to, in a case where it is determined according to the environment data that the vehicle is in a target driving scene, output prompt information for prompting the driver to perform scene description and / or driving instruction description by voice.
[0112] In some embodiments, the speech data is speech data of the driver within a preset time period after the prompt information is output.
[0113] As to the apparatus in the above embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and will not be described here in detail.
[0114] The present disclosure also provides a computer readable storage medium having stored thereon computer program instructions which, when executed by a processor, implement the steps of the vehicle control method provided by the present disclosure.
[0115] Figure 4 is a block diagram of a vehicle 400 according to an exemplary embodiment. For example, the vehicle 400 can be a hybrid vehicle, or a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles.
[0116] Referring to Figure 4 The vehicle 400 can include various subsystems, such as an infotainment system 410, a perception system 420, a decision control system 430, a drive system 440, and a computing platform 450. The vehicle 400 can include more or fewer subsystems, and each subsystem can include multiple components. In addition, each subsystem and each component of the vehicle 400 can be interconnected by wired or wireless means.
[0117] In some embodiments, the infotainment system 410 can include a communication system, an entertainment system, a navigation system, and the like.
[0118] The perception system 420 can include several sensors for sensing information of the environment around the vehicle 400. For example, the perception system 420 can include a global positioning system (which can be a GPS system, a Beidou system, or other positioning systems), an inertial measurement unit (IMU), a laser radar, a millimeter wave radar, an ultrasonic radar, and a camera.
[0119] The decision control system 430 can include a computing system, a vehicle controller, a steering system, a throttle, and a braking system.
[0120] The drive system 440 can include components that provide power motion for the vehicle 400. In one embodiment, the drive system 440 can include an engine, an energy source, a transmission system, and wheels. The engine can be one or a combination of an internal combustion engine, an electric motor, an air compression engine. The engine can convert energy provided by the energy source into mechanical energy.
[0121] Part or all of the functions of the vehicle 400 are controlled by the computing platform 450. The computing platform 450 can include at least one processor 451 and a memory 452, and the processor 451 can execute instructions 453 stored in the memory 452.
[0122] The processor 451 can be any conventional processor, such as a commercial available CPU. The processor can also include a Graphic Process Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof.
[0123] The memory 452 can be implemented by any type of volatile or nonvolatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read only memory (PROM), read only memory (ROM), magnetic storage, flash memory, magnetic or optical disk.
[0124] In addition to the instructions 453, the memory 452 can also store data, such as road map, route information, the position, direction, speed of the vehicle, etc. The data stored in the memory 452 can be used by the computing platform 450.
[0125] In the embodiments of the present disclosure, the processor 451 can execute the instructions 453 to complete all or part of the steps of the vehicle control method described above.
[0126] In another exemplary embodiment, a computer program product is also provided, which contains a computer program executable by a programmable device, the computer program having code portions for performing the vehicle control method described above when executed by the programmable device.
[0127] Likewise, although the present disclosure has been described and illustrated with respect to one or more implementations, equivalent alterations and modifications will become apparent to those skilled in the art that do not depart from the true spirit and scope of the disclosure. The present disclosure includes all such modifications and alterations and is limited only by the scope of the following claims. In particular regard to the various functions performed by the above described components (e.g., elements, resources, etc.), the terms (including a reference to a "means") used to describe such components are intended to correspond, unless otherwise indicated, to any component which performs the described function (e.g., a functional equivalent), even though not structurally equivalent to the disclosed structure. In addition, although a particular feature of the disclosure can have been disclosed with respect to only one of several implementations, other implementations can include the particular feature. For example, the disclosure can be implemented with respect to other implementations that incorporate the particular feature, and that implement other features as disclosed herein, and each of the various implementations have a reasonable expectation of support. Furthermore, to the extent that the terms "includes", "including", "has", "have", "having", or variants thereof are used in either the detailed description or the claims, such terms are intended to be inclusive in a manner similar to the term "comprising".
[0128] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features of the disclosure disclosed herein. It is intended that the present disclosure be considered as including any variations, uses, or adaptations of the disclosure following, in general, the principles of the disclosure and including such steps, compositions, components, and / or elements known in the art to be appropriate. It is specifically intended that the present disclosure include all such modifications and alterations in the application, processes, and procedures as fall within the usual scope of the technology and are appreciated by those skilled in the art. The specification and examples are to be considered exemplary only, with the true scope and spirit of the disclosure being indicated by the following claims.
[0129] It is to be understood that the present disclosure is not limited to the precise construction described and shown herein and that changes can be made in various details without departing from the scope of the disclosure. The scope of the present disclosure is limited only by the claims that follow.
Claims
1. A vehicle control method characterized by, The method comprises: acquiring environment data collected by a vehicle-mounted sensor; in a case where it is determined according to the environment data that the vehicle is in a target driving scene, outputting prompt information for prompting a driver to describe a scene and a driving instruction by voice; acquiring a scene image collected by a vehicle-mounted camera and voice data of the driver, the intention type expressed by the voice data including a scene description intention and a driving instruction intention, and the target driving scene including a narrow road meeting scene, a narrow road overtaking scene, a mountain road meeting scene, a rural road meeting scene and a community meeting scene; generating a target driving control instruction according to the scene image and the voice data; automatically driving the vehicle according to the target driving control instruction; the generating of the target driving control instruction according to the scene image and the voice data comprises: performing scene extraction processing on text data corresponding to the voice data of the intention type being the scene description intention by a large language model to obtain first scene description text; performing scene extraction processing on the scene image and the first scene description text by a visual language model to obtain second scene description text; generating a target driving control instruction according to the second scene description text.
2. The method of claim 1, wherein, the generating of the target driving control instruction according to the scene image and the voice data comprises: performing voice recognition on the voice data to obtain text data corresponding to the voice data; performing intention recognition on the text data to obtain an intention type corresponding to the text data; generating a target driving control instruction according to the scene image, the text data and the intention type.
3. The method of claim 2, wherein, the generating of the target driving control instruction according to the scene image, the text data and the intention type comprises: in a case where the intention type is a driving instruction intention, performing control instruction extraction processing on the text data by a large language model to obtain a first sub-driving control instruction; performing scene extraction processing on the scene image by a visual language model to obtain third scene description text; generating a second sub-driving control instruction according to the third scene description text; fusing the first sub-driving control instruction and the second sub-driving control instruction to obtain the target driving control instruction.
4. The method of claim 1, wherein, The voice data is voice data of the driver within a preset time period after the prompt information is output.
5. A vehicle control device characterized by comprising: The device comprises: a first acquisition module configured to acquire environment data collected by a vehicle-mounted sensor; a second acquisition module configured to, in a case where it is determined according to the environment data that the vehicle is in a target driving scene, output prompt information for prompting a driver to describe a scene and a driving instruction by voice, acquire a scene image collected by a vehicle-mounted camera and voice data of the driver, the intention type expressed by the voice data including a scene description intention and a driving instruction intention, and the target driving scene including a narrow road meeting scene, a narrow road overtaking scene, a mountain road meeting scene, a rural road meeting scene and a community meeting scene; The generating module is configured to generate a target driving control instruction according to the scene image and the voice data; The control module is configured to automatically drive control the vehicle according to the target driving control instruction; The generating module is further configured to perform scene extraction processing on text data corresponding to voice data of an intent type of the scene description intent by a large language model to obtain first scene description text; perform scene extraction processing on the scene image and the first scene description text by a visual language model to obtain second scene description text; and generate a target driving control instruction according to the second scene description text.
6. A vehicle characterized by comprising: Comprise: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the steps of the method of any one of claims 1 to 4 when invoking the executable instructions stored on the memory.
7. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 4.
8. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 4. The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 4.
Citation Information
Patent Citations
Method and device for controlling vehicle running and vehicle
CN113226886A