Vehicle-mounted visual and voice collaborative interaction method and device based on multi-modal large model
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2026-04-08
- Publication Date
- 2026-08-04
AI Technical Summary
[0005]本发明提供一种基于多模态大模型的车载视觉和语音协同交互方法和装置,用以解决现有技术中车载人机交互方式智能化程度较低的缺陷,实现提高车载人机交互方式的智能化程度
[0019] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the in-vehicle vision and voice collaborative interaction method based on a multimodal large model as described above.
Smart Images

Figure CN122511244A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automatic control technology, and in particular to a method and apparatus for vehicle-mounted vision and voice collaborative interaction based on a multimodal large model. Background Technology
[0002] With the rapid development of intelligent vehicles, autonomous driving assistance systems, and in-vehicle infotainment systems, the functional complexity and information carrying capacity of in-vehicle systems continue to increase. Modern in-vehicle systems typically integrate multiple functions such as navigation, voice assistants, multimedia entertainment, online services, and third-party applications, placing higher demands on human-computer interaction methods.
[0003] Currently, in-vehicle human-machine interaction mainly relies on the following technical paths: 1) Touch-based graphical interface interaction, where users need to manually complete operations through multi-level menus, which can be distracting in driving scenarios; 2) Voice interaction based on speech recognition and natural language understanding, which requires users to clearly and completely input target information (such as the complete destination name or video name) before the system can execute, making it highly dependent on user input; 3) Function triggering mechanisms based on fixed rules or scene templates, such as fixed voice commands and shortcut scene buttons, which have limited flexibility and are difficult to adapt to open and diverse user needs.
[0004] Therefore, the level of intelligence in existing in-vehicle human-machine interaction methods is relatively low. Summary of the Invention
[0005] This invention provides a method and device for in-vehicle vision and voice collaborative interaction based on a multimodal large model, which addresses the shortcomings of low intelligence in existing in-vehicle human-machine interaction methods and improves the intelligence level of in-vehicle human-machine interaction methods.
[0006] This invention provides a method for in-vehicle vision and voice collaborative interaction based on a multimodal large model, comprising: Receives voice commands input by the user through the vehicle-mounted voice acquisition module, wherein the voice commands are used to represent the user's operational needs and intentions; In response to the voice command, acquire the target image captured by the vehicle-mounted camera device; A sequence of operation instructions is determined, wherein the sequence of operation instructions is determined by a multimodal large model based on the demand type and the target content, wherein the demand type is the type corresponding to the operation demand intent determined based on the text information corresponding to the voice instruction, and the target content is the content in the target image related to the demand type; The sequence of operation instructions is parsed and executed to complete at least one set of graphical user interface (GUI) operations corresponding to the intended operation requirement.
[0007] According to the present invention, a method for in-vehicle vision and voice collaborative interaction based on a multimodal large model is provided, wherein determining the sequence of operation instructions includes: The text information and the target image are sent to the server, whereby the text information and the target image are used to instruct the server to determine the demand type and the target content through the multimodal large model, and to determine the operation instruction sequence based on the demand type and the target content; Receive the sequence of operation instructions sent by the server.
[0008] According to the present invention, a method for in-vehicle vision and voice collaborative interaction based on a multimodal large model is provided, wherein determining the sequence of operation instructions includes: The text information and the target image are input into the multimodal large model. The demand type and the target content are determined by the multimodal large model, and the operation instruction sequence is determined based on the demand type and the target content.
[0009] According to the present invention, an in-vehicle vision and voice collaborative interaction method based on a multimodal large model is provided, wherein the target content is determined through the multimodal large model, including: The target image is visually understood using the multimodal large model to obtain at least one information element in the target image; Based on the required type, at least one target information element related to the intention to complete the operation requirement is selected from at least one of the information elements; The target information element is determined as the target content.
[0010] According to the present invention, a method for in-vehicle vision and voice collaborative interaction based on a multimodal large model is provided, wherein acquiring the target image captured by the in-vehicle camera device includes: The target image is obtained by acquiring an image of an external carrier containing the target information captured by an in-vehicle camera device. The external carrier includes a display screen of a mobile terminal, printed material, or the surface of an object containing readable information.
[0011] According to the present invention, a method for in-vehicle vision and voice collaborative interaction based on a multimodal large model is provided, wherein acquiring the target image captured by the in-vehicle camera device includes: The in-vehicle camera captures images containing the target's direction, and determines the user's spatial direction information based on the images. The target direction includes the user's gesture direction or gaze direction. Images of the external environment are captured by an external vehicle-mounted camera, and at least one candidate information carrier located outside the vehicle is identified based on the images. Based on the spatial pointing information and the spatial location information of each of the candidate information carriers, the target information carrier pointed to by the user is determined from at least one of the candidate information carriers; The vehicle-mounted camera outside the vehicle acquires an image containing the target information carrier, and identifies the image as the target image.
[0012] According to the present invention, a method for in-vehicle vision and voice collaborative interaction based on a multimodal large model is provided, wherein acquiring the target image captured by the in-vehicle camera device includes: Acquire the original images captured by the vehicle-mounted camera device; The original image is preprocessed to obtain the target image. The preprocessing includes at least one of the following: resolution adjustment, image cropping, and image format adjustment.
[0013] According to the present invention, an in-vehicle vision and voice collaborative interaction method based on a multimodal large model is provided, wherein the demand type includes navigation type or multimedia information playback type.
[0014] According to the present invention, a vehicle-mounted vision and voice collaborative interaction method based on a multimodal large model is provided, wherein parsing and executing the operation instruction sequence includes: Based on the aforementioned requirement type and target content, an operation confirmation message is generated; The operation confirmation information is output through the in-vehicle interactive interface; Upon receiving a confirmation instruction for the operation confirmation information, the operation instruction sequence is parsed and executed.
[0015] This invention also provides an in-vehicle vision and voice collaborative interaction system based on a multimodal large model, comprising: The vehicle-mounted voice acquisition module is used to acquire voice commands input by the user, and the voice commands are used to represent the user's operational needs and intentions. Vehicle-mounted camera equipment, used to capture images of the target; The vehicle controller is used to determine an operation command sequence, which is determined by a multimodal large model based on demand type and target content. The demand type is the type corresponding to the operation demand intent determined based on the text information corresponding to the voice command, and the target content is the content in the target image related to the demand type. The vehicle controller is also used to parse and execute the sequence of operation instructions to complete at least one set of user interface (GUI) operations corresponding to the intended operation requirement.
[0016] The present invention also provides an in-vehicle vision and voice collaborative interaction device based on a multimodal large model, comprising: The receiving module is used to receive voice commands input by the user through the vehicle-mounted voice acquisition module, wherein the voice commands are used to represent the user's operational needs and intentions. The acquisition module is used to acquire the target image captured by the vehicle-mounted camera device in response to the voice command; The determination module is used to determine the sequence of operation instructions. The sequence of operation instructions is determined by a multimodal large model based on the demand type and target content. The demand type is the type corresponding to the operation demand intention determined based on the text information corresponding to the voice instruction. The target content is the content in the target image related to the demand type. The processing module is used to parse and execute the sequence of operation instructions to complete at least one set of graphical user interface (GUI) operations corresponding to the intended operation requirement.
[0017] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the in-vehicle vision and voice collaborative interaction method based on a multimodal large model as described above.
[0018] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the in-vehicle vision and voice collaborative interaction method based on a multimodal large model as described above.
[0019] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the in-vehicle vision and voice collaborative interaction method based on a multimodal large model as described above.
[0020] The present invention provides a vehicle-mounted visual and voice collaborative interaction method and device based on a multimodal large model. This method receives voice commands input by a user through a vehicle-mounted voice acquisition module. The voice commands represent the user's operational intent. In response to the voice commands, the method acquires a target image captured by a vehicle-mounted camera, determines an operation command sequence, and establishes this sequence through a multimodal large model based on the demand type and target content. The demand type is determined based on the text information corresponding to the voice command, and the target content is the content in the target image related to the demand type. The method then parses and executes the operation command sequence to complete at least one set of GUI operations corresponding to the operational intent. Because it can trigger and coordinate external visual information through voice, and the multimodal large model performs unified semantic understanding and operation planning to form the operation command sequence, it effectively enables the direct transmission of existing external information to the vehicle system. This eliminates the need for additional manual operation or multiple voice confirmations by the user, automatically responding to user needs. This effectively breaks down the barriers between external information and the vehicle system, reduces user operation steps, and improves the versatility and intelligence of the vehicle-mounted human-machine interaction method in various scenarios. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0022] Figure 1 This is a flowchart illustrating the in-vehicle vision and voice collaborative interaction method based on a multimodal large model provided in an embodiment of the present invention.
[0023] Figure 2 This is a schematic diagram of the structure of an in-vehicle vision and voice collaborative interaction system based on a multimodal large model, provided in an embodiment of the present invention.
[0024] Figure 3 This is a schematic diagram of the structure of an in-vehicle vision and voice collaborative interaction device based on a multimodal large model, provided in an embodiment of the present invention.
[0025] Figure 4 This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0027] Currently, in-vehicle human-machine interaction mainly relies on the following technological paths: First, touch-based graphical interface interaction, where users complete tasks by clicking, swiping, and other operations on the in-vehicle display. For example, navigating to a destination or playing a movie requires voice input or manual input of the detailed destination name or video title. This method has advantages in terms of functionality, but in driving scenarios, it often requires multi-level menu operations, resulting in numerous operation steps and distraction from the driver's attention.
[0028] Second, there's voice interaction based on speech recognition and natural language understanding. In-vehicle voice assistants use Automatic Speech Recognition (ASR) and Natural Language Understanding (NLU) technologies to interpret user commands and perform operations such as navigation and audio / video playback. With this approach, when a user wants to navigate to a destination or play specific video or music content, they typically need to explicitly input the complete and specific destination name or content name via voice. The system then matches and executes the input. If the input information is incomplete or ambiguous, the system struggles to respond correctly. This method is highly dependent on the standardization and completeness of the user's input.
[0029] Thirdly, there are rule-based or scenario template-based function triggering mechanisms. Some in-vehicle systems trigger specific functions through preset rules or scenario templates, such as fixed voice commands or shortcut scene buttons. This method is efficient in specific scenarios, but its flexibility is limited and it is difficult to adapt to open and diverse user needs.
[0030] In real-world usage scenarios, user needs often originate not only from within the vehicle system but are also closely related to the external environment or external information sources. When in the car, users typically encounter or pay attention to the following types of information: (1) Information presented by external physical carriers or display carriers, such as roadside billboards, shop signs, bus stop signs, paper maps, or flyers. (2) Information located outside or inside the car but not belonging to the vehicle system itself, such as addresses displayed on mobile phone screens, articles or movie recommendations in in-car magazines, and product packaging or business cards placed inside the car. (3) Information that exists in image form but has not yet been input into the vehicle system in a structured manner, such as restaurant menu photos taken by users with their mobile phones, QR codes on street posters, and name signs on building facades.
[0031] The information mentioned above often explicitly includes the user's desired content, such as the destination name and audio / video content identifiers. However, existing in-vehicle systems typically cannot directly utilize this type of information. If users wish to complete navigation or audio / video playback operations on the vehicle's infotainment system based on this information, they usually still need to re-enter the corresponding content via voice or manual input.
[0032] In summary, the level of intelligence in existing in-vehicle human-machine interaction technologies remains relatively low.
[0033] In view of the above-mentioned problems, this invention proposes an in-vehicle vision and voice collaborative interaction method based on a multimodal large model. In this method, information from the external environment or external information sources can be fully utilized, and combined with concise voice commands, multimodal fusion input and multimodal joint understanding can be performed. The system performs unified semantic reasoning on the above-mentioned unstructured information through the multimodal large model, automatically maps and generates an executable sequence of in-vehicle graphical interface operation commands, thereby completing the understanding of user needs and the automatic operation response of the in-vehicle graphical interface. Users no longer need to completely and accurately repeat or input target content information through voice or manual means. This can effectively break down the barriers between external information and the in-vehicle system, reduce user operation steps, and build an end-to-end workflow for demand triggering, understanding and automatic operation response, thereby improving the intelligence level of in-vehicle human-machine interaction.
[0034] The embodiments of the present invention can be applied to in-vehicle human-machine interaction scenarios, such as navigation scenarios based on external physical identifiers, multimedia and information query or navigation scenarios based on mobile terminals or printed materials, or navigation scenarios based on points of interest (POI) in the external environment of the vehicle.
[0035] The subject executing this method can be an electronic device such as a vehicle-mounted system, terminal equipment, computer, server, server cluster, or a specially designed in-vehicle vision and voice collaborative interaction device based on a multimodal large model. It can also be an in-vehicle vision and voice collaborative interaction device based on a multimodal large model installed in the electronic device. The in-vehicle vision and voice collaborative interaction device based on a multimodal large model can be implemented through software, hardware, or a combination of both.
[0036] Figure 1 This is a flowchart illustrating the in-vehicle vision and voice collaborative interaction method based on a multimodal large model provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes: Step 101: Receive voice commands input by the user through the vehicle voice acquisition module. The voice commands are used to represent the user's operational needs and intentions.
[0037] In this step, the user's voice command is used to trigger the request processing flow. The voice command is used to express the user's current unfinished operation request, such as navigation, playing audio or video content, etc.
[0038] It should be noted that the voice command input by the user does not need to contain complete and explicit target information, such as the full destination name or content name. It only needs to serve to trigger the multimodal demand understanding process. For example, the user can input "navigate to here," "I want to go here," "play this movie," etc. This voice command is vague and conversational; it only indicates the user's intended action.
[0039] Since the voice commands input by users do not need to contain complete and precise destination or content names, users only need to use vague and directional natural language (such as "I want to go here" or "Just play this"), which greatly reduces the interaction threshold and memory burden, and can improve the naturalness and efficiency of in-vehicle human-machine interaction.
[0040] Step 102: In response to the voice command, acquire the target image captured by the vehicle-mounted camera.
[0041] In this step, the vehicle-mounted camera equipment may include an in-vehicle camera for monitoring the driver or passengers, an external camera for surround view or driving, etc. The specific form and location of the vehicle-mounted camera equipment are not limited in this embodiment of the invention, as long as the target image can be acquired.
[0042] Upon receiving a voice command, the vehicle will activate and control the in-vehicle camera to capture images of the target location. For example, a user might say "Navigate here" while simultaneously displaying the map location on their phone screen to the in-vehicle camera. At this point, the in-vehicle camera can capture an image of the phone screen containing the map location.
[0043] The collected target images contain target content related to the user's operational needs and intentions, such as destination names, video titles, or song titles.
[0044] It should be noted that the acquisition of the target image is triggered by voice commands, rather than by random environmental shooting. Through this triggering sequence and logic, it is possible to bind the user's vague voice commands with a specific and clear visual information source, providing a foundation for realizing the user's operational needs and intentions in the future.
[0045] Furthermore, by acquiring target images from in-vehicle cameras, users can obtain the necessary information visually without having to manually input the destination or content name. This breaks down the barriers between external information and the vehicle's infotainment system, eliminating the tedious steps of repeatedly stating information between the device and the interface. Consequently, this significantly reduces the user's operational burden and cognitive load, effectively mitigating the risk of distraction caused by complex interactions in driving scenarios, thereby improving interaction efficiency and driving safety.
[0046] Step 103: Determine the sequence of operation instructions. The sequence of operation instructions is determined by a multimodal large model based on the demand type and target content. The demand type is the type of operation demand intent determined based on the text information corresponding to the voice command. The target content is the content in the target image that is related to the demand type.
[0047] With the development of deep learning and large-scale model technology, Visual-Language Models (VLMs) are gradually gaining the ability to perform unified semantic modeling and joint reasoning of images and natural language. These models are no longer limited to simple object recognition or keyword matching, but can understand the user's high-level intent within the joint context of vision and language. Based on this, VLMs can determine the type of need corresponding to the operational intent based on the text information corresponding to the voice command, and determine the target content related to the need type from the target image.
[0048] In this context, "demand type" can be understood as the executable function category corresponding to a user's intended action within the vehicle's infotainment system. For example, demand types can include navigation type or multimedia information playback type. For instance, if the text information is "Go here," "Please navigate here," or "I want to go here," the demand type determined based on the text information is "navigation type." If the text information is "Play this," "Please play this movie," or "I want to watch this," the demand type determined based on the text information is "multimedia information playback type," and so on.
[0049] The target content can be understood as the core information entity parsed and extracted from the target image, directly related to the determined demand type, and usable by the vehicle's infotainment system to perform specific operations. It is the key data identified by the multimodal large model in the image that directly satisfies the user's operational intent. For example, if the target image contains the address information of restaurant A, and the demand type is "navigation type," then the determined target content is "the address information of restaurant A." If the target image is a poster of movie B displayed on the user's mobile phone screen, and the demand type is "playing multimedia information type," then the determined target content is information such as the identification (ID) or name of movie B.
[0050] Furthermore, a multimodal large model can be used to determine the sequence of operation instructions based on the type of demand and the target content. This sequence of operation instructions includes action language descriptions, such as entering "Restaurant A" in the search box and clicking the "Start Navigation" button, as well as structured instructions directly recognized and executed by the vehicle's infotainment system, such as CLICK(box=[0.233,0.525, 0.354,0.759], element_info="Start Navigation"). By combining natural language and structured commands, the task corresponding to the intended operation can be completed more accurately and efficiently.
[0051] Step 104: Parse and execute the sequence of operation instructions to complete at least one set of graphical user interface (GUI) operations corresponding to the intended operation requirements.
[0052] In this step, after acquiring the sequence of operation instructions, the vehicle's infotainment system parses it. Since this sequence includes both action language descriptions and structured instructions, the system uses the precise semantic intent from the action language interpreter and extracts the parameters from the structured instructions. The system then converts the parsed results into specific function calls or message passes that can be executed at the vehicle's underlying level, driving the corresponding GUI components to complete the operation. This dual parsing and execution mechanism allows for a more efficient and accurate transformation of user needs into actual user interface interactions.
[0053] After parsing is complete, the system will automatically complete the corresponding interface navigation, function calls, and operation execution according to the logic and order defined by the operation instruction sequence, in order to complete at least one set of graphical user interface (GUI) operations corresponding to the operation requirements. Taking navigating to restaurant A as an example, the car system will automatically launch the navigation application, enter the address of restaurant A in the search box, and automatically click the "Start Navigation" button without the user touching the screen.
[0054] The in-vehicle vision and voice collaborative interaction method based on a multimodal large model provided in this invention receives voice commands input by the user through the in-vehicle voice acquisition module. These voice commands represent the user's operational intent. In response to the voice commands, the method acquires the target image captured by the in-vehicle camera device, determines the sequence of operation commands, and establishes this sequence through the multimodal large model based on the demand type and target content. The demand type is the type of operational intent determined based on the text information corresponding to the voice command, and the target content is the content in the target image related to the demand type. The method parses and executes the sequence of operation commands to complete at least one set of GUI operations corresponding to the operational intent. Because it can trigger and coordinate external visual information through voice, and the multimodal large model performs unified semantic understanding and operation planning to form the sequence of operation commands, it can effectively achieve the direct transmission of existing external information to the in-vehicle system. This eliminates the need for additional manual operation or multiple rounds of voice confirmation by the user, thus effectively breaking down the barriers between external information and the in-vehicle system, reducing user operation steps, and improving the versatility and intelligence of in-vehicle human-machine interaction in various scenarios.
[0055] For example, in one possible implementation, when determining the sequence of operation instructions, voice instructions and target images can be sent to the server. The voice instructions and target images are used to instruct the server to determine the demand type and target content through a multimodal big model, and to determine the sequence of operation instructions based on the demand type and target content, and to receive the sequence of operation instructions sent by the server.
[0056] Specifically, to reduce the hardware computing load and memory consumption of the vehicle's infotainment system, the multimodal large model can be deployed to a cloud server. In this way, the vehicle's infotainment system itself does not undertake complex model calculation tasks, but instead uploads voice commands and target images to the server through the vehicle's network communication module.
[0057] After receiving voice commands and target images, the server converts the voice commands into text information using a voice conversion model. The text information and target image are then input into a multimodal big data model. This model, as the core cognitive and reasoning unit, performs a joint understanding of the text information and target image within a unified semantic space. Specifically, based on the content contained in the target image and the intended needs expressed in the text information, the multimodal big data model performs the following reasoning process: understanding the user's actual need type, such as navigation or multimedia playback; parsing the target content related to the need type from the target image; and combining this with the functional structure of the in-vehicle system to infer one or more sets of in-vehicle GUI operation steps required to fulfill that need type. Through this reasoning process, a mapping from the user's high-level intentions to specific in-vehicle system operations can be achieved.
[0058] The multimodal large model outputs the results of inference in a structured manner, which may include the user's demand type, target content, and operation instruction sequence.
[0059] For example, the server can transmit the demand type, target content, and operation instruction sequence to the vehicle terminal through the communication interface, or it can only transmit the operation instruction sequence to the vehicle terminal.
[0060] After receiving the sequence of operation instructions from the server, the vehicle's infotainment system will proceed to parse and execute the sequence of operation instructions.
[0061] It should be noted that a voice conversion model can also be deployed on the vehicle's infotainment system. The collected voice commands are processed by the in-vehicle system's built-in speech recognition engine using Automatic Speech Recognition (ASR) to convert them into text information. The recognized text information and the target image are then sent to the server. In this way, the server no longer performs speech recognition; instead, it directly inputs the received text information and target image into a multimodal large model.
[0062] In this embodiment, the multimodal large model can be deployed on a cloud server. This allows for the processing of text information and target images, as well as the generation of operation instruction sequences, on the server, reducing the hardware computing load on the vehicle's infotainment system.
[0063] In another possible implementation, when determining the sequence of operation instructions, the text information corresponding to the voice instructions and the target image can be input into the multimodal large model. The demand type and target content are determined through the multimodal large model, and the sequence of operation instructions is determined based on the demand type and target content.
[0064] Specifically, to improve the real-time performance and reliability of human-computer interaction, the multimodal large model can be deployed locally on the vehicle's infotainment system. After completing the text conversion of voice commands and the acquisition of target images, the vehicle's infotainment system can directly input the converted text information and the acquired target images into the locally deployed multimodal large model without relying on an external network connection. The locally deployed multimodal large model performs the same inference task as the multimodal large model deployed on the server in the aforementioned embodiment to determine the sequence of operation commands.
[0065] In this embodiment, a multimodal large model can be deployed locally on the vehicle's infotainment system. This locally deployed model processes text information and target images, and generates operation command sequences, thereby improving the real-time performance of human-computer interaction. Furthermore, it can provide stable and reliable human-computer interaction services even in environments with poor network signal or no network access.
[0066] For example, when determining target content through a multimodal large model, the target image can be visually understood through the multimodal large model to obtain at least one information element in the target image. Based on the demand type, at least one target information element related to the intention to complete the operation demand is selected from the at least one information element, thereby determining the target information element as the target content.
[0067] Specifically, after the target image is input into the multimodal large model, the model performs a comprehensive visual understanding of the target image, obtaining multiple discrete information elements contained within it. These information elements can be text, icons, or special symbols such as barcodes or QR codes. For example, for a target image of a restaurant flyer, the multimodal large model may identify multiple information elements such as the "XX Restaurant" logo, "Address: No. 1, XX Road", "Telephone: 12345678", and "Signature Dish: YYY".
[0068] Furthermore, the multimodal big model will combine the demand type with the analysis of the correlation between each information element and the demand type, in order to filter these identified information elements and obtain at least one target information element related to the intention to complete the operation. For example, if the demand type is "navigation type", the multimodal big model will evaluate and filter out elements that are directly related to the demand type "navigation type", such as location, address, and name, and filter out elements that are not related to the user's intention to complete the operation (such as "telephone" or "signature dish").
[0069] The selected target information elements can be identified as the target content. For example, in a navigation scenario, the address "No. 1, XX Road" selected from the target image of a restaurant flyer would be used as the final target content.
[0070] In this embodiment, a multimodal large model is used to perform visual understanding of the target image, and at least one information element obtained after visual understanding is filtered to obtain at least one target information element related to the intention to complete the operation requirement. This target information element is determined as the target content. In this way, the accuracy and robustness of target content extraction can be improved. Even if the target image contains multiple irrelevant information interferences, the core information can be effectively focused on, thereby ensuring the accuracy of the subsequently generated operation instruction sequence.
[0071] For example, in one possible implementation, when acquiring the target image captured by the vehicle-mounted camera, an image of an external carrier containing the target information can be acquired from the vehicle-mounted camera to obtain the target image. The external carrier includes the display screen of a mobile terminal, printed materials, or the surface of an object containing readable information.
[0072] Specifically, in this scenario, users can actively display external objects containing target information to the in-vehicle camera. For example, users can point the screen of a mobile device such as a smartphone or tablet towards the in-vehicle camera, which may display a map location, movie title, or song title. They can also point printed materials such as books, magazines, brochures, business cards, or paper maps towards the in-vehicle camera, or any object with text, graphics, or QR codes on its surface towards the in-vehicle camera.
[0073] In this embodiment, the user actively and intuitively demonstrates the action (such as pointing a mobile phone screen or printed material towards the camera), enabling the in-vehicle camera to directly and clearly capture the target image of the external information carrier, thus completing the input of visual information. This method eliminates the need for the user to repeatedly switch between the external information carrier and the vehicle's infotainment system, repeatedly inputting or confirming target content, reducing the operational burden and greatly improving the naturalness and intuitiveness of human-computer interaction, enhancing user-friendliness in driving scenarios. Simultaneously, due to the close-range image acquisition, the captured target image quality is high, ensuring the accuracy and reliability of subsequent multimodal large-scale model recognition.
[0074] In another possible implementation, when acquiring the target image captured by the in-vehicle camera, an image containing the target direction can also be acquired by the in-vehicle camera, and the user's spatial direction information can be determined based on the image. The target direction includes the user's gesture direction or gaze direction. An external in-vehicle camera can be used to acquire images of the external environment, and at least one candidate information carrier located outside the vehicle can be identified based on the external environment images. Based on the spatial direction information and the spatial position information of each candidate information carrier, the target information carrier that the user is pointing to can be determined from at least one candidate information carrier. An image containing the target information carrier can be acquired by the external in-vehicle camera, and the image can be identified as the target image.
[0075] Specifically, images containing the user's gestures or gaze can be captured by in-vehicle cameras. The gaze direction can be determined using eye-tracking technology. Based on this image, a three-dimensional spatial pointing vector, i.e., spatial pointing information, emanating from the user (usually an eye or hand reference point) can be calculated using gesture recognition algorithms, eye-tracking algorithms, or a fusion of both. This allows the determination of which direction the user is pointing outside the vehicle.
[0076] The vehicle-mounted camera will synchronously or continuously collect images of the external environment and use target detection algorithms to analyze these images, identify all carriers that may contain valid information, such as billboards or road signs, and determine these carriers as candidate information carriers. It can also estimate the three-dimensional coordinates of each candidate information carrier in the vehicle coordinate system, i.e., spatial location information.
[0077] Furthermore, the spatial directional information representing the user's intention direction can be geometrically calculated and matched with the spatial location information of all candidate information carriers, and the candidate information carrier with the smallest spatial distance from the user's directional direction can be determined as the target information carrier.
[0078] Control or invoke the vehicle-mounted camera outside the vehicle to acquire an image containing the target information carrier and use it as the target image.
[0079] For example, when a user sees a billboard on the roadside, they can point to it with their hand and simultaneously issue a voice command, such as "Navigate here." The in-vehicle camera system captures an image containing the user's gesture and determines the spatial direction information of the finger's pointing finger. This spatial direction information can be understood as a virtual ray pointing from the user out of the vehicle. Simultaneously, the external in-vehicle camera system captures images of the external environment and identifies roadside billboards, road signs, and other objects as candidate information carriers.
[0080] The calculated spatial pointing information is matched with the spatial location information of each candidate information carrier (such as multiple billboards and road signs). The billboard that best matches the user's gesture pointing direction is identified as the target information carrier, and an image containing the billboard is captured by an external vehicle-mounted camera as the target image.
[0081] In this embodiment, by utilizing the collaborative work of in-vehicle and external vehicle-mounted cameras, precise capture of the user's gestures or gaze and real-time perception of complex external environmental information are achieved, and the two are accurately correlated and matched within a unified spatial coordinate system. This collaborative working mode of the in-vehicle and external vision systems can improve safety and convenience in driving scenarios. Furthermore, the above method can effectively achieve the direct transmission of existing external information to the vehicle's infotainment system, significantly reducing additional user steps and improving interaction efficiency in driving scenarios.
[0082] It should be noted that for a certain information carrier outside the vehicle, the user can also take a picture of the information carrier with a mobile terminal and then show the resulting image to the in-vehicle camera device so that the in-vehicle camera device can capture a target image containing the information carrier.
[0083] For example, based on the above embodiments, when acquiring the target image captured by the vehicle-mounted camera device, the original image captured by the vehicle-mounted camera device may be acquired first, and the original image may be preprocessed to obtain the target image. The preprocessing includes at least one of the following: resolution adjustment, image cropping, and image format adjustment.
[0084] Specifically, the raw images captured by vehicle-mounted camera equipment may vary in resolution, size, presence of useless background information (such as the target information carrier not being centered), or data format due to factors such as camera specifications, lighting conditions, and the user's viewing angle and distance. If directly used for subsequent multimodal large-scale model inference, it may affect the model's recognition accuracy and processing efficiency. Therefore, after acquiring the raw images, preprocessing is necessary to optimize image quality and adapt to the input requirements of multimodal large-scale models, resulting in a target image that meets the requirements.
[0085] Preprocessing includes at least one of resolution adjustment, image cropping, and image format adjustment. For example, resolution adjustment can uniformly scale or sample the original image to the standard input size required for model training, ensuring stable model processing. Image cropping can extract key regions containing target information from the original image. For example, when a user displays a mobile phone screen, cropping the portion outside the screen borders makes the image content more focused on the text or images on the screen. Image format adjustment can convert the original image from its original format (such as YUV, RAW, etc.) captured by the camera to a common format supported by the multimodal large model (such as RGB, JPEG, PNG).
[0086] In this embodiment, by preprocessing the original image, the data quality and consistency input to the multimodal large model can be significantly improved, enabling the multimodal large model to more accurately and efficiently parse the target content from the target image, thereby providing a foundation for the subsequent generation of precise operation instruction sequences.
[0087] For example, in order to improve the safety and reliability of human-computer interaction, when parsing and executing the sequence of operation instructions, operation confirmation information can be generated based on the demand type and target content, and the operation confirmation information can be output through the vehicle interaction interface. When a confirmation instruction for the operation confirmation information is received, the sequence of operation instructions can be parsed and executed.
[0088] Specifically, after obtaining the sequence of operation instructions, the vehicle's infotainment system will not execute the sequence immediately. Instead, it will first generate an operation confirmation message based on the request type and target content, and then output this confirmation message to the user through the in-vehicle interface. For example, if the request type is navigation and the target content is "Restaurant A", the system will generate a message such as "Please confirm whether you want to navigate to Restaurant A".
[0089] When the system detects a confirmation command from the user, such as a positive voice response like "confirm" or "okay," or after the user taps the "confirm" button on the screen, the system will parse and execute the sequence of operation instructions. If the user does not issue a confirmation command, or issues a negative or cancel command, the entire operation process will be terminated.
[0090] In this embodiment, by outputting operation confirmation information, and only upon receiving a confirmation instruction for the operation confirmation information, the steps of parsing and executing the operation instruction sequence are performed. This effectively prevents erroneous operations caused by model recognition errors or user accidental triggering, thereby improving the safety and reliability of human-computer interaction.
[0091] Figure 2 This is a schematic diagram of the structure of the in-vehicle vision and voice collaborative interaction system based on a multimodal large model provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the system 200 includes: The vehicle-mounted voice acquisition module 11 is used to acquire voice commands input by the user, and the voice commands are used to represent the user's operational needs and intentions. Vehicle-mounted camera device 12 is used to acquire target images; The vehicle controller 13 is used to determine an operation command sequence, which is determined by a multimodal large model based on the demand type and target content. The demand type is the type corresponding to the operation demand intention determined based on the text information corresponding to the voice command, and the target content is the content in the target image related to the demand type. The vehicle controller 13 is also used to parse and execute the sequence of operation instructions to complete at least one set of user interface (GUI) operations corresponding to the intended operation requirement.
[0092] The system in this embodiment can be used in any of the methods in the side embodiment of the vehicle vision and voice collaborative interaction method based on multimodal large model. Its specific implementation process and technical effects are similar to those in the side embodiment of the vehicle vision and voice collaborative interaction method based on multimodal large model. For details, please refer to the detailed description in the side embodiment of the vehicle vision and voice collaborative interaction method based on multimodal large model, which will not be repeated here.
[0093] The following describes the in-vehicle vision and voice collaborative interaction device based on a multimodal large model provided by the present invention. The in-vehicle vision and voice collaborative interaction device based on a multimodal large model described below can be referred to in correspondence with the in-vehicle vision and voice collaborative interaction method based on a multimodal large model described above.
[0094] Figure 3 This is a schematic diagram of the structure of the in-vehicle vision and voice collaborative interaction device based on a multimodal large model provided in an embodiment of the present invention, as shown below. Figure 3 As shown, the in-vehicle vision and voice collaborative interaction device 300 based on a multimodal large model includes: The receiving module 21 is used to receive voice commands input by the user through the vehicle voice acquisition module, wherein the voice commands are used to represent the user's operational needs and intentions. Acquisition module 22 is used to acquire the target image captured by the vehicle-mounted camera device in response to the voice command; The determination module 23 is used to determine the operation instruction sequence, which is determined by a multimodal large model based on the demand type and target content. The demand type is the type corresponding to the operation demand intention determined based on the text information corresponding to the voice instruction, and the target content is the content in the target image related to the demand type. The processing module 24 is used to parse and execute the sequence of operation instructions to complete at least one set of graphical user interface (GUI) operations corresponding to the intended operation requirement.
[0095] In one example embodiment, the determining module 23 is specifically used for: The text information and the target image are sent to the server, whereby the text information and the target image are used to instruct the server to determine the demand type and the target content through the multimodal large model, and to determine the operation instruction sequence based on the demand type and the target content; Receive the sequence of operation instructions sent by the server.
[0096] In one example embodiment, the determining module 23 is specifically used for: The text information and the target image are input into the multimodal large model. The demand type and the target content are determined by the multimodal large model, and the operation instruction sequence is determined based on the demand type and the target content.
[0097] In one example embodiment, the determining module 23 is specifically used for: The target image is visually understood using the multimodal large model to obtain at least one information element in the target image; Based on the required type, at least one target information element related to the intention to complete the operation requirement is selected from at least one of the information elements; The target information element is determined as the target content.
[0098] In one example embodiment, the acquisition module 22 is specifically used for: The target image is obtained by acquiring an image of an external carrier containing the target information captured by an in-vehicle camera device. The external carrier includes a display screen of a mobile terminal, printed material, or the surface of an object containing readable information.
[0099] In one example embodiment, the acquisition module 22 is specifically used for: The in-vehicle camera captures images containing the target's direction, and determines the user's spatial direction information based on the images. The target direction includes the user's gesture direction or gaze direction. Images of the external environment are captured by an external vehicle-mounted camera, and at least one candidate information carrier located outside the vehicle is identified based on the images. Based on the spatial pointing information and the spatial location information of each of the candidate information carriers, the target information carrier pointed to by the user is determined from at least one of the candidate information carriers; The vehicle-mounted camera outside the vehicle acquires an image containing the target information carrier, and identifies the image as the target image.
[0100] In one example embodiment, the acquisition module 22 is specifically used for: Acquire the original images captured by the vehicle-mounted camera device; The original image is preprocessed to obtain the target image. The preprocessing includes at least one of the following: resolution adjustment, image cropping, and image format adjustment.
[0101] In one example embodiment, the demand type includes a navigation type or a multimedia information playback type.
[0102] In one example embodiment, processing module 24 is specifically used for: Based on the aforementioned requirement type and target content, an operation confirmation message is generated; The operation confirmation information is output through the in-vehicle interactive interface; Upon receiving a confirmation instruction for the operation confirmation information, the operation instruction sequence is parsed and executed.
[0103] The apparatus of this embodiment can be used in any embodiment of the method in the side embodiment of the vehicle vision and voice collaborative interaction method based on multimodal large model. Its specific implementation process and technical effects are similar to those in the side embodiment of the vehicle vision and voice collaborative interaction method based on multimodal large model. For details, please refer to the detailed description in the side embodiment of the vehicle vision and voice collaborative interaction method based on multimodal large model, which will not be repeated here.
[0104] Figure 4 This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of the present invention, such as... Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a vehicle-mounted vision and voice collaborative interaction method based on a multimodal large model. The method includes: receiving a voice command input by a user through a vehicle-mounted voice acquisition module, the voice command representing the user's operational intention; in response to the voice command, acquiring a target image acquired by a vehicle-mounted camera device; determining an operation command sequence, the operation command sequence being determined by the multimodal large model based on the demand type and target content, the demand type being the type corresponding to the operational intention determined based on the text information corresponding to the voice command, and the target content being the content in the target image related to the demand type; parsing and executing the operation command sequence to complete at least one set of graphical user interface (GUI) operations corresponding to the operational intention.
[0105] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0106] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the in-vehicle vision and voice collaborative interaction method based on a multimodal large model provided by the above methods. The method includes: receiving a voice command input by a user through an in-vehicle voice acquisition module, the voice command being used to represent the user's operational intention; in response to the voice command, acquiring a target image acquired by an in-vehicle camera device; determining an operation command sequence, the operation command sequence being determined by a multimodal large model based on a demand type and target content, the demand type being the type corresponding to the operational intention determined based on the text information corresponding to the voice command, and the target content being the content in the target image related to the demand type; parsing and executing the operation command sequence to complete at least one set of graphical user interface (GUI) operations corresponding to the operational intention.
[0107] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the in-vehicle vision and voice collaborative interaction method based on a multimodal large model provided by the above methods. The method includes: receiving a voice command input by a user through an in-vehicle voice acquisition module, the voice command representing the user's operational intent; in response to the voice command, acquiring a target image acquired by an in-vehicle camera device; determining an operation command sequence, the operation command sequence being determined by a multimodal large model based on a demand type and target content, the demand type being the type corresponding to the operational intent determined based on text information corresponding to the voice command, and the target content being content in the target image related to the demand type; parsing and executing the operation command sequence to complete at least one set of graphical user interface (GUI) operations corresponding to the operational intent.
[0108] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0109] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for in-vehicle vision and voice collaborative interaction based on a multimodal large model, characterized in that, include: Receives voice commands input by the user through the vehicle-mounted voice acquisition module, wherein the voice commands are used to represent the user's operational needs and intentions; In response to the voice command, acquire the target image captured by the vehicle-mounted camera device; A sequence of operation instructions is determined, wherein the sequence of operation instructions is determined by a multimodal large model based on the demand type and the target content, wherein the demand type is the type corresponding to the operation demand intent determined based on the text information corresponding to the voice instruction, and the target content is the content in the target image related to the demand type; The sequence of operation instructions is parsed and executed to complete at least one set of graphical user interface (GUI) operations corresponding to the intended operation requirement.
2. The in-vehicle vision and voice collaborative interaction method based on a multimodal large model according to claim 1, characterized in that, The determination of the operation instruction sequence includes: The voice command and the target image are sent to the server, and the voice command and the target image are used to instruct the server to determine the demand type and the target content through the multimodal big model, and to determine the operation command sequence based on the demand type and the target content; Receive the sequence of operation instructions sent by the server.
3. The in-vehicle vision and voice collaborative interaction method based on a multimodal large model according to claim 1, characterized in that, The determination of the operation instruction sequence includes: The text information corresponding to the voice command and the target image are input into the multimodal large model. The demand type and the target content are determined by the multimodal large model, and the operation command sequence is determined based on the demand type and the target content.
4. The in-vehicle vision and voice collaborative interaction method based on a multimodal large model according to claim 1, characterized in that, Determining the target content through the multimodal large model includes: The target image is visually understood using the multimodal large model to obtain at least one information element in the target image; Based on the required type, at least one target information element related to the intention to complete the operation requirement is selected from at least one of the information elements; The target information element is determined as the target content.
5. The in-vehicle vision and voice collaborative interaction method based on a multimodal large model according to claim 1, characterized in that, The acquisition of the target image captured by the vehicle-mounted camera includes: The target image is obtained by acquiring an image of an external carrier containing target information captured by an in-vehicle camera device. The external carrier includes a display screen of a mobile terminal, printed material, or the surface of an object containing readable information.
6. The in-vehicle vision and voice collaborative interaction method based on a multimodal large model according to claim 1, characterized in that, The acquisition of the target image captured by the vehicle-mounted camera includes: The in-vehicle camera captures images containing the target's direction, and determines the user's spatial direction information based on the images. The target direction includes the user's gesture direction or gaze direction. Images of the external environment are captured by an external vehicle-mounted camera, and at least one candidate information carrier located outside the vehicle is identified based on the images. Based on the spatial pointing information and the spatial location information of each of the candidate information carriers, the target information carrier pointed to by the user is determined from at least one of the candidate information carriers; The vehicle-mounted camera outside the vehicle acquires an image containing the target information carrier, and identifies the image as the target image.
7. The in-vehicle vision and voice collaborative interaction method based on a multimodal large model according to any one of claims 1-6, characterized in that, The required types include navigation type or multimedia information playback type.
8. The in-vehicle vision and voice collaborative interaction method based on a multimodal large model according to any one of claims 1-6, characterized in that, The parsing and execution of the operation instruction sequence includes: Based on the aforementioned requirement type and target content, an operation confirmation message is generated; The operation confirmation information is output through the in-vehicle interactive interface; Upon receiving a confirmation instruction for the operation confirmation information, the operation instruction sequence is parsed and executed.
9. A vehicle-mounted vision and voice collaborative interaction system based on a multimodal large model, characterized in that, include: The vehicle-mounted voice acquisition module is used to acquire voice commands input by the user, and the voice commands are used to represent the user's operational needs and intentions. Vehicle-mounted camera equipment, used to capture images of the target; The vehicle controller is used to determine an operation command sequence, which is determined by a multimodal large model based on demand type and target content. The demand type is the type corresponding to the operation demand intent determined based on the text information corresponding to the voice command, and the target content is the content in the target image related to the demand type. The vehicle controller is also used to parse and execute the sequence of operation instructions to complete at least one set of user interface (GUI) operations corresponding to the intended operation requirement.
10. A vehicle-mounted vision and voice collaborative interaction device based on a multimodal large model, characterized in that, include: The receiving module is used to receive voice commands input by the user through the vehicle-mounted voice acquisition module, wherein the voice commands are used to represent the user's operational needs and intentions. The acquisition module is used to acquire the target image captured by the vehicle-mounted camera device in response to the voice command; The determination module is used to determine the sequence of operation instructions. The sequence of operation instructions is determined by a multimodal large model based on the demand type and target content. The demand type is the type corresponding to the operation demand intention determined based on the text information corresponding to the voice instruction. The target content is the content in the target image related to the demand type. The processing module is used to parse and execute the sequence of operation instructions to complete at least one set of graphical user interface (GUI) operations corresponding to the intended operation requirement.