Speech interaction method, system and device based on multi-modal large model and storage medium

By combining multimodal large models with speech, text, and visual information to infer user intent, the problem of in-vehicle voice interaction technology being unable to understand natural language commands and lacking visual information assistance is solved, achieving higher speech recognition accuracy and user experience.

CN122116895APending Publication Date: 2026-05-29HUIZHOU DESAY SV AUTOMOTIVE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUIZHOU DESAY SV AUTOMOTIVE
Filing Date
2025-12-26
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing in-vehicle voice interaction technologies cannot understand natural language commands that are beyond their scope and lack visual information assistance, resulting in insufficient voice recognition accuracy in complex in-vehicle environments and low accuracy in acquiring and parsing visual information from the interface, which seriously affects the user experience.

Method used

A voice interaction method based on a multimodal large model is adopted. By receiving user language commands, recognizing and converting them into speech text, and combining them with the visual information of the current interface of the vehicle system, semantic understanding and information acquisition are performed to realize user intent reasoning and output the results.

Benefits of technology

Without modifying system code or expanding the preset command library, it dynamically adapts to diverse user voice needs, improves semantic understanding accuracy, reduces development and maintenance costs, and avoids interaction failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116895A_ABST
    Figure CN122116895A_ABST
Patent Text Reader

Abstract

The embodiment of the application relates to the technical field of data processing, and discloses a voice interaction method, system and device based on a multi-modal large model and a storage medium, the method comprising: receiving a language instruction of a user; converting the language instruction into voice text; determining whether the voice text is a fixed instruction, if yes, directly executing, otherwise, executing the next step; obtaining visual information of a current interface, the multi-modal large model performing semantic understanding and information acquisition according to the voice text and the visual information, realizing user intention reasoning, and outputting the reasoning result; and displaying the output result. For non-fixed instructions, the multi-modal large model performs semantic understanding and information acquisition according to the voice text and the visual information, and realizes user intention reasoning. The reasoning ability of the multi-modal large model is used to dynamically adapt to diversified and personalized voice demands of users without modifying system codes or expanding a preset instruction library.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of data processing technology, specifically to a voice interaction method, system, device, and storage medium based on a multimodal large model. Background Technology

[0002] With the development of vehicle networking technology, voice interaction has become the core interaction method of in-vehicle systems and is widely used in scenarios such as music playback, navigation control, and information query.

[0003] Existing in-vehicle voice interaction technology mainly relies on a preset command set and application interface docking mode. The in-vehicle system can only recognize predefined fixed semantic commands and needs to dock with third-party applications in advance to provide corresponding services.

[0004] However, existing in-vehicle voice interaction technologies cannot understand natural language commands that are beyond their scope, resulting in insufficient voice recognition accuracy in complex in-vehicle environments and affecting the continuity of interaction.

[0005] In addition, existing in-vehicle voice interaction technology lacks visual information assistance. The in-vehicle system cannot autonomously infer answers by combining the current interface context, such as song information on the music playback interface or content cover of the video interface. This results in low accuracy in acquiring and parsing visual information from the interface, making it impossible to provide reliable input data for multimodal large models and seriously affecting the user experience. Summary of the Invention

[0006] To address the problems of existing in-vehicle voice interaction technologies being unable to understand natural language commands beyond their scope and lacking visual information assistance, resulting in insufficient voice recognition accuracy in complex in-vehicle environments and low accuracy in acquiring and parsing interface visual information, which seriously affects user experience, this invention provides a voice interaction method, system, device, and storage medium based on a multimodal large model.

[0007] According to an embodiment of the present invention, a voice interaction method based on a multimodal large model is provided, comprising the following steps: Receive user language commands; The voice command is recognized to convert the language command into speech text; Determine whether the spoken text belongs to a pre-trained fixed instruction. If so, execute the corresponding spoken instruction directly; otherwise, proceed to the next step. The system acquires the visual information of the current interface, inputs the speech text and the visual information into the multimodal large model, and the multimodal large model performs semantic understanding and information acquisition based on the speech text and the visual information to realize user intent reasoning and output the reasoning results. The output results will be displayed.

[0008] In some optional implementations, visual information of the current interface is obtained, and the speech text and visual information are input into a multimodal large model. The multimodal large model performs user intent inference based on the speech text and visual information, and outputs the inference result, specifically including: Capture screenshots of the current interface of the vehicle's infotainment system and camera footage, and process them to obtain visual information about the current interface; The spoken text and the visual information are input into a multimodal large model; The multimodal large model performs fusion processing on the speech text and the visual information; Semantic understanding and information acquisition are performed on the fused speech text and visual information to obtain user intent inference; The inference results are output according to a preset protocol.

[0009] In some optional implementations, user language commands are received, specifically including: The system receives the user's voice commands via an in-vehicle microphone or voice output device.

[0010] In some optional implementations, the output results are displayed, specifically including: The output results are read aloud via voice. And / or, display the output results via text or images.

[0011] According to another objective of embodiments of the present invention, a voice interaction system based on a multimodal large model is provided, comprising: The image recognition module is used to capture screenshots of the current interface of the vehicle system and camera images, and process them to obtain the visual information of the current interface. The voice interaction module is used to receive the user's language commands, recognize the voice commands, convert the language commands into speech text, and determine whether the speech text belongs to a fixed command that has been trained. If so, the corresponding voice command is executed directly; otherwise, the inference result output by the multimodal large model module is executed. A multimodal large model module is used to receive the speech text and the visual information, perform semantic understanding and information acquisition on the speech text and the visual information, realize user intent reasoning, and output the reasoning results to the voice interaction module; and The interactive display module is used to display the inference results output by the multimodal large model module executed by the voice interaction module according to the driving scenario.

[0012] Some optional implementations also include: an external interface scheduling module; The external interface scheduling module is used to interface with third-party applications to convert the inference results output by the multimodal large model module into operation signals that can be executed by the third-party applications.

[0013] Some optional implementations also include: a security management module; The safety management module is used to select different display methods for the inference results output by the multimodal large model module according to the current driving scenario.

[0014] In some optional implementations, the multimodal large model module includes an input module, an understanding module, and an output module; The input module is used to receive the voice text output by the voice interaction module and the visual information output by the image recognition module, and to perform fusion processing on the voice text and the visual information; The understanding module is used to perform semantic understanding and information acquisition on the fused speech text and visual information in order to obtain user intent reasoning; The output module is used to convert the inference results into structured data and output them according to a preset protocol.

[0015] According to another objective of the present invention, a computer device is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction that causes the processor to perform an operation of a voice interaction method based on a multimodal large model as described above.

[0016] According to another objective of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the voice interaction method based on a multimodal large model as described above.

[0017] Compared with the prior art, the present invention has the following advantages: This invention provides a voice interaction method based on a multimodal large model. For non-fixed commands, the multimodal large model performs semantic understanding and information acquisition based on speech, text, and visual information, enabling user intent reasoning. Without modifying system code or expanding the preset command library, the reasoning capability of the multimodal large model can dynamically adapt to the diverse and personalized voice needs of users. This breaks the limitation of traditional voice interaction, which can only recognize preset commands, allowing the voice interaction system to understand natural language commands beyond its scope, and significantly reducing the development and maintenance costs of the voice interaction system.

[0018] In addition, by using a multimodal large model to perform semantic understanding and information acquisition based on speech, text and visual information, user intent reasoning is realized, which solves the problems of traditional voice interaction lacking context support and being unable to resolve ambiguous references. Visual information provides context anchors for semantic understanding, improving the accuracy of semantic understanding and avoiding interaction failures caused by ambiguous instructions.

[0019] The above description is merely an overview of the technical solutions of the embodiments of the present invention. In order to better understand the technical means of the embodiments of the present invention and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0020] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 The diagram shows a flowchart of a voice interaction method based on a multimodal large model provided by an embodiment of the present invention.

[0021] Figure 2 The diagram illustrates the implementation flow of the inference result of a voice interaction method based on a multimodal large model provided by an embodiment of the present invention.

[0022] Figure 3 The diagram shows a structural block diagram of a voice interaction system based on a multimodal large model provided by an embodiment of the present invention.

[0023] Figure 4 A structural block diagram of a computer device provided by an embodiment of the present invention is shown. Detailed Implementation

[0024] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein.

[0025] To address the problems of existing in-vehicle voice interaction technologies being unable to understand natural language commands beyond their scope and lacking visual information assistance, resulting in insufficient voice recognition accuracy in complex in-vehicle environments and low accuracy in acquiring and parsing interface visual information, which seriously affects user experience, this invention provides a voice interaction method based on a multimodal large model.

[0026] This invention provides a voice interaction method based on a multimodal large model, such as... Figure 1 As shown, it includes the following steps: S10. Receive the user's language instructions.

[0027] This step is specifically as follows: The system receives the user's voice commands via an in-vehicle microphone or voice output device.

[0028] In this step, the vehicle microphone or voice output device is used for voice input, thus enabling the vehicle microphone or voice output device to capture the natural language commands issued by the user.

[0029] For example, when a user issues a natural language command such as "Who sings this song?" or "I want to watch a video of Three Little Pigs", the vehicle's microphone or voice output device can capture the corresponding voice command.

[0030] S20. Recognize the voice command to convert the language command into speech text.

[0031] This step is specifically as follows: Based on the above step S10, after the vehicle microphone or voice output device receives the voice command to be issued, it performs noise reduction, accent correction and other processing on the voice command, and converts the voice command into standard voice text to ensure the accuracy of subsequent understanding.

[0032] For example, after a car microphone or voice output device receives a user's natural language command such as "Who sings this song?" or "I want to watch a video of Three Little Pigs", it converts it into a voice-text format of "Who sings this song?" or "I want to watch a video of Three Little Pigs".

[0033] S30. Determine whether the spoken text belongs to a pre-trained fixed instruction. If so, execute the corresponding spoken instruction directly; otherwise, proceed to the next step.

[0034] This step is specifically as follows: The voice interaction system queries the built-in fixed command library to determine whether the voice text obtained in step S20 is a pre-trained fixed command, such as "open navigation" or "open music". If so, the corresponding voice command is executed directly. That is, if the voice text obtained in step S20 is determined to be a pre-trained fixed command, the corresponding operation is executed directly, and the process ends.

[0035] If it is determined that the speech text obtained in step S20 is not a fixed instruction from the training, such as "Who sings this song?" or "I want to watch the Three Little Pigs video", then proceed to the next step, that is, enter the multimodal large model understanding process.

[0036] S40. Obtain the visual information of the current interface, input the speech text and the visual information into the multimodal large model, the multimodal large model performs semantic understanding and information acquisition based on the speech text and the visual information, realizes user intent reasoning, and outputs the reasoning results.

[0037] This step is specifically as follows: When step S30 determines that the voice text obtained in step S20 is not a fixed instruction for training, it captures the visual information of the current interface of the vehicle system according to the voice instruction, inputs the voice text and visual information into the multimodal large model, and establishes a relationship between the voice text and visual information to realize semantic understanding and information acquisition of the voice text and visual information, thereby realizing user intent reasoning and outputting the reasoning results.

[0038] For example, the multimodal big data model establishes an association between the voice text "Who sings this song?" and the visual information "Anti-Hero" and "Midnights" corresponding to the current music interface. By combining contextual learning, it performs semantic understanding and information retrieval on the voice text and visual information, understands the user's intent "to obtain information about the singer of the currently playing song", and combines the multimodal big data model knowledge base or external retrieval to infer that the singer is Taylor Swift, expands the album-related information, and outputs the inference result.

[0039] S50. Display the output results.

[0040] This step is specifically as follows: The results are output via voice broadcast; And / or, display the output results via text or images.

[0041] That is, this step enables the display of the inference results output in step S40 through voice broadcast or text display.

[0042] This embodiment provides a voice interaction method based on a multimodal large model. For non-fixed commands, the multimodal large model performs semantic understanding and information acquisition based on voice text and visual information, realizing user intent reasoning. Without modifying the system code or expanding the preset command library, the reasoning capability of the multimodal large model can dynamically adapt to the diverse and personalized voice needs of users, breaking the limitation of traditional voice interaction that can only recognize preset commands. This allows the voice interaction system to understand natural language commands beyond the scope, significantly reducing the development and maintenance costs of the voice interaction system.

[0043] In addition, by using a multimodal large model to perform semantic understanding and information acquisition based on speech, text and visual information, user intent reasoning is realized, which solves the problems of traditional voice interaction lacking context support and being unable to resolve ambiguous references. Visual information provides context anchors for semantic understanding, improving the accuracy of semantic understanding and avoiding interaction failures caused by ambiguous instructions.

[0044] This embodiment, as a preferred embodiment, optimizes the specific implementation process of step S40 in the voice interaction method based on a multimodal large model provided in this embodiment: acquiring the visual information of the current interface, inputting the voice text and the visual information into the multimodal large model, and the multimodal large model performing semantic understanding and information acquisition based on the voice text and the visual information to realize user intent reasoning and outputting the reasoning results.

[0045] Before this step, the training of a large multimodal model was first implemented.

[0046] The training of this multimodal large model involves: pre-training for in-vehicle scenarios based on a general large model, incorporating data from in-vehicle application scenarios such as music, video, and navigation to improve the ability to understand scenarios; and the training process of this multimodal large model also supports continuous optimization through user interaction data to dynamically expand the scope of semantic understanding, adapting to new scenarios and new instructions without reconstructing the multimodal large model.

[0047] The multimodal large model in this embodiment has the capabilities of context learning, incremental training, and complex instruction decomposition, which can decompose the user's ambiguous and complex natural language instructions into simple task chains that can be executed by the voice interaction system.

[0048] In this embodiment, step S40 involves acquiring the visual information of the current interface, inputting the speech text and the visual information into a multimodal large model, and then the multimodal large model performs user intent inference based on the speech text and the visual information, and outputs the inference result, such as... Figure 2 As shown, it specifically includes: S410. Capture a screenshot of the current interface of the vehicle's infotainment system and the camera image, and process them to obtain visual information of the current interface.

[0049] This step is specifically as follows: Based on voice commands, screenshots of the vehicle's current interface and camera images are captured, and interface elements, text information, and visual features are extracted. The information from multiple images is then segmented, labeled, and understood separately to obtain the visual information of the current interface.

[0050] S420. Input the speech text and the visual information into the multimodal large model.

[0051] In this step, the multimodal large model enables the simultaneous reception of speech, text, and visual information.

[0052] S430. The multimodal large model performs fusion processing on the speech text and the visual information.

[0053] In this step, the multimodal large model implements semantic classification and labeling of text and icons in visual information, and semantically aligns them with keywords in speech text to establish a correspondence between text references and visual entities.

[0054] S440. Perform semantic understanding and information acquisition on the fused speech text and visual information to obtain user intent reasoning.

[0055] This step is specifically as follows: By training a multimodal fusion algorithm and combining it with context, the algorithm can uncover the deep connections between speech, text, and visual information, thereby inferring user intent and solving the problems of traditional voice interaction lacking context and being unable to understand ambiguous references.

[0056] S450. Output the inference result according to a preset protocol.

[0057] This step is specifically as follows: After converting the unstructured inference results output by the multimodal large model into standardized and structured data, the data is output according to a preset protocol.

[0058] In some optional embodiments, the present invention also discloses a voice interaction system based on a multimodal large model.

[0059] This embodiment discloses a voice interaction system based on a multimodal large model, such as... Figure 3 As shown, it includes: an image recognition module 100, a voice interaction module 200, a multimodal large model module 300, and an interactive display module 400.

[0060] in: The image recognition module 100 is used to capture screenshots of the current interface of the vehicle system and camera images, and process them to obtain visual information of the current interface.

[0061] In this embodiment, the image recognition module 100 serves as a visual perception entry point, responsible for capturing screenshots of the vehicle interface and camera images, extracting interface elements, text information, and visual features, and accurately outputting structured visual information data according to voice commands.

[0062] The image recognition module 100 supports two methods of acquiring visual information: screen capture and application programming interface (API) calls. It can segment and annotate multi-image information and understand it separately, ensuring the accuracy and relevance of the output visual information data and providing reliable visual input for multimodal large models.

[0063] In addition, the image recognition module 100 is compatible with all in-vehicle visualization scenarios, such as music playback interfaces, video application interfaces, and navigation interfaces, solving the problem of lack of contextual support in voice interaction.

[0064] In this embodiment, the process of the image recognition module 100 acquiring visual information is described in detail in step S40 of the above-mentioned voice interaction system based on a multimodal large model. This embodiment will not repeat the description of the process.

[0065] The voice interaction module 200 is used to receive the user's language commands, recognize the voice commands, convert the language commands into speech text, and determine whether the speech text belongs to a fixed command that has been trained. If so, the corresponding voice command is executed directly; otherwise, the inference result output by the multimodal large model module is executed.

[0066] In this embodiment, the voice interaction module 200 receives user voice input and completes front-end processing and back-end implementation of voice recognition, text conversion, command judgment and action execution. It is the core hub connecting the user and the voice interaction system.

[0067] In the process of speech recognition, the voice interaction module 200 adopts advanced noise reduction and accent correction algorithms to adapt to the complex in-vehicle environment and improve the accuracy of speech-to-text conversion. In the process of instruction judgment, it quickly recognizes and executes regular instructions through a built-in pre-trained fixed instruction library, while non-fixed instructions trigger the multimodal large model module. In the process of action execution, it receives structured instructions output by the multimodal large model and triggers actions such as voice broadcasting and application operation, connecting the understanding and execution stages.

[0068] In this embodiment, the process of the voice interaction module 200 acquiring and processing the user's language commands is described in detail in the specific implementation process of steps S10-S30 in the above-mentioned voice interaction system based on a multimodal large model. This embodiment will not repeat the description.

[0069] The multimodal large model module 300 is used to receive the speech text and the visual information, perform semantic understanding and information acquisition on the speech text and the visual information, realize user intent reasoning, and output the reasoning results to the speech interaction module.

[0070] In this embodiment, the multimodal large model module 300 has the capabilities of context learning, incremental training, and complex instruction decomposition, which can decompose the user's fuzzy and complex natural language instructions into simple task chains that the system can execute.

[0071] In this embodiment, the reasoning process of the multimodal large model module 300 regarding the user's intent is described in detail in the specific implementation process of step S40 in the above-mentioned voice interaction system based on a multimodal large model, and will not be repeated in this embodiment.

[0072] The interactive display module 400 is used to display the inference results output by the multimodal large model module executed by the voice interaction module according to the driving scenario.

[0073] In this embodiment, the interactive display module 400 displays the interactive results according to the characteristics of the vehicle scenario, and supports multiple forms of feedback such as voice broadcast and / or text or image display to meet the information acquisition needs of users in different scenarios.

[0074] The interactive display module 40 can provide secondary interaction options based on the user's interaction status, and supports two secondary interaction triggering methods: voice and touch, which improves the continuity and flexibility of the interaction.

[0075] In this embodiment, the process of the interactive display module 400 displaying the reasoning results is described in detail in the specific implementation process of step S50 in the above-mentioned voice interaction system based on a multimodal large model. This embodiment will not repeat the description.

[0076] This embodiment is a preferred embodiment. A voice interaction system based on a multimodal large model also includes: an external interface scheduling module 500 and a security management module 600.

[0077] The external interface scheduling module 500 is used to interface with third-party applications to convert the inference results output by the multimodal large model module into operation signals that can be executed by the third-party applications.

[0078] In this embodiment, the external interface scheduling module 500 serves as the interface hub between the voice interaction system and third-party applications, enabling the conversion of abstract instructions generated by the multimodal large model into executable operation signals for third-party applications, without requiring applications to pre-customize interface interfaces.

[0079] The external interface scheduling module 500 supports universal interface adaptation with mainstream in-vehicle applications, and can convert commands such as "play a certain album" and "open a certain video" into API call signals of the application, realizing seamless conversion from multimodal large model inference results to application operations and reducing the cost of multi-application integration.

[0080] The safety management module 600 is used to select different display methods for the inference results output by the multimodal large model module according to the current driving scenario.

[0081] In this embodiment, the safety management module 600 controls interactive behavior based on driving status levels, balancing the convenience of interaction with driving safety, which is a dedicated safety design for in-vehicle scenarios.

[0082] The safety management module 600 is categorized into the following scenarios: during driving, it primarily uses concise voice prompts to avoid distracting the driver with complex text or images; in parking or low-speed scenarios, it can display detailed information. The safety management module 600 also performs risk assessments on vehicle control-related commands. Actively interactive vehicle control commands are not executed directly but require secondary confirmation from the user to avoid misoperation. Complex operations and content that may affect driving safety are automatically masked during driving, retaining only core information.

[0083] This embodiment is a preferred embodiment, and the multimodal large model module 300 includes an input module 310, an understanding module 320, and an output module 330.

[0084] The input module 310 is used to receive the voice text output by the voice interaction module 200 and the visual information output by the image recognition module 100, and to perform fusion processing on the voice text and visual information.

[0085] In this embodiment, the input module 310 receives the voice and text output by the voice-text interaction module 200 and the visual information output by the image recognition module 100, and establishes a connection through a cross-modal attention mechanism to achieve the fusion processing of voice and text and visual information.

[0086] The understanding module 320 is used to perform semantic understanding and information acquisition on the fused speech, text and visual information in order to obtain user intent reasoning.

[0087] In this embodiment, the understanding module 320 parses the user intent and infers the user intent by combining the multimodal large model knowledge base or external retrieval.

[0088] The output module 330 is used to convert the inference results into structured data and output them according to a preset protocol.

[0089] In this embodiment, the output module 330 converts the reasoning results into structured data and outputs it according to a preset protocol.

[0090] In some alternative embodiments, the present invention also discloses a computer device.

[0091] Figure 4 A schematic diagram of the structure of a computer device provided in an embodiment of the present invention is shown.

[0092] like Figure 4 As shown, the computer device includes: a processor 710, a memory 720, a communication interface 730, and a communication bus 740. The processor 710, the memory 720, and the communication interface 730 communicate with each other through the communication bus 740.

[0093] In this embodiment, the memory 720 stores at least one executable instruction, which causes the processor 720 to perform operations of a voice interaction method based on a multimodal large model as described in any of the above embodiments; the communication interface 730 is used to communicate with other network elements such as clients or other servers. The processor 710 is used to execute program 750, specifically performing the relevant steps in the above embodiments of the voice interaction method based on a multimodal large model.

[0094] Specifically, program 750 may include program code, which includes computer-executable instructions.

[0095] The processor 710 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The computer device includes one or more processors, which may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.

[0096] Memory 720 is used to store program 750. Memory 720 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0097] Specifically, program 750 can be called by processor 710 to cause the computer device to perform the operation of a voice interaction method based on a multimodal large model as described in any of the above embodiments.

[0098] In some optional embodiments, the present invention also discloses a computer-readable storage medium storing at least one executable instruction that, when executed on a computer device, causes the computer device to perform the steps of a voice interaction method based on a multimodal large model in any of the above method embodiments.

[0099] The specific implementation process of the voice interaction method based on a multimodal large model described in this embodiment can be found in any of the above method embodiments, and will not be repeated in this embodiment.

[0100] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. Similarly, for the sake of brevity and to aid in understanding one or more aspects of the invention, in the description of exemplary embodiments of the invention above, various features of the embodiments are sometimes grouped together in a single embodiment, figure, or description thereof. The claims, which follow the detailed description, are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of the invention.

[0101] Those skilled in the art will understand that the modules in the device of the embodiment can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiment can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components, except that at least some of such features and / or processes or units are mutually exclusive.

[0102] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.

Claims

1. A voice interaction method based on a multimodal large model, characterized in that, Includes the following steps: Receive user language commands; The voice command is recognized to convert the language command into speech text; Determine whether the spoken text belongs to a pre-trained fixed instruction. If so, execute the corresponding spoken instruction directly; otherwise, proceed to the next step. The system acquires the visual information of the current interface, inputs the speech text and the visual information into the multimodal large model, and the multimodal large model performs semantic understanding and information acquisition based on the speech text and the visual information to realize user intent reasoning and output the reasoning results. The output results will be displayed.

2. The voice interaction method based on a multimodal large model according to claim 1, characterized in that, The system acquires visual information from the current interface, inputs the speech text and visual information into a multimodal large model, and the multimodal large model performs user intent inference based on the speech text and visual information, and outputs the inference result, specifically including: Capture screenshots of the current interface of the vehicle's infotainment system and camera footage, and process them to obtain visual information about the current interface; The spoken text and the visual information are input into a multimodal large model; The multimodal large model performs fusion processing on the speech text and the visual information; Semantic understanding and information acquisition are performed on the fused speech text and visual information to obtain user intent inference; The inference results are output according to a preset protocol.

3. The voice interaction method based on a multimodal large model according to claim 1, characterized in that, Receiving user language commands, specifically including: The system receives the user's voice commands via an in-vehicle microphone or voice output device.

4. The voice interaction method based on a multimodal large model according to claim 1, characterized in that, The output results will be displayed, specifically including: The output results will be read aloud via voice. And / or, display the output results via text or images.

5. A voice interaction system based on a multimodal large model, characterized in that, include: The image recognition module is used to capture screenshots of the current interface of the vehicle system and camera images, and process them to obtain the visual information of the current interface. The voice interaction module is used to receive the user's language commands, recognize the voice commands, convert the language commands into speech text, and determine whether the speech text belongs to a fixed command that has been trained. If so, the corresponding voice command is executed directly; otherwise, the inference result output by the multimodal large model module is executed. The multimodal large model module is used to receive the speech text and the visual information, perform semantic understanding and information acquisition on the speech text and the visual information, realize user intent reasoning, and output the reasoning results to the speech interaction module; as well as The interactive display module is used to display the inference results output by the multimodal large model module executed by the voice interaction module according to the driving scenario.

6. A voice interaction system based on a multimodal large model according to claim 5, characterized in that, Also includes: External interface scheduling module; The external interface scheduling module is used to interface with third-party applications to convert the inference results output by the multimodal large model module into operation signals that can be executed by the third-party applications.

7. A voice interaction system based on a multimodal large model according to claim 5, characterized in that, Also includes: Security management module; The safety management module is used to select different display methods for the inference results output by the multimodal large model module according to the current driving scenario.

8. A voice interaction system based on a multimodal large model according to claim 5, characterized in that, The multimodal large model module includes an input module, an understanding module, and an output module; The input module is used to receive the voice text output by the voice interaction module and the visual information output by the image recognition module, and to perform fusion processing on the voice text and the visual information; The understanding module is used to perform semantic understanding and information acquisition on the fused speech text and visual information in order to obtain user intent reasoning; The output module is used to convert the inference results into structured data and output them according to a preset protocol.

9. A computer device, characterized in that, include: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction that causes the processor to perform an operation of a voice interaction method based on a multimodal large model as described in any one of claims 1-4.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute claim 1. The voice interaction method based on a multimodal large model as described in any one of the 4.