Voice interaction method and device, equipment and storage medium

By acquiring audio signals in the vehicle terminal and using a large language model to identify their type, and outputting corresponding vehicle control tasks or prompts, the problem of high model complexity and long processing time in the prior art is solved, and efficient and accurate vehicle control task execution is achieved.

CN120954403APending Publication Date: 2025-11-14GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511172665.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In existing technologies, vehicle terminals need to first determine the type of voice, and only execute vehicle control tasks if the type is an instruction type. This results in high model complexity, excessive processing time, and low efficiency.

Method used

By acquiring audio signals collected by the vehicle terminal, the target prompt words are determined based on the audio signals. The type of audio signal is identified using a large language model. The vehicle control task is output under executable command audio signals, otherwise the target prompt information is output. This reduces the model for voice type recognition and lowers the model complexity.

Benefits of technology

It achieves high efficiency and accuracy in vehicle control tasks, reduces the space occupied by the model used to process audio signals, lowers model complexity, and ensures the processing accuracy of vehicle control tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954403A_ABST
    Figure CN120954403A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a voice interaction method and device, equipment and a storage medium, and the method is applied to a vehicle-mounted terminal, and comprises the steps: obtaining an audio signal collected by the vehicle-mounted terminal; a target prompt word is determined based on the audio signal, the type of the audio signal is obtained according to the target prompt word, and the type of the audio signal comprises an executable instruction audio signal; when the audio signal is an executable instruction audio signal, outputting a vehicle control task corresponding to the audio signal, the vehicle control task being determined based on the target prompt word; under the condition that the audio signal is not the executable instruction audio signal, target prompt information is output, and the target prompt information is used for indicating that the audio signal does not have the corresponding vehicle control task. The obtained vehicle control task has both efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to vehicle technology, and includes, but is not limited to, a voice interaction method, device, equipment, and storage medium. Background Technology

[0002] During user-vehicle interaction, voice is usually required. However, since the voice given by the user may sometimes be a command and sometimes just casual conversation, the vehicle needs to determine the type of the voice before executing the corresponding task determination logic if the voice is a command.

[0003] In related technologies, during the interaction between a user and the vehicle's in-vehicle terminal, the in-vehicle terminal can first determine the type of voice, and only if the type of voice is a command type will it determine the vehicle control task that needs to be executed. Summary of the Invention

[0004] The voice interaction method, apparatus, device, and storage medium provided in this application embodiment are implemented as follows:

[0005] One aspect of this application provides a voice interaction method applied to an in-vehicle terminal, including:

[0006] Acquire audio signals collected by the vehicle-mounted terminal;

[0007] The target prompt word is determined based on the audio signal, and the type of audio signal is obtained based on the target prompt word. The type of audio signal includes: executable instruction audio signal;

[0008] When the audio signal is an executable instruction audio signal, the vehicle control task corresponding to the audio signal is output, and the vehicle control task is determined based on the target prompt word;

[0009] If the audio signal is not an executable instruction audio signal, output a target prompt message. The target prompt message is used to indicate that there is no corresponding vehicle control task for the audio signal.

[0010] Another aspect of this application embodiment provides a voice interaction device applied to an in-vehicle terminal, including: an acquisition module, a determination module, and an execution module;

[0011] The acquisition module is used to acquire audio signals collected by the vehicle terminal;

[0012] The determination module is used to determine the target prompt word based on the audio signal, and to obtain the type of audio signal according to the target prompt word. The type of audio signal includes: executable instruction audio signal;

[0013] The execution module is used to output the vehicle control task corresponding to the audio signal when the audio signal is an executable instruction audio signal. The vehicle control task is determined based on the target prompt word.

[0014] The execution module is also used to output target prompt information when the audio signal is not an executable instruction audio signal. The target prompt information is used to indicate that there is no corresponding vehicle control task for the audio signal.

[0015] The computer device provided in this application includes a memory and a processor. The memory stores a computer program that can run on the processor, and the processor executes the program to implement the method of this application.

[0016] The computer-readable storage medium provided in this application embodiment stores a computer program thereon, which, when executed by a processor, implements the method provided in this application embodiment.

[0017] The voice interaction method, apparatus, device, and storage medium provided in this application embodiment can acquire audio signals collected by an in-vehicle terminal; determine target prompt words based on the audio signals; and obtain the type of audio signals based on the target prompt words. The types of audio signals include: executable instruction audio signals; when the audio signal is an executable instruction audio signal, output a vehicle control task corresponding to the audio signal, the vehicle control task being determined based on the target prompt words; when the audio signal is not an executable instruction audio signal, output target prompt information, the target prompt information indicating that there is no corresponding vehicle control task for the audio signal. Specifically, after acquiring the audio signal, the target prompt words can be determined based on the audio signal, thereby determining the vehicle control task corresponding to the audio signal. Then, based on the determined type of the audio signal, the execution of the vehicle control task or the output of target prompt information can be determined. By synchronously identifying the type of the audio signal and the corresponding vehicle control task, the space occupied by the model used to process the audio signal can be reduced. Since no additional model is set up to identify the type of the audio signal, parameter redundancy in the processing process is also reduced. This reduces model complexity while ensuring the processing accuracy of the vehicle control task, resulting in a vehicle control task that is both efficient and accurate. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1This is a schematic diagram illustrating the application scenario provided in the embodiments of this application;

[0020] Figure 2 This is a flowchart illustrating the related technologies provided in the embodiments of this application;

[0021] Figure 3 This is a flowchart illustrating the voice interaction method provided in the embodiments of this application;

[0022] Figure 4 This is a schematic diagram of the process for determining the type of audio signal provided in the embodiments of this application;

[0023] Figure 5 This is a schematic diagram of the process for generating target prompt words provided in the embodiments of this application;

[0024] Figure 6 This is another schematic diagram of the process for generating target prompt words provided in the embodiments of this application;

[0025] Figure 7 This is a flowchart illustrating the output target prompt information provided in the embodiments of this application;

[0026] Figure 8 This is a schematic diagram of the process for executing the vehicle control task provided in the embodiments of this application;

[0027] Figure 9 This is a schematic diagram of the overall process of the voice interaction method provided in the embodiments of this application;

[0028] Figure 10 This is a schematic diagram of the structure of the voice interaction device provided in the embodiments of this application;

[0029] Figure 11 This is a schematic diagram of the structure of the computer device provided in the embodiments of this application. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the specific technical solutions of this application will be further described in detail below with reference to the accompanying drawings of the embodiments of this application. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.

[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0032] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0033] It should be noted that the terms "first, second, third" used in the embodiments of this application are used to distinguish similar or different objects and do not represent a specific order of objects. It can be understood that "first, second, third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0034] To more clearly explain the voice interaction method provided in the embodiments of this application, a possible application scenario in the embodiments of this application will be explained in detail below.

[0035] Figure 1 This is a schematic diagram of the application scenario provided in the embodiments of this application. Please refer to it. Figure 1 This scenario may include: an in-vehicle terminal 100, which may be, for example, a smart vehicle system. The in-vehicle terminal 100 is installed in the smart cockpit and can have intelligent dialogue with the user in the smart cockpit. For example, it can generate corresponding answers based on the user's questions, or execute corresponding vehicle control tasks based on the user's instructions.

[0036] Specifically, generating corresponding answers to user questions can be achieved through chat-related functions; executing corresponding vehicle control tasks based on user commands can be done by determining the user's intent based on their voice, thereby executing the corresponding vehicle control tasks.

[0037] The vehicle terminal 100 can execute corresponding vehicle control tasks according to the user's instructions. For example, if the user instructs to open a certain door, the vehicle terminal 100 can execute the task of opening the corresponding door based on the instruction.

[0038] In this scenario, one user or multiple users can interact with the vehicle terminal 100. There are no specific restrictions. For example, the user in the driver's seat can interact with the vehicle terminal 100, and the user in the passenger seat can also interact with the vehicle terminal 100.

[0039] During the interaction between the user and the vehicle terminal 100, communication can be achieved through language. For example, the user can ask a question, the vehicle terminal 100 can collect the corresponding audio signal to obtain a voice command, and then perform voice recognition on the voice command to determine the user's intention. Subsequently, based on the large language model, the vehicle control task to be performed or the response to be given can be determined, etc. No specific restrictions are imposed here.

[0040] It should be noted that during the interaction between the user and the vehicle, the voice provided by the user can be casual conversation or a command to perform vehicle control tasks. In order to achieve voice recognition, in related technologies, during the interaction between the user and the vehicle's on-board terminal, the on-board terminal can first determine the type of the voice, and only if the voice type is a command type will it determine the vehicle control task to be performed.

[0041] The following explains one implementation process implemented in the relevant technology.

[0042] Figure 2 The following is a flowchart illustrating the related technologies provided in the embodiments of this application. Please refer to it. Figure 2 The method includes:

[0043] First, we use the first neural network model, which is... Figure 2 Model 1, as shown, determines the type of speech, and when the speech type is an instruction type, it uses a second neural network model, that is... Figure 2 Model 2 in the model is used to determine the vehicle control task corresponding to the voice, thereby completing the process of determining the vehicle control task.

[0044] In actual implementation, determining the type of speech requires the use of a neural network model. Similarly, determining the vehicle control task also typically requires a neural network model. This may result in the vehicle terminal storing two different models, one for determining the speech type and the other for determining the vehicle control task. This leads to high model complexity in the entire vehicle terminal and excessively long speech processing time, resulting in low efficiency.

[0045] To address the aforementioned problems in related technologies, this application provides a voice interaction method, the actual implementation of which will be explained below.

[0046] Figure 3 This is a flowchart illustrating the voice interaction method provided in the embodiments of this application. Please refer to... Figure 3 Voice interaction methods include:

[0047] S310: Acquire audio signals collected by the vehicle terminal.

[0048] It should be noted that the subject of this method can be the aforementioned vehicle-mounted terminal, through which multiple audio signals can be collected.

[0049] In one embodiment, the audio signal can be a voice command given by the user, such as "Open the driver's side window," "Close the passenger side door," "Turn up the media playback volume," or "Turn down the air conditioning temperature," which requires the in-vehicle terminal to perform a corresponding vehicle control task. Alternatively, the audio signal can also be casual conversation information given by the user, such as "What's the outdoor temperature today?" or "Sing me a song," which requires the in-vehicle terminal to respond to the user's question.

[0050] In one embodiment, an audio acquisition device, such as a microphone or radio, can be installed in the smart cabin. No specific limitation is made here. The audio acquisition device can communicate with the vehicle terminal, and the vehicle terminal can obtain the audio signal provided by the user through the audio acquisition device.

[0051] For example, while driving, a user can verbally request something, such as "Turn on navigation." This request is captured by the audio capture device and sent to the in-vehicle terminal as an audio signal. When resting in the car, a user can chat with the in-vehicle terminal, such as saying, "The weather is nice today." This chat is also captured by the audio capture device and sent to the in-vehicle terminal as an audio signal.

[0052] It should be noted that each time the vehicle terminal collects an audio signal, it can execute the voice interaction method provided in this application embodiment once, thereby realizing the recognition and processing of the audio signal, and then determining whether to execute the vehicle control task or generate the corresponding voice response.

[0053] S320: Determine the target prompt word based on the audio signal, and obtain the type of audio signal based on the target prompt word.

[0054] It should be noted that target prompts can be determined based on audio signals. These target prompts can be generated based on audio signals or extracted from the text information corresponding to the audio signals; no specific restrictions are imposed here.

[0055] In one embodiment, the vehicle terminal can determine the target prompt word corresponding to the audio signal by searching, and then determine the type of the audio signal based on the target prompt word.

[0056] The type of audio signal can be obtained by inputting the target prompt word into a neural network model, such as a large language model, or by determining the type of audio signal based on the degree of matching between the target prompt word and the audio signal. No specific restrictions are imposed here.

[0057] Optionally, the type of audio signal includes: executable instruction audio signal.

[0058] It should be noted that, as explained above, audio signals can include various types such as command type and chat type. Among them, executable command audio signals can be one type of command, which can refer to the audio signal corresponding to the command that the vehicle can execute.

[0059] For example, "Open the left window" and "Activate automatic parking mode" are both command-type audio signals. If the vehicle has the ability to automatically open the left window, then the audio signal "Open the left window" can be determined to be an executable command audio signal. If the vehicle does not have the ability to activate automatic parking mode, then the audio signal "Activate automatic parking mode" can be determined to be an executable command audio signal.

[0060] The type of audio signal can be determined in the above way, that is, whether the type of audio signal is an executable instruction audio signal.

[0061] S330: When the audio signal is an executable instruction audio signal, output the vehicle control task corresponding to the audio signal.

[0062] It should be noted that if the audio signal is determined to be an executable command audio signal, the vehicle corresponding to that audio signal can be output to perform the task.

[0063] Among them, the vehicle control task is determined based on the target cue words.

[0064] For example, the vehicle control task can be obtained by inputting the target prompt word and the audio signal into a large language model; or, the vehicle control task corresponding to the target prompt word can be determined based on the degree of matching between the audio signal and the target prompt word, etc., without specific limitations.

[0065] In one embodiment, the vehicle control task can be determined by recognizing the audio signal.

[0066] For example, for the audio signal "open the driver's side door", the audio signal can be identified and it can be determined that the audio signal is a task that requires opening the driver's side door.

[0067] Among them, the vehicle control task can be a specific task for the vehicle where the on-board terminal is located, or it can be controlling a certain unit of the vehicle to perform a certain task, such as controlling the opening of the passenger side door of the vehicle.

[0068] It should be noted that vehicle control tasks can be output in a specified format, for example:<API,Araguments> In this context, API can be the corresponding execution unit, representing a unit that performs the task, such as a car window, a car door, an air conditioner, or a display, without specific limitations. Araguments can be the corresponding execution parameters, representing parameter values ​​for performing the task. For example, execution parameters for a switch class can be represented by 0 / 1 to indicate whether it is closed or open; execution parameters for an adjustment class can be set by a specific numerical value.

[0069] It should be noted that after outputting the vehicle control task corresponding to the audio signal, the corresponding execution unit can be controlled to execute the vehicle control task.

[0070] For example, vehicle control tasks that can be implemented by the in-vehicle terminal, such as volume control and media playback, can be executed by the in-vehicle terminal as an execution unit. For instance, the in-vehicle terminal can increase the playback volume or switch the playback media. For vehicle control tasks that cannot be implemented by the in-vehicle terminal, such as door control and window control, the corresponding execution unit can execute them. For instance, the in-vehicle terminal can send the corresponding vehicle control task to the door control unit or window control unit, which can then control the opening or closing of the doors or windows.

[0071] S340: If the audio signal is not an executable instruction audio signal, output the target prompt message.

[0072] It should be noted that if it is determined that the audio signal is not an executable instruction audio signal, a target prompt message can be output.

[0073] The target prompt information is used to indicate that there is no corresponding vehicle control task for the audio signal.

[0074] In one embodiment, an audio signal that is not an executable instruction may include various situations, such as: the audio signal is an instruction type audio signal, but the vehicle cannot execute it, and the audio signal is not an instruction type audio signal, such as: an unrecognizable audio signal, a chat-type audio signal, an audio signal containing harmful information, etc., without specific limitations.

[0075] In these cases, instead of outputting a vehicle control task, a target prompt message will be output, which may include the fact that the audio signal does not correspond to a vehicle control task.

[0076] In addition, the target prompt information can also include other content. Different types of audio signals can include different content. For example, for chatty audio signals, chatty responses can be generated; for audio signals containing harmful information, prompts for harmful information can be generated, such as harmful words, phrases, etc. There are no specific restrictions here; for unrecognizable audio signals, prompts for re-entry can be given.

[0077] The voice interaction method provided in this application embodiment can acquire audio signals collected by an in-vehicle terminal; determine target prompt words based on the audio signals; and obtain the type of audio signals based on the target prompt words. The types of audio signals include: executable instruction audio signals; when the audio signal is an executable instruction audio signal, output a vehicle control task corresponding to the audio signal, the vehicle control task being determined based on the target prompt words; when the audio signal is not an executable instruction audio signal, output target prompt information, the target prompt information indicating that there is no corresponding vehicle control task for the audio signal. Specifically, after acquiring the audio signal, the target prompt words can be determined based on the audio signal, thereby determining the vehicle control task corresponding to the audio signal. Then, based on the determined type of the audio signal, the execution of the vehicle control task or the output of target prompt information can be determined. By synchronously identifying the type of the audio signal and the corresponding vehicle control task, the space occupied by the model used to process the audio signal can be reduced. Since no additional model is set up to identify the type of the audio signal, parameter redundancy in the processing process is also reduced. This reduces model complexity while ensuring the processing accuracy of the vehicle control task, resulting in a vehicle control task that is both efficient and accurate.

[0078] The following explains one possible implementation process for determining the type of audio signal provided in the embodiments of this application.

[0079] Figure 4 This is a flowchart illustrating the process of determining the type of audio signal provided in the embodiments of this application. Please refer to... Figure 4 The target prompt word is determined based on the audio signal, and the type of audio signal is obtained based on the target prompt word, including:

[0080] S410: Determine target prompt words based on audio signals.

[0081] It should be noted that target prompts can be obtained through retrieval. Different retrieval methods can be used for different types of audio signals. Alternatively, to ensure the accuracy of the retrieval, multiple retrieval methods can be combined to determine the target prompts based on the audio signal.

[0082] In one embodiment, the target prompt word can be a single prompt word, or it can be multiple prompt words. For multiple prompt words, they can be prompt words obtained using one retrieval method, or they can be multiple prompt words obtained using multiple retrieval methods.

[0083] For example, the target cue word could be something like "car window" or "open," or related words like "media volume" or "increase." There are no specific restrictions here, and the target cue word can be determined based on the actual audio signal.

[0084] It should be noted that there may be cases where the target prompt word cannot be determined. In such cases, the type of audio signal can be defined as an unrecognizable audio signal.

[0085] S420: Input the target prompt word and audio signal into the large language model, and determine the type of audio signal through the large language model.

[0086] In one embodiment, after obtaining the target prompt word, the target prompt word and the audio signal can be input into a large language model. The large language model can be a pre-trained model. The input of the large language model can be the target prompt word and the audio signal. The output of the large language model can be the type of the audio signal. If the type of the audio signal is an executable instruction audio signal, the output content also includes the vehicle control task corresponding to the audio signal.

[0087] It should be noted that the executable instruction audio signal can be a standard signal, such as an instruction that tells the vehicle to perform a vehicle control task that the vehicle can perform.

[0088] It should be noted that a Large Language Model (LLM) can be a pre-trained model with prior knowledge. It can determine the type of audio signal based on the recognition result of the audio signal, and it can also identify the vehicle control task corresponding to the audio signal based on the target prompt word, such as opening the car window, turning on the air conditioner, and turning down the media volume, etc.

[0089] The voice interaction method provided in this application embodiment can determine target prompt words based on audio signals; the target prompt words and audio signals are input into a large language model, and the type of audio signal is determined by the large language model. Specifically, processing the target prompt words through the large language model enables rapid and accurate recognition of audio signals, thereby improving the efficiency of determining the type of audio signal and vehicle control tasks.

[0090] The following explains one feasible implementation process for generating target prompts provided in the embodiments of this application.

[0091] Figure 5 This is a schematic diagram of the process for generating target prompt words provided in the embodiments of this application. Please refer to... Figure 5 Determining target prompts based on audio signals, including:

[0092] S510: Perform retrieval processing on the audio signal to determine multiple prompt words corresponding to the audio signal.

[0093] In one embodiment, the retrieval of audio signals may include multiple methods, such as using an AC automaton, using a search enhancement generative algorithm (RAG) for retrieval, or combining the above two retrieval methods to determine multiple prompt words corresponding to the audio signal.

[0094] Alternatively, if only the AC automaton is used for retrieval, one or more prompt words can be obtained. These prompt words can be concatenated to obtain the target prompt word.

[0095] If you only use the search enhancement to generate RAG search, you can also get one or more suggestion words. You can then combine these suggestion words to get the target suggestion words.

[0096] If you use the two methods described above to search, you can get multiple suggestions. You can then combine these suggestions to get the target suggestion.

[0097] S520: Generate target prompt words based on multiple prompt words.

[0098] In one embodiment, the multiple prompt words obtained can be discrete prompt words. Before inputting them into the large language model, these discrete prompt words can be concatenated to obtain a concatenated prompt word. The concatenated prompt word is then input into the large language model as the target prompt word, thereby realizing the recognition of audio signal type and vehicle control task.

[0099] The voice interaction method provided in this application embodiment can perform retrieval processing on audio signals to determine multiple prompt words corresponding to the audio signals; and generate target prompt words based on the multiple prompt words. Specifically, different retrieval processing methods can yield multiple different prompt words, and then concatenating these multiple prompt words can yield an accurate target prompt word, thereby improving the accuracy of determining the type of audio signal.

[0100] The following is a detailed explanation of one of the optional implementation processes for generating target prompts provided in the embodiments of this application.

[0101] Figure 6 For another flowchart illustrating the generation of target prompt words provided in this application embodiment, please refer to... Figure 6 The audio signal is processed to identify multiple prompt words corresponding to the audio signal, including:

[0102] S610: Identify keywords in the target voice command and determine the target vehicle information corresponding to the keywords based on multi-pattern matching.

[0103] It should be noted that a generalized lexicon can be pre-built in the large language model of the in-vehicle terminal. This lexicon can combine all entity generalization words and entity words into a "generalized lexicon". This lexicon basically contains all entity words and generalization words, and is composed of a key1-value1 mapping, denoted as mapping 1, where key1 is the entity generalization word and value1 is the original entity word. It also has another key2-value2 mapping, denoted as mapping 2, where key2 is the vehicle entity and value2 is vehicle knowledge.

[0104] Assuming the user inputs an audio signal of "kinetic energy recovery standard mode", the AC automaton can determine the generalized words in it based on the above generalized vocabulary, such as "kinetic energy", "kinetic energy recovery", "standard mode", etc.

[0105] The retrieved entities are duplicates, such as "kinetic energy" and "kinetic energy recovery". Furthermore, the original entities corresponding to "kinetic energy", "kinetic energy recovery", and "standard mode" all contain "X-pedal driving mode" and "one-pedal mode", so deduplication is necessary.

[0106] Since both "kinetic energy" and "kinetic energy recovery" contain "kinetic energy," the longest entity is retained. This is because "kinetic energy" itself may contain more original entities and is not the entity that best matches the user's audio signal. To select the longest entity string, it is possible to match the user's command to the greatest extent possible.

[0107] After deduplication, the remaining entries are "Kinetic Energy Recovery" and "Standard Mode". Since there is no repetition at the character level, they are directly mapped to the original entities, resulting in: ["X-pedal Driving Mode, One-Pedal Mode"] and ["Driving Mode", "Standard Mode", "Energy Recovery Level", "Ejection Mode", "Comfort Driving", "Custom", "Driver Mode", "X-pedal Driving Mode, One-Pedal Mode"]. Both of these contain "X-pedal Driving Mode, One-Pedal Mode", which is the vehicle entity that best matches the user command. The common substring of both is retained, and this is also the unique entity that matches the audio signal.

[0108] If the deduplicated entity contains multiple entities, the original entities of all entities can be mapped to obtain the vehicle entity. These entities will be used as the mapping between entity and vehicle knowledge.

[0109] Once the vehicle entity is obtained, it can be mapped to vehicle knowledge. All vehicle knowledge is input into the embedd layer of the large language model to obtain vehicle knowledge Embedd, where vehicle knowledge can be the target vehicle information mentioned above.

[0110] S620: Determine the first prompt word based on the target vehicle information.

[0111] It should be noted that each target vehicle information can correspond to a prompt word. By determining the vehicle knowledge in the above way, that is, after determining the target vehicle information, the first prompt word corresponding to the target vehicle information can be determined.

[0112] In one embodiment, the first prompt word includes multiple prompts.

[0113] S630: The second suggestion word was obtained by generating RAG through search enhancement.

[0114] It should be noted that relevant target interface parameters can be retrieved through RAG, and the second prompt word can be determined based on the target interface parameters.

[0115] In one embodiment, after receiving the audio signal "adjust the seat position forward by 30%", the audio signal can be input into the RAG retrieval model. The RAG model retrieves similar query results and similar interface information, wherein:

[0116] The similarity query results are constructed based on the training data of the large language model. This part of the data already exists, and only positive and negative pairs need to be constructed. The similar interface information is retrieved based on the description of the interface. The retrieved results are mapped to the corresponding interface information. This part of the data is also constructed based on the training data of the large model. The positive and negative pairs are mapped to the current interface and other different interfaces.

[0117] Optionally, the audio signal can be converted into a vector, and the BGE model can be used to retrieve query result vectors and interface information description vectors that are similar to the audio signal from the query result vector library and interface information description vector library, respectively, and map them to similar query results and similar interface information.

[0118] When new requirements are iterated, some data examples can be added directly, and the large speech model can obtain the results based on the examples and interface information. Only the RAG small model needs to be iterated, without the need for frequent iterations of the large language model.

[0119] Alternatively, the two methods mentioned above can be combined, using AC automata and RAG retrieval methods respectively to determine the corresponding prompt words.

[0120] In one embodiment, generating a target prompt word based on multiple prompt words includes:

[0121] S640: Concatenate the first prompt word and the second prompt word to obtain the target prompt word.

[0122] It should be noted that after obtaining multiple prompt words, the first prompt word and the second prompt word can be concatenated, that is, both prompt words can be used as input prompt words for the target voice command.

[0123] The target prompt can be the result of combining the first prompt and the second prompt. For example, if the first prompt includes "car window" and "open", and the second prompt includes "media volume" and "turn up", then the target prompt can include all the content included in the above two prompts.

[0124] The voice interaction method provided in this application embodiment can identify keywords in the target voice command and determine the target vehicle information corresponding to the keywords based on multi-pattern matching; determine a first prompt word based on the target vehicle information; retrieve a second prompt word through retrieval enhancement (RAG); and concatenate the first and second prompt words to obtain the target prompt word. The above methods enable AC automaton retrieval and RAG retrieval, thereby obtaining two different prompt words, namely the first and second prompt words. Concatenation of these two different prompt words accurately yields the target prompt word input into the large language model, improving the accuracy of the large language model processing.

[0125] The following explains one possible implementation process of the output target prompt information provided in the embodiments of this application.

[0126] Figure 7 This is a flowchart illustrating the output target prompt information provided in the embodiments of this application. Please refer to... Figure 7The types of audio signals also include: unclear audio signals, harmful audio signals, or audio signals with non-executable instructions.

[0127] Unclear audio signals can be audio signals that the vehicle terminal cannot recognize, such as: being unable to determine the target prompt word, or being unable to determine the text information corresponding to the audio signal.

[0128] Harmful audio signals can be identified by recognizing the audio signal and determining that the corresponding text information contains harmful words. The harmful words can be multiple pre-defined words, without specific restrictions, and can be set according to actual needs.

[0129] The non-executable instruction audio signal can be an instruction-type audio signal, but the vehicle terminal does not have the ability to execute the instruction. For example, for a vehicle without an automatic parking function, the automatic parking instruction is the non-executable instruction audio signal.

[0130] In the process of outputting prompt information for the various types of audio signals mentioned above, in addition to the prompt information for indicating that the audio signal does not have a corresponding vehicle control task, it may also include prompt information corresponding to each type of audio signal that is not activated.

[0131] In one embodiment, outputting target prompt information includes: when the type of audio signal is an unclear audio signal, outputting first prompt information, the first prompt information being used to indicate that the audio signal cannot be recognized.

[0132] It should be noted that for unclear audio signals, since they cannot be recognized by the vehicle terminal, the vehicle terminal does not know the content corresponding to the audio signal and therefore cannot make a corresponding response. Therefore, the corresponding first prompt message can be used to indicate that the audio signal cannot be recognized.

[0133] For example, it can output corresponding prompt voice, such as preset voices like "Please say it again" or "Sorry, I didn't hear you" to respond to the audio signal.

[0134] In one embodiment, outputting target prompt information includes: when the type of the audio signal is a harmful audio signal, outputting second prompt information, the second prompt information being used to indicate that the audio signal contains harmful information.

[0135] It should be noted that for harmful audio signals, the corresponding text information usually contains sensitive words, and responses cannot be made based on such text. Therefore, the vehicle terminal cannot make a corresponding response. Thus, the corresponding second prompt information can be used to indicate that the audio signal contains harmful information.

[0136] For example, it can output corresponding prompt voice, such as preset voices like "unable to recognize" or "sensitive information exists" to respond to the audio signal.

[0137] In one embodiment, outputting target prompt information includes: when the type of the audio signal is an unexecutable instruction audio signal, outputting third prompt information, the third prompt information being used to indicate that the vehicle control task corresponding to the audio signal cannot be executed.

[0138] It should be noted that for non-executable command audio signals, these audio signals can be command-type audio signals that can be recognized by the in-vehicle terminal, and the in-vehicle terminal can also determine the corresponding vehicle control task. However, the in-vehicle terminal itself does not have the conditions to execute the vehicle control task, for example, it does not have the corresponding control interface. Therefore, the in-vehicle terminal cannot execute the vehicle control task corresponding to the audio signal.

[0139] For example, a corresponding prompt voice can be output, such as a preset voice message like "Please execute manually" or "Failed to call the interface" to respond to the audio signal.

[0140] In one embodiment, the type of the above-mentioned audio signal also includes: a chat-type audio signal.

[0141] For chatty audio signals, responses can be generated based on a large language model.

[0142] It should be noted that the chat-type audio signal may be a chat between users in the cabin and other users, or it may be a chat between a user and the in-vehicle terminal. The in-vehicle terminal can choose to respond to the chat based on the actual situation. For example, if it is determined to be a chat between a user and the in-vehicle terminal, a response can be generated for that chat. If it is determined to be a chat between a user and other users, no corresponding chat response can be generated.

[0143] In the voice interaction method provided in this application embodiment, when the audio signal type is an unclear audio signal, a first prompt message is output, indicating that the audio signal cannot be recognized. When the audio signal type is a harmful audio signal, a second prompt message is output, indicating that the audio signal contains harmful information. When the audio signal type is an audio signal for an unexecutable command, a third prompt message is output, indicating that the vehicle control task corresponding to the audio signal cannot be executed. By generating different types of prompt messages for different types of audio signals, the processing differences for different audio signals can be improved, increasing the comprehensiveness and relevance of the responses.

[0144] The following explains the implementation process for the executable instruction audio signal, specifically the vehicle control task corresponding to the output audio signal.

[0145] Figure 8 This is a flowchart illustrating the execution of vehicle control tasks provided in the embodiments of this application. Please refer to... Figure 8 Output vehicle control tasks corresponding to the audio signals, including:

[0146] S810: Filter out target interface parameters from preset interface parameters based on target prompt words.

[0147] The target interface parameters are used to indicate the execution unit and execution parameters of the vehicle control task.

[0148] It should be noted that the target interface parameters corresponding to the audio signal can be determined based on the target prompt words.

[0149] The target interface parameters are used to indicate the execution unit and execution parameters of the vehicle control task.

[0150] The execution unit is the vehicle device that performs the vehicle control task, and the execution parameters are the specific parameter adjustments made by the execution unit during the execution of the vehicle control task.

[0151] It should be noted that different methods can be used to determine the above target interface parameters for different types of audio signals.

[0152] S820: Outputs vehicle control tasks corresponding to audio signals based on target interface parameters.

[0153] It should be noted that after determining the target interface parameters (API) and arguments, the aforementioned vehicle control task can be obtained.<API,Araguments> It can output the vehicle control task and control the corresponding execution unit to execute the vehicle control task according to the execution parameters.

[0154] In the voice interaction method provided in this application embodiment, target interface parameters can be filtered from preset interface parameters based on target prompt words; and vehicle control tasks corresponding to audio signals can be output according to the target interface parameters. Specifically, determining the target interface parameters through target prompt words and then determining the vehicle control task based on the target interface parameters allows for accurate and rapid acquisition and execution of the vehicle control task.

[0155] The following will explain the voice interaction method provided in the embodiments of this application through an overall implementation process.

[0156] Figure 9 This is a schematic diagram of the overall process of the voice interaction method provided in the embodiments of this application. Please refer to it. Figure 9 The method includes:

[0157] First, the type of audio signal can be determined by searching using an A / C automaton. Keyword searches are performed using a basic vocabulary and a generalized vocabulary to identify the corresponding keywords. Then, the target vehicle information can be determined based on the keywords, and the first prompt word can be obtained based on the target vehicle information.

[0158] Alternatively, a search can be performed using RAG retrieval to obtain a second prompt word. After obtaining the second prompt word, the first prompt word and the second prompt word can be concatenated and input into a large language model. The large language model can also input the aforementioned audio signal, and the type of the audio signal can be determined after recognition by the large language model.

[0159] In the case where the type of audio signal is an executable instruction audio signal, it can be further determined whether the executable instruction audio signal is a switch type or an adjustment type, so that the vehicle control task can be output and executed.

[0160] If the audio signal type is not an executable audio signal, a target prompt message can be output.

[0161] The voice interaction method provided in this application embodiment can break the traditional two-stage separation mode and construct a single-stage multi-task model. It integrates instruction classification (idle chat / harmful / standard instructions) and slot parsing (API / Arguments generation) functions. It combines vehicle knowledge retrieved in real time by RAG retrieval and AC automaton, and reduces parameter redundancy by sharing the underlying feature extraction layer, thereby reducing model complexity while maintaining task processing accuracy. In addition, the results of RAG retrieval and AC automaton retrieval can be used as decision anchors. The model prioritizes judging whether the query matches the standard instruction pattern in the vehicle function knowledge base: for queries with high matching degree, API parameters are directly generated; for queries with low matching degree, they are classified as idle chat / harmful based on semantic features or incomplete slots are returned. This achieves end-to-end lightweight reasoning of "retrieval-judgment-generation", balancing efficiency and accuracy.

[0162] It should be understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the above flowcharts may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0163] Based on the foregoing embodiments, this application provides a voice interaction device, which includes various modules and units included in each module, and can be implemented by a processor; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), microprocessor (MPU), digital signal processor (DSP) or field programmable gate array (FPGA), etc.

[0164] Figure 10 This is a schematic diagram of the structure of the voice interaction device provided in the embodiments of this application. Please refer to... Figure 10 In another aspect of the embodiments of this application, a voice interaction device is also provided, which is applied to an in-vehicle terminal, including: an acquisition module 1010, a determination module 1020 and an execution module 1030;

[0165] The acquisition module 1010 is used to acquire audio signals collected by the vehicle terminal;

[0166] The determining module 1020 is used to determine the target prompt word based on the audio signal, and to obtain the type of audio signal according to the target prompt word. The type of audio signal includes: executable instruction audio signal;

[0167] The execution module 1030 is used to output a vehicle control task corresponding to the audio signal when the audio signal is an executable instruction audio signal. The vehicle control task is determined based on the target prompt word.

[0168] The execution module 1030 is also used to output target prompt information when the audio signal is not an executable instruction audio signal. The target prompt information is used to indicate that there is no corresponding vehicle control task for the audio signal.

[0169] In one embodiment, the determining module 1020 is specifically used to determine the target prompt word based on the audio signal; input the target prompt word and the audio signal into a large language model, and determine the type of the audio signal through the large language model.

[0170] In one embodiment, the determining module 1020 is specifically used to perform retrieval processing on the audio signal, determine multiple prompt words corresponding to the audio signal, and generate target prompt words based on the multiple prompt words.

[0171] In one embodiment, the determining module 1020 is specifically used to determine the keywords in the target voice command and determine the target vehicle information corresponding to the keywords based on multi-pattern matching; determine the first prompt word based on the target vehicle information; and retrieve the second prompt word by generating RAG through retrieval enhancement.

[0172] In one embodiment, the determining module 1020 is specifically used to concatenate the first prompt word and the second prompt word to obtain the target prompt word.

[0173] In one embodiment, the audio signal type further includes: unclear audio signal, harmful audio signal, or non-executable instruction audio signal; the execution module 1030 is specifically configured to output a first prompt message when the audio signal type is an unclear audio signal, the first prompt message indicating that the audio signal cannot be recognized; output a second prompt message when the audio signal type is a harmful audio signal, the second prompt message indicating that the audio signal contains harmful information; and output a third prompt message when the audio signal type is a non-executable instruction audio signal, the third prompt message indicating that the vehicle control task corresponding to the audio signal cannot be executed.

[0174] In one embodiment, the execution module 1030 is specifically used to filter target interface parameters from preset interface parameters based on target prompt words. The target interface parameters are used to indicate the execution unit and execution parameters of the vehicle control task. The vehicle control task corresponding to the audio signal is output according to the target interface parameters.

[0175] The voice interaction device provided in this application embodiment can acquire audio signals collected by an in-vehicle terminal; determine target prompt words based on the audio signals; and obtain the type of audio signals based on the target prompt words. The types of audio signals include: executable command audio signals; when the audio signal is an executable command audio signal, output a vehicle control task corresponding to the audio signal, the vehicle control task being determined based on the target prompt words; when the audio signal is not an executable command audio signal, output target prompt information, the target prompt information indicating that there is no corresponding vehicle control task for the audio signal. Specifically, after acquiring the audio signal, the target prompt words can be determined based on the audio signal, thereby determining the vehicle control task corresponding to the audio signal. Then, based on the determined type of the audio signal, the execution of the vehicle control task or the output of target prompt information can be determined. By synchronously identifying the type of the audio signal and the corresponding vehicle control task, the space occupied by the model used to process the audio signal can be reduced. Since no additional model is set up to identify the type of the audio signal, parameter redundancy in the processing process is also reduced. This reduces model complexity while ensuring the processing accuracy of the vehicle control task, resulting in a vehicle control task that is both efficient and accurate.

[0176] The descriptions of the above device embodiments are similar to those of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0177] It should be noted that, in the embodiments of this application... Figure 10 The module division of the voice interaction device shown is illustrative and represents only one logical functional division; in actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, exist as separate physical units, or have two or more units integrated into one unit. The integrated units can be implemented in hardware, as software functional units, or a combination of both.

[0178] It should be noted that, in the embodiments of this application, if the above-described methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.

[0179] Figure 11 This is a schematic diagram of the structure of the computer device provided in the embodiments of this application. Please refer to... Figure 11 This application provides a computer device, which may be the aforementioned vehicle-mounted terminal, or a vehicle-mounted system, smart cockpit, etc., including the vehicle-mounted terminal. No specific limitations are imposed here, and its internal structure diagram may be as follows. Figure 11 As shown. The computer device includes a processor 1120, memory, and a network interface 1140 connected via a system bus 1110. The processor 1120 provides computing and control capabilities. The memory includes a non-volatile storage medium 1131 and internal memory 1132. The non-volatile storage medium 1131 stores an operating system, computer programs, and a database. The internal memory 1132 provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium 1131. The database is used to store data. The network interface 1140 is used to communicate with external terminals via a network connection. When the computer program is executed by the processor 1120, it implements the aforementioned methods.

[0180] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method provided in the above embodiments.

[0181] This application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the steps in the method provided in the above-described method embodiments.

[0182] Those skilled in the art will understand that Figure 11The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0183] In one embodiment, the voice interaction device provided in this application can be implemented as a computer program, and the computer program can be implemented in such a way as... Figure 11 The device operates on the computer device shown. The memory of the computer device can store the various program modules that make up the above-described apparatus. The computer program, composed of the various program modules, causes the processor to execute the steps of the methods in the various embodiments of this application described in this specification.

[0184] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium, storage medium, and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0185] It should be understood that the phrases "one embodiment," "an embodiment," or "some embodiments" mentioned throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment," "in one embodiment," or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. The descriptions of the various embodiments above tend to emphasize the differences between the various embodiments; their similarities or commonalities can be referred to mutually, and for the sake of brevity, they will not be repeated here.

[0186] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three kinds of relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist simultaneously, and object B exists alone.

[0187] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0188] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or modules can be electrical, mechanical, or other forms.

[0189] The modules described above as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules. They may be located in one place or distributed across multiple network units. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.

[0190] In addition, each functional module in the various embodiments of this application can be integrated into one processing unit, or each module can be a separate unit, or two or more modules can be integrated into one unit; the integrated modules can be implemented in hardware or in the form of hardware plus software functional units.

[0191] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0192] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0193] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0194] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0195] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0196] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A voice interaction method, characterized in that, Applications in vehicle-mounted terminals include: Acquire the audio signal collected by the vehicle-mounted terminal; Based on the audio signal, a target prompt word is determined, and the type of the audio signal is obtained according to the target prompt word. The type of the audio signal includes: an executable instruction audio signal. When the audio signal is an executable instruction audio signal, a vehicle control task corresponding to the audio signal is output, and the vehicle control task is determined based on the target prompt word; If the audio signal is not an executable instruction audio signal, a target prompt message is output, which indicates that the audio signal does not correspond to a vehicle control task.

2. The method according to claim 1, characterized in that, The step of determining the target prompt word based on the audio signal and obtaining the type of the audio signal based on the target prompt word includes: The target prompt word is determined based on the audio signal; The target prompt word and the audio signal are input into a large language model, and the type of the audio signal is determined by the large language model.

3. The method according to claim 2, characterized in that, The step of determining the target prompt word based on the audio signal includes: The audio signal is processed to determine multiple prompt words corresponding to the audio signal; The target prompt word is generated based on the multiple prompt words.

4. The method according to claim 3, characterized in that, The step of retrieving and processing the audio signal to determine multiple prompt words corresponding to the audio signal includes: The keywords in the target voice command are identified, and the target vehicle information corresponding to the keywords is determined based on multi-pattern matching; The first prompt word is determined based on the target vehicle information; The second suggestion word was obtained by generating RAG through search enhancement.

5. The method according to claim 4, characterized in that, The process of generating the target prompt word based on the plurality of prompt words includes: The first prompt word and the second prompt word are concatenated to obtain the target prompt word.

6. The method according to claim 1, characterized in that, The types of audio signals also include: unclear audio signals, harmful audio signals, or audio signals with non-executable instructions; The output target prompt information includes: If the type of the audio signal is an unclear audio signal, a first prompt message is output, which is used to indicate that the audio signal cannot be recognized; If the type of the audio signal is a harmful audio signal, a second prompt message is output, which is used to indicate that the audio signal contains harmful information; If the type of the audio signal is an unexecutable instruction audio signal, a third prompt message is output, which indicates that the vehicle control task corresponding to the audio signal cannot be executed.

7. The method according to claim 1, characterized in that, The vehicle control task corresponding to the output audio signal includes: Based on the target prompt words, target interface parameters are filtered from preset interface parameters. The target interface parameters are used to indicate the execution unit and execution parameters of the vehicle control task. The vehicle control task corresponding to the audio signal is output based on the target interface parameters.

8. A voice interaction device, characterized in that, Applied to in-vehicle terminals, it includes: an acquisition module, a determination module, and an execution module; The acquisition module is used to acquire the audio signal collected by the vehicle terminal; The determining module is configured to determine a target prompt word based on the audio signal, and to obtain the type of the audio signal according to the target prompt word. The type of the audio signal includes: an executable instruction audio signal. The execution module is configured to output a vehicle control task corresponding to the audio signal when the audio signal is an executable instruction audio signal, wherein the vehicle control task is determined based on the target prompt word; The execution module is further configured to output target prompt information when the audio signal is not an executable instruction audio signal, the target prompt information being used to indicate that the audio signal does not correspond to a vehicle control task.

9. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Speech recognition method and system based on AC automaton hot word enhancement

    CN114187902A

  • Voice interaction method, server and readable storage medium

    CN117877478A

  • Cabin voice instruction processing method, apparatus and device, and medium

    CN118212921A

  • Model training method of large language model, server and storage medium

    CN120183389A