Voice interaction method and device, equipment and storage medium
By using preset prompt word recognition and target prompt word generation, the problem of slow response to simple voice commands is solved, enabling rapid response and efficient execution of vehicle control tasks.
Patent Information
- Application Number
- CN202511172533.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-04
AI Technical Summary
In existing technologies, when users interact with vehicles, simple voice commands are slow to respond, resulting in low voice interaction efficiency and the in-vehicle terminal being unable to execute vehicle control tasks in a timely manner.
The system recognizes target voice commands by using preset prompt words. If the recognition is successful, the system directly outputs and executes the vehicle control task. If the recognition fails, the system generates target prompt words and outputs the vehicle control task based on them.
It improves the response speed and execution efficiency of vehicle control tasks, especially the rapid processing of simple voice commands, and reduces the process time for determining vehicle control tasks.
Smart Images

Figure CN120895035A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to vehicle interaction technology, and relate to but are not limited to a voice interaction method and device, equipment, and storage medium. BACKGROUND
[0002] In the process of user interaction with the vehicle, it is usually necessary to request the vehicle terminal to perform the corresponding vehicle control task through the voice instruction mode, for example: opening the door, automatic parking, etc.
[0003] In the related art, in the process of user interaction with the vehicle terminal, there may be a situation that the user gives a simple instruction, but the vehicle terminal responds slowly, which leads to low efficiency of voice interaction and the vehicle terminal cannot execute the vehicle control task in time. SUMMARY
[0004] The voice interaction method and device, equipment, and storage medium provided by the embodiments of the present application are implemented as follows:
[0005] In one aspect of the embodiments of the present application, a voice interaction method is provided, applied to a vehicle terminal, comprising:
[0006] obtaining a target voice instruction collected by the vehicle terminal;
[0007] identifying the target voice instruction according to a preset prompt word;
[0008] in the case of successful identification of the target voice instruction, outputting a vehicle control task corresponding to the target voice instruction and executing the vehicle control task;
[0009] in the case of failed identification of the target voice instruction, determining a target prompt word based on the target voice instruction, outputting the vehicle control task corresponding to the target voice instruction according to the target prompt word, and executing the vehicle control task.
[0010] In another aspect of the embodiments of the present application, a voice interaction device is also provided, applied to a vehicle terminal, comprising: an acquisition module, an identification module, and an execution module;
[0011] The acquisition module is configured to obtain a target voice instruction collected by the vehicle terminal;
[0012] The identification module is configured to identify the target voice instruction according to a preset prompt word;
[0013] The execution module is configured to, in the case of successful identification of the target voice instruction, output a vehicle control task corresponding to the target voice instruction and execute the vehicle control task;
[0014] The execution module is further configured to determine a target prompt word based on the target voice instruction in a case where the target voice instruction recognition fails, and output a vehicle control task corresponding to the target voice instruction according to the target prompt word and execute the vehicle control task.
[0015] The computer device provided in the embodiments of the present application includes a memory and a processor. The memory stores a computer program that can run on the processor. The processor implements the method of the embodiments of the present application when executing the program.
[0016] The computer readable storage medium provided in the embodiments of the present application stores a computer program. The computer program is executed by a processor to implement the method provided in the embodiments of the present application.
[0017] The voice interaction method and device, equipment and storage medium provided in the embodiments of the present application can acquire a target voice instruction collected by a vehicle terminal, recognize the target voice instruction according to a preset prompt word, output a vehicle control task corresponding to the target voice instruction and execute the vehicle control task in a case where the target voice instruction recognition is successful, and determine a target prompt word based on the target voice instruction in a case where the target voice instruction recognition fails, output a vehicle control task corresponding to the target voice instruction according to the target prompt word and execute the vehicle control task. After the target voice instruction is acquired, the target voice instruction can be recognized based on the preset prompt word. If the target voice instruction can be recognized successfully through the preset prompt word, the vehicle control task can be output and executed in a faster manner, the inference time for determining the vehicle control task can be reduced, the generation of the target prompt word is not needed in the case where the preset prompt word recognition is successful, the process for determining the vehicle control task is reduced, and thus the vehicle control task can be determined more efficiently and quickly. In addition, after the vehicle control task is determined more quickly, the response speed of the vehicle control task can be improved, and the vehicle control task can be executed more quickly. BRIEF DESCRIPTION OF DRAWINGS
[0018] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort based on these drawings.
[0019] Figure 1 The application scenario provided in the embodiments of the present application is shown in the figure.
[0020] Figure 2 The flowchart of the voice interaction method provided in the embodiments of the present application is shown in the figure.
[0021] Figure 3A relationship diagram for identifying a target voice instruction provided in an embodiment of the present application;
[0022] Figure 4 Another relationship diagram for identifying a target voice instruction provided in an embodiment of the present application;
[0023] Figure 5 Another flow diagram of a voice interaction method provided in an embodiment of the present application;
[0024] Figure 6 A flow diagram of generating a target prompt word provided in an embodiment of the present application;
[0025] Figure 7 A processing flow diagram of an adjustment type instruction provided in an embodiment of the present application;
[0026] Figure 8 A processing flow diagram of a switch type instruction provided in an embodiment of the present application;
[0027] Figure 9 A processing flow diagram of an uncertain instruction provided in an embodiment of the present application;
[0028] Figure 10 Still another flow diagram of a voice interaction method provided in an embodiment of the present application;
[0029] Figure 11 A whole flow diagram of a voice interaction method provided in an embodiment of the present application;
[0030] Figure 12 A structure diagram of a voice interaction device provided in an embodiment of the present application;
[0031] Figure 13 A structure diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0032] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the specific technical solutions of the present application will be further described below with reference to the accompanying drawings of the embodiments of the present application. The following embodiments are used to explain the present application, but are not used to limit the scope of the present application.
[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description herein is for describing the embodiments of the present application only and is not intended to limit the present application.
[0034] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments, but it is understood that "some embodiments" can be the same subset or a different subset of all possible embodiments, and can be combined with each other as long as there is no conflict.
[0035] It should be noted that the terms "first", "second", "third" involved in the embodiments of the present application are used to distinguish similar or different objects, and do not represent a specific order of the objects. Understandably, "first", "second", "third" can be interchanged in a specific order or sequence as long as it is allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0036] In order to more clearly explain the voice interaction method provided in the embodiments of the present application, a possible application scenario in the embodiments of the present application will be explained in detail below.
[0037] Figure 1 For the application scenario provided in the embodiments of the present application, please refer to Figure 1 The scenario can include a vehicle terminal 100, for example, the vehicle terminal 100 can be an intelligent vehicle machine. The vehicle terminal 100 is arranged in an intelligent cabin and can have an intelligent conversation with a user in the intelligent cabin, for example, it can generate a corresponding answer based on a user's question, or execute a corresponding vehicle control task based on a user's instruction. The voice interaction method provided in the embodiments of the present application is mainly for the scenario where the vehicle terminal 100 determines and executes a corresponding vehicle control task when a user gives a voice instruction.
[0038] The vehicle terminal 100 can execute a corresponding vehicle control task according to a user's instruction, for example, if the user instructs to open a certain door, the vehicle terminal can execute the task of opening the corresponding door based on the instruction.
[0039] In this scenario, one user can interact with the vehicle terminal 100, or multiple users can interact with the vehicle terminal 100, which is not limited specifically herein, for example, a user on the main driver's seat can interact with the vehicle terminal 100, and a user on the co-driver's seat can also interact with the vehicle terminal 100.
[0040] During the process of the user interacting with the vehicle terminal 100, it can be achieved through language communication, for example, the user speaks a corresponding question, the vehicle terminal 100 collects a corresponding audio signal to obtain a voice instruction, and then the voice instruction can be recognized to determine the user's intention, and then a large language model can be used to determine the vehicle control task to be executed, which is not limited specifically herein.
[0041] In the related art, in the process of executing a vehicle task, a plurality of preset steps need to be executed for each voice instruction given by a user, so as to realize recognition of the voice instruction.
[0042] The method is suitable for relatively complex voice instructions. However, in an actual application scenario, a user is usually driving a car and will not give a relatively complex voice instruction. The voice instruction given is usually a relatively simple instruction, such as "open the window" and "automatic parking". However, for a simple instruction, the plurality of preset steps still need to be executed, which will result in slow response of the vehicle terminal and further result in low efficiency of voice interaction and the vehicle terminal being unable to execute a vehicle control task in time.
[0043] To solve the above problems in the related art, a voice interaction method is provided in the embodiments of the present application. The actual implementation process of the method is explained as follows.
[0044] Figure 2 For a flowchart of the voice interaction method provided in the embodiments of the present application, please refer to Figure 2 The method comprises the following steps.
[0045] S210: Obtain a target voice instruction collected by a vehicle terminal.
[0046] It should be noted that the execution subject of the method can be the vehicle terminal described above. A plurality of voice instructions can be collected by the vehicle terminal, and any one of the voice instructions can be used as a target voice instruction.
[0047] Optionally, the target voice instruction can be a voice instruction given by a user, such as "open the window of the main driver", "close the door of the co-driver", "increase the volume of media playing", "lower the temperature of the air conditioner", and the like, which are voice instructions requiring the vehicle terminal to execute a corresponding vehicle control task.
[0048] In an embodiment, an audio collection device, such as a microphone, a radio, and the like, can be arranged in an intelligent vehicle cabin, which is not specifically limited herein. The audio collection device can be in communication connection with the vehicle terminal. The vehicle terminal can obtain a target voice instruction provided by a user through the audio collection device.
[0049] For example, a user can speak a request, such as "open the navigation", during driving a car. The request can be collected by the audio collection device and sent to the vehicle terminal as a target voice instruction collected by the vehicle terminal.
[0050] In an embodiment, the vehicle terminal can collect various types of voice, such as casual voice, inquiry voice, etc. The vehicle terminal can first perform voice recognition on these voices to determine whether the voice is a voice instruction. If so, the voice can be used as the target voice instruction. If not, the voice can be ignored, or other preset functions can be performed, which are not limited here.
[0051] S220: Recognize the target voice instruction according to the preset prompt word.
[0052] It should be noted that the preset prompt word can be one or more words configured in advance. The target voice instruction can be determined based on the preset prompt word.
[0053] In an embodiment, the target voice instruction can be recognized by a large language model. In this case, the preset prompt word and the target voice instruction can be input into the large language model. If the target voice instruction can be recognized by the large language model based on the preset prompt word, it can be determined that the target voice instruction can be recognized according to the preset prompt word. If the target voice instruction cannot be recognized by the large language model based on the preset prompt word, it can be determined that the target voice instruction cannot be recognized according to the preset prompt word.
[0054] In an embodiment, the target voice instruction can be recognized by matching words. In this case, the target voice instruction can be first converted into text information by audio-to-text conversion. Then it can be determined whether there is a match between the preset prompt word and the words in the text information. If there is a match between the preset prompt word and the words in the text information, it can be determined that the target voice instruction can be recognized according to the preset prompt word. If all preset prompt words do not match the words in the text information, it can be determined that the target voice instruction cannot be recognized according to the preset prompt word.
[0055] Optionally, any one of the above two methods can be used to recognize the target voice instruction in actual implementation, so that it can be determined whether the target voice instruction is successfully recognized. Different methods can be used to determine the vehicle control task corresponding to the target voice instruction for the target voice instruction successfully recognized and the target voice instruction failed to be recognized.
[0056] It should be noted that identifying the target voice instruction refers to identifying the vehicle control task corresponding to the target voice instruction, for example, for the target voice instruction "open the main driver door", the target voice instruction is identified through the preset prompt word, and it can be determined that the target voice instruction corresponds to the vehicle control task; correspondingly, for the target voice instruction "lower the right glass", the target voice instruction cannot be identified through the preset prompt word, and the vehicle control task corresponding to the target voice instruction is not determined.
[0057] S230: In the case that the target voice instruction is successfully identified, output the vehicle control task corresponding to the target voice instruction and execute the vehicle control task.
[0058] In an embodiment, in the case that the target voice instruction is successfully identified, the vehicle control task corresponding to the target instruction can be identified, so that the vehicle control task corresponding to the target instruction can be output, and the vehicle control task can be executed.
[0059] The vehicle control task can be a specific task for the vehicle in which the vehicle terminal is located, and can be a task of controlling a certain unit of the vehicle to execute a certain task, for example, a task of controlling the co-driver door of the vehicle to open.
[0060] It should be noted that the vehicle control task can be output in a specified format, for example, <API, Araguments>, where API can be a corresponding execution unit, used to represent a unit that executes the task, for example, a certain window, a certain door, an air conditioner or a display, etc., which is not limited here. Araguments can be a corresponding execution parameter, used to represent a parameter value for executing the task, for example, for an on-off type execution parameter, 0 / 1 can be used to represent off or on; for an adjustment type execution parameter, a specific numerical value can be set.
[0061] It should be noted that after outputting the vehicle control task corresponding to the target voice instruction, the corresponding execution unit can be controlled to execute the vehicle control task.
[0062] In an embodiment, for the case that the vehicle control task corresponding to the target voice instruction can be identified, no additional work needs to be performed, the identified vehicle control task can be output, and the task can be executed.
[0063] S240: In the case that the target voice instruction is not successfully identified, determining a target prompt word based on the target voice instruction, outputting the vehicle control task corresponding to the target voice instruction according to the target prompt word and executing the vehicle control task.
[0064] In an embodiment, in the case that the target voice instruction recognition fails, the vehicle control task corresponding to the target instruction cannot be recognized, so the vehicle control task can be further determined by determining the target prompt word.
[0065] The target prompt word corresponding to the target voice instruction can be determined by searching, matching, etc., and the target prompt word can be generated based on the target voice instruction or extracted from the text information corresponding to the target voice instruction, which is not limited here.
[0066] After obtaining the target prompt word, the vehicle control task can be determined based on the target prompt word, for example: the target prompt word, the preset prompt word and the target voice instruction can be input into a large language model, and the vehicle control task is output by the large language model.
[0067] Alternatively, the target prompt word and the target voice instruction can also be input into a large language model, and the vehicle control task is output by the large language model.
[0068] Alternatively, the target voice instruction can also be recognized by the aforementioned word matching method, the text information corresponding to the target voice instruction can be obtained, and then it can be determined whether there is a match between the target prompt word and the word in the text information. If there is a match between the target prompt word and the word in the text information, it can be determined that the target voice instruction can be recognized according to the target prompt word.
[0069] In the above manner, in the case that the target voice instruction recognition fails, the vehicle control task corresponding to the target voice instruction can be further inferred by generating the target prompt word, and then the vehicle control task can be executed.
[0070] In the voice interaction method provided in the embodiments of the present application, a target voice instruction collected by a vehicle terminal can be acquired; the target voice instruction is recognized according to a preset prompt word; in the case that the target voice instruction is recognized successfully, a vehicle control task corresponding to the target voice instruction is output and the vehicle control task is executed; in the case that the target voice instruction is not recognized successfully, a target prompt word is determined based on the target voice instruction, and the vehicle control task corresponding to the target voice instruction is output according to the target prompt word and the vehicle control task is executed. After the target voice instruction is acquired, the target voice instruction can be recognized based on the preset prompt word. If the target voice instruction can be recognized successfully through the preset prompt word, the vehicle control task can be output and executed in a faster manner, the inference time for determining the vehicle control task can be reduced, and in the case that the preset prompt word is recognized successfully, the generation of the target prompt word is not needed, the process for determining the vehicle control task is reduced, so that the vehicle control task can be determined more efficiently and quickly. In addition, after the vehicle control task is determined more quickly, the response speed of the vehicle control task can be improved, and the vehicle control task can be executed more quickly.
[0071] It should be noted that in the application scenario of the vehicle terminal, the user is usually in the state of driving a car, and in this state, the user usually does not give a relatively complex voice instruction, but a relatively simple voice instruction, for example, a voice instruction that can be recognized through a preset prompt word. For such a relatively simple voice instruction, the generation of the target prompt word can not be needed, the process for determining the vehicle control task is reduced, so that the vehicle control task can be determined more efficiently and quickly, and it is suitable for executing the corresponding vehicle control task based on the voice instruction of the user more quickly in the case that the vehicle is driving.
[0072] In an embodiment, compared with the same processing process for all voice instructions in the related art, different processing processes can be used to determine the vehicle control task for the two types of voice instructions that are recognized successfully and not recognized successfully, the diversity and flexibility of determining the vehicle control task can be increased, the vehicle control task can be obtained for different voice instructions in multiple ways, and especially for the voice instructions that are recognized successfully, the vehicle control task can be determined in a more efficient manner, the response time of the vehicle terminal is saved, and the corresponding vehicle control task can be output and executed more quickly.
[0073] The following explains one feasible implementation process of recognizing the target voice instruction provided in the embodiments of the present application.
[0074] Figure 3 For the relationship diagram of recognizing the target voice instruction provided in the embodiments of the present application, please refer to Figure 3According to the preset prompt word, the target voice instruction is recognized, including: inputting the preset prompt word and the target voice instruction into a large language model to recognize the target voice instruction.
[0075] It should be noted that the large language model (LLM) can be a pre-trained model with prior knowledge. In the case of a relatively simple target voice instruction, the corresponding vehicle control task of the target voice instruction can be recognized based on the preset prompt word, such as opening the window, turning on the air conditioner, and reducing the media volume, etc.
[0076] Optionally, after the vehicle control task is recognized by the above-mentioned manner, the vehicle control task can be output.
[0077] For reference Figure 3 The input of the large language model can be the target voice instruction and the preset prompt word. If the recognition is successful, the output of the large language model can be obtained, that is, the above-mentioned vehicle control task; if the recognition fails, the large language model will wait for the generation of the target prompt word and will not directly output the vehicle control task temporarily.
[0078] The following explains another possible implementation process of recognizing the target voice instruction provided in the embodiments of the present application.
[0079] Figure 4 For another relationship diagram of recognizing the target voice instruction provided in the embodiments of the present application, please refer to Figure 4 According to the preset prompt word, the target voice instruction is recognized, including: matching the preset prompt word with the text information corresponding to the target voice instruction to recognize the target voice instruction.
[0080] It should be noted that the preset prompt word can include multiple words, and the text information corresponding to the target voice instruction can be a sentence. The multiple words can be matched with each sentence. If the words in the preset prompt word are included in the sentence, it can be determined that the matching is successful, that is, the recognition is successful. Conversely, if no word in the preset prompt word is matched in the sentence, it can be determined that the matching fails, that is, the recognition fails.
[0081] For reference Figure 4 Each preset prompt word can be matched with the text information corresponding to the target voice instruction in turn, so as to realize the recognition of the target voice instruction. If any one of the preset prompt words is matched successfully, it can be determined that the recognition is successful. If none of the preset prompt words is matched successfully, it can be determined that the recognition fails.
[0082] In the voice interaction method provided in the embodiments of the present application, the preset prompt word and the target voice instruction can be input into the large language model to identify the target voice instruction, or the text information corresponding to the preset prompt word and the target voice instruction can be matched to identify the target voice instruction. Through the above two ways, the identification of the target voice instruction can be quickly and accurately realized, so that the efficiency of determining the vehicle control task can be improved.
[0083] The following explains another possible implementation process of the voice interaction method provided in the embodiments of the present application.
[0084] Figure 5 For another flowchart of the voice interaction method provided in the embodiments of the present application, please refer to Figure 5 Before determining the target prompt word based on the target voice instruction, the method further includes:
[0085] S510: Determine the type of the target voice instruction.
[0086] It should be noted that the type of the target voice instruction can be identified by the large language model, for example: for the target voice instruction that can identify the corresponding vehicle control task, the vehicle control task can be output; for the target voice instruction that cannot identify the corresponding vehicle control task, the type of the target voice instruction can be determined.
[0087] For example, if the text information corresponding to the target voice instruction includes a word indicating the type, or the type of the target voice instruction is determined in the process of semantic understanding, the type of the target voice instruction can be determined.
[0088] The type of the target voice instruction includes any one of the following: adjustment type instruction, switch type instruction, and uncertain instruction.
[0089] It should be noted that the adjustment type instruction can be an instruction for adjusting the numerical value of a parameter, for example: adjusting the temperature of an air conditioner to a certain temperature value, increasing the media volume to a certain value, etc.; the switch type instruction can be an instruction for adjusting the state of an execution unit, for example: opening a certain door, closing a certain door, opening a certain window, closing a certain window, etc., which is not limited here.
[0090] For the target voice instruction whose type cannot be identified by the large language model, the type of the target voice instruction can be divided into an uncertain instruction.
[0091] Determining the target prompt word based on the target voice instruction includes:
[0092] S520: Determine the target prompt word according to the type of the target voice instruction.
[0093] In an embodiment, after determining the type of the target voice instruction, a corresponding manner can be selected to determine the target prompt word based on the type of the target voice instruction.
[0094] In an embodiment, the target prompt word can be determined by first determining the target interface parameter.
[0095] In an embodiment, the type of the target voice instruction can be determined, and the target prompt word can be determined according to the type of the target voice instruction. In this way, the target prompt word can be determined in different manners based on different types, thereby improving the flexibility of determining the target prompt word.
[0096] The following explains a feasible implementation process of generating the target prompt word in the voice interaction method provided in an embodiment.
[0097] Figure 6 For a flowchart of generating the target prompt word provided in an embodiment, please refer to Figure 6 The target prompt word can be determined according to the type of the target voice instruction, including:
[0098] S610: The target interface parameter corresponding to the target voice instruction is determined according to the type of the target voice instruction.
[0099] It should be noted that the target interface parameter corresponding to the target voice instruction can be determined according to the type of the target voice instruction.
[0100] The target interface parameter is used to indicate the execution unit and the execution parameter of the vehicle control task.
[0101] The execution unit is a vehicle device that executes the vehicle control task, and the execution parameter is a specific parameter adjustment implemented by the execution unit in the process of executing the vehicle control task.
[0102] It should be noted that the target interface parameter can be determined in different manners for different types of target voice instructions.
[0103] S620: The target prompt word is generated based on the target interface parameter.
[0104] In an embodiment, after obtaining the target interface parameter, the target prompt word can be generated, where the target interface parameter can be used as the target prompt word, or the target prompt word corresponding to the target interface parameter can be generated based on the target interface parameter, and the like, which is not limited here.
[0105] It should be noted that the target prompt word can be one or more words generated based on the target interface parameter, for example, if the target interface parameter is "car window A" and "open", the corresponding target prompt word can be "driver's window" and "open", and the relationship between the target prompt word and the target interface parameter is not limited here, and can be set according to actual needs.
[0106] In the voice interaction method provided by the embodiments of the present application, the target interface parameter corresponding to the target voice instruction can be determined according to the type of the target voice instruction, and the target prompt word can be generated based on the target interface parameter. Different ways can be used to determine the target interface parameter based on different types of target voice instructions, so as to improve the adaptability of determining the target interface parameter. For different types of target voice instructions, the determination method of the target interface parameter corresponding to the target voice instruction of the type can be used to obtain the target interface parameter.
[0107] The implementation process of determining the target interface parameter for the above three types of target voice instructions will be explained below.
[0108] Figure 7 For the processing flow of the adjustment type instruction provided in the embodiments of the present application, please refer to Figure 7 According to the type of the target voice instruction, the target interface parameter corresponding to the target voice instruction is determined, including: in the case that the type of the target voice instruction is an adjustment type instruction, the target interface parameter is retrieved through RAG retrieval.
[0109] It should be noted that in the case that the target voice instruction is an adjustment type instruction, the related target interface parameter can be retrieved through RAG, and the target prompt word determined through the target interface parameter can be input into the large language model, so as to determine the vehicle control task.
[0110] For example, if the target voice instruction is "voice broadcast volume down", in the case that the preset prompt word is not recognized successfully, the large language model can determine that the target voice instruction is an adjustment type instruction.
[0111] After determining that it is an adjustment type instruction, the retrieval of the adjustment type instruction can be performed, for example, the retrieval of the target interface parameter is performed using the above-mentioned RAG retrieval method, and the target interface parameter can be obtained through retrieval.
[0112] The thinking logic of the large language model is as follows:
[0113] Firstly, it can be understood that the target voice instruction is that the user thinks that the current voice broadcast volume is too large and needs to be reduced; further, the voice broadcast volume on the vehicle can be determined by searching for related knowledge, and the voice broadcast volume can be adjusted relatively, and the corresponding target interface parameter can be determined, for example: the voice broadcast volume can be adjusted through an interface, and the specific adjustment mode can be to reduce, and the target interface parameter can be obtained.
[0114] In this process, if it is determined that the target voice instruction is an adjustment type instruction, RAG retrieval can be performed, and then the target interface parameter can be determined, and the target prompt word can be obtained.
[0115] Figure 8 For the processing flowchart of the switch type instruction provided in the embodiments of the present application, please refer to Figure 8 According to the type of the target voice instruction, the target interface parameter corresponding to the target voice instruction is determined, including: in the case that the type of the target voice instruction is a switch type instruction, determining the key word in the target voice instruction, and determining the target vehicle information corresponding to the key word based on multi-mode matching; determining the target interface parameter according to the target vehicle information.
[0116] It should be noted that a generalization word library can be pre-constructed in the large language model of the vehicle terminal, and all entity generalization words and entity words can be combined into the "generalization word library". The word library basically contains all entity words and generalization words of the switch type, and the form is key1-value1 mapping, denoted as mapping1, wherein key1 is an entity generalization word, and value1 is an original entity word. Meanwhile, it also has another key2-value2 mapping, denoted as mapping2, wherein key2 is a vehicle entity, and value2 is vehicle knowledge.
[0117] Suppose the target voice instruction input by the user is "kinetic energy recovery standard mode", the generalization words therein can be determined by the AC automatic machine according to the above generalization word library, such as "kinetic energy", "kinetic energy recovery" and "standard mode".
[0118] These retrieved entity parts are repetitive, such as "kinetic energy" and "kinetic energy recovery". At the same time, the original entities corresponding to "kinetic energy", "kinetic energy recovery" and "standard mode" all contain "X-pedal driving mode, single pedal mode", so de-duplication needs to be performed.
[0119] Since "kinetic energy" and "kinetic energy recovery" both contain "kinetic energy", the longest entity is retained, because "kinetic energy" itself may contain more original entities, which are not the most matched entities with the target voice instruction of the user. In order to select the longest entity string, the user instruction can be matched to the maximum extent.
[0120] After deduplication, the remaining "energy recovery" "standard mode", both have no repetition at the character level, so it is directly mapped to the original entity, resulting in: ["X-pedal driving mode, single pedal mode"] and ["driving mode", "standard mode", "energy recovery level", "ejection mode", "comfortable ride", "custom", "driver mode", "X-pedal driving mode, single pedal mode"], both of which contain "X-pedal driving mode, single pedal mode", which is the most matched vehicle entity with the user's instruction, and the common substring of both is retained, which is the only entity, consistent with the target voice instruction.
[0121] If the deduplicated entity contains multiple entities, the original entity of all entities can be mapped to obtain the vehicle entity, and these entities will be used as the mapping of entity->vehicle knowledge.
[0122] When the vehicle entity is obtained, it can be mapped to the vehicle knowledge, and all vehicle knowledge is input into the embedd layer of the large language model to obtain the vehicle knowledge Embedd.
[0123] It should be noted that in the case of a target voice instruction as a switch instruction, the relevant information can be retrieved by calling an AC automaton, for example: determining the keywords in the target voice instruction, and determining the target vehicle information corresponding to the keywords based on multi-mode matching, and then determining the target interface parameter according to the target vehicle information, and inputting the target prompt word determined by the target interface parameter into the large language model, so as to determine the vehicle control task.
[0124] Example: if the target voice instruction is "automatic parking", in the case where the preset prompt word is not recognized successfully, the large language model can determine that the target voice instruction is a switch instruction.
[0125] After determining as a switch instruction, the retrieval of the switch instruction can be performed, for example, the target interface parameter is retrieved using the above-mentioned AC automaton, and the corresponding keywords can be obtained through retrieval, and the corresponding target vehicle information can be determined through the keywords, and then the above-mentioned target interface parameter can be determined based on the target vehicle information.
[0126] The thinking logic of the large language model is as follows:
[0127] Firstly, it can be understood that the target voice instruction is that the user wants to turn on the automatic parking function; and then the target vehicle information can be determined through the AC automaton, for example, it can be determined that the automatic parking is a function on the vehicle, and when turning on / off, a parameter can be supplemented to realize the automatic parking, and then the above-mentioned target interface parameter can be obtained.
[0128] That is to say, for the switch type instruction, the AC automatic machine mode can be used for retrieval; for the adjustment type instruction, the RAG retrieval mode can be used for retrieval, so as to determine the target interface parameter based on a more adaptive mode.
[0129] Figure 9 For the processing flow of the uncertain instruction provided in the embodiment of the application, please refer to Figure 9 , according to the type of the target voice instruction, the target interface parameter corresponding to the target voice instruction is determined, including: in the case that the type of the target voice instruction is an uncertain instruction, the keyword in the target voice instruction is determined, and the target vehicle information corresponding to the keyword is determined based on multi-mode matching; the first interface parameter is determined according to the target vehicle information; the RAG retrieval is performed through retrieval enhancement to obtain the second interface parameter; and the target interface parameter is obtained based on the first interface parameter and the second interface parameter.
[0130] It should be noted that for the uncertain type target voice instruction, the AC automatic machine and the RAG retrieval mode can be combined to determine the corresponding interface parameter, if the target voice instruction is of the switch type, the first interface parameter can be obtained, and the second interface parameter cannot be obtained; if the target voice instruction is of the adjustment type, the second interface parameter can be obtained, and the first interface parameter cannot be obtained.
[0131] For the case where the first interface parameter exists and the second interface parameter does not exist, the first interface parameter can be used as the target interface parameter; for the case where the second interface parameter exists and the first interface parameter does not exist, the second interface parameter can be used as the target interface parameter.
[0132] For example, if the target voice instruction is "screen switching to automatic", in the case that the preset prompt word is not recognized successfully, the target voice instruction can be determined as an uncertain instruction through the large language model.
[0133] For the uncertain instruction, the processes shown in Figure 7 and Figure 8 can be used for retrieval to finally determine the target interface parameter.
[0134] The thinking logic of the large language model is as follows:
[0135] Firstly, it can be understood that the target voice instruction is that the user wants to switch the screen to automatic mode; then the target vehicle information can be determined through various retrieval methods, for example, it can be determined that the screen appearance includes multiple modes, supplementary parameters, appearance, etc., and then the target interface parameter can be obtained.
[0136] Through the above method, the target interface parameter corresponding to the target voice instruction can be accurately determined.
[0137] The following explains another possible implementation of the voice interaction method provided in the embodiments of the present application.
[0138] Figure 10 For another flowchart of the voice interaction method provided in the embodiments of the present application, please refer to Figure 10 According to the target prompt word, the vehicle control task corresponding to the target voice instruction is output and the vehicle control task is executed, including:
[0139] S1010: The target prompt word is spliced with the preset prompt word to obtain a spliced prompt word.
[0140] It should be noted that after obtaining the target prompt word, the target prompt word and the preset prompt word can be spliced, that is, both of the two prompt words can be used as the input prompt word of the target voice instruction.
[0141] The spliced prompt word can be the result of splicing the target prompt word and the preset prompt word, for example, the preset prompt word includes "window" and "open", and the target prompt word includes "media volume" and "turn up", and the spliced prompt word can include all the contents included in the above two prompt words.
[0142] S1020: The spliced prompt word and the target voice instruction are input into the large language model to obtain the vehicle control task corresponding to the target voice instruction and execute the vehicle control task.
[0143] It should be noted that for the target voice instruction that fails to be recognized by the preset prompt word, the large language model can be in a state of waiting for processing, and after obtaining the above spliced prompt word, the spliced prompt word and the target voice instruction can be re-input into the large language model, and the vehicle control task corresponding to the target voice instruction can be obtained after being recognized by the large language model, and then the vehicle control task can be executed.
[0144] In the voice interaction method provided in the embodiments of the present application, the target prompt word can be spliced with the preset prompt word to obtain a spliced prompt word, and the spliced prompt word and the target voice instruction can be input into the large language model to obtain the vehicle control task corresponding to the target voice instruction and execute the vehicle control task. Among them, by splicing the prompt word, the large language model can more accurately determine the vehicle control task corresponding to the target voice instruction.
[0145] The following explains the voice interaction method provided in the embodiments of the present application through an overall implementation process.
[0146] Figure 11 For the overall flowchart of the voice interaction method provided in the embodiments of the present application, please refer to Figure 11 The method includes:
[0147] Firstly, the preset prompt word and the target voice instruction can be input into the large language model, if the large language model can identify the corresponding vehicle control task, the vehicle control task can be output, and the vehicle control task can be executed.
[0148] If the large language model cannot identify the corresponding vehicle control task, the type of the target voice instruction can be determined, which can include on-off type instruction, adjustment type instruction and uncertain instruction.
[0149] For on-off type instruction and uncertain instruction, keyword retrieval can be performed by AC automatic machine, keyword retrieval can be performed by the above basic word library and generalization word library, so as to determine the corresponding keyword, and then the target vehicle information can be determined according to the keyword, and the target interface parameter can be determined based on the target vehicle information, and then the target prompt word can be obtained, after obtaining the target prompt word, the target prompt word and the preset prompt word can be spliced and input into the large language model, since it is the same large language model, the target voice instruction does not need to be input repeatedly, the vehicle control task can be output by the large language model after identification, and the vehicle control task can be executed.
[0150] For adjustment type instruction and uncertain instruction, RAG retrieval can be used for retrieval, so as to determine the target interface parameter, and then the target prompt word can be obtained, after obtaining the target prompt word, the target prompt word and the preset prompt word can be spliced and input into the large language model, since it is the same large language model, the target voice instruction does not need to be input repeatedly, the vehicle control task can be output by the large language model after identification, and the vehicle control task can be executed.
[0151] It should be noted that in the voice interaction method provided in the embodiments of the application, a dynamic task perception intelligent decision mechanism is adopted, an adaptive calling strategy based on task complexity is proposed, different ways can be used to identify target voice instructions with different difficulty levels, a "perception-decision-execution" closed loop intelligence is constructed, and the time consumed for identifying simple instructions is saved. In addition, an asynchronous collaborative lightweight knowledge empowerment architecture is also adopted, in the process of retrieval to generate the target prompt word, a non-blocking knowledge acquisition channel is adopted, when the large language model determines that external knowledge assistance is needed, RAG retrieval or AC automatic machine calling is triggered by the asynchronous message mechanism, and the self enters the elastic processing state of "reasoning-waiting".
[0152] It should be understood that although each step in the above flowcharts is shown in sequence according to the direction of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless explicitly stated herein, there is no strict order limitation for the execution of these steps, and these steps can be executed in other orders. Moreover, at least part of the steps in the above flowcharts can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be alternately executed with at least part of other steps or sub-steps or stages of other steps.
[0153] Based on the foregoing embodiments, the embodiments of the present application provide a voice interaction device, which includes the modules included therein and the units included in the modules, and can be implemented by a processor; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0154] Figure 12 For the structural schematic diagram of the voice interaction device provided in the embodiments of the present application, please refer to Figure 12 Another aspect of the embodiments of the present application also provides a voice interaction device applied to a vehicle terminal, which includes an acquisition module 1210, an identification module 1220, and an execution module 1230.
[0155] The acquisition module 1210 is configured to acquire a target voice instruction collected by the vehicle terminal.
[0156] The identification module 1220 is configured to identify the target voice instruction according to a preset prompt word.
[0157] The execution module 1230 is configured to, in the case that the target voice instruction is successfully identified, output a vehicle control task corresponding to the target voice instruction and execute the vehicle control task.
[0158] The execution module 1230 is further configured to, in the case that the target voice instruction is not successfully identified, determine a target prompt word based on the target voice instruction, output the vehicle control task corresponding to the target voice instruction according to the target prompt word, and execute the vehicle control task.
[0159] In an embodiment, the identification module 1220 is specifically configured to input the preset prompt word and the target voice instruction into a large language model to identify the target voice instruction; or match text information corresponding to the preset prompt word and the target voice instruction to identify the target voice instruction.
[0160] In an embodiment, the execution module 1230 is specifically configured to determine a type of the target voice instruction, the type of the target voice instruction comprising any one of the following: a regulation type instruction, a switch type instruction, and an uncertain instruction; and determine the target prompt word according to the type of the target voice instruction.
[0161] In an embodiment, the execution module 1230 is specifically configured to determine a target interface parameter corresponding to the target voice instruction according to the type of the target voice instruction, the target interface parameter being used to indicate an execution unit and an execution parameter of a vehicle control task; and generate the target prompt word based on the target interface parameter.
[0162] In an embodiment, the execution module 1230 is specifically configured to, when the type of the target voice instruction is the regulation type instruction, retrieve the target interface parameter by searching a retrieval augmented generation (RAG).
[0163] In an embodiment, the execution module 1230 is specifically configured to, when the type of the target voice instruction is the switch type instruction, determine a keyword in the target voice instruction, determine target vehicle information corresponding to the keyword based on multi-mode matching, and determine the target interface parameter according to the target vehicle information.
[0164] In an embodiment, the execution module 1230 is specifically configured to, when the type of the target voice instruction is the uncertain instruction, determine a keyword in the target voice instruction, determine target vehicle information corresponding to the keyword based on multi-mode matching, determine a first interface parameter according to the target vehicle information, retrieve a second interface parameter by searching a retrieval augmented generation (RAG), and obtain the target interface parameter based on the first interface parameter and the second interface parameter.
[0165] In an embodiment, the execution module 1230 is specifically configured to splice the target prompt word with a preset prompt word to obtain a spliced prompt word, input the spliced prompt word and the target voice instruction into a large language model to obtain a vehicle control task corresponding to the target voice instruction, and execute the vehicle control task.
[0166] The voice interaction device provided in this application embodiment can acquire target voice commands collected by an in-vehicle terminal; recognize the target voice commands based on preset prompt words; if the target voice command is successfully recognized, output and execute the vehicle control task corresponding to the target voice command; if the target voice command recognition fails, determine the target prompt word based on the target voice command, output and execute the vehicle control task corresponding to the target voice command based on the target prompt word. Specifically, after acquiring the target voice command, it can be recognized based on preset prompt words. If the recognition is successful using the preset prompt words, the vehicle control task can be output and executed more quickly, reducing the inference time for determining the vehicle control task. If the preset prompt word is successfully recognized, there is no need to generate the target prompt word, reducing the process of determining the vehicle control task, thus allowing for more efficient and faster determination of the vehicle control task. Furthermore, faster determination of the vehicle control task improves the response speed and allows for faster execution of the vehicle control task.
[0167] The descriptions of the above device embodiments are similar to those of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0168] It should be noted that, in the embodiments of this application... Figure 12 The module division of the voice interaction device shown is illustrative and represents only one logical functional division; in actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, exist as separate physical units, or have two or more units integrated into one unit. The integrated units can be implemented in hardware, as software functional units, or a combination of both.
[0169] It should be noted that, in the embodiments of this application, if the above-described methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.
[0170] Figure 13 For the structural schematic diagram of the computer device provided in the embodiments of the present application, please refer to Figure 13 The computer device provided in the embodiments of the present application can be the vehicle-mounted terminal described above, or a car machine, an intelligent vehicle cabin, etc. that includes the vehicle-mounted terminal, and is not specifically limited here, and its internal structure diagram can be as shown in Figure 13 The computer device includes a processor 1320, a memory and a network interface 1340 connected through a system bus 1310. The processor 1320 of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium 1331 and an internal memory 1332. The non-volatile storage medium 1331 stores an operating system, a computer program and a database. The internal memory 1332 provides an environment for the operating system and the computer program in the non-volatile storage medium 1331 to run. The database of the computer device is configured to store data. The network interface 1340 of the computer device is configured to communicate with external terminals through network connection. The computer program is executed by the processor 1320 to implement the above method.
[0171] The computer readable storage medium provided in the embodiments of the present application stores a computer program, and the computer program is executed by the processor to implement the steps in the method provided in the above embodiments.
[0172] The computer program product provided in the embodiments of the present application includes instructions, and when the computer program product is run on a computer, the computer is caused to execute the steps in the method provided in the above method embodiments.
[0173] Those skilled in the art can understand that Figure 13 The structure shown in the above embodiments is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0174] In one embodiment, the voice interaction device provided in the present application can be implemented in the form of a computer program, which can run on a computer device as shown in Figure 13 The memory of the computer device can store various program modules constituting the above device. The computer program constituted by the various program modules causes the processor to execute the steps in the method of each embodiment of the present application described in the specification.
[0175] It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium, storage medium and device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0176] It should be understood that the "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiments are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" or "in some embodiments" appearing throughout the specification does not necessarily mean the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of each process does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The sequence number of the above embodiments of the present application is only for description, not representing the advantages and disadvantages of the embodiments. The above description of each embodiment tends to emphasize the differences between each embodiment, and the same or similar parts can be referred to each other, and for the sake of brevity, this paper will not be repeated here.
[0177] The term "and / or" in this paper is only a description of the association relationship between the associated objects, which means that there can be three kinds of relationships, for example, object A and / or object B, which can represent the following three cases: object A exists alone, object A and object B exist together, and object B exists alone.
[0178] It should be noted that in this paper, the term "include", "contain" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or includes elements inherent to such process, method, article or device. Without more limitations, the element defined by the sentence "including a…" does not exclude the presence of other identical elements in the process, method, article or device including the element.
[0179] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other manners. The embodiments described above are merely exemplary, for example, the division of the modules is only a logical function division, and there can be another division manner for the actual implementation, for example, a plurality of modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the various components shown or discussed can be indirect coupling or communication connection through some interfaces, devices or modules, and can be electrical, mechanical or other forms.
[0180] The modules described above as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules; they can be located in one place, or distributed on multiple network units; and some or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments.
[0181] In addition, each functional module in each embodiment of the present application can be integrated into one processing unit, or each module can be a separate unit, or two or more modules can be integrated into one unit; the integrated module can be realized in the form of hardware or hardware plus software functional unit.
[0182] Those of ordinary skill in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by a program instructing related hardware, and the foregoing program can be stored in a computer readable storage medium, and the program executes the steps of the above-mentioned method embodiments when executed; and the foregoing storage medium includes mobile storage devices, read-only memories (ROM), magnetic disks or optical disks and various storage media that can store program codes.
[0183] Alternatively, the integrated units of the present application, if implemented in the form of software functional modules and sold or used as independent products, can also be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes a plurality of instructions for causing an electronic device to execute all or part of the methods described in the embodiments of the present application. The foregoing storage medium includes mobile storage devices, ROM, magnetic disks or optical disks and various storage media that can store program codes.
[0184] The methods disclosed in the several method embodiments provided in the present application can be combined arbitrarily without conflict, to obtain new method embodiments.
[0185] The features disclosed in several product embodiments provided by the present application can be arbitrarily combined, without conflict, to obtain new product embodiments.
[0186] The features disclosed in several method or device embodiments provided by the present application can be arbitrarily combined, without conflict, to obtain new method embodiments or device embodiments.
[0187] The above description is merely illustrative of the application, and the scope of the application is not limited thereto. Any variations and modifications of the application, which would occur to those skilled in the art, are to be considered within the scope of the application. Therefore, the scope of the application is to be determined by the claims.
Claims
1. A voice interaction method, characterized in that, Applications in vehicle-mounted terminals include: Acquire the target voice command collected by the vehicle terminal; The target voice command is identified based on preset prompt words; If the target voice command is successfully recognized, the vehicle control task corresponding to the target voice command is output and the vehicle control task is executed. If the target voice command recognition fails, a target prompt word is determined based on the target voice command, and a vehicle control task corresponding to the target voice command is output and executed according to the target prompt word.
2. The method according to claim 1, characterized in that, The step of recognizing the target voice command based on preset prompt words includes: The preset prompt words and the target voice command are input into a large language model to recognize the target voice command; or... The preset prompt words are matched with the text information corresponding to the target voice command to identify the target voice command.
3. The method according to claim 2, characterized in that, Before determining the target prompt word based on the target voice command, the method further includes: Determine the type of the target voice command, wherein the type of the target voice command includes any one of the following: adjustment command, switch command, and uncertain command; Determining the target prompt word based on the target voice command includes: The target prompt word is determined based on the type of the target voice command.
4. The method according to claim 3, characterized in that, Determining the target prompt word based on the type of the target voice command includes: The target interface parameters corresponding to the target voice command are determined according to the type of the target voice command. The target interface parameters are used to indicate the execution unit and execution parameters of the vehicle control task. The target prompt word is generated based on the target interface parameters.
5. The method according to claim 4, characterized in that, Determining the target interface parameters corresponding to the target voice command based on the type of the target voice command includes: When the target voice command is an adjustment command, the target interface parameters are retrieved by using the Retrieval Enhancement Generation RAG.
6. The method according to claim 4, characterized in that, Determining the target interface parameters corresponding to the target voice command based on the type of the target voice command includes: When the type of the target voice command is a switch command, the keywords in the target voice command are determined, and the target vehicle information corresponding to the keywords is determined based on multi-pattern matching; The target interface parameters are determined based on the target vehicle information.
7. The method according to claim 4, characterized in that, Determining the target interface parameters corresponding to the target voice command based on the type of the target voice command includes: When the type of the target voice command is an uncertain command, the keywords in the target voice command are determined, and the target vehicle information corresponding to the keywords is determined based on multi-pattern matching; Determine the first interface parameters based on the target vehicle information; The second interface parameters were retrieved by enhancing the RAG retrieval process. The target interface parameters are obtained based on the first interface parameters and the second interface parameters.
8. The method according to claim 1, characterized in that, The step of outputting and executing the vehicle control task corresponding to the target voice command based on the target prompt word includes: The target prompt word is concatenated with the preset prompt word to obtain the concatenated prompt word; The concatenated prompt words and the target voice command are input into the large language model to obtain the vehicle control task corresponding to the target voice command and execute the vehicle control task.
9. A voice interaction device, characterized in that, Applied to in-vehicle terminals, it includes: an acquisition module, an identification module, and an execution module; The acquisition module is used to acquire the target voice command collected by the vehicle terminal; The recognition module is used to recognize the target voice command based on preset prompt words; The execution module is used to output a vehicle control task corresponding to the target voice command and execute the vehicle control task when the target voice command is successfully recognized. The execution module is further configured to, in the event that the target voice command recognition fails, determine a target prompt word based on the target voice command, output a vehicle control task corresponding to the target voice command according to the target prompt word, and execute the vehicle control task.
10. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Vehicle control method and system and vehicle
CN112435660A
Control instruction determination method and device, storage medium and electronic device
CN116072113A
Equipment control method and device based on large language model
CN118155610A
Voice instruction processing method and device and electronic equipment
CN118782032A
Vehicle control method, server and computer readable storage medium
CN120108392A
Cited By
Audio processing method and device, computer equipment and storage medium
CN122157682A
Audio processing method and device, computer device and storage medium
CN122157682B