Voice control method and device
By combining speech text and vehicle state, using inference model and reinforcement learning algorithm optimized speech control method, the problem of low recognition accuracy and poor stability in vehicle cockpit functional control is solved, and more efficient and accurate user intention recognition and operation execution is achieved.
Patent Information
- Application Number
- CN202510492567.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-04
Smart Images

Figure CN120260569A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of vehicles, and particularly to a voice control method and device. Background Art
[0002] With the development of science and technology, it is possible to control some in-vehicle cockpit functions through voice. For example, a driver can use voice commands to control basic functions such as air-conditioning adjustment, music playback, or phone calls. The application of this technology liberates the driver's hands to a great extent and helps to improve driving safety. However, voice recognition technology can only trigger a preset set of in-vehicle cockpit functions through preset fixed languages, resulting in poor generalization ability and a narrow function coverage range.
[0003] With the development of large language models (LLMs), in current technology, more complex user languages are analyzed and understood through large language models to control in-vehicle cockpit functions. For example, various languages with similar semantics can be used to trigger corresponding in-vehicle cockpit functions, which can improve the generalization ability and thus expand the coverage of in-vehicle cockpit functions.
[0004] However, in current technology, when using a large language model to recognize user speech to control corresponding in-vehicle cockpit functions, there are problems such as low recognition accuracy and unstable recognition results. Summary of the Invention
[0005] Based on the above problems, the embodiments of this application provide a voice control method and device, which combine the user input voice and vehicle state, comprehensively recognize the user intention, and convert the user intention into a structured instruction, thereby achieving a deep understanding of the user intention, improving the recognition accuracy, and ensuring the stability of the recognition result.
[0006] The embodiments of this application disclose the following technical solutions:
[0007] In a first aspect, this application provides a voice control method, including:
[0008] When receiving the voice text input by the user, obtain the vehicle state;
[0009] Based on the voice text and the vehicle state, recognize the user intention;
[0010] When the user intention represents a next action, convert the next action represented by the user intention into a structured instruction; wherein, the structured instruction is used to trigger the corresponding operation of the in-vehicle cockpit function module;
[0011] Send the structured instruction to the corresponding in-vehicle cockpit function module so that the in-vehicle cockpit function module responds to the structured instruction and performs the operation corresponding to the structured instruction.
[0012] Optionally, identifying the user intention based on the voice text and the vehicle state includes:
[0013] Input the voice text and the vehicle state into the inference model so that the inference model outputs the user intention based on the voice text and the vehicle state.
[0014] Optionally, the method further includes:
[0015] Store the voice text and vehicle state input into the inference model, and the user intention output by the inference model as a set of training samples;
[0016] Use a preset reward function to evaluate the training samples and obtain the corresponding reward value;
[0017] Optimize the inference model using a reinforcement learning algorithm based on the corresponding reward value.
[0018] Optionally, the preset reward function includes: a first reward, a second reward, a third reward, and a fourth reward; wherein, the first reward is used to evaluate whether the inference model accurately outputs the user intention, the second reward is used to evaluate whether the user intention output by the inference model considers the safety state of the vehicle itself, the third state is used to evaluate the efficiency of the inference model in outputting the user intention based on the voice text and the vehicle state, and the fourth reward is used to evaluate whether the user intention output by the inference model considers the user's habits.
[0019] Optionally, when the user intention represents the next action, converting the next action represented by the user intention into a structured instruction includes:
[0020] Input the user intention into the instruction following model so that the instruction following model responds to the next action represented by the user intention and converts the next action represented by the user intention into a structured instruction.
[0021] Optionally, when the user intention represents the next action, converting the next action represented by the user intention into a structured instruction includes:
[0022] When the user intention represents the next action, obtain the instruction text of the next action in a preset format based on the user intention; wherein, the instruction text of the next action is used to represent the in-vehicle cockpit function module corresponding to the next action and the operation performed by the corresponding in-vehicle cockpit function module.
[0023] Convert the instruction text of the next action into a corresponding structured instruction; wherein, the structured instruction is a function that can be recognized and executed by the corresponding in-vehicle cockpit function module.
[0024] Optionally, before sending the structured instruction to the corresponding in-vehicle cockpit function module, the method further includes:
[0025] Check whether the parameters in the structured instruction are complete;
[0026] When the parameters in the structured instruction are incomplete, generate and output a prompt message based on the structured instruction to prompt that the parameters in the structured instruction are incomplete;
[0027] The sending the structured instruction to the corresponding in-vehicle cockpit function module includes:
[0028] When the parameters in the structured instruction are complete, send the structured instruction to the corresponding in-vehicle cockpit function module.
[0029] Optionally, the method further includes:
[0030] Receive the execution result returned by the in-vehicle cockpit function module;
[0031] Update the vehicle state based on the execution result to obtain an updated vehicle state;
[0032] Modify the voice text input by the user according to the updated vehicle state to obtain an updated voice text;
[0033] Based on the updated voice text and the updated vehicle state, determine whether all user intents corresponding to the voice text input by the user are completed;
[0034] When not all user intents corresponding to the voice text input by the user are completed, the identifying the user intent based on the voice text and the vehicle state includes:
[0035] Re-identify the user intent based on the updated voice text and the updated vehicle state.
[0036] Optionally, the updating the vehicle state based on the execution result to obtain an updated vehicle state includes:
[0037] Update the vehicle state based on the execution result through the inference model to obtain an updated vehicle state;
[0038] Modifying the voice text input by the user according to the updated vehicle state to obtain an updated voice text, including:
[0039] Modifying the voice text input by the user according to the updated vehicle state through the instruction following model to obtain an updated voice text;
[0040] Judging whether all user intents corresponding to the voice text input by the user are completed based on the updated voice text and the updated vehicle state, including:
[0041] Judging whether all user intents corresponding to the voice text input by the user are completed based on the updated voice text and the updated vehicle state through the inference model.
[0042] In a second aspect, the present application provides a voice control device, including:
[0043] An acquisition module, configured to acquire a vehicle state when receiving a voice text input by a user;
[0044] An identification module, configured to identify a user intent based on the voice text and the vehicle state;
[0045] A structuring module, configured to convert the next action represented by the user intent into a structured instruction when the user intent represents a next action; wherein, the structured instruction is used to trigger a corresponding in-vehicle cockpit function module to execute a corresponding operation;
[0046] A sending module, configured to send the structured instruction to a corresponding in-vehicle cockpit function module, so that the in-vehicle cockpit function module responds to the structured instruction and executes an operation corresponding to the structured instruction.
[0047] Compared with the prior art, the present application has the following beneficial effects: comprehensively considering the voice text and the vehicle state, identifying the user intent, converting the user intent into a structured instruction, and sending it to a corresponding in-vehicle cockpit function module to complete a corresponding operation. Since the voice text and the vehicle state are comprehensively considered, a deep understanding of the user intent is achieved, the accuracy of identification is improved, and the stability of the identification result is ensured. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0049] Figure 1 A flowchart showing the process of a voice control method provided by an embodiment of the present application;
[0050] Figure 2 A flowchart showing the optimization process of an inference model provided by an embodiment of the present application;
[0051] Figure 3 A flowchart showing the process of generating and sending structured instructions provided by an embodiment of the present application;
[0052] Figure 4 A flowchart showing the process of another voice control method provided by an embodiment of the present application;
[0053] Figure 5 A structural diagram of a voice control device provided by an embodiment of the present application. Detailed implementation manners
[0054] As described above, in the current technology, a large language model is used to analyze and understand more complex user voices, so as to voice-control in-vehicle cockpit functions. A large language model refers to a language processing model constructed using deep learning technology and having a large number of parameters.
[0055] However, since the original design intention of the large language model is general language processing, its performance in the professional scenario of in-vehicle cockpit function control cannot meet the expectations, especially for specific technical field terms or precise operations, such as: guidance, path calculation, etc. Therefore, the large language model in the current technology has the problem of low recognition accuracy; secondly, the large language model is a probability-based model, and there may be instability in the speech recognition results, that is, for the same input user language, the recognition results output by the large language model may vary greatly. For example: for the same user voice, sometimes it cannot be recognized, and sometimes the recognition results of the two times are completely different.
[0056] The present application provides a voice control method, including: when receiving the voice text input by the user, obtaining the vehicle state; based on the voice text and the vehicle state, identifying the user intention; when the user intention represents the next action, converting the next action represented by the user intention into a structured instruction; sending the structured instruction to the corresponding in-vehicle cockpit function module, so that the in-vehicle cockpit function module responds to the structured instruction and executes the operation corresponding to the structured instruction. The embodiments of the present application comprehensively consider the voice text and the vehicle state, identify the user intention, realize the in-depth understanding of the user intention, improve the accuracy of the user intention recognition and ensure the stability of the recognition result, thereby improving the accuracy of the voice control.
[0057] Furthermore, by converting the user's intention into a structured instruction and sending it to the corresponding in-vehicle cockpit function module to complete the corresponding operation, the generalization of voice input for triggering operations is improved, and the accuracy of the triggered operations / functions is ensured, thereby enhancing the stability of voice control.
[0058] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0059] Embodiment 1:
[0060] The following combines Figures 1-3 , and details a voice control method provided by the embodiments of this application. This method is applied to a cloud server that communicates or interacts with a vehicle machine system (including: in-vehicle cockpit function modules).
[0061] Among them, the cloud server that communicates or interacts with the vehicle machine system generally refers to the backend infrastructure that supports vehicle networking functions and provides data processing, storage, and various services. The cloud server communicates with each in-vehicle cockpit function module through vehicle networking to achieve various functions and services.
[0062] As Figure 1 shown, a language control method provided by the embodiments of this application includes the following steps:
[0063] S101. When receiving the voice text input by the user, obtain the vehicle state.
[0064] Among them, the vehicle state (CarState) refers to various information provided by the vehicle machine system about the current operating state of the vehicle. The vehicle state includes but is not limited to: vehicle speed, which refers to the current driving speed of the vehicle; mileage, which refers to the total distance the vehicle has traveled, usually divided into two recording methods: short mileage and long mileage; energy situation, for a fuel vehicle, it refers to the remaining fuel in the fuel tank, and for an electric vehicle, it refers to the remaining battery power and battery life; seat belt state, which refers to the information on whether the seat belt is fastened; headlight state, which refers to the working state of the front headlights, taillights, turn signals, etc.; the display content on the navigation device, for example: the display information of the in-vehicle navigation function (such as the current location, destination direction, estimated arrival time, traffic status update, etc.).
[0065] Specifically, receive and recognize the user's input voice to obtain the voice text input by the user; when recognizing / receiving the voice text input by the user, obtain the current vehicle state of the vehicle from the in-vehicle system of the vehicle.
[0066] In a possible implementation, for receiving and recognizing the user's input voice, multiple microphones (i.e., microphone array) are usually equipped on the in-vehicle system to capture the user's voice. After the voice input by the user is collected through the microphone array of the in-vehicle system, the voice input by the user is recognized through Automatic Speech Recognition (ASR) technology to obtain the voice text input by the user.
[0067] Exemplarily, the user says "I want to go home". The voice "I want to go home" is collected through the microphone array, and through automatic language recognition, the voice text "I want to go home" is obtained.
[0068] S102. Identify the user's intention based on the voice text input by the user and the vehicle state.
[0069] Among them, the user's intention is the specific requirement of the in-vehicle cockpit function expressed by the user through voice, such as a specific operation request or function requirement. Exemplarily, the user's intention can be to call the home function of the navigation application, turn on the air conditioner, etc.
[0070] In a possible implementation, input the voice text input by the user and the vehicle state into the reasoning model; the reasoning model outputs the user's intention based on the voice text and the vehicle state. Further, after splicing the voice text input by the user and the vehicle state, input it into the reasoning model.
[0071] Among them, the reasoning model is a large language model with deep reasoning ability and specially optimized, which is used to output the corresponding user intention based on the input voice text and vehicle state.
[0072] Exemplarily, the voice text input by the user is "I want to go home, I want to listen to a song by Jay Chou", and the vehicle state includes: the navigation function is in the on state, and the music function is in the on state. Input the above voice text and vehicle state into the reasoning model. The reasoning model first obtains the corresponding initial user intentions "turn on the navigation function, navigate home, turn on the music function, retrieve and play songs by Jay Chou" based on the voice text "I want to go home, I want to listen to a song by Jay Chou"; the reasoning model combines the vehicle state and knows that the navigation function and music function are in the on state and no longer need to be turned on, so it determines that all the corresponding user intentions are "navigate home, retrieve and play songs by Jay Chou".
[0073] Further, when the user intention inferred by the inference model for the speech text and vehicle state is used to represent an action, it represents the next action corresponding to the speech text, rather than all actions corresponding to the speech text. Exemplarily, when the inference model combines the vehicle state and the speech text and determines that all corresponding user intentions are "navigate home, retrieve and play Jay Chou's songs", the user intention output by the inference model is "navigate home", which represents the next action, and the next action can be "call the navigation function to go home" or "call the navigation's go-home function".
[0074] In a possible implementation, during the use of the inference model, the reinforcement learning algorithm is simultaneously used to optimize the inference model, improving the inference model's ability to understand and respond to the input speech text and vehicle state, enabling it to more accurately output the user intention, and as time goes by and more data accumulates, the performance of the inference model will gradually improve, thus achieving more efficient and accurate operations.
[0075] For ease of understanding, the following combines Figure 2 , and details the optimization process of the inference model.
[0076] S201. Store the speech text and vehicle state input into the inference model, as well as the user intention output by the inference model, as a set of training samples.
[0077] Specifically, when the inference model receives the speech text and vehicle state, while outputting the corresponding user intention, store the speech text, vehicle state, and user intention as a set of training samples.
[0078] S202. Use a preset reward function to evaluate the training samples and obtain the corresponding reward value.
[0079] Specifically, for each set of stored training samples, use the preset reward function to evaluate the training samples and obtain the corresponding reward value.
[0080] Among them, the reward value reflects the effect of the inference model in inferring the user intention. When the reward value is positive, it indicates positive feedback, that is, the inference model's understanding of the vehicle state and speech text and the finally output user intention are correct; when the reward value is negative, it indicates negative feedback, that is, the inference model's understanding of the vehicle state and speech text and the finally output user intention are incorrect.
[0081] Among them, the preset reward function includes: the first reward, the second reward, the third reward, and the fourth reward.
[0082] Among them, the first reward can also be called the accuracy reward, which is used to evaluate whether the inference model accurately outputs (identifies) the user's intention. Specifically, the inference model should accurately identify and output the user's intention. When the inference model can correctly identify and output the user's intention, the reward value of the first reward is positive; otherwise, it is negative. Exemplarily, when the window state in the vehicle state is the fully closed state, and the voice text input by the user is "open the left front window", if the inference model identifies the function of the user's intention as "open the window", a positive reward is given to the first reward, and if the parameter of the user's intention "left front window" is identified, a positive reward is given to the first reward; correspondingly, when the parameter of the user's intention identified by the inference model is not "left front window", a negative reward is given to the first reward.
[0083] Among them, the second reward can also be called the safety reward, which is used to evaluate whether the user's intention output (identified) by the inference model takes into account the safety state of the vehicle itself. Specifically, the inference model should consider which functions and operations can be performed under different safety states of the vehicle to ensure the safety of the vehicle itself. When the function / operation corresponding to the user's intention identified by the inference model can ensure the safety of the vehicle itself, the reward value of the second reward is positive; otherwise, it is negative. Exemplarily, when the vehicle is traveling at high speed, to ensure the safety of the vehicle itself, it should be ensured that the vehicle's windows are in the closed state. When the user's intention identified by the inference model is that the window cannot be opened, a positive reward is given to the second reward; when the user's intention identified by the inference model is to open the window, a negative feedback is given to the second reward.
[0084] Among them, the third reward can also be called the inference efficiency reward, which is used to evaluate the efficiency of the inference model in outputting (identifying) the user's intention based on the voice text and the vehicle state. Specifically, the inference model should identify the user's intention as quickly as possible and with less consumption. When the inference model can quickly identify the user's intention, the reward value of the third reward is positive; otherwise, it is negative. Exemplarily, the efficiency of the inference model in outputting the user's intention based on the voice text and the vehicle state is characterized by the token used by the inference model (the basic unit for the inference model to consume resources, serving as a unit for measuring efficiency). When the token is less than the preset token, a positive feedback is given to the third reward; when the token is greater than the preset token, a negative reward is given to the third reward.
[0085] Among them, the fourth reward can also be called a habit adaptation reward, which is used to evaluate whether the user intention output by the inference model takes into account the user's habits. Specifically, when the inference model outputs the user intention based on the speech text and the vehicle state, it can also be optimized according to the user's habits, where the user's habits can be obtained from the local storage of the vehicle's in-vehicle system. When the user intention output by the inference model takes into account the user's habits, the reward value of the fourth reward is positive; otherwise, it is negative. Exemplarily, the user's habit is to often eat Sichuan cuisine, the speech text is "I'm hungry", and the vehicle state includes: the navigation function is in the open state. When the user intention output by the inference model based on the speech text and the vehicle state is "Search for Sichuan restaurants in the navigation", a positive reward is given to the fourth reward; when the user intention output by the inference model is "Search for Cantonese restaurants in the navigation" and does not take into account the user's habits, a negative reward is given to the fourth reward.
[0086] For ease of understanding, the preset reward function in the embodiments of the present application is introduced below in conjunction with formula (1).
[0087] R 推理 = w1×R1 + w2×R2 + w3×R3 + w4×R4 (1)
[0088] Among them, R 推理 is the preset reward function, and R 推理 can also be labeled as R(s, a, s'), where s is the current vehicle state, a is the user intention output by the inference model, and s' is the next vehicle state; R1, R2, R3, and R4 are the first reward, the second reward, the third reward, and the fourth reward respectively; w1, w2, w3, and w4 are the weights corresponding to the first reward, the second reward, the third reward, and the fourth reward respectively, and w1 + w2 + w3 + w4 = 1.
[0089] S203. Optimize the inference model using a reinforcement learning algorithm based on the corresponding reward value.
[0090] Among them, reinforcement learning (RL) is a machine learning method that enables a computer to autonomously learn in an uncertain environment through a trial-and-error approach. Reinforcement learning algorithms include: Q-learning, SARSA (State-Action-Reward-State-Action), Deep Q-Network (DQN), etc.
[0091] Specifically, based on the corresponding reward value, the parameters of the inference model are adjusted using a reinforcement learning algorithm so that a better user intention can be output when encountering a similar input / situation in the future.
[0092] The above combination Figure 2The optimization process of the inference model is introduced in detail. Next, continue to combine with Figure 1 introduce a voice control method provided by an embodiment of the present application.
[0093] S103. When the user intention represents the next action, convert the next action represented by the user intention into a structured instruction.
[0094] Among them, the structured instruction is used to trigger the in-vehicle cockpit function module to execute the corresponding operation, and the corresponding operation to be executed is related to the next action represented by the user intention. For example: if the user intention is "navigate home", the structured instruction is used to trigger the navigation function module to execute the operation of navigating home. It can be simply understood that: the next action represented by the user intention is the colloquial version or the description that is easy for the user to understand of the corresponding operation to be executed by the subsequent in-vehicle cockpit function module.
[0095] In a possible implementation manner, determine whether the user intention represents the next action or the voice input by the user is for chatting; when it is determined that the user intention represents the next action, convert the next action represented by the user intention into a structured instruction; when it is determined that the user intention represents that the voice input by the user is for chatting, generate a corresponding reply message based on the voice text input by the user. Among them, the reply message is a reply to the voice text input by the user.
[0096] Exemplarily, if the user intention is "lower the air conditioner temperature to 20°C", then the user intention represents the next action; the conversion of the user intention into a structured instruction is a structured instruction used to trigger the air conditioner module to execute the operation of lowering the temperature to 20°C.
[0097] Exemplarily, if the user intention is "ask the current time" and there is no need to trigger the execution of the corresponding operation on the in-vehicle cockpit function module, then the user intention represents that the voice input by the user is for chatting; the corresponding reply message generated based on the voice text input by the user can be "It is 3:15 pm now. If you have a schedule, I can help you check the traffic conditions".
[0098] In a possible implementation manner, send the user intention to the instruction following model; the instruction following model, in response to the user intention representing the next action, converts the next action represented by the user intention into a structured instruction.
[0099] Specifically, when the instruction following model receives the user intention, it analyzes the user intention to determine whether the user intention represents the next action or the voice input by the user is for chatting; when it is determined that the user intention represents the next action, convert the next action represented by the user intention into a structured instruction.
[0100] S104. Send the structured instruction to the corresponding in-vehicle cockpit function module so that the in-vehicle cockpit function module can execute the operation corresponding to the structured instruction in response to the structured instruction.
[0101] Among them, the in-vehicle cockpit functions generally refer to a series of technologies and features installed inside the vehicle, aiming to enhance the driving experience, safety, and passenger comfort; these functions can cover a wide range of technologies from entertainment systems to driver assistance systems. Exemplarily, the in-vehicle cockpit functions include: infotainment functions (such as navigation functions, Bluetooth connection functions, audio functions), driver assistance functions, dashboard display functions, environmental perception functions, etc.
[0102] Exemplarily, send the structured instruction of navigating home to the navigation function module so that the navigation function module can execute the function of navigating home in response to the structured instruction.
[0103] For ease of understanding, the following will introduce in detail the process of generating and sending the structured instruction in combination with Figure 3 ,
[0104] S301. When the user intention represents the next action, based on the user intention, obtain the instruction text of the next action in a preset format.
[0105] Among them, the instruction text of the next action is used to represent the in-vehicle cockpit function module corresponding to the next action and the operation executed by the corresponding in-vehicle cockpit function module.
[0106] Specifically, when the user intention represents the next action, through the instruction following model, obtain the instruction text of the next action in a preset format based on the user intention.
[0107] In a possible implementation manner, the preset format is stored in the instruction following model in the form of an instruction following prompt word. Exemplarily, the preset format is: [In-vehicle cockpit function module: Operation executed by the in-vehicle cockpit function module].
[0108] For example: If the user intention is "call navigation to home", then the corresponding instruction text of the next action in the preset format is: [Navigation function: Navigate home].
[0109] S302. Convert the instruction text of the next action into the corresponding structured instruction.
[0110] Among them, the structured instruction is a function that can be recognized and executed by the corresponding in-vehicle cockpit function module.
[0111] Specifically, the instruction following model matches the corresponding function based on the instruction text of the next action and fills in the relevant parameters to obtain the structured instruction.
[0112] S303. Check whether the parameters in the structured instruction are complete.
[0113] Specifically, the instruction following model checks whether the parameters in the structured instruction are complete. Exemplarily, if the instruction text for the next action is: [Music function: Retrieve the music of singer A], then based on the instruction text for the next action, a corresponding structured instruction a is generated. Checking the structured instruction a reveals that the parameter of singer A is missing, that is, the structured instruction lacks the relevant parameters of singer A and cannot be executed.
[0114] When the parameters in the structured instruction are complete, proceed to S304.
[0115] When the parameters in the structured instruction are incomplete, proceed to S305.
[0116] S304. Send the structured instruction to the corresponding in-vehicle cockpit function module.
[0117] S305. Based on the structured instruction, generate and output a prompt message to indicate that the parameters in the structured instruction are incomplete.
[0118] In a possible implementation, when the parameters in the structured instruction are incomplete, a prompt message is generated and output, and the prompt message also carries the parameter information that needs to be supplemented.
[0119] Exemplarily, if the instruction text for the next action is: [Music function: Retrieve the music of singer A], and the parameter of "singer A" is missing in the generated corresponding structured instruction a, then the prompt message carries the information to prompt the input or addition of the "singer" parameter.
[0120] Generally speaking, in the embodiments of the present application, when receiving the voice text input by the user, the vehicle state is obtained, and the voice text and the vehicle state are sent to the inference model; the inference model, based on the voice text and the vehicle state, identifies the user intention and sends the identified user intention to the instruction following model; the instruction following model, in response to the user intention representing the next action, converts the next action represented by the user intention into a structured instruction, and sends the structured instruction to the corresponding in-vehicle cockpit function module through the vehicle communication function. By combining the inference model and the instruction following model, different models handle partial tasks (i.e., the inference model identifies the user intention, and the instruction following model generates the structured instruction), realizing the transparency of the intermediate processing process, thereby enhancing the interpretability of the overall voice control function, effectively overcoming the limitations of a single model in multi-task processing, and providing a clearer and more reliable basis for complex task execution. Further, both the inference model and the instruction following model can be independently trained to improve their effects in handling their respective tasks.
[0121] A voice control method provided by an embodiment of the present application includes: when receiving a voice text input by a user, obtaining the vehicle state; based on the voice text and the vehicle state, identifying the user intention; when the user intention represents a next action, converting the next action represented by the user intention into a structured instruction; and sending the structured instruction to a corresponding in-vehicle cockpit function module, so that the in-vehicle cockpit function module responds to the structured instruction and executes the operation corresponding to the structured instruction. The embodiment of the present application comprehensively considers the voice text and the vehicle state, identifies the user intention, realizes a deep understanding of the user intention, improves the accuracy of user intention recognition and ensures the stability of the recognition result, thereby improving the accuracy of voice control.
[0122] Furthermore, by converting the user intention into a structured instruction and sending it to a corresponding in-vehicle cockpit function module to complete the corresponding operation, the generalization of voice input for triggering operations is improved, and the accuracy of the triggered operation / function is ensured, thereby improving the stability of voice control.
[0123] Furthermore, comprehensively considering the voice text and the vehicle state, identifying the user intention is achieved through an inference model. During the use of the inference model, the voice text, the vehicle state, and the user intention are used as training samples, and the reward value of the training samples is calculated based on a preset reward function. Based on the reward value, the inference model is continuously optimized using a reinforcement learning algorithm. This improves the ability of the inference model to understand and respond to the input voice text and vehicle state, enabling it to more accurately output the user intention. And as time goes by and more data accumulates, the performance of the inference model will gradually improve, thereby achieving more efficient and accurate operations.
[0124] Embodiment 2:
[0125] The following combines Figure 4 , and details another voice control method provided by an embodiment of the present application.
[0126] As Figure 4 shown, a voice control method provided by an embodiment of the present application includes the following steps:
[0127] S401. When receiving a voice text input by a user, obtain the vehicle state.
[0128] S402. Based on the voice text input by the user and the vehicle state, identify the user intention.
[0129] S403. When the user intention represents a next action, convert the next action represented by the user intention into a structured instruction.
[0130] S404. Send the structured instruction to the corresponding in-vehicle cockpit function module so that the in-vehicle cockpit function module responds to the structured instruction and performs the operation corresponding to the structured instruction.
[0131] It should be noted that the above S401 - S404 is the same as S101 - S104 in Embodiment 1. Therefore, for the specific implementation details of S401 - S404, please refer to S101 - S104 in Embodiment 1 and will not be elaborated here.
[0132] S405. Receive the execution result returned by the in-vehicle cockpit function module.
[0133] Among them, the execution result includes: instruction execution status (such as success, failure), the status of the function module after execution, etc.
[0134] Exemplarily, for the structured instruction of "open the window", the execution result returned by the window control function module received may include: successfully opening the window, and successfully opening the front left window, etc.; or being unable to open the window, and the reason for being unable to open the window (such as: mechanical failure).
[0135] Exemplarily, for the structured instruction of "retrieve the songs of singer A", the execution result returned by the music function module received may include: successfully retrieving the songs of singer A, and the list of songs of singer A retrieved, etc.
[0136] S406. Update the vehicle state based on the execution result to obtain the updated vehicle state.
[0137] Exemplarily, the current vehicle state is: all windows are in the closed state, and the structured instruction is the structured instruction of "open the front left window". The execution result returned by the window control function module received is that the front left window has been opened; then based on the execution result, modify the state of the windows in the current vehicle state from being in the closed state to the front left window being in the open state and the remaining windows being in the closed state to obtain the updated vehicle state.
[0138] In a possible implementation, through the inference model, update the vehicle state based on the execution result to obtain the updated vehicle state. That is, input the execution result into the inference model so that the inference model updates the vehicle state based on the execution result to obtain the updated vehicle state, and outputs the updated vehicle state to the instruction following model.
[0139] S407. Modify the user input voice text according to the updated vehicle state to obtain the updated voice text.
[0140] Specifically, based on the updated vehicle state and the structured instructions that have been issued, it is possible to determine whether the current user intention has been completed. Therefore, according to the updated vehicle state, the voice text input by the user is modified to obtain the updated voice text, thereby preventing the same user intention from being repeatedly executed / completed.
[0141] Exemplarily, the voice text input by the user is "Increase the volume, open the window". Assume that the first recognized user intention is "Increase the volume". Subsequently, a structured instruction is generated based on the user intention and sent to the volume control function module, and the operation of increasing the volume is executed and completed. And the vehicle state is updated based on the operation of increasing the volume, so that the vehicle state can represent that the operation of increasing the volume has been completed. And based on the updated vehicle state, the voice text input by the user is modified to "Open the window" for subsequent execution of the "Open the window" operation. However, if the vehicle state is not updated and the voice text input by the user is not modified based on the updated vehicle state, then the user intention recognized subsequently based on the voice text input by the user and the vehicle state may still be "Increase the volume", and the operation of increasing the volume will be repeatedly executed, resulting in a dead loop in the entire process and unable to meet all user intentions.
[0142] In a possible implementation, through an instruction following model, according to the updated vehicle state, the voice text input by the user is modified to obtain the updated voice text. That is, the inference model outputs the updated vehicle state to the instruction following model, and the instruction following model modifies the voice text input by the user according to the updated vehicle state to obtain the updated voice text, and outputs the updated voice text to the inference model.
[0143] Specifically, the instruction following model modifies the voice text input by the user according to the updated vehicle state and the last output structured instruction (i.e., the operation executed by the in-vehicle cockpit function module triggered by the structured instruction) to obtain the updated voice text.
[0144] In the embodiments of the present application, the vehicle state is updated through the execution result, and the voice text input by the user is updated according to the updated vehicle state, so as to improve the stability and accuracy of recognizing the user intention subsequently. And the double - helix update can prevent repeated functions from being called multiple times, thereby ensuring the accuracy and stability of the completion of the user intention.
[0145] S408. Based on the updated voice text and the updated vehicle state, determine whether all user intentions corresponding to the voice text input by the user are completed.
[0146] Specifically, when the updated speech text is empty, or based on the updated speech text and the updated vehicle state, the user intention representing the next action cannot be recognized, it is considered that all user intentions corresponding to the user input speech text are completed.
[0147] In a possible implementation, through an inference model, based on the updated speech text and the updated vehicle state, it is determined whether all user intentions corresponding to the user input speech text are completed. That is, the inference model determines whether all user intentions corresponding to the user input speech text are completed based on the updated vehicle state and the updated speech text output by the received instruction following model.
[0148] When all user intentions corresponding to the user input speech text are not completed, S402 - S408 are repeated until all user intentions corresponding to the user input speech text are completed.
[0149] When all user intentions corresponding to the user input speech text are completed, S409 is performed.
[0150] S409: Generate and output a response message based on the execution results returned by all received in - vehicle cockpit function modules.
[0151] Among them, the response message is used to prompt that all user intentions corresponding to the user input speech text have been completed, as well as the execution results of each in - vehicle cockpit function module.
[0152] Exemplarily, assuming that all user intentions are: open the window, lower the air - conditioner temperature to 20°C, and the window control function module and the air - conditioner function module both correctly execute the corresponding operations, then the response message can be: "The vehicle has been opened and the air - conditioner temperature has been adjusted to 20°C. Do you have any other requests?"
[0153] In a possible implementation, through an inference model, based on the execution results returned by all received in - vehicle cockpit function modules, a response message is generated and output.
[0154] Furthermore, in the embodiments of the present application, by combining an inference model and an instruction following model, different models handle partial tasks. The inference model not only recognizes user intentions, but also updates the vehicle state and determines whether the user intentions are completed; the instruction following model not only generates structured instructions, but also updates the user input speech text. Further effectively overcomes the limitations of a single model in multi - task processing, and provides a clearer and more reliable basis for the execution of more complex tasks.
[0155] A voice control method provided by an embodiment of the present application includes: when receiving a voice text input by a user, obtaining the vehicle state; based on the voice text input by the user and the vehicle state, identifying the user intention; when the user intention represents a next action, converting the next action represented by the user intention into a structured instruction; sending the structured instruction to a corresponding in-vehicle cockpit function module, so that the in-vehicle cockpit function module responds to the structured instruction and executes the operation corresponding to the structured instruction; receiving an execution result returned by the in-vehicle cockpit function module; based on the execution result, updating the vehicle state to obtain an updated vehicle state; modifying the voice text input by the user according to the updated vehicle state to obtain an updated voice text; based on the updated voice text and the updated vehicle state, determining whether all user intentions corresponding to the voice text input by the user are completed; when all user intentions corresponding to the voice text input by the user are completed, repeating the above process; when all user intentions corresponding to the voice text input by the user are completed, generating and outputting a response message based on all received execution results returned by the in-vehicle cockpit function module. The embodiment of the present application comprehensively considers the voice text and the vehicle state, identifies the user intention, realizes a deep understanding of the user intention, improves the accuracy of user intention recognition and ensures the stability of the recognition result, thereby improving the accuracy of voice control.
[0156] Further, by converting the user intention into a structured instruction and sending it to the corresponding in-vehicle cockpit function module to complete the corresponding operation, the generalization of the voice input for triggering the operation is improved, and the accuracy of the triggered operation / function is ensured, thereby improving the stability of voice control.
[0157] Further, the embodiment of the present application updates in a double helix manner, that is, updating the vehicle state through the execution result and updating the voice text input by the user according to the updated vehicle state, so as to improve the stability and accuracy of identifying the user intention in the future; and it can prevent repeated functions from being called multiple times, thereby ensuring the accuracy and stability of the completion of the user intention.
[0158] Further, when all user intentions corresponding to the voice text input by the user are completed, a response message is generated and output based on all received execution results returned by the in-vehicle cockpit function module. It can effectively prompt that the current in-vehicle cockpit function module has completed all user intentions, and the response message can also carry the execution results of each in-vehicle cockpit function module to prompt the user of the execution results of each in-vehicle cockpit function module.
[0159] Embodiment Three:
[0160] The following combines Figure 5 , and details a voice control device provided by an embodiment of the present application.
[0161] AsFigure 5 As shown in Figure 5 , a language control device provided by an embodiment of the present application includes the following modules:
[0162] An acquisition module 501, configured to acquire the vehicle state when receiving a voice text input by a user;
[0163] An identification module 502, configured to identify the user intention based on the voice text and the vehicle state;
[0164] A structuring module 503, configured to convert the next action represented by the user intention into a structured instruction when the user intention represents a next action; wherein, the structured instruction is used to trigger the in-vehicle cockpit function module to execute a corresponding operation;
[0165] A sending module 504, configured to send the structured instruction to the corresponding in-vehicle cockpit function module, so that the in-vehicle cockpit function module responds to the structured instruction and executes the operation corresponding to the structured instruction.
[0166] In a possible implementation manner, the identification module 502 is specifically configured to input the voice text and the vehicle state into an inference model, so that the inference model outputs the user intention based on the voice text and the vehicle state.
[0167] In a possible implementation manner, the device further includes: a sample storage module, a reward calculation module, and a reinforcement learning module.
[0168] The sample storage module is configured to store the voice text and the vehicle state input into the inference model, and the user intention output by the inference model as a set of training samples.
[0169] The reward calculation module is configured to evaluate the training samples using a preset reward function to obtain a corresponding reward value.
[0170] The reinforcement learning module is configured to optimize the inference model based on the corresponding reward value using a reinforcement learning algorithm.
[0171] In a possible implementation manner, the structuring module 503 is specifically configured to input the user intention into an instruction following model, so that the instruction following model converts the next action represented by the user intention into a structured instruction in response to the next action represented by the user intention. That is, through the instruction following model, when the user intention represents a next action, the next action represented by the user intention is converted into a structured instruction.
[0172] In a possible implementation, the structured module 503 is specifically configured to, when the user intention represents the next action, obtain an instruction text of the next action in a preset format based on the user intention; wherein, the instruction text of the next action is used to represent the in-vehicle cockpit function module corresponding to the next action and the operations performed by the corresponding in-vehicle cockpit function module; convert the instruction text of the next action into a corresponding structured instruction; wherein, the structured instruction is a function that can be recognized and executed by the corresponding in-vehicle cockpit function module.
[0173] In a possible implementation, the device further includes: a parameter check module and a prompt output module.
[0174] The parameter check module is used to check whether the parameters in the structured instruction are complete; the prompt output module is used to, when the parameters in the structured instruction are incomplete, generate and output a prompt message based on the structured instruction to prompt that the parameters in the structured instruction are incomplete.
[0175] The sending module 504 is specifically configured to, when the parameters in the structured instruction are complete, send the structured instruction to the corresponding in-vehicle cockpit function module.
[0176] In a possible implementation, the device further includes: an execution result receiving module, a vehicle state update module, a voice text update module, and an intention completion judgment module.
[0177] The execution result receiving module is used to receive the execution result returned by the in-vehicle cockpit function module;
[0178] The vehicle state update module is used to update the vehicle state based on the execution result to obtain the updated vehicle state;
[0179] The voice text update module is used to modify the voice text input by the user according to the updated vehicle state to obtain the updated voice text;
[0180] The intention completion judgment module is used to judge whether all the user intentions corresponding to the voice text input by the user are completed based on the updated voice text and the updated vehicle state;
[0181] The recognition module 502 is specifically configured to, when not all the user intentions corresponding to the voice text input by the user are completed, re-recognize the user intention based on the updated voice text and the updated vehicle state.
[0182] In a possible implementation, the vehicle status update module is specifically configured to update the vehicle status based on the execution result through an inference model to obtain the updated vehicle status; the voice text update module is specifically configured to modify the voice text input by the user according to the updated vehicle status through an instruction following model to obtain the updated voice text; the intention completion judgment module is specifically configured to judge whether all user intentions corresponding to the voice text input by the user are completed based on the updated voice text and the updated vehicle status through an inference model.
[0183] In a possible implementation, the device further includes: a response information output module, configured to generate and output response information based on the execution results returned by all received in-vehicle cockpit function modules when all user intentions corresponding to the voice text input by the user are completed.
[0184] In a possible implementation, the device further includes: a reply information output module, configured to generate a corresponding reply information based on the voice text input by the user when the user intention indicates that the voice input by the user is for chatting; wherein, the reply information is a reply to the voice text input by the user.
[0185] A voice control device provided by an embodiment of the present application includes: an acquisition module 501, configured to acquire the vehicle status when receiving the voice text input by the user; an identification module 502, configured to identify the user intention based on the voice text and the vehicle status; a structuring module 503, configured to convert the next action represented by the user intention into a structured instruction when the user intention represents the next action; a sending module 504, configured to send the structured instruction to the corresponding in-vehicle cockpit function module, so that the in-vehicle cockpit function module responds to the structured instruction and executes the operation corresponding to the structured instruction. The embodiment of the present application comprehensively considers the voice text and the vehicle status, identifies the user intention, and converts the user intention into a structured instruction, thereby realizing a deep understanding of the user intention, improving the accuracy of recognition and ensuring the stability of the recognition result.
[0186] Further, the identification of the user intention by comprehensively considering the voice text and the vehicle status is implemented through an inference model. During the use of the inference model, the voice text, the vehicle status, and the user intention are used as training samples, and the reward value of the training samples is calculated based on a preset reward function. Based on the reward value, the inference model is continuously optimized using a reinforcement learning algorithm. This improves the ability of the inference model to understand and respond to the input voice text and vehicle status, enabling it to more accurately output the user intention. And as time goes by and more data accumulates, the performance of the inference model will be gradually improved, thereby achieving more efficient and accurate operations.
[0187] It should be noted that the embodiments in this specification are all described in a progressive manner. For the identical or similar parts among the embodiments, reference can be made to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the method and apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the relevant content. The method and apparatus embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components referred to as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0188] As described above, this is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A voice control method, characterized in that, Including: When receiving the speech text input by the user, obtain the vehicle state; Based on the speech text and the vehicle state, identify the user intention; When the user intention represents the next action, convert the next action represented by the user intention into a structured instruction; wherein, the structured instruction is used to trigger the in-vehicle cockpit function module to execute the corresponding operation; Send the structured instruction to the corresponding in-vehicle cockpit function module, so that the in-vehicle cockpit function module responds to the structured instruction and executes the operation corresponding to the structured instruction.
2. The method according to claim 1, wherein The identifying the user intention based on the speech text and the vehicle state includes: Input the speech text and the vehicle state into the inference model, so that the inference model outputs the user intention based on the speech text and the vehicle state.
3. The method according to claim 2, wherein The method further includes: Store the speech text and vehicle state input into the inference model, and the user intention output by the inference model as a set of training samples; Use a preset reward function to evaluate the training samples and obtain the corresponding reward value; Based on the corresponding reward value, use a reinforcement learning algorithm to optimize the inference model.
4. The method according to claim 3, wherein The preset reward function includes: a first reward, a second reward, a third reward, and a fourth reward; wherein, the first reward is used to evaluate whether the inference model accurately outputs the user intention, the second reward is used to evaluate whether the user intention output by the inference model considers the safety state of the vehicle itself, the third state is used to evaluate the efficiency of the inference model to output the user intention based on the speech text and the vehicle state, and the fourth reward is used to evaluate whether the user intention output by the inference model considers the user's habits.
5. The method according to claim 2, wherein The converting the next action represented by the user intention into a structured instruction when the user intention represents the next action includes: Input the user intention into the instruction following model, so that the instruction following model converts the next action represented by the user intention into a structured instruction in response to the user intention representing the next action.
6. The method according to claim 1, wherein The converting the next action represented by the user intention into a structured instruction when the user intention represents the next action includes: When the user intention represents the next action, based on the user intention, obtain the instruction text of the next action in a preset format; wherein, the instruction text of the next action is used to represent the in-vehicle cockpit function module corresponding to the next action and the operation executed by the corresponding in-vehicle cockpit function module; Convert the instruction text of the next action into the corresponding structured instruction; wherein, the structured instruction is a function that can be recognized and executed by the corresponding in-vehicle cockpit function module.
7. The method according to claim 6, characterized in that, Before sending the structured instruction to the corresponding in-vehicle cockpit function module, the method further includes: Check whether the parameters in the structured instruction are complete; When the parameters in the structured instruction are incomplete, generate and output a prompt message based on the structured instruction to prompt that the parameters in the structured instruction are incomplete; Sending the structured instruction to the corresponding in-vehicle cockpit function module includes: When the parameters in the structured instruction are complete, sending the structured instruction to the corresponding in-vehicle cockpit function module.
8. The method according to claim 5, characterized in that The method further includes: Receiving the execution result returned by the in-vehicle cockpit function module; Based on the execution result, updating the vehicle state to obtain an updated vehicle state; According to the updated vehicle state, modifying the voice text input by the user to obtain an updated voice text; Based on the updated voice text and the updated vehicle state, determining whether all user intents corresponding to the voice text input by the user are completed; When not all user intents corresponding to the voice text input by the user are completed, the identifying the user intent based on the voice text and the vehicle state includes: Based on the updated voice text and the updated vehicle state, re-identifying the user intent.
9. The method according to claim 8, wherein The updating the vehicle state based on the execution result to obtain an updated vehicle state includes: Through the inference model, updating the vehicle state based on the execution result to obtain an updated vehicle state; The modifying the voice text input by the user according to the updated vehicle state to obtain an updated voice text includes: Through the instruction following model, modifying the voice text input by the user according to the updated vehicle state to obtain an updated voice text; The determining whether all user intents corresponding to the voice text input by the user are completed based on the updated voice text and the updated vehicle state includes: Through the inference model, determining whether all user intents corresponding to the voice text input by the user are completed based on the updated voice text and the updated vehicle state.
10. A voice control device, characterized in that, It includes: An acquisition module, configured to acquire the vehicle state when receiving the voice text input by the user; An identification module, configured to identify the user intent based on the voice text and the vehicle state; A structuring module, configured to convert the next action represented by the user intent into a structured instruction when the user intent represents a next action; wherein, the structured instruction is used to trigger the corresponding in-vehicle cockpit function module to execute the corresponding operation; A sending module, configured to send the structured instruction to the corresponding in-vehicle cockpit function module, so that the in-vehicle cockpit function module responds to the structured instruction and executes the operation corresponding to the structured instruction.
Citation Information
Cited By
Voice interaction control method and system for electric scooter
CN120998196A
Cabin interaction method and vehicle
CN122526428A