Voice processing method and device, medium and equipment
By using two voice recognition models with parallel processing in the voice interaction system, the problem of long voice reply delay in traditional systems is solved, and the user experience is improved.
Patent Information
- Application Number
- CN202510275968.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-05-27
AI Technical Summary
After receiving the user's voice control command, the traditional voice interaction system outputs voice reply delay is long, which seriously affects the user experience.
By using two different voice recognition models to process the user's voice control instructions in parallel, target response information is generated and target operations are performed, and the reply delay is reduced.
It effectively reduces the delay in replying to user voice commands and improves the voice interaction experience of the on-board system.
Smart Images

Figure CN120048261A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of intelligent automobile technology, and in particular to a voice processing method, device, medium and equipment. Background Art
[0002] With the continuous development of vehicle technology, the vehicle's voice interaction system has been widely used in daily life. Usually, users can trigger the vehicle to play music, navigate, open windows, etc. through the voice interaction system.
[0003] However, in the process of implementing the present disclosure, the inventors found that there are at least the following problems in the prior art: since the traditional voice interaction system is a voice interaction link pipeline, after receiving the user's voice control command, the voice response output based on the voice interaction system is delayed for a long time, which seriously affects the user experience.
[0004] Thus, there is an urgent need for a voice processing method to solve the problem of long delay in replying in traditional voice interaction systems. Summary of the invention
[0005] In order to solve the above technical problems, the present disclosure provides a voice processing method to solve the problem of long delay in replying in traditional voice interaction systems.
[0006] On the one hand, the present disclosure provides a voice processing method, which includes: determining to obtain a user's voice control instruction; processing the voice control instruction based on a first voice recognition model to generate target response information corresponding to the voice control instruction; processing the voice control instruction based on a second voice recognition model to obtain a target operation corresponding to the voice control instruction; executing the target operation corresponding to the voice control instruction; wherein the target response information is generated and output before executing the target operation.
[0007] On the other hand, the present disclosure provides a voice processing device, including: a determination module, used to obtain a user's voice control instructions; a processing module, used to process the voice control instructions based on a first voice recognition model, and generate target response information corresponding to the voice control instructions; the processing module is also used to process the voice control instructions based on a second voice recognition model to obtain a target operation corresponding to the voice control instructions; an execution module, used to execute the target operation corresponding to the voice control instructions; wherein the target response information is generated and output before executing the target operation.
[0008] On the other hand, the present disclosure proposes a computer program product, which, when an instruction processor in the computer program product executes, performs the speech processing method described in the embodiment of the first aspect of the present disclosure.
[0009] On the other hand, the present disclosure proposes an electronic device, comprising: a processor; a memory for storing executable instructions of the processor; the processor is used to read the executable instructions from the memory and execute the instructions to implement the speech processing method described in the first aspect above.
[0010] The disclosed embodiment provides a voice processing method, which can process the voice control instruction based on a first voice recognition model after collecting the user's voice control instruction, generate target response information corresponding to the voice control instruction, and process the voice control instruction based on a second voice recognition model to obtain the target operation corresponding to the voice control instruction, and execute the target operation. That is, the disclosed embodiment is based on two voice recognition models with different voice control instruction processing speeds to process the collected voice control instructions in parallel at the same time. Because the voice recognition model with a faster voice control instruction processing speed can generate and output the target response information earlier before executing the target operation corresponding to the voice control instruction, the delay in replying to the user's voice instruction is reduced, thereby improving the voice interaction experience of the vehicle-mounted system. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is a structural diagram of an in-vehicle speech processing model provided by an exemplary embodiment of the present disclosure.
[0012] Figure 2 It is a flowchart of a speech processing method provided by an exemplary embodiment of the present disclosure.
[0013] Figure 3 It is a flowchart of a speech processing method provided by another exemplary embodiment of the present disclosure.
[0014] Figure 4 It is a flowchart of a speech processing method provided by another exemplary embodiment of the present disclosure.
[0015] Figure 5 It is a structural diagram of a speech processing device provided by yet another exemplary embodiment of the present disclosure.
[0016] Figure 6 is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0017] To explain the present disclosure, example embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all of the embodiments. It should be understood that the present disclosure is not limited to the example embodiments.
[0018] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure unless specifically stated otherwise.
[0019] Application Overview
[0020] In the field of intelligent vehicle technology, drivers or passengers can interact with vehicles through voice interaction systems to control vehicles to perform certain operations. For example, the driver sends a voice control command "please open the window". After the vehicle receives the voice control command, it processes the voice control command through the voice interaction system to execute the "open the window" operation corresponding to the voice control command.
[0021] The traditional voice interaction system is a pipeline voice interaction link, which may include a voice recognition module, a natural language understanding module, a natural language response module, and a voice synthesis module. The voice recognition module may be used to convert the received voice control command into text information, the natural language understanding module may perform semantic analysis on the text information to determine the intention of the user's voice command, the natural language response module may generate text response information based on the user's intention information, and the voice synthesis module may be used to convert the text response information into voice response information; after that, the voice response information may be output through the vehicle's speakers.
[0022] Since the interaction of large text models has achieved disruptive effects in many fields, large text models can be connected in series in the voice interaction link to replace the natural language understanding module and the natural language generation module, thereby greatly improving the understanding ability of the voice interaction system and providing more natural responses. However, due to the slow running speed and long delay of large text models, directly introducing large text models in the voice interaction link will increase the delay of the voice interaction link, resulting in a longer delay in outputting voice responses, seriously affecting the user experience. In this way, after receiving the user's voice control command, how to ensure that the voice control command is responded to with a short delay becomes an urgent problem to be solved.
[0023] Based on the above problems, the embodiment of the present disclosure provides a voice processing method. After the user's voice control command is collected through the audio sensor of the vehicle, the voice control command can be processed based on the first voice recognition model to generate the target response information corresponding to the voice control command, and the voice control command can be processed based on the second voice recognition model to obtain the target operation corresponding to the voice control command, and the target operation is executed. That is, in the embodiment of the present disclosure, two voice recognition models with different processing speeds for voice control commands are used to process the same collected voice control command in parallel at the same time. Because the target response information can be generated and output earlier by the voice recognition model with a faster processing speed for voice control commands before executing the target operation corresponding to the voice control command, the delay in responding to the user's voice command is reduced, thereby improving the voice interaction experience of the vehicle-mounted system.
[0024] Exemplary Systems
[0025] Figure 1 It is a structural diagram of an in-vehicle speech processing model provided by an exemplary embodiment of the present disclosure.
[0026] For example, Figure 1 As shown, the above-mentioned in-vehicle speech processing model includes a speech acquisition module 110, a first speech recognition module 120, a second speech recognition module 130 and an execution module 140; wherein the second speech recognition module 130 and the execution module 140 are connected.
[0027] The voice collection module 110 can be used to obtain the user's voice control instructions collected by the vehicle's audio sensor, and transmit the voice control instructions to the first voice recognition module 120 and the second voice recognition module 130 simultaneously.
[0028] The first voice recognition module 120 may be used to process the voice control instruction after receiving the voice control instruction and generate target response information corresponding to the voice control instruction.
[0029] In some examples, the first speech recognition module 120 includes a first speech recognition model, which can process the voice control instruction and generate target response information corresponding to the voice control instruction.
[0030] The second voice recognition module 130 may be used to process the voice control instruction after receiving the voice control instruction to obtain a target operation corresponding to the voice control instruction.
[0031] In some examples, the second speech recognition module 130 includes a second speech recognition model, which includes multiple sub-models. The voice control instructions need to be processed serially based on the multiple sub-models to obtain the target operation corresponding to the voice control instructions.
[0032] In the embodiment of the present disclosure, target response information is generated and output before the target operation is performed.
[0033] The execution module 140 may be used to execute the target operation corresponding to the voice control instruction.
[0034] The disclosed embodiments provide a voice processing method, which can process the voice control instructions simultaneously based on two parallel voice recognition models when voice control instructions input by the user are collected. Since the time required for the first voice recognition model of the two voice recognition models to process the voice control instructions is less than the time required for the second voice recognition model to process the voice control instructions, the first voice recognition model with a faster processing speed for the voice control instructions can generate and output target response information, thereby reducing the delay in responding to the voice control instructions and improving the user's voice interaction experience.
[0035] Exemplary Methods
[0036] Figure 2 It is a flowchart of a speech processing method provided by an exemplary embodiment of the present disclosure.
[0037] Exemplarily, the above method may be executed by a vehicle, or by a chip in a vehicle.
[0038] like Figure 2 As shown, the above method may include the following steps:
[0039] Step 201, obtaining a user's voice control instruction.
[0040] In some embodiments, the voice control instructions input by the user in the vehicle can be collected in real time through the audio sensor, that is, the voice control instructions are the voice instructions collected by the audio sensor in real time; or, the voice control instructions input by the user in the vehicle can be collected periodically through the audio sensor, that is, the voice control instructions are the voice instructions input by the user collected by the audio sensor in a collection cycle closest to the current moment.
[0041] In some examples, the audio sensor may be a vehicle-mounted microphone or other sensor that can collect voice commands.
[0042] In other embodiments, after acquiring the voice control instruction collected by the audio sensor of the vehicle, the voice control instruction may be preprocessed, and then the preprocessed voice control instruction may be subsequently processed. For example, the preprocessing method may include performing noise reduction processing on the voice control instruction.
[0043] Step 202: Process the voice control instruction based on the first speech recognition model to generate target response information corresponding to the voice control instruction.
[0044] The above-mentioned first speech recognition model can be an end-to-end speech large model, which can support end-to-end processing from speech to speech, that is, the first speech recognition model does not need to convert the speech information into text information before processing, but can directly recognize and process the speech information and obtain the corresponding speech reply, thereby greatly simplifying the processing flow of the speech information and improving the processing efficiency, that is, the first speech recognition model has a fast processing speed and low latency. In this way, when processing the voice control instructions received from the user input based on the first speech recognition model, the voice control instructions can be directly processed, instead of first converting the voice control instructions into corresponding text information, and then processing the converted text information to generate the target response information corresponding to the voice control instructions. For example, the first speech recognition model can be GPT-4o.
[0045] In some embodiments, the target response information is information for replying to the voice control instruction. When the first voice recognition model recognizes that the voice control instruction belongs to the first type of scenario, the output target response information can be used to indicate the execution progress of the voice control instruction, that is, when the target response information is output, the user can determine that the vehicle has received the voice control instruction input by the user and is processing the voice control instruction based on the target response information; when the first voice recognition model recognizes that the voice control instruction belongs to the second type of scenario, the output target response information can be used to indicate the specific content corresponding to the voice control instruction.
[0046] In some examples, the first category of scenes may include vehicle control scenes, navigation scenes, or entertainment scenes, etc.; the second category of scenes may include question-and-answer scenes or chat scenes.
[0047] For example, when the voice control command is "open the car window", the voice control command belongs to the first type of scenario; for another example, when the voice control command is "please tell a joke", the voice control command belongs to the second type of scenario.
[0048] In some examples, the target response information may include voice response information and display response information, and the embodiment of the present disclosure does not limit the target response message. When the target response message also includes a display response message, the display response information may be a text response message, a picture response information, or a video response message.
[0049] Exemplarily, the first speech recognition model is GPT-4o for exemplary description. When a voice control command of "open the window" is received, the voice control command can be input into GPT-4o, so that the voice control command can be processed based on GPT-4o to obtain the target response information of "helping you operate", so that the user knows based on the target response information that the vehicle has received the voice control command input by the user and is processing the voice control command; for another example, when the received voice control command is "please tell a joke", the voice control command can be input into GPT-4o, so that the voice control command can be processed based on GPT-4o to obtain the specific content of the joke.
[0050] In some examples, when the first speech recognition model is GPT-4o, the first speech recognition model processes the voice control instruction, and the delay time for generating target response information corresponding to the voice control instruction is 200ms.
[0051] Step 203: Process the voice control instruction based on the second speech recognition model to obtain a target operation corresponding to the voice control instruction.
[0052] Wherein, the target response information is generated and output before executing the target operation.
[0053] The above-mentioned second speech recognition model can be a traditional speech processing model, which usually includes multiple sub-models. When the traditional speech processing model processes the voice control instructions input by the user, one of the sub-models is required to convert the voice control instructions into text information, and then the text information is processed in turn by other sub-models. This processing method has an intermediate conversion step, which will cause conversion errors and delays. Therefore, the traditional speech processing model has low efficiency and poor accuracy in processing speech. In this way, when the received voice control instructions are processed based on the second speech recognition module, the voice control instructions are serially processed based on multiple sub-models to obtain the target operation corresponding to the voice control instructions.
[0054] In some embodiments, the target response information is associated with the target operation. The target response information can be used to indicate the operation progress of the target operation, so different target response information has different corresponding target operations.
[0055] Exemplarily, when the target operation is to close the window, the corresponding response information is "the window is closing", and the target response information is used to indicate the progress of the operation of closing the window; when the target operation is to open the window, the corresponding target response information is "the window is opening", and the target response information is used to indicate the progress of the operation of opening the window.
[0056] Step 204: Execute the target operation corresponding to the voice control instruction.
[0057] In some embodiments, executing a target operation corresponding to a voice control instruction may specifically include: first determining a target control object based on the voice control instruction, and then executing a corresponding target operation on the target control object through a control unit inside the vehicle.
[0058] For example, the voice control command is "please open the window" and the corresponding target operation is to open the window. The vehicle can determine the target control object as the window based on the voice control command "please open the window", and then control the window to move through the window controller to achieve the purpose of opening the window.
[0059] The disclosed embodiment provides a voice processing method, which can process the voice control command based on a first voice recognition model to generate target response information corresponding to the voice control command after the voice control command input by the user is collected through the audio sensor of the vehicle, and process the voice control command based on a second voice recognition model to obtain the target operation corresponding to the voice control command, and execute the target operation. That is, in the disclosed embodiment, two voice recognition models with different processing speeds for voice control commands are used to process the same collected voice control command in parallel at the same time. Because the target response information can be generated and output earlier by the voice recognition model with a faster processing speed for voice control commands before executing the target operation corresponding to the voice control command, the delay in responding to the user's voice command is reduced, thereby improving the voice interaction experience of the vehicle-mounted system.
[0060] like Figure 3 As shown in the above Figure 2 Based on the illustrated embodiment, step 202 may include the following steps:
[0061] Step 2021: Process the voice control instruction based on the first speech recognition model to obtain target response information corresponding to the voice control instruction.
[0062] For the relevant explanation of the first speech recognition model, reference may be made to the description in the above embodiment, which is not limited in the embodiments of the present disclosure.
[0063] In some examples, the target response information may include voice response information and display response information, and the embodiment of the present disclosure does not limit the target response message. When the target response message also includes a display response message, the display response information may be a text response message, a picture response information, or a video response message.
[0064] Step 2022, determine the type of target response information.
[0065] In some embodiments, the type of the target response information may include at least one of the following: text type, voice type, video type, etc.
[0066] In some embodiments, the type of the target response information may be determined based on the specific content of the target response information.
[0067] Exemplarily, if the target response information only includes text content, the type of the target response information is a text type; or, if the target response information only includes an image, the type of the target response information is an image type; for another example, if the target response information includes an image and a text description, the type of the target response information includes both a text type and an image type.
[0068] Step 2023: output the target response information based on the type of the target response information.
[0069] In some embodiments, the output mode of the target response information is related to the type of the target response information, that is, the output mode of the target response information is determined based on the type of the target response information. Different types of target response information correspond to different types of output modes.
[0070] Exemplarily, when the target response information is of text type, the target response information of text type can be displayed on the display interface of the vehicle's central control screen; when the target response information is of voice type, the target response information of voice type can be played through the vehicle's built-in speakers.
[0071] The voice processing method provided by the embodiment of the present disclosure, on the premise of directly recognizing and processing the voice control instructions based on the first voice recognition model to obtain the target response information, can output different types of target response information based on the type of the target response information, thereby enriching the interaction mode and content and improving the interaction experience.
[0072] In some embodiments, the above step 2021 may include the following steps:
[0073] Step 2021a, determining the current state information of the target control object corresponding to the voice control instruction.
[0074] In some embodiments, the target control object corresponding to the voice control instruction can be determined by identifying keywords in the voice control instruction.
[0075] For example, when the voice control command is "open the car window", it can be recognized that the keyword in the voice control command is the car window, so the target control object corresponding to the voice control command "open the car window" is the car window.
[0076] In some embodiments, various states of different vehicle components can be detected by various sensors of the vehicle, and the state information of different vehicle components can be recorded in a state table. Various sensors of the vehicle monitor the states of different vehicle components in real time, and update the state table when the states of various vehicle components change, that is, the state table records the current state information of various vehicle components in real time, so the current state information of the target control object corresponding to the voice control instruction can be queried from the state table. The above state table can be in the form of a large model text prompt.
[0077] For example, the target control object is a car window. When it is determined that the target control object corresponding to the voice control instruction is a car window, the state information corresponding to the car window can be queried from the real-time state table, that is, the current state information of the car window.
[0078] Step 2021b: Process the voice control instruction and the current state information based on the first speech recognition model to obtain target response information corresponding to the voice control instruction.
[0079] In some embodiments, when a vehicle receives a voice control command input by a user, when the current state information of the target control object corresponding to the voice control command is different, the target response information corresponding to the obtained voice control command is also different.
[0080] For example, the voice control command "open the window" is used as an example for illustration. When the user inputs the voice control command "open the window", the target control object corresponding to the voice control command "open the window" is the window; when the current state information corresponding to the window is queried through the state table that the window is not open, the voice control command and the current window state information are processed based on the first voice recognition model, and the corresponding target response information is "the window is open"; or, when the current state information corresponding to the window is queried through the state table that the window is open, the voice control command and the current window state information are processed based on the first voice recognition model, and the corresponding target response information is "the window is already in the open state".
[0081] The voice processing method provided by the embodiment of the present disclosure can obtain target response information by combining the voice control instruction and the current state information of the target control object corresponding to the voice control instruction when receiving the voice control instruction. Therefore, the obtained target response information is more accurate, thereby improving the accuracy of voice interaction.
[0082] like Figure 4 As shown in the above Figure 2 Based on the illustrated embodiment, step 203 may include the following steps:
[0083] Step 2031: Convert the voice control instruction into text information based on the second speech recognition sub-model in the second speech recognition model.
[0084] In some examples, the second speech recognition submodel converts speech information into text information by recognizing and understanding speech information. Thus, when receiving a speech control instruction input by a user, the speech control instruction can be converted into corresponding text information based on the second speech recognition submodel. For example, the second speech recognition submodel can be an automatic speech recognition model (Automatic Speech Recognition, ASR).
[0085] For example, the second speech recognition submodel is ASR. When the user inputs a voice control command of "how is the weather today", the voice control command is received by the vehicle microphone, and the voice control command is converted into the corresponding text information "how is the weather today" based on the ASR model.
[0086] In some examples, when the second speech recognition sub-model is ASR, the delay time for the second speech recognition sub-model to convert the voice control instruction into text information is 600ms.
[0087] Step 2032: Perform semantic analysis on the text information corresponding to the voice control instruction based on the natural language understanding sub-model in the second speech recognition model to obtain the target control type and user intention information.
[0088] In some examples, the natural language understanding sub-model extracts key information from the text information by analyzing the grammar and semantics of the text, and determines the intent type corresponding to the voice control instruction based on the key information, uses the intent type as the target control type, and determines the user intent information based on the key information. In this way, after receiving the text information generated by the second speech recognition sub-model, the natural language understanding sub-model can perform semantic analysis on the text information, extract key information from the text information, and obtain the target control type and user intent information. For example, the natural language understanding sub-model can be a semantic analysis model (Natural Language Understanding, NLU).
[0089] In some examples, when the natural language understanding sub-model is NLU, the delay time for the natural language understanding sub-model to perform semantic analysis on the text information corresponding to the voice control instruction is 100ms.
[0090] In some examples, the target control type may be any of the following: vehicle control type, entertainment type, navigation type, information query type, question-and-answer chat type, etc. The user intention information may be key information identified from text information corresponding to the voice control instruction.
[0091] For example, the natural language understanding sub-model is used as NLU for example. When the text information corresponding to the voice control command is "How is the weather today?", the text information "How is the weather today?" can be voice analyzed based on NLU, and key information including "today", "weather" and "how" can be extracted, and the target control type can be determined as the weather query type in the information query type based on the above key information, and the user intention information can be determined as today's weather conditions based on the key information.
[0092] Step 2033: Determine the target operation corresponding to the voice control instruction based on the target control type and the user intention information.
[0093] In some embodiments, when the target control type belongs to a control type not supported by the second speech recognition model, no subsequent response will be made. However, when the target control type belongs to a control type supported by the second speech recognition model, the corresponding target operation can be determined based on the user intent information.
[0094] Exemplarily, when the target control type is a certain type of chat question and answer type, since the second voice recognition model does not support this type of chat question and answer, no subsequent response will be made; when the target control type is a vehicle control type (such as opening the window), since the second voice recognition model supports this vehicle control type, it can perform the corresponding window opening operation based on the user's intention to open the window.
[0095] The speech processing method provided by the embodiment of the present disclosure can convert the voice control instruction into text information based on the second speech recognition submodel in the second speech recognition model, and perform semantic analysis on the text information based on the natural language understanding submodel in the second speech recognition model to obtain the target control type and user intent information, and then determine the target operation corresponding to the voice control instruction based on the target control type and user intent information. Therefore, the target operation corresponding to the voice control instruction can be executed more accurately, thereby enabling the user to interact with the vehicle through voice commands to accurately control various functions of the vehicle.
[0096] In some embodiments, the above step 2033 may include the following steps:
[0097] Step 2033a, determining the inclusion relationship between the target control type and the preset control type.
[0098] Among them, the preset control type is the control type supported by the second speech recognition model.
[0099] In some examples, the preset control type may include at least one of the following: vehicle control type, entertainment type, navigation type, information query type, etc. The target control type may be any of the following: vehicle control type, entertainment type, navigation type, information query type, question-and-answer chat type, etc.
[0100] For example, the preset control type includes at least one of the following: vehicle control type, entertainment type, navigation type and information query type. When the target control type is a chat question and answer type, the target control type is not included in the preset control type; when the target control type is a vehicle control type, the target control type is included in the preset control type.
[0101] Step 2033b, in response to the inclusion relationship that the preset control type includes the target control type, determining the target operation corresponding to the voice control instruction according to the user intention information.
[0102] In some embodiments, when the preset control type includes a target control type, the user intention information can be determined as the target operation corresponding to the voice control instruction.
[0103] For example, the target control type is the vehicle control type, and the user intention information is opening the window. When it is determined that the preset control type includes the vehicle control type, opening the window can be determined as the target operation.
[0104] In some embodiments, determining the target operation corresponding to the voice control instruction according to the user intention information in step 2033b above may specifically include the following steps:
[0105] (a) Determine the current state information of the target control object corresponding to the voice control instruction.
[0106] In some embodiments, the target control object corresponding to the voice control instruction can be determined by identifying keywords in the voice control instruction.
[0107] Exemplarily, when the voice control command is "turn on the air conditioner", the keyword in the voice control command is recognized as air conditioner, so the target control object corresponding to the voice control command "turn on the air conditioner" is the air conditioner.
[0108] In some embodiments, various states of different vehicle components can be detected by various sensors of the vehicle, and the state information of different vehicle components can be recorded in a state table. Various sensors of the vehicle monitor the states of different vehicle components in real time, and update the state table when the states of various vehicle components change, that is, the state table records the current state information of various vehicle components in real time, so the current state information of the target control object corresponding to the voice control instruction can be queried from the state table. The above state table can be in the form of a large model text prompt.
[0109] For example, the target control object is a car window. When it is determined that the target control object corresponding to the voice control instruction is a car window, the state information corresponding to the car window can be queried from the real-time state table, that is, the current state information of the car window.
[0110] (b) Determine the target operation based on the current state information and user intention information.
[0111] When the current state information of the target control object matches the user intention information, the target operation can be determined based on the user intention information; conversely, when the current state information of the target control object does not match the user intention information, the target operation cannot be determined, that is, no operation needs to be performed at this time.
[0112] In some embodiments, the step of determining the target operation according to the current state information of the target control object and the user intention information may specifically include the following steps:
[0113] Determine the matching relationship between the current state information of the target control object and the user intention information;
[0114] In response to the matching relationship being that the current state information of the target control object matches the user intention information, a target operation is determined according to the user intention information.
[0115] In some examples, the matching relationship between the current state information of the target control object and the user intention information includes two types: (1) the current state information of the target control object matches the user intention information; (2) the current state information of the target control object does not match the user intention information.
[0116] For example, the target control object is an air conditioner. When the user intention information is to turn on the air conditioner, when the current state information of the air conditioner is that the air conditioner is in the off state, it means that the current state information matches the user intention information, and at this time, it can be determined according to the user intention information that the target operation is to turn on the air conditioner; on the contrary, when the current state information of the air conditioner is that the air conditioner is in the on state, it means that the current state information does not match the user intention information, and no operation is performed at this time.
[0117] The voice processing method provided by the embodiment of the present disclosure can determine the inclusion relationship between the target control type and the preset control type, and when the preset control type includes the target control type, determine the current state information of the target control object corresponding to the voice control instruction, and then determine the target operation in combination with the current state information and the user intention information. Therefore, when it is determined that the current state information of the target control object matches the user intention information, it can determine and execute the corresponding target operation according to the user intention information to realize the interaction between the user and the vehicle, and when it is determined that the current state information of the target control object does not match the user intention information, no response is made, thereby avoiding the execution of unreasonable target operations when the current state information of the target control object does not match the user intention information, thereby improving the accuracy of voice interaction.
[0118] In some other embodiments, the above step 2033 may include the following steps:
[0119] Step 2033c: in response to the target control type being the first preset control type among the preset control types, determining initial query information corresponding to the user intention information.
[0120] Among them, the preset control type is the control type supported by the second speech recognition model.
[0121] In some examples, the first preset type is an information query type. When the target control type is a weather query type, the target control type belongs to the information query type, that is, the target control type is the first preset type.
[0122] In some embodiments, determining the initial query information corresponding to the user intent information may include: (1) sending a request message to the corresponding information interface to obtain the initial query information corresponding to the user intent information; (2) when a corresponding application (e.g., a weather application) is installed on the vehicle-mounted device, the initial query information corresponding to the user intent information may be obtained from the application.
[0123] Exemplarily, the target control type is a weather query type, and the first preset type is an information query type. Since the weather query type is an information query type, a request message can be sent to an information interface of a weather service provider (such as Baidu Weather), so that the information interface can return corresponding weather data based on the parameters such as the city to be queried and the date included in the request message, and the weather data can include temperature, humidity, air quality, etc.
[0124] In some embodiments, the initial query information may be incoherent information including symbols, images, etc. The initial query information may be the latest query information or historical query information, which is not limited in the embodiments of the present disclosure and may be determined according to actual conditions.
[0125] Step 2033d: Process the initial query information based on the natural language response sub-model in the second speech recognition model to obtain text query information.
[0126] In some examples, the second speech recognition submodel can generate a coherent text that conforms to the user's speech habits by analyzing the input data. When the initial query information is obtained, since the initial query information includes incoherent information such as symbols and images, the initial query information can be processed based on the natural language response submodel to obtain a natural language response that conforms to grammatical and semantic rules, that is, text query information. For example, the natural language response submodel can be a natural language generation model (Natural Language Generation, NLG).
[0127] For example, the initial query information is weather data. After obtaining the latest weather data, the NLG model can process the weather data and generate text query information that conforms to grammatical and semantic rules, such as "Today, Beijing is sunny and the temperature is between 10-20 degrees Celsius."
[0128] In some examples, when the natural language response sub-model is NLG, the delay time for the natural language response sub-model to process the initial query information to obtain the text query information is 100ms.
[0129] Step 2033e: convert the text query information into speech query information based on the speech synthesis sub-model in the second speech recognition model.
[0130] In some examples, the speech synthesis sub-model can be used to convert text information into speech information so that the user can receive the speech information through hearing. After receiving the text query information generated by the natural language response sub-model, the text query information can be converted into speech query information based on the speech synthesis sub-model. For example, the speech synthesis sub-model can be a speech synthesis model (Text To Speech, TTS).
[0131] In some examples, when the speech synthesis sub-model is TTS, the delay time for the speech synthesis sub-model to convert text query information into speech query information is 300ms.
[0132] Step 2033f, determining the target operation based on the output type of the voice query information.
[0133] In some embodiments, since the voice query information is of audio type, the output type of the voice query type is to play the voice query information through a speaker of the vehicle, that is, the target operation is to play the voice query information through the speaker.
[0134] For example, the voice query information is "the weather is sunny, the temperature is moderate, and it is suitable for outdoor activities". The voice "the weather is sunny, the temperature is moderate, and it is suitable for outdoor activities" is played through the speaker of the vehicle system so that the user can know the weather information.
[0135] The voice processing method provided by the embodiment of the present disclosure, when the target control type corresponding to the received voice control instruction is the first preset control type, can first determine the initial query information corresponding to the user intention information, and process the initial query information based on the natural language response sub-model to obtain text query information, and then convert the text query information into voice query information based on the voice synthesis sub-model. Therefore, it is possible to determine the target operation based on the output type of the voice query information, so that when the user inputs a voice control instruction of the information query type, the corresponding queried information can be provided to the user by voice.
[0136] Exemplary Devices
[0137] Figure 5 The present invention provides a schematic diagram of the structure of a speech processing device provided in an embodiment of the present invention. The speech processing device can be set in an electronic device such as a terminal device or a server, or in an object such as a vehicle, to execute the speech processing method of any of the above embodiments of the present invention.
[0138] like Figure 5 As shown, the device 300 may include:
[0139] The determination module 301 may be used to obtain a user's voice control instruction.
[0140] The processing module 302 may be configured to process the voice control instruction based on the first voice recognition model to generate target response information corresponding to the voice control instruction.
[0141] The processing module 302 may also be used to process the voice control instruction based on the second speech recognition model to obtain a target operation corresponding to the voice control instruction.
[0142] The execution module 303 may be used to execute the target operation corresponding to the voice control instruction; wherein, the target response information is generated and outputted before executing the target operation.
[0143] The disclosed embodiment provides a voice processing device, which can process the voice control instruction based on a first voice recognition model to generate target response information corresponding to the voice control instruction after collecting the voice control instruction input by the user through the audio sensor of the vehicle, and process the voice control instruction based on a second voice recognition model to obtain the target operation corresponding to the voice control instruction, and execute the target operation. That is, in the disclosed embodiment, two voice recognition models with different processing speeds for voice control instructions are used to process the collected voice control instructions in parallel at the same time. Because the voice recognition model with a faster processing speed for voice control instructions can generate and output the target response information earlier before executing the target operation corresponding to the voice control instruction, the delay in replying to the user's voice instruction is reduced, thereby improving the voice interaction experience of the vehicle-mounted system.
[0144] In one possible implementation, the processing module 302 may be specifically used to: process the voice control instruction based on the first speech recognition model to obtain target response information corresponding to the voice control instruction; determine the type of the target response information; and output the target response information based on the type of the target response information.
[0145] In one possible implementation, the processing module 302 can be specifically used to: determine the current state information of the target control object corresponding to the voice control instruction; process the voice control instruction and the current state information based on the first speech recognition model to obtain the target response information corresponding to the voice control instruction.
[0146] In one possible implementation, the processing module 302 can be specifically used to: convert the voice control instruction into text information based on the second speech recognition submodel in the second speech recognition model; perform semantic analysis on the text information corresponding to the voice control instruction based on the natural language understanding submodel in the second speech recognition model to obtain the target control type and user intent information; and determine the target operation corresponding to the voice control instruction based on the target control type and user intent information.
[0147] In one possible implementation, the processing module 302 can be specifically used to: determine the inclusion relationship between the target control type and the preset control type; wherein the preset control type is a control type supported by the second speech recognition model; in response to the inclusion relationship being that the preset control type includes the target control type, determine the target operation corresponding to the voice control instruction according to the user intention information.
[0148] In a possible implementation, the processing module 302 may be specifically used to: determine the current state information of the target control object corresponding to the voice control instruction; and determine the target operation according to the current state information and the user intention information.
[0149] In a possible implementation, the processing module 302 may be specifically used to: determine a matching relationship between current state information and user intent information; in response to the matching relationship being a match between the current state information and the user intent information, determine a target operation according to the user intent information.
[0150] In one possible implementation, the processing module 302 can be specifically used to: in response to the target control type being the first preset control type among the preset control types, determine the initial query information corresponding to the user intention information; wherein the preset control type is a control type supported by the second speech recognition model; process the initial query information based on the natural language response submodel in the second speech recognition model to obtain text query information; convert the text query information into speech query information based on the speech synthesis submodel in the second speech recognition model; and determine the target operation based on the output type of the speech query information.
[0151] The beneficial technical effects corresponding to the exemplary embodiment of the present device can be found in the corresponding beneficial technical effects of the above exemplary method section, which will not be repeated here.
[0152] Exemplary Electronic Devices
[0153] Figure 6 A structural diagram of an electronic device provided in an embodiment of the present disclosure includes at least one processor 61 and a memory 62.
[0154] The processor 61 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 60 to perform desired functions.
[0155] The memory 62 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 61 may execute one or more computer program instructions to implement the speech processing method and / or other desired functions of the various embodiments of the present disclosure described above.
[0156] In one example, the electronic device 60 may further include: an input device 63 and an output device 64, and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0157] The input device 63 may include various sensors, including but not limited to: a distance sensor for detecting the distance between the target object and the vehicle; an image sensor for collecting information about the vehicle's surrounding environment. In some examples, the input device may also include a pressure sensor for detecting seat pressure, determining whether there is a passenger, and the position of the passenger; a temperature sensor for monitoring the temperature in the cabin; a humidity sensor for monitoring the humidity in the cabin to assist in adjusting the interior environment; an air quality sensor for monitoring the air quality in the vehicle, such as carbon dioxide, volatile organic compounds (VOCs), etc.; a light sensor for detecting the light intensity inside and outside the vehicle; an acceleration sensor for detecting changes in the acceleration of the vehicle; a distance sensor for detecting the distance between the vehicle and other objects; a touch screen sensor for interaction with the vehicle's infotainment system; a biometric sensor, such as fingerprint recognition, facial recognition, etc.; a heart rate monitor for monitoring the driver's heart rate; a sound sensor for voice recognition and interaction to achieve voice control functions; a seat sensor for monitoring the use of the seat, such as whether the seat is occupied and the body shape of the passenger; a wireless communication sensor, such as Bluetooth, Wi-Fi, etc., for connecting with smart devices to achieve data transmission and remote control. In addition to the examples given above, the input device may also include more or fewer sensors, which will not be described in detail here.
[0158] The output device 64 can output various information or signals to other hardware or devices, which may include displays, vehicle audio, seats, windows, steering wheels, etc., as well as communication networks and remote output devices connected thereto, etc. The displays may include a plurality of different display screens such as a driver's display screen, a co-driver's display screen, and a rear display screen, and the vehicle audio may include a plurality of speakers arranged at different positions in the vehicle cabin, and the different display screens or speakers may work independently.
[0159] Of course, to simplify, Figure 6 Only some of the components related to the present disclosure in the electronic device 60 are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application situations, the electronic device 60 may also include any other appropriate components.
[0160] Exemplary computer program products and computer-readable storage media
[0161] In addition to the above-mentioned methods and devices, embodiments of the present disclosure may also provide a computer program product, including computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the speech processing method of various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section.
[0162] The computer program product may be written in any combination of one or more programming languages to write program code for performing the operations of the disclosed embodiments, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0163] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enables the processor to execute the steps of the speech processing method of various embodiments of the present disclosure described in the above “Exemplary Method” section.
[0164] Computer readable storage media can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium is, for example, but not limited to, a system, device or device including electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0165] The basic principles of the present disclosure are described above in conjunction with specific embodiments. However, the advantages, strengths, effects, etc. mentioned in the present disclosure are only examples and not limitations, and cannot be considered as necessary for each embodiment of the present disclosure. In addition, the specific details disclosed above are only for the purpose of illustration and ease of understanding, rather than limitation, and the above details do not limit the present disclosure to being implemented by adopting the above specific details.
[0166] Those skilled in the art may make various changes and modifications to the present disclosure without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present disclosure claims and their equivalents, the present disclosure is also intended to include these modifications and variations.
Claims
1. A speech processing method, comprising: Obtain the user's voice control instructions; Processing the voice control instruction based on the first speech recognition model to generate target response information corresponding to the voice control instruction; Processing the voice control instruction based on a second speech recognition model to obtain a target operation corresponding to the voice control instruction; Execute the target operation corresponding to the voice control instruction; wherein, before executing the target operation, generate and output the target response information.
2. The method according to claim 1, wherein: The processing of the voice control instruction based on the first voice recognition model to generate target response information corresponding to the voice control instruction includes: Processing the voice control instruction based on the first speech recognition model to obtain the target response information corresponding to the voice control instruction; Determining the type of the target response information; The target response information is output based on the type of the target response information.
3. The method according to claim 2, wherein: The processing of the voice control instruction based on the first voice recognition model to obtain the target response information corresponding to the voice control instruction includes: Determine current state information of a target control object corresponding to the voice control instruction; The voice control instruction and the current state information are processed based on the first speech recognition model to obtain the target response information corresponding to the voice control instruction.
4. The method according to claim 1, wherein: The processing of the voice control instruction based on the second voice recognition model to obtain a target operation corresponding to the voice control instruction includes: Converting the voice control instruction into text information based on a second voice recognition sub-model in the second voice recognition model; Performing semantic analysis on the text information corresponding to the voice control instruction based on the natural language understanding sub-model in the second voice recognition model to obtain target control type and user intention information; Based on the target control type and the user intention information, a target operation corresponding to the voice control instruction is determined.
5. The method according to claim 4, wherein: The determining, based on the target control type and the user intention information, a target operation corresponding to the voice control instruction includes: Determine an inclusion relationship between the target control type and a preset control type; wherein the preset control type is a control type supported by the second speech recognition model; In response to the inclusion relationship that the preset control type includes the target control type, a target operation corresponding to the voice control instruction is determined according to the user intention information.
6. The method according to claim 5, wherein: The determining, according to the user intention information, a target operation corresponding to the voice control instruction includes: Determine current state information of a target control object corresponding to the voice control instruction; The target operation is determined according to the current state information and the user intention information.
7. The method according to claim 6, wherein: The determining the target operation according to the current state information and the user intention information includes: Determining a matching relationship between the current state information and the user intention information; In response to the matching relationship being a match between the current state information and the user intent information, the target operation is determined according to the user intent information.
8. The method according to claim 4, wherein: The determining, based on the target control type and the user intention information, a target operation corresponding to the voice control instruction includes: In response to the target control type being a first preset control type among preset control types, determining initial query information corresponding to the user intention information; wherein the preset control type is a control type supported by the second speech recognition model; Processing the initial query information based on the natural language response sub-model in the second speech recognition model to obtain text query information; Converting the text query information into speech query information based on the speech synthesis sub-model in the second speech recognition model; A target operation is determined based on the output type of the voice query information.
9. A speech processing device, comprising: A determination module, used to obtain a user's voice control command; A processing module, configured to process the voice control instruction based on a first speech recognition model to generate target response information corresponding to the voice control instruction; The processing module is further used to process the voice control instruction based on the second speech recognition model to obtain a target operation corresponding to the voice control instruction; An execution module is used to execute the target operation corresponding to the voice control instruction; wherein, before executing the target operation, the target response information is generated and output.
10. A computer-readable storage medium storing a computer program, wherein the computer program is used to execute the speech processing method according to any one of claims 1 to 8.
11. An electronic device, comprising: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the speech processing method described in any one of claims 1-8 above.
Citation Information
Cited By
Voice processing method and device
CN120895042A