Voice interaction method, vehicle and computer readable storage medium

By dynamically adjusting the pause duration and combining semantic analysis and user habit identifiers to optimize in-vehicle voice interaction, the problem of rigid interaction modes in existing technologies has been solved, improving the efficiency of voice interaction and user experience.

CN116168700BActive Publication Date: 2025-12-09GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310167856.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-23
Publication Date
2025-12-09
Estimated Expiration
2043-02-23

AI Technical Summary

Technical Problem

In existing in-vehicle voice technology, the interaction between users and voice assistants is a question-and-answer format, resulting in a rigid interaction mode, poor fluency and convenience, and a poor user experience.

Method used

By receiving the first voice request, determining the pause duration, and generating a response after receiving a follow-up voice request within the pause duration, the pause time is dynamically adjusted to reduce the number of interaction rounds. The response time is optimized by combining semantic analysis and user habit identifiers.

Benefits of technology

It improves the efficiency and user experience of voice interaction, reduces unnecessary waiting time and interaction costs, and enables a more flexible interaction mode.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116168700B_ABST
    Figure CN116168700B_ABST
Patent Text Reader

Abstract

The application discloses a voice interaction method, comprising: receiving a first voice request; determining a pause duration for replying according to the first voice request; if a subsequent voice request is received within the pause duration, generating a reply according to the first voice request and the subsequent voice request to complete the voice interaction. According to the application, the pause duration for replying can be determined according to the received first voice request, and the reply can be generated according to the first voice request and the subsequent voice request received within the pause duration, and finally the voice interaction is completed. In the voice interaction method of the application, a plurality of voice requests of a user only need one reply of a voice assistant, the number of turns of a dialogue in the voice interaction process is reduced, the waste of interaction cost caused by multiple interactions is avoided, and the user interaction experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent voice, in particular to a voice interaction method, a vehicle and a computer readable storage medium. BACKGROUND

[0002] At present, the vehicle-mounted voice technology can support the user to interact in the vehicle cabin through voice, for example, to control vehicle parts or interact with components in the vehicle system user interface. Generally, the current interaction mode of the user and the voice assistant is a question and answer mode, that is, the voice assistant can reply to the voice request of the user and then the next round of interaction can be performed. The interaction mode is relatively rigid as a whole, the fluency and convenience of voice interaction are poor, and the user interaction experience is not good. SUMMARY

[0003] The present application provides a voice interaction method, a vehicle and a computer readable storage medium.

[0004] The voice interaction method of the present application comprises:

[0005] receiving a first voice request;

[0006] determining a pause duration for reply according to the first voice request;

[0007] if a subsequent voice request is received within the pause duration, generating a reply according to the first voice request and the subsequent voice request to complete the voice interaction.

[0008] In this way, the present application can determine the pause duration for reply according to the received first voice request, and generate a reply according to the first voice request and the subsequent voice request received within the pause duration to finally complete the voice interaction. In the voice interaction method of the present application, the voice assistant can reply to multiple voice requests of the user that meet the conditions at one time, reduce the number of turns of the dialogue in the voice interaction process, avoid the waste of interaction cost caused by multiple interactions, make the interaction mode more flexible, and improve the user interaction experience.

[0009] The method further comprises:

[0010] performing semantic analysis on the first voice request to obtain a semantic identifier;

[0011] confirming a user habit identifier of the first voice request according to a history record of voice requests;

[0012] determining the pause duration according to the semantic identifier and the user habit identifier.

[0013] Therefore, the pause duration can be determined according to the semantic identification of the first voice request and the user habit identification, and a response can be made after the pause duration, so that the waste of interaction cost caused by multiple interactions is avoided.

[0014] The pause duration is determined according to the semantic identification and the user habit identification.

[0015] If it is confirmed according to the semantic identification that the semantic of the first voice request is complete and it is confirmed according to the user habit identification that the first voice request conforms to the user expression habit, it is determined to make a response after pausing for a first duration.

[0016] Therefore, the pause duration before the response to the first voice request can be determined according to the semantic completeness of the first voice request and the degree of conformity to the user expression habit, so that meaningless time delay of multiple interactions is avoided, and the user interaction experience is improved.

[0017] The pause duration is determined according to the semantic identification and the user habit identification.

[0018] If it is confirmed according to the semantic identification that the semantic of the first voice request is complete or it is confirmed according to the user habit identification that the first voice request conforms to the user expression habit, it is determined to make a response after pausing for a second duration.

[0019] Therefore, the pause duration before the response to the first voice request can be determined according to the semantic completeness of the first voice request and the degree of conformity to the user expression habit, so that meaningless time delay of multiple interactions is avoided, and the user interaction experience is improved.

[0020] The pause duration is determined according to the semantic identification and the user habit identification.

[0021] If it is confirmed according to the semantic identification that the semantic of the first voice request is not complete and it is confirmed according to the user habit identification that the first voice request does not conform to the user expression habit, it is determined to make a response after pausing for a third duration.

[0022] Therefore, the pause duration before the response to the first voice request can be determined according to the semantic completeness of the first voice request and the degree of conformity to the user expression habit, so that the long waiting time of the voice assistant when the semantic is not clear is avoided, and the user interaction experience is improved.

[0023] The method further includes:

[0024] If it is confirmed according to the semantic identification that the first voice request meets the preset rule, it is determined to make a response after pausing for a fourth duration.

[0025] Therefore, the pause duration before the reply to the first voice request is determined according to whether the preset rule is met, so that the user has sufficient thinking time to improve the voice request, the voice interaction efficiency is prevented from being affected by the fast reply, and the user interaction experience is improved.

[0026] The method further includes:

[0027] If no subsequent voice request is received within the pause duration, a reply is generated according to the first voice request after the pause ends, and the voice interaction is ended.

[0028] Therefore, when no subsequent voice request is received after the pause duration ends, a reply is generated according to the first voice request, so that meaningless time delay in the voice interaction process is avoided, and the user interaction experience is improved.

[0029] The method includes:

[0030] The voice request is received according to the silence identifier.

[0031] Therefore, whether the subsequent voice request is received within the pause duration can be confirmed according to the silence identifier, so that the reply method for different voice requests is determined, and the user interaction experience is improved.

[0032] The vehicle of the present application includes a processor and a memory, the memory stores a computer program, and the computer program is executed by the processor to implement the above method.

[0033] The computer readable storage medium of the present application stores a computer program, and when the computer program is executed by one or more processors, the above method is implemented.

[0034] Additional aspects and advantages of the embodiments of the present application will be in part apparent and in part pointed out hereinafter in the description of the embodiments of the present application, some of which will be apparent from the description of the embodiments of the present application, or will be learned from the practice of the embodiments of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0035] The above and / or additional aspects and advantages of the present application will become apparent and more readily appreciated from the following description of the embodiments, taken in conjunction with the accompanying drawings in which:

[0036] Figure 1 is a flowchart of the voice interaction method of the present application;

[0037] Figure 2 is a flowchart of the voice interaction method of the present application;

[0038] Figure 3 is a flowchart of the voice interaction method of the present application;

[0039] Figure 4 is a flowchart of a voice interaction method of the present application;

[0040] Figure 5 is a flowchart of a voice interaction method of the present application;

[0041] Figure 6 is a flowchart of a voice interaction method of the present application;

[0042] Figure 7 is a flowchart of a voice interaction method of the present application. DETAILED DESCRIPTION

[0043] Embodiments of the present application are described in detail below with reference to the accompanying drawings, in which the same or similar components have the same or similar reference numerals throughout the several views, and the embodiments described below are examples for explaining the present application and are not intended to limit the present application.

[0044] At present, the vehicle-mounted voice technology can support the user to interact in the vehicle cabin through voice, for example, to control vehicle parts or interact with components in the vehicle system user interface. Generally, the current user and voice assistant interaction mode is a question and answer mode, that is, the voice assistant can reply to the current user's voice request and then proceed to the next round of interaction. For example, in the following vehicle intelligent cabin use scenario, the user and voice assistant interaction process is as follows:

[0045] Step 1: the user says "hello XX (wake-up keyword)", trying to wake up the voice assistant;

[0046] Step 2: the voice assistant replies "OK", responding after being woken up;

[0047] Step 3: the user issues a voice request again: "turn on the air conditioner";

[0048] Step 4: the voice assistant replies "OK, the air conditioner has been turned on for you" after waiting for the user to finish "turn on the air conditioner";

[0049] Step 5: the user continues to issue a voice request: "turn on the sports mode";

[0050] Step 6: the voice assistant replies "OK, the sports mode has been turned on for you" after waiting for the user to finish "turn on the sports mode".

[0051] According to the above reply process, in the related art, if multiple voice requests issued by the user need to be replied to, the voice interaction round must be no less than the number of voice requests issued by the user.

[0052] In the related art, whether a voice request in a current turn is ended is mainly determined based on voice activity detection (VAD). In actual operation, the voice assistant usually needs to wait for the voice request to end and then wait for a fixed time, for example, 500 ms, before replying. In addition, the voice request issued by the user needs to be subjected to voice recognition, natural language understanding, etc. In addition, the above process may have network delay, etc. The user may need to wait for more than 2 seconds after issuing the voice request before receiving the reply.

[0053] In addition, if the voice request issued by the user cannot be effectively recognized or the corresponding instruction cannot be effectively executed, the voice assistant needs to be re-awakened. This will additionally increase the interaction turn and interaction time in the interaction. The overall process efficiency of the voice interaction is low, and the user experience is poor.

[0054] Based on the above possible problems, please refer to Figure 1 The present application provides a voice interaction method, comprising:

[0055] 01: receiving a first voice request;

[0056] 02: determining a pause duration for replying according to the first voice request;

[0057] 03: if a subsequent voice request is received within the pause duration, generating a reply according to the first voice request and the subsequent voice request to complete the voice interaction.

[0058] The present application also provides a vehicle comprising a memory and a processor. The voice interaction method of the present application can be implemented by the vehicle of the present application. Specifically, the memory stores a computer program, and the processor is configured to receive a first voice request, determine a pause duration for replying according to the first voice request, and if a subsequent voice request is received within the pause duration, generate a reply according to the first voice request and the subsequent voice request to complete the voice interaction.

[0059] In a voice interaction, in terms of timing, the voice request issued by the user first is the first voice request. The voice request issued later than the first voice request is the subsequent voice request. The pause duration is the waiting time for determining whether the voice request is ended. After the pause is ended, a reply is made. In some examples, the pause duration can be determined according to the relevant information of the first voice request.

[0060] The pause duration is the time interval between the first voice request issued by the user and the consecutive voice request corresponding thereto. When the time interval between the two voice requests is less than the pause duration, the voice request issued earlier in time is the first voice request, and the voice request issued later can be regarded as the consecutive voice request corresponding to the first voice request, and the voice assistant only needs to make one round of reply to the above two voice requests.

[0061] The length of the pause duration can be dynamically adjusted by the voice assistant through analysis of user habits, semantics and the like. The voice interaction method of the present application can dynamically adjust the fixed pause duration in the related art voice activity detection (VAD), and in some scenarios, the fixed pause duration can be shortened from 500-1000 milliseconds to 100-200 milliseconds, reducing the user waiting time and improving the speed of single-turn dialogue and reply execution efficiency.

[0062] After the user issues the first voice request, if the consecutive voice request is received within the pause duration, according to the definition of the above pause duration, the voice assistant needs to generate a reply in combination with the first voice request and the consecutive voice request. Conversely, if no consecutive voice request is received within the pause duration, the voice assistant can directly generate a reply according to the first voice request. Finally, the process of voice interaction is completed.

[0063] In the above example of the vehicle intelligent cockpit use scenario, the interaction process after the voice assistant reduces the reply round is as follows:

[0064] Step 1: The user says "Hello XX (wake-up keyword)", trying to wake up the voice assistant;

[0065] The voice assistant determines whether the user is trying to wake up himself and whether the user's voice request is complete. When the voice assistant finds that the user is trying to wake up himself and the voice request is not complete for the time being, it waits for the user to issue a voice request again, and no reply will be made to the user during the waiting period.

[0066] Step 2: The user issues a voice request again to the voice assistant: "Turn on the air conditioner";

[0067] The voice assistant continuously determines whether the user's voice request is complete and simultaneously identifies the instructive sentence therein. For example, after the user issues the voice requests "Hello" and "Turn on the air conditioner", the voice assistant determines whether the user continues to issue a voice request. If the user is accustomed to issuing the voice request "Exercise mode" after issuing the voice request "Turn on the air conditioner", the voice assistant will continue to wait for the user to issue a voice request again.

[0068] Step 3: The user issues a voice request again to the voice assistant: "Exercise mode";

[0069] Step 4: The voice assistant replies: "OK, the air conditioner is turned on, and the sports mode is started"

[0070] According to the above reply process, the voice interaction method of the present application can reduce the interaction process of 6 steps required by the related technology to 4 steps. The voice assistant can uniformly reply to multiple voice requests of the user after the user finishes issuing the voice request. The multiple voice requests of the user only require one reply of the voice assistant, which can greatly improve the reply efficiency of the voice assistant and the efficiency of voice interaction.

[0071] In summary, according to the received first voice request, the present application determines the pause duration for reply, and generates a reply according to the first voice request and the subsequent voice request received within the pause duration, and finally completes the voice interaction. In the voice interaction method of the present application, for multiple voice requests of the user that meet the conditions, the voice assistant can reply at one time, reduce the number of turns of the dialogue in the voice interaction process, and avoid the waste of interaction cost caused by multiple interactions. The interaction mode is more flexible.

[0072] Please refer to Figure 2 , the method further comprises:

[0073] 04: performing semantic analysis on the first voice request to obtain a semantic identifier;

[0074] 05: confirming a user habit identifier of the first voice request according to a history record of voice requests;

[0075] 06: determining a pause duration according to the semantic identifier and the user habit identifier.

[0076] The processor is configured to perform semantic analysis on the first voice request to obtain a semantic identifier, confirm a user habit identifier of the first voice request according to a history record of voice requests, and determine a pause duration according to the semantic identifier and the user habit identifier.

[0077] Specifically, the semantic analysis can be performed on the first voice request to determine whether the first voice request is a complete and valid sentence. For example, in the vehicle intelligent cockpit scenario, the semantic of the voice request "navigate to" is not complete, while the semantic of "navigate to the train station" is complete.

[0078] The process of performing semantic analysis on the first voice request and obtaining the corresponding semantic identifier is the core mechanism. When the user's voice request is a complete and valid sentence, it is determined that the voice request is ended. As shown in Figure 3 , the specific steps are as follows:

[0079] Step 1: acoustic feature extraction and acoustic model are used to recognize the pronunciation sequence; at the same time, a decoder function is used to recognize the words and syntax structure of the voice request.

[0080] Step 2: Deterministic recognition of the voice request using pronunciation sequence information to obtain the deterministic confidence of the voice request.

[0081] This method extracts pronunciation information in the process of the user issuing a voice request, such as a pinyin sequence, and calculates the deterministic information in the user's voice request, including the clarity of speech and whether the speech hesitates, through confidence analysis. The closer the confidence is to 1, the clearer the speaker's voice request is, and the less hesitant it is.

[0082] Step 3: Text confidence analysis of the voice request using a word graph.

[0083] The word graph is the result of combining pronunciation sequence and language model, and is used to calculate the text confidence of voice recognition. The higher the confidence, the more accurate the recognition result text is, and the more complete the text representation is.

[0084] Step 4: Syntax structure analysis of the text obtained after the voice request is recognized by the decoder to obtain the instruction confidence of the voice request.

[0085] Syntax analysis of the voice request can obtain the sentence structure result, such as subject-predicate structure and verb-object structure. The sentence structure after syntax analysis is input to the voice request instruction analysis for confidence scoring. The higher the confidence, the more explicit the user's voice request intent is, and the higher the credibility is.

[0086] Step 5: Threshold determination based on the deterministic confidence, text confidence, and instruction confidence of the voice request obtained in the above steps.

[0087] When the above three confidences of the voice request simultaneously meet the threshold value, it can be determined that the voice request has ended. For example, set the threshold value to 0.8 or 0.9, etc. When the above three confidences simultaneously meet the set threshold value, the semantic identification can indicate that the user's semantic is complete.

[0088] Compared with the traditional single natural language understanding method, the above process of obtaining voice identification introduces intermediate information in the process of understanding the content of the voice request. The joint determination mechanism of word graph and syntax analysis significantly improves the accuracy.

[0089] In addition, after semantic analysis of the first voice request to obtain the semantic identification, it is also necessary to determine whether the content said by the user is complete according to the habitual expressions of the user in the historical use of the voice request. For example, in the historical record, the user will say "turn on the air conditioner", "navigate to the company" or "go to the company", and "play a song" before driving. When the user says "navigate to", it hits the historical record "navigate to the company", and is judged as incomplete content. Therefore, when the user supplements the name of the place to which the navigation is needed, it can be judged that the voice request is complete. The user habit identification of the first voice request can be confirmed according to whether the voice request hits the historical record of the user habit and the semantic completeness.

[0090] For the first voice request issued by the user, the semantic identification obtained through semantic analysis and the user habit identification obtained by comparing the historical record of the voice request are used to determine the pause time length of the first voice request.

[0091] In this way, the pause time length can be determined according to the semantic identification and the user habit identification of the first voice request, and a reply can be made after the pause time length, so as to avoid the waste of interaction cost caused by multiple interactions.

[0092] Step 06 includes:

[0093] If it is confirmed according to the semantic identification that the semantic of the first voice request is complete and according to the user habit identification that the first voice request conforms to the user expression habit, it is determined to reply after pausing for a first time length.

[0094] The processor is configured to determine to reply after pausing for a first time length if it is confirmed according to the semantic identification that the semantic of the first voice request is complete and according to the user habit identification that the first voice request conforms to the user expression habit.

[0095] Specifically, the semantic analysis of the first voice request obtains the semantic identification, it is determined that the semantic of the first voice request is complete, and according to the historical record of the voice request, it is confirmed that the user habit identification of the first voice request conforms to the user expression habit, so the first voice request is in a completed state with high probability. In this case, the pause time can be determined as a relatively short time, such as about 100 milliseconds, which is referred to as a first time length. The above-mentioned first time length can be set to 90 milliseconds, 100 milliseconds, 110 milliseconds, etc., which is not limited here.

[0096] In one example, the user issues a first voice request "navigate to the company", the first voice request is analyzed by words and syntax, it is determined that the user pronounces clearly when issuing the voice request, and the certainty confidence is close to 1. The sentence is a typical verb-object structure, and the instruction confidence is high. And according to the word graph analysis, the text of the first voice request is accurate, complete and has high text confidence. Finally, it is determined that the semantic of the first voice request "navigate to the company" is complete.

[0097] Further, the "navigate to company" hits the history record of the user's common voice request before driving, and confirms that the first voice request meets the user's expression habit.

[0098] Since the semantic of the above first voice request "navigate to company" is complete, and the first voice request meets the user's expression habit according to the user habit identification, it is determined to reply after a pause of a first duration, such as 100 milliseconds.

[0099] In this way, the pause duration before replying to the first voice request can be determined according to the semantic completeness of the first voice request and the degree of fit with the user's expression habit, avoiding meaningless time delay of multiple interactions and improving user interaction experience.

[0100] Step 06 includes:

[0101] If the semantic of the first voice request is complete according to the semantic identification or the first voice request meets the user's expression habit according to the user habit identification, it is determined to reply after a pause of a second duration.

[0102] The processor is configured to determine to reply after a pause of a second duration if the semantic of the first voice request is complete according to the semantic identification or the first voice request meets the user's expression habit according to the user habit identification.

[0103] Specifically, the first voice request is subjected to semantic analysis to obtain a semantic identification, and it is determined that the semantic of the first voice request is complete. Or according to the history record of the voice request, it is confirmed that the user habit identification of the first voice request meets the user's expression habit. Only one of the two conditions of the semantic completeness of the first voice request and the meeting of the user's expression habit is required, and then the first voice request can be determined to be in a completed state. In this case, the pause time, such as about 200 milliseconds, can be determined as a second duration. The above first duration can be set to 190 milliseconds, 200 milliseconds, 210 milliseconds, etc., which is not specifically limited here. Since the first voice request before the second duration may have incomplete semantic or not meet the user's habit, the understanding of the voice assistant to the first voice request requires more time than the understanding of the first voice request before the first duration. Therefore, the second duration should be set longer than the first duration.

[0104] In one example, the user issues a first voice request "navigate to the hospital", the first voice request is analyzed in terms of words and syntax, and it is determined that the user utters the voice request clearly and the certainty confidence is close to 1. The sentence is a typical verb-object structure, and the instruction confidence is high. According to the word graph analysis, the first voice request text is accurate and complete, and the text confidence is high. Finally, it is determined that the semantic of the first voice request "navigate to the hospital" is complete. Although "navigate to the hospital" does not hit the historical record in the common voice request before driving, it is still determined that the first voice request has been completed, and the pause time, i.e., the second time length, is determined, such as 200 milliseconds. After pausing for the second time length, the reply content can be "navigation has been started for you, and the nearest hospital from the current distance is input as the destination".

[0105] In another example, the user issues a first voice request "go to the company", the first voice request is analyzed in terms of words and syntax, and it is determined that the user utters the voice request clearly, but the sentence structure is not complete, and the instruction confidence and the text confidence do not reach the threshold. However, "go to the company" hits the historical record of the common voice request "go to the company" of the user before driving, and it is determined that the first voice request conforms to the user's expression habit, and it is still determined that the first voice request has been completed, and the pause time, i.e., the second time length, is determined, such as 200 milliseconds.

[0106] In this way, the pause time before the reply to the first voice request can be determined according to the semantic completeness of the first voice request and the degree of fit with the user's expression habit, avoiding meaningless time delay of multiple interactions and improving the user interaction experience.

[0107] Step 06 comprises:

[0108] If it is determined according to the semantic identification that the semantic of the first voice request is not complete and according to the user habit identification that the first voice request does not conform to the user's expression habit, it is determined to reply after pausing for a third time length.

[0109] The processor is configured to determine to reply after pausing for a third time length if it is determined according to the semantic identification that the semantic of the first voice request is not complete and according to the user habit identification that the first voice request does not conform to the user's expression habit.

[0110] Specifically, semantic analysis is performed on the first voice request to obtain a semantic identifier, it is determined that the semantic of the first voice request is incomplete, and according to the historical record of the voice request, it is confirmed that the user habit identifier of the first voice request does not conform to the user expression habit, and then a pause time, such as about 500 milliseconds, is determined, which is referred to as a third time length. The third time length can be set to 490 milliseconds, 500 milliseconds, 510 milliseconds, etc., which is not specifically limited here. Because the first voice request before the third time length has the conditions of incomplete semantics and not meeting the user habit, the understanding of the voice assistant to the first voice request takes more time than the understanding of the first voice request before the first time length and the second time length. Therefore, the third time length should be set to be longer than the first time length and the second time length.

[0111] In one example, the user issues a first voice request "quite good", and the first voice request is subjected to word and syntax analysis to obtain that the user pronounces clearly when issuing the voice request, but the sentence only has an adjective, the sentence structure is incomplete, and the instruction confidence and the text confidence do not reach the threshold.

[0112] Further, "quite good" does not hit the historical record in the commonly used voice request before driving. Accordingly, it is determined to reply after pausing for a third time length. The third time length can be set, such as 500 milliseconds.

[0113] In this way, the pause time before replying to the first voice request can be determined according to the semantic integrity of the first voice request and the degree of fit with the user's expression habit, avoiding the long waiting time of the voice assistant when the semantics is unclear, and improving the user interaction experience.

[0114] Step 06 includes:

[0115] If it is confirmed according to the semantic identifier that the first voice request meets the preset rule, it is determined to reply after pausing for a fourth time length.

[0116] The processor is configured to, in a case where the user-defined element exists in the custom vocabulary, concatenate and combine a plurality of corresponding flow nodes to determine a target flow node.

[0117] Specifically, when the user issues a voice request, the sentence may be paused due to thinking time, which cannot guarantee the fluency of the voice request while guaranteeing the completeness of the voice request. For the pause that may occur in the user voice request, a preset rule can be set. The preset rule can be a special rule set for "common hesitation words" in the voice request, such as "navigate to", "I want to listen", "open", etc. The number and type of preset rules are not limited here

[0118] The user pauses after the aforementioned hesitation words because they haven't yet selected the specific instruction. If the voice assistant responds immediately after a short delay, the user experience will be poor. To ensure the content after the hesitation word pause is effective, a longer pause time, such as about one second, can be determined, referred to as the fourth duration. If the user still doesn't issue another voice request after the fourth duration, the response will directly address the first voice request, i.e., the hesitation word.

[0119] In one example, the user issues the first voice request, "Navigate to...". Semantic analysis of the first voice request yields a semantic identifier that indicates it conforms to the rule of "common hesitation words". Based on this, the pause time, i.e., the fourth duration, is determined, such as 1 second.

[0120] The fourth duration mentioned above can be set to 0.9 seconds, 1 second, 1.5 seconds or longer to ensure that users have sufficient time to think. The specific value of the fourth duration is not limited here.

[0121] In this way, the pause duration before responding to the first voice request can be determined based on whether the first voice request meets the preset rules, ensuring that the user has sufficient time to think and complete the voice request, avoiding excessively fast responses that affect the efficiency of voice interaction, and improving the user interaction experience.

[0122] Please see Figure 4 The methods also include:

[0123] 07: If no voice continuation request is received within the pause duration, a response will be generated based on the first voice request after the pause ends and the voice interaction will end.

[0124] The processor is used to generate a response based on the first voice request and end the voice interaction after the pause ends if no voice continuation request is received within the pause duration.

[0125] Specifically, for the user's first voice request, considering the four semantic integrity criteria and the degree of conformity with the user's expression habits, and taking into account the result of the silence determination, if the voice assistant still does not receive a continuation voice request after the pause duration ends, it generates a response based on the specific content of the first voice request, thus ending the voice interaction. The response determination process for the user's voice request is as follows: Figure 5 As shown.

[0126] When the first voice request is semantically complete and conforms to the user's expression habits, a response is determined after a pause of the first duration. In one example, the user makes a voice request "Navigate to the company," which is semantically complete and conforms to the user's expression habits. The response could be "Navigation has been enabled for you, and you have entered the address set as 'the company' as your destination."

[0127] When the semantic of the first voice request is complete or conforms to the user's expression habit, it is determined to reply after a pause of a second duration. In one example, the user issues a voice request "go to the company", the semantic of which is not complete but conforms to the user's expression habit, and the reply content can be "I have started navigation for you and set the address of 'company' as the destination".

[0128] When the semantic of the first voice request is not complete and does not conform to the user's expression habit, it is determined to reply after a pause of a third duration. In one example, the user issues a voice request "very good", the semantic of which is not complete and does not conform to the user's expression habit, and it is determined to reply after a pause of a third duration, such as "I don't quite understand your meaning".

[0129] When the first voice request meets the preset rule, it is determined to reply after a pause of a fourth duration. In one example, the user issues a first voice request "navigate to", and the semantic identification shows that it meets the rule of "common hesitation word". Accordingly, it is determined to reply after a fourth duration, such as "please confirm the destination you navigate to".

[0130] In this way, when the pause duration ends and no subsequent voice request is received, a reply is generated according to the first voice request, avoiding meaningless delay in the voice interaction process and improving the user interaction experience.

[0131] Please refer to Figure 6 , the method further comprises:

[0132] 08: confirming whether a subsequent voice request is received within the pause duration according to the silence identification.

[0133] The processor is configured to confirm whether a subsequent voice request is received within the pause duration according to the silence identification.

[0134] Specifically, the voice interaction method of the present application needs to determine whether a voice request is received again within a certain time after the first voice request is issued, to determine whether the voice request is the subsequent voice request of the first voice request. If the time interval between the voice request received again and the first voice request is long, the voice assistant needs to be awakened again through a wake-up type voice request.

[0135] The determination process uses VAD (Voice Activity Detection) frame-level silence identification to identify whether the user's audio is silent frame by frame. Within the pause duration, the silence identification is combined with the above-mentioned semantic identification and user habit identification to realize the process of reducing interaction rounds, such as Figure 7As shown, the specific principle is that if the mute identifier shows that the user audio of the current frame is mute, it is judged that the current frame does not receive the continuous speech request, and the next frame is jumped to continue to listen to the audio without immediately replying, and the interaction is reduced by one round. If the mute identifier shows that the user audio of the current frame is not mute, it is judged that the continuous speech request is received within the pause duration. In the subsequent speech interaction process, the speech request can be regarded as a continuous speech request processing, and after a reply is obtained, the speech interaction process is ended.

[0136] In this way, whether the continuous speech request is received within the pause duration can be confirmed according to the mute identifier, so as to determine the reply method for different speech requests, and improve the user interaction experience.

[0137] The computer readable storage medium of the present application stores a computer program, which, when executed by one or more processors, implements the above method.

[0138] In the description of the present application, the description of the terms "above", "specifically", "further", and the like means that the specific features, structures, materials or characteristics described in combination with the embodiments or examples are contained in at least one embodiment or example of the present application. In the present application, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present application and the features of the different embodiments or examples without contradiction.

[0139] Any process or method descriptions in flow charts or otherwise described herein represents an example of executable request code that includes one or more steps for achieving a particular result. The scope of preferred embodiments of the present application includes additional implementation in which the functions described in the illustrated or discussed order are not performed in the order shown or discussed, including functions performed in a substantially simultaneous manner or in reverse order according to the functions involved, which should be understood by those skilled in the art of the embodiments to which the present application belongs.

[0140] Although the embodiments of the present application have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application.

Claims

1. A vehicle voice interaction method, characterized by, The method comprises: receiving a first voice request; determining a pause duration for replying according to the first voice request; if a subsequent voice request is received within the pause duration, generating a reply according to the first voice request and the subsequent voice request to complete the voice interaction; The method further comprises: performing semantic analysis on the first voice request to obtain a semantic identification; confirming a user habit identification of the first voice request according to a history record of voice requests; determining the pause duration according to the semantic identification and the user habit identification.

2. The voice interaction method of claim 1, wherein, The determination of the pause duration according to the semantic identification and the user habit identification comprises: if it is confirmed according to the semantic identification that the semantic of the first voice request is complete and it is confirmed according to the user habit identification that the first voice request conforms to the user expression habit, determining to reply after a first pause duration.

3. The voice interaction method of claim 1, wherein, The determination of the pause duration according to the semantic identification and the user habit identification comprises: if it is confirmed according to the semantic identification that the semantic of the first voice request is complete or it is confirmed according to the user habit identification that the first voice request conforms to the user expression habit, determining to reply after a second pause duration.

4. The voice interaction method of claim 1, wherein, The determination of the pause duration according to the semantic identification and the user habit identification comprises: if it is confirmed according to the semantic identification that the semantic of the first voice request is not complete and it is confirmed according to the user habit identification that the first voice request does not conform to the user expression habit, determining to reply after a third pause duration.

5. The voice interaction method of claim 1, wherein, The method further comprises: if it is confirmed according to the semantic identification that the first voice request meets a preset rule, determining to reply after a fourth pause duration.

6. The voice interaction method of claim 1, wherein, The method further comprises: if no subsequent voice request is received within the pause duration, generating a reply according to the first voice request after the pause ends and ending the voice interaction.

7. The voice interaction method of any of claims 1-6, wherein, The method comprises: confirming whether the subsequent voice request is received within the pause duration according to a silence identification.

8. A vehicle characterized by comprising: The vehicle comprises a memory and a processor, the memory stores a computer program, and the computer program is executed by the processor to implement the method of any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and when the computer program is executed by one or more processors, the method of any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Voice interaction method and device

    CN109377998A

  • Speech recognition method and device, electronic equipment and storage medium

    CN114582333A