Voice interaction method, voice interaction device, server and readable storage medium
By receiving and supplementing voice requests in the vehicle cabin, the problem of repeated wake-up due to cross-voice zone command inheritance in multi-user scenarios is solved, enabling intelligent response and greater freedom of expression for all users in the vehicle.
Patent Information
- Application Number
- CN202211091712.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-07
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2042-09-07
AI Technical Summary
When inheriting commands across different voice zones in multi-user scenarios, users in different voice zones need to repeatedly wake up the voice assistant and cannot express themselves concisely, which makes it impossible to cater to all users in the vehicle and limits the freedom of expression.
By receiving voice requests from the vehicle cabin and combining the functional operations of the first voice request to complete the current voice request, and issuing commands for adjustment within a reasonable time range without having to repeatedly wake up the voice assistant, it supports intelligent response from users in different voice zones throughout the vehicle.
This allows users in different audio zones throughout the vehicle to speak without having to repeatedly wake up the voice assistant within a reasonable timeframe, giving users greater freedom of expression, conforming to the language habits of multi-person conversations, and improving the intelligent response capability of the vehicle's infotainment system.
Smart Images

Figure CN115527533B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vehicle technology, and in particular to a voice interaction method, a voice interaction device, a server, and a readable storage medium. Background Technology
[0002] In related technologies, when inheriting commands across voice zones in multi-user scenarios, users in different voice zones need to repeatedly wake up the voice assistant when issuing commands. Moreover, after being woken up, the assistant does not support the user's concise and abbreviated expressions, which makes it impossible to cater to the use of all users in the vehicle. Furthermore, the freedom of expression of users in different voice zones is severely limited. Summary of the Invention
[0003] This invention aims to at least solve one of the technical problems existing in the prior art. To this end, one objective of this invention is to propose a voice interaction method that enables users in different voice zones throughout the vehicle to adjust the same function with voice zone attributes within a reasonable time frame without repeatedly waking up the voice assistant, thus achieving intelligent response of the vehicle's infotainment system to users in multi-user scenarios.
[0004] According to the voice interaction method of the present invention, the method includes: receiving a current voice request from a user in a vehicle cabin; when it is determined that the number of rounds of the current voice request is less than a target value and the time interval from the first voice request is less than a target duration, supplementing the current voice request with the functional operation of the first voice request; wherein the first voice request is used to indicate a preset function; and when it is determined that the supplemented current voice request is unambiguous, issuing a control command to the vehicle so that the vehicle executes the operation corresponding to the supplemented current voice request according to the control command and performs voice interaction with the user.
[0005] According to the voice interaction method of the present invention, when users in different voice zones of the vehicle adjust the same function with voice zone attributes within a reasonable time range, they can directly issue commands for adjustment without repeatedly waking up the voice assistant, omitting the control name and voice zone information. Furthermore, it can complete the current voice request, giving users greater freedom of expression and conforming more to the language habits of multi-person dialogue with contextual reference, thus realizing the vehicle system's intelligent response to users in multi-person scenarios.
[0006] According to the voice interaction method of the present invention, the determination that the completed voice request is unambiguous includes any of the following cases: determining that the default audio region corresponding to the first voice request is all audio regions; determining that the content of the current voice request has a specified audio region; determining that the content of the current voice request does not have a specified audio region, and the sound source audio region of the current voice request is the same as the specified audio region in the content of the first voice request.
[0007] Therefore, several unambiguous scenarios for the completed current-round voice request are listed to facilitate the issuance of control commands to the vehicle, enabling the vehicle to perform the operation corresponding to the completed current-round voice request and interact with the user via voice.
[0008] According to the voice interaction method of the present invention, issuing control commands to the vehicle includes: if it is determined that the completed current voice request is executable, issuing control commands to the vehicle so that the vehicle performs an operation corresponding to the completed current voice request and interacts with the user via voice according to the control commands; if it is determined that the completed current voice request is not executable, issuing control commands to the vehicle so that the vehicle interacts with the user via voice according to the control commands.
[0009] Therefore, if the completed voice request is executable, it facilitates the execution of the corresponding operation and voice interaction with the user to meet the user's needs. If the completed voice request is not executable, it facilitates voice interaction between the control command and the user to remind the user, thereby improving driving safety.
[0010] According to the voice interaction method of the present invention, after the current voice request is completed by combining the function operation content of the first round of voice request, the method includes: if it is determined that the completed current voice request is ambiguous, determining a voice broadcast to indicate the situation; and issuing the voice broadcast.
[0011] Therefore, if it is determined that the completed voice request in this round is ambiguous, a voice broadcast to indicate the situation is determined and issued to confirm with the user a second time.
[0012] According to the voice interaction method of the present invention, determining the voice broadcast for indicating the situation includes: when the sound source region of the current voice request can perform the functional operation of the first voice request, obtaining a first broadcast based on the sound source region of the current voice request and the functional operation of the first voice request, wherein the first broadcast is used to indicate an unambiguous voice request.
[0013] Therefore, by issuing the first broadcast, an unambiguous voice request can be indicated to guide the user's speech, thereby educating the user while guiding them to complete the instruction.
[0014] According to the voice interaction method of the present invention, determining the voice broadcast for indicating the situation includes: when the sound source region of the current voice request cannot perform the function operation of the first voice request, obtaining a second broadcast based on the sound source region of the current voice request and the function operation of the first voice request, wherein the second broadcast is used to indicate that the sound source region of the current voice request cannot perform the function operation of the first voice request.
[0015] Thus, a second broadcast can be obtained to indicate that the sound source region of the current voice request cannot perform the function operation of the first voice request, so as to remind the user of the sound source region of the current voice request that the operation cannot be performed through the second broadcast.
[0016] According to the voice interaction method of the present invention, the method further includes: if it is determined that the voice request is casual conversation, it is not included in the round.
[0017] Therefore, if the voice request is determined to be casual conversation, it will be rejected and not included in the round, thus making the inheritance scope reasonable. This ensures that false calls are not caused by an excessively large effective scope, while also taking into account the actual use scenario to ensure the effectiveness of use in multi-person casual conversation.
[0018] The present invention also proposes a voice interaction device.
[0019] According to the voice interaction device of the present invention, applied to a vehicle, the vehicle including multiple audio output channels, the device includes: a receiving module for receiving a current voice request from a user in the vehicle cabin; and a processing module for supplementing the current voice request by combining the functional operation of the first voice request, provided that the number of rounds of the current voice request is less than a target value and the time interval from the first voice request is less than a target duration; wherein the first voice request is used to instruct a preset function sending module to issue a control command to the vehicle, provided that the supplemented current voice request is unambiguous, so that the vehicle executes the operation corresponding to the supplemented current voice request and interacts with the user via voice according to the control command.
[0020] According to the voice interaction device of the present invention, when users in different voice zones of the vehicle adjust the same function with voice zone attributes within a reasonable time range, there is no need to repeatedly wake up the voice assistant. The control name and voice zone information can be omitted and the command can be issued directly for adjustment. Moreover, the voice request in this round can be completed, so that the user has greater freedom of expression and is more in line with the language habits when there is context reference in multi-person dialogue. This realizes the intelligent response of the vehicle system to the user in multi-person scenarios.
[0021] The present invention also proposes a server.
[0022] According to the server of the present invention, the server includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, it implements any of the methods described above.
[0023] According to the server of the present invention, when users in different audio zones of the vehicle adjust the same function with audio zone attributes within a reasonable time range, there is no need to repeatedly wake up the voice assistant. The control name and audio zone information can be omitted and the command can be issued directly for adjustment. Moreover, the server can complete the current voice request, so that the user has greater freedom of expression and is more in line with the language habits of multi-person dialogue with context reference. This realizes the intelligent response of the vehicle system to the user in multi-person scenarios.
[0024] The present invention also proposes a non-volatile computer-readable storage medium for computer programs.
[0025] The non-volatile computer-readable storage medium of the computer program according to the present invention implements any of the methods described above when the computer program is executed by one or more processors.
[0026] The non-volatile computer-readable storage medium of the computer program according to the present invention enables users in different voice zones throughout the vehicle to adjust the same function with voice zone attributes within a reasonable time range without repeatedly waking up the voice assistant. The control name and voice zone information can be omitted to directly issue commands for adjustment, and the voice request can be completed in the current round, so that the user's expression has greater freedom and is more in line with the language habits of multi-person dialogue with context reference, realizing the vehicle system's intelligent response to users in multi-person scenarios.
[0027] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0028] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the voice interaction method in conjunction with the following drawings, wherein:
[0029] Figure 1 This is one of the flowcharts of the voice interaction method of the present invention;
[0030] Figure 2 This is the second flowchart of the voice interaction method of the present invention;
[0031] Figure 3 This is the third flowchart of the voice interaction method of the present invention;
[0032] Figure 4 This is the fourth flowchart of the voice interaction method of the present invention;
[0033] Figure 5 This is the fifth flowchart of the voice interaction method of the present invention;
[0034] Figure 6 This is the sixth flowchart of the voice interaction method of the present invention;
[0035] Figure 7 This is the seventh flowchart of the voice interaction method of the present invention;
[0036] Figure 8 This is the eighth flowchart of the voice interaction method of the present invention;
[0037] Figure 9 This is a schematic diagram showing the connection status of the receiving module, processing module and sending module of the voice interaction device of the present invention.
[0038] Figure 10 This is a schematic diagram illustrating the connection state between the non-volatile computer-readable storage medium and the processor in the interaction method of the present invention. Detailed Implementation
[0039] The technical solutions of this invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described voice interaction method is only a part of, not all, of this invention. All other methods obtained by those skilled in the art based on the methods of this invention without inventive effort are within the scope of protection of this invention.
[0040] The terms "first," "second," etc., used in the specification and claims of this invention are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the methods of the invention can be implemented in sequences other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0041] The following description, in conjunction with the accompanying drawings, details the voice interaction method, voice interaction device, server, and readable storage medium provided by the present invention through specific application scenarios.
[0042] The following is for reference. Figures 1-10 The voice interaction method according to the present invention is described.
[0043] like Figure 1 As shown, the voice interaction method according to the present invention includes:
[0044] Step 10: Receive the current voice request from the user in the vehicle cabin.
[0045] Step 20: If it is determined that the number of rounds of the current voice request is less than the target value, and the time interval from the first voice request is less than the target duration, complete the current voice request by combining the functional operation of the first voice request; wherein, the first voice request is used to indicate the preset function.
[0046] Step 30: If it is confirmed that the completed voice request is unambiguous, issue a control command to the vehicle so that the vehicle can perform the operation corresponding to the completed voice request and interact with the user via voice.
[0047] It should be noted that this interaction method is applied to a vehicle, meaning that the interaction method of this invention can be implemented by the vehicle of this invention. The vehicle includes a memory and a processor. The memory stores a computer program, and the processor receives the current round of voice requests from the user in the vehicle cabin. If it is determined that the number of rounds of the current voice request is less than a target value, and the time interval from the first voice request is less than a target duration, the processor completes the current voice request by combining the functional operation of the first voice request. The first voice request is used to indicate a preset function. If it is determined that the completed current voice request is unambiguous, a control command is issued to the vehicle so that the vehicle can execute the operation corresponding to the completed current voice request and interact with the user via voice.
[0048] With technological advancements, voice interaction has been increasingly widely applied across various fields. Currently, most technologies lack specific strategies for cross-voice region command inheritance in multi-user scenarios. Therefore, users in different voice regions need to repeatedly wake up the voice assistant when issuing commands, and even after waking up, concise, abbreviated expressions from the user are not supported. While some technologies support wake-up-free "I also want" expressions for the passenger voice region, this doesn't cater to all users in the vehicle, and the freedom of expression remains severely limited.
[0049] In this invention, the case of cross-voice region instruction inheritance in a multi-person scenario is addressed, that is, in a scenario where, after the first round of voice request instructs the preset function, a voice request from a user in another voice region is received for the current round (i.e., the next round or the following round).
[0050] It should be noted that "this round of voice request" in the above description refers to the voice request after the first round of voice request, that is, the current round of voice request is issued after the first round of voice request.
[0051] If the voice request in this round falls within the first round of cross-voice zone inheritance support, it will be included in the cross-voice zone integration module; otherwise, it will not be included. Here, a voice zone refers to a location within the vehicle's cabin, such as the driver's seat, passenger seat, and rear seats, which correspond to the driver's voice zone, passenger's voice zone, and rear seat voice zone, respectively. The scope of cross-voice zone inheritance support includes: air conditioning switch, temperature, seat ventilation, seat heating, screen brightness, windows, seat position, seat back, seat cushion, lumbar support, seat massage, seat rhythm, reading lights, etc.
[0052] The scope of cross-region inheritance support is divided according to functional type:
[0053] The default sound zones correspond one-to-one with the sound source sound zones: windows, seat position, seat back, seat cushion, lumbar support, seat ventilation, seat heating, seat massage, seat rhythm, and reading light.
[0054] The default sound zone does not correspond one-to-one with the sound source sound zone: screen brightness (passenger seat - passenger seat screen, driver seat + rear seats - driver seat screen);
[0055] The default audio range is all: air conditioner on / off and temperature.
[0056] Step 20 involves supplementing the current voice request with the functional operations of the first voice request, provided that the number of rounds of the current voice request is less than the target value and the time interval from the first voice request is less than the target duration. This allows the user of the current voice request to omit the control name and audio region information and directly issue commands for adjustment, giving the user greater freedom of expression and conforming more to the language habits of multi-person dialogue with contextual reference.
[0057] Understandably, referring to the context maintenance rules, it is necessary to determine whether subsequent voice requests in this round should enter the inheritance module. For example, if the current cross-voice region module has not timed out and the number of inheritance rounds is less than or equal to the target value (the maximum number of inheritance rounds for the cross-voice region module): if the number of rounds of the current voice request is less than the target value, and it is determined that the time interval from the first voice request is less than the target duration, then the number of inheritance rounds will be increased by one, and the current voice request will be completed by combining the functional operation of the first voice request.
[0058] It should be noted that the context maintenance rule can be the number of inheritable rounds for cross-sound region modules. For example, the number of inheritable rounds for cross-sound region modules is 4, that is, the target value is 4. The round counting rule is as follows: in addition to the rounds that are successfully inherited, other rule instructions and ambiguous rounds are also included in the inheritance rounds, but idle chat is not included in the inheritance rounds. The context storage duration is 30 seconds, that is, the time interval between the current round of voice request and the previous round of voice request is 30 seconds.
[0059] Therefore, by limiting the number of inheritance rounds and the effective time, and clarifying the content requirements for inclusion in the inheritance rounds, the scope of inheritance can be rationalized. This ensures that there are no false recalls due to an excessively large effective scope, while also taking into account actual usage scenarios to ensure the effectiveness of use in casual conversations among multiple people.
[0060] The wording for the succession round includes, but is not limited to, "I want it too," "Turn it down a bit," "Turn it up a bit," or "Turn it to the front." If the voice request in this round does not conform to the wording for the succession round, it will not be rewritten, and the operation and response will be consistent with issuing the voice request separately.
[0061] For the second-round strategy description of "I want it too": If the first round gives a complete instruction and the second round instruction is expressed as "I want it too", then the same operation as the first round is performed on the same function in the XX voice zone, or the operation is performed according to the description.
[0062] Note: Generalized expressions of "I want one too" include:
[0063] The first type: XX also / also: XX includes "I / my / here / here / driver's seat / passenger's seat / front row / ...", where "I" is equivalent to the designated vocal range.
[0064] The second type: plus one: including common expressions such as "plus one / plus two".
[0065] The third type: also + action: mainly refers to the omission of pitch range and function, such as "also raise the pitch a little".
[0066] Subsequent inheritance explanation: When the "I want it too" condition is met in the next round:
[0067] If the third round of instructions is in a different pitch range than "I want to" it will skip the pitch range of "I want to" and the pitch range of the previous round to make ambiguous judgments.
[0068] If the third round is in the same pitch range as (I also want), then the corresponding operation is performed based on the pitch range of (I also want).
[0069] For example, when the registers of "I want it too" in the subsequent rounds and the second round are different:
[0070] (First round) Driver: "Open the car window."
[0071] Action+TTS: "The driver's side window is opening."
[0072] (Second round) Co-pilot: "Me too."
[0073] Action+TTS: "The passenger window is opening."
[0074] (Third round) Driver: "Drive halfway."
[0075] Action+TTS: "The driver's side window is halfway open."
[0076] Or, if the register of "I want it too" is the same in the subsequent round and the second round:
[0077] (First round) Driver: "Open the car window."
[0078] Action+TTS: "The driver's side window is opening."
[0079] (Second round) Co-pilot: "Add one."
[0080] Action+TTS: "The passenger window is opening."
[0081] (Third round) Co-pilot: "Turn it off."
[0082] Action+TTS: "The passenger window is closing."
[0083] It should be noted that the above description of the inheritable rounds of cross-region modules and the duration of context storage is for illustrative purposes only and does not represent any limitation on this.
[0084] In this process, the function operation of the first round of voice request is combined to complete the current round of voice request. For example, if the current round of voice request is "open the car window", the voice region is omitted, or if the current round of voice request is "I also open", both the voice region and the control name are omitted. Therefore, for functions with cross-voice region attributes, in order to avoid adjusting functions that are inconsistent with the user's target voice region due to accidental operation, it is necessary to introduce an ambiguity strategy.
[0085] Next, step 30: If it is confirmed that the completed voice request is unambiguous, a control command is issued to the vehicle so that the vehicle can perform the operation corresponding to the completed voice request and interact with the user via voice.
[0086] Definition of ambiguity:
[0087] The default audio range is all functions (air conditioner on / off, temperature): When the audio range is not specified in this round of voice requests, regardless of whether it is consistent with the audio range in the first round, no ambiguity is judged, and the default is all audio ranges and executed directly.
[0088] Other (default voice zone corresponds one-to-one with the source voice zone, default voice zone does not correspond one-to-one with the source voice zone): When the voice request in this round does not specify a voice zone and the voice zone specified in the first round is inconsistent with that specified in this round, ambiguity arises. The operation is confirmed through ambiguous statements and no operation is performed.
[0089] Ambiguous statements:
[0090] Ambiguous wording regarding the executable nature of the command: If you want to use $(intentName, corresponding function operation name), you can simply say $(intentName, corresponding function operation name);
[0091] Ambiguous wording regarding the instruction not being executable: Do you mean $(intentName, corresponding function operation name)? (The corresponding one does not support TTS).
[0092] Therefore, by introducing an ambiguity strategy, it is possible to avoid adjusting functions that are inconsistent with the user's target audio range due to accidental operation.
[0093] According to the voice interaction method of the present invention, when users in different voice zones of the vehicle adjust the same function with voice zone attributes within a reasonable time range, they can directly issue commands for adjustment without repeatedly waking up the voice assistant, omitting the control name and voice zone information. Furthermore, it can complete the current voice request, giving users greater freedom of expression and conforming more to the language habits of multi-person dialogue with contextual reference, thus realizing the vehicle system's intelligent response to users in multi-person scenarios.
[0094] Step 30 determines that the completed voice request for this round is unambiguous, including any of the following cases:
[0095] The default audio region for the first round of voice requests is set to all audio regions.
[0096] For example, if the default audio range for the first round of voice requests is determined to be all audio ranges (such as air conditioner on / off, temperature, etc.), and the default audio range for the second round is also all audio ranges:
[0097] (First round) Driver: "Turn on the driver's seat air conditioning."
[0098] Action+TTS: "The driver's air conditioning is on."
[0099] (Next round) Co-pilot: "Turn it off."
[0100] Action+TTS: "The air conditioner is off."
[0101] In other words, the default voice zone is all functions (such as air conditioner on / off and temperature): when the voice request in this round does not specify the voice zone, regardless of whether it is consistent with the voice zone in the first round, it will not be judged as ambiguous, and will default to all voice zones and be executed directly.
[0102] The content of this round of voice requests is determined to have a specified audio range.
[0103] In other words, when a specific audio range is requested in the current round of voice, no ambiguity is considered; the specified audio range is assumed to be the one requested in the current round of voice and the request is executed directly.
[0104] It is determined that the content of this round of voice requests does not have the specified audio region, and the audio region of the sound source in this round of voice requests is the same as the specified audio region in the content of the first round of voice requests.
[0105] In other words, if no audio region is specified in the current voice request, and the audio region of the sound source in the current voice request is the same as the specified audio region in the content of the first voice request, no ambiguity is determined, and the specified audio region in the content of the first voice request is assumed and executed directly.
[0106] For example, if the sound source region of this round of voice request is the same as the specified region in the content of the first round of voice request:
[0107] (First round) Driver: "Adjust the passenger seat forward."
[0108] Action+TTS: "The passenger seat is being adjusted forward."
[0109] (Second round) Co-pilot: "Move to the front."
[0110] Action+TTS: "Adjust the passenger seat to the furthest forward position."
[0111] Therefore, several unambiguous scenarios for the completed current-round voice request are listed to facilitate the issuance of control commands to the vehicle, enabling the vehicle to perform the operation corresponding to the completed current-round voice request and interact with the user via voice.
[0112] See Figure 2 Step 30 involves issuing control commands to the vehicle, including:
[0113] Step 31: If it is determined that the completed voice request for this round is executable, a control command is issued to the vehicle so that the vehicle can perform the operation corresponding to the completed voice request for this round and interact with the user via voice according to the control command.
[0114] In other words, if it is determined that the completed voice request is unambiguous, it is then determined whether the completed voice request is executable. If it is executable, the vehicle executes the operation corresponding to the completed voice request according to the control command and interacts with the user via voice.
[0115] For example:
[0116] (Initial voice request) Driver: "Open the window";
[0117] Action+TTS: "The driver's side window is opening."
[0118] (Second round, i.e., this round of voice request) Co-pilot: "Co-pilot, drive to 60%."
[0119] Action+TTS: "The passenger window is being opened to 60%."
[0120] Therefore, if the completed voice request is executable, it is convenient to execute the corresponding operation and interact with the user via voice to meet the user's needs.
[0121] See Figure 3 Step 30 involves issuing control commands to the vehicle, including:
[0122] Step 32: If it is determined that the completed voice request for this round is not executable, a control command is issued to the vehicle so that the vehicle can interact with the user via voice according to the control command.
[0123] In other words, if it is determined that the completed voice request is unambiguous, it is then determined whether the completed voice request is executable. If it is not executable, the vehicle will interact with the user via voice according to the control instructions.
[0124] For example:
[0125] (First round) Co-pilot: "Adjust the seat forward."
[0126] Action+TTS: "The passenger seat is being adjusted forward."
[0127] (Next round, i.e., this round of voice request) Driver: "I'll move this a bit back."
[0128] TTS: "Voice adjustment of the driver's seat position is not supported while driving. Please adjust it manually."
[0129] Therefore, if the completed voice request is not executable, it facilitates voice interaction between the control command and the user to remind the user, which helps improve safety during driving.
[0130] See Figure 4 and Figure 5 Step 20 includes: After completing the current round of voice requests by combining the content of the first round of voice request function operations, the method includes:
[0131] Step 21: If it is determined that the completed voice request in this round is ambiguous, determine the voice broadcast used to indicate the situation;
[0132] Step 22: Issue a voice broadcast.
[0133] In other words, if it is determined that there is ambiguity in the current round of voice requests after completion, a voice broadcast to indicate the situation is determined and issued to obtain secondary confirmation from the user.
[0134] Further, see Figure 6Step 20 determines the voice announcement used to indicate the situation, including:
[0135] Step 211: If the sound source region of the current voice request can perform the function operation of the first voice request, a first broadcast is obtained based on the sound source region of the current voice request and the function operation of the first voice request. The first broadcast is used to indicate an unambiguous voice request.
[0136] In other words, if it is determined that the current round of voice requests is ambiguous after completion, the function operation of the first round of voice requests is substituted into the sound source area of the current round of voice requests. If it can be executed, the first broadcast is obtained to indicate the unambiguous voice requests, so as to remind the user of the sound source area of the current round of voice requests.
[0137] For example:
[0138] Example 1: When the pitch range of the subsequent rounds is the same as that of the first round:
[0139] (First round) Driver's seat: "Turn on the driver's seat heater."
[0140] Action+TTS: "Driver's seat heating is on."
[0141] (Second round) Co-pilot: "Turn to the highest gear."
[0142] (First broadcast) TTS: "If you want to (adjust the passenger seat heating), you can just say (adjust the passenger seat heating)."
[0143] (Third round) Driver: "Shift to second gear."
[0144] Action+TTS: "Driver's seat heating is now set to level two."
[0145] Example 2: When the pitch range of the subsequent rounds differs from that of the first round:
[0146] (First round) Driver: "Open the car window."
[0147] Action+TTS: "The driver's side window is opening."
[0148] (Next round) Co-pilot: "Turn it off."
[0149] (First broadcast) TTS: "If you want to close the passenger window, you can just say (close the passenger window)."
[0150] (Third round) Left rear: "Open halfway."
[0151] TTS: "If you want to (open the left rear window), you can just say (open the left rear window)."
[0152] Therefore, the first broadcast can indicate an unambiguous voice request to guide the user's speech, thus educating the user while guiding them to complete the instruction.
[0153] See Figure 7 Step 20 determines the voice announcement used to indicate the situation, including:
[0154] Step 212: If the sound source region of the current voice request cannot perform the function operation of the first voice request, a second broadcast is obtained based on the sound source region of the current voice request and the function operation of the first voice request. The second broadcast is used to indicate that the sound source region of the current voice request cannot perform the function operation of the first voice request.
[0155] In other words, if it is determined that the current round of voice requests is ambiguous after completion, the function operation of the first round of voice requests is substituted into the sound source area of the current round of voice requests. If it cannot be executed, then based on the sound source area of the current round of voice requests and the function operation of the first round of voice requests, a second broadcast is obtained to indicate that the sound source area of the current round of voice requests cannot execute the function operation of the first round of voice requests, so as to remind the user of the sound source area of the current round of voice requests that the operation cannot be executed.
[0156] For example: (First round) Driver: "Turn on the seat ventilation."
[0157] Action+TTS: "Driver's seat ventilation is on."
[0158] (Second round, i.e., the second round of voice request) Left back: "Turn to the highest setting".
[0159] (Second broadcast) TTS: "Do you want to (adjust the left rear seat ventilation)? The rear seats do not support ventilation function."
[0160] (Third round) Driver: "Shift to second gear."
[0161] Action+TTS: "Driver's seat ventilation is now set to level two."
[0162] Therefore, a second announcement is made to remind users in the audio range of the voice source request that the operation is not feasible.
[0163] The method also includes:
[0164] If the voice request is determined to be casual conversation, it will not be counted in the round.
[0165] In other words, if the voice request is determined to be casual conversation, it will be rejected and not included in the round, thereby making the inheritance scope reasonable. This ensures that false calls are not caused by an excessively large effective scope, while also taking into account the actual use scenario to ensure the effectiveness of use in multi-person casual conversation.
[0166] The following is combined with Figure 8 A specific implementation of the voice interaction method of the present invention is described below:
[0167] First, the first sentence is entered into the cross-regional inheritance module, and the timer t starts. The number of inheritance rounds n=2. Then, a new command is input, which is the voice request for this round.
[0168] If the new instruction does not satisfy: t≤timeout (the time interval between the current round of voice request and the first round of voice request is less than the target duration), that is, n≤4 (the number of rounds of voice request is less than the target value), then the current round of processing ends.
[0169] If the new instruction satisfies: t≤timeout (the time interval between the current voice request and the first voice request is less than the target duration), that is, n≤4 (the number of rounds of the current voice request is less than the target value), then it is determined whether the new instruction is chat.
[0170] If the new instruction is determined to be casual conversation, it is rejected and the current round of processing ends.
[0171] If the new instruction is determined not to be idle chat, the inheritance round is increased to n = n + 1, and then it is determined whether the new instruction is an inheritable instruction.
[0172] If the new instruction is determined not to be an inheritable instruction, then it is determined whether the instruction is executable. If it is executable, then the Action + corresponding TTS is executed, and then the current round of processing ends.
[0173] If the new instruction is determined to be an inheritable instruction, then it is determined whether the instruction is ambiguous. If the instruction is determined to be ambiguous, then the ambiguity is broadcast and the current round of processing ends.
[0174] If the instruction is determined to be unambiguous, then determine whether the instruction is executable. If the instruction is determined to be unexecutable, then execute the corresponding exception reporting text message (TTS) and then end the current round of processing.
[0175] If the instruction is determined to be executable, then execute the Action + corresponding TTS, and then end the current round of processing.
[0176] The present invention also proposes a voice interaction device.
[0177] According to the present invention, a voice interaction device is applied to a vehicle, the vehicle including multiple audio output channels, such as... Figure 9 As shown, the device includes: a receiving module 301, a processing module 302, and a transmitting module 303.
[0178] Specifically, the receiving module 301 is used to receive the current round of voice requests from the user in the vehicle cabin; the processing module 302 is used to complete the current round of voice requests by combining the functional operation of the first round of voice requests when it is determined that the number of rounds of the current round of voice requests is less than a target value and the time interval from the first round of voice requests is less than a target duration; wherein, the first round of voice requests is used to indicate a preset function; the sending module 303 is used to issue control commands to the vehicle when it is determined that the completed current round of voice requests is unambiguous, so that the vehicle can perform the operation corresponding to the completed current round of voice requests according to the control commands and conduct voice interaction with the user.
[0179] According to the voice interaction device of the present invention, when users in different voice zones of the vehicle adjust the same function with voice zone attributes within a reasonable time range, there is no need to repeatedly wake up the voice assistant. The control name and voice zone information can be omitted and the command can be issued directly for adjustment. Moreover, the voice request in this round can be completed, so that the user has greater freedom of expression and is more in line with the language habits when there is context reference in multi-person dialogue. This realizes the intelligent response of the vehicle system to the user in multi-person scenarios.
[0180] In this invention, when the processing module 302 determines that the current round of voice request after completion is unambiguous, it includes any of the following situations: determining that the default audio region corresponding to the first round of voice request is all audio regions; determining that the content of the current round of voice request has a specified audio region; determining that the content of the current round of voice request does not have a specified audio region, and the audio region of the sound source of the current round of voice request is the same as the specified audio region in the content of the first round of voice request.
[0181] In this invention, the processing module 302 is further configured to issue control commands to the vehicle, including: if it is determined that the completed current voice request is executable, issuing control commands to the vehicle so that the vehicle can perform the operation corresponding to the completed current voice request and interact with the user via voice according to the control commands; if it is determined that the completed current voice request is not executable, issuing control commands to the vehicle so that the vehicle can interact with the user via voice according to the control commands.
[0182] In this invention, the processing module 302 is further configured to determine a voice broadcast to indicate the situation if it is determined that the current round of voice request after completion is ambiguous; and to issue the voice broadcast.
[0183] In this invention, the processing module 302 is further configured to, when the sound source region of the current voice request can perform the functional operation of the first voice request, obtain a first broadcast based on the sound source region of the current voice request and the functional operation of the first voice request, wherein the first broadcast is used to indicate an unambiguous voice request.
[0184] In this invention, the processing module 302 is further configured to, when the sound source region of the current voice request cannot perform the function operation of the first voice request, obtain a second broadcast based on the sound source region of the current voice request and the function operation of the first voice request, wherein the second broadcast is used to indicate that the sound source region of the current voice request cannot perform the function operation of the first voice request.
[0185] In this invention, the processing module 302 is also used to exclude the voice request from the round if it is determined that the voice request is casual conversation.
[0186] The voice interaction device in this invention can be a device, or a component, integrated circuit, or chip in a mobile terminal. For example, the mobile terminal can be a mobile phone, tablet computer, laptop computer, PDA, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc.
[0187] The voice interaction device in this invention can be a device with an operating system. This operating system can be QNX, Linux, Android, WinCE, or other possible operating systems; this invention does not impose specific limitations.
[0188] The voice interaction device provided by this invention can achieve Figures 1 to 8 The various processes in the method can achieve the same technical effect, and to avoid repetition, they will not be described in detail here.
[0189] The present invention also proposes a server.
[0190] According to the server of the present invention, the server includes a memory and a processor 200, such as Figure 10 As shown, the memory stores a computer program 101. When the computer program 101 is executed by the processor 200, it implements any of the above methods and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0191] The present invention also proposes a non-volatile computer-readable storage medium 100 for computer programs.
[0192] The non-volatile computer-readable storage medium 100 of the present invention implements any of the above methods when the computer program 101 is executed by one or more processors 200, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0193] The processor 200 is the processor 200 in the electronic device described above. The readable storage medium includes the computer-readable storage medium 100, such as a computer read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.
[0194] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of the present invention is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0195] Through the above description of the embodiments, those skilled in the art can clearly understand that the above methods can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the present invention.
[0196] The method of the present invention has been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other modifications under the guidance of the present invention without departing from the spirit and scope of the claims, and all of these modifications are within the protection scope of the present invention.
Claims
1. A voice interaction method, characterized in that, include: Receive the current voice request from the user in the vehicle cabin; If it is determined that the number of rounds of the current voice request is less than the target value, and the time interval from the first voice request is less than the target duration, the control names and / or audio region information omitted in the current voice request are supplemented by combining the functional operation of the first voice request; wherein, the first voice request is used to indicate a preset function; the round counting rule is: in addition to the rounds that are successfully inherited, other rule instructions and ambiguous rounds are also included in the inheritance rounds, and idle chat is not included in the inheritance rounds; If it is determined that the completed voice request is unambiguous, a control command is issued to the vehicle so that the vehicle can perform the operation corresponding to the completed voice request and interact with the user via voice according to the control command; the determination that the completed voice request is unambiguous includes any of the following situations: The default audio region corresponding to the first round of voice requests is determined to be all audio regions; when the audio region is not specified in the current round of voice requests, it is not considered ambiguous regardless of whether it is consistent with the audio region of the first round. It is determined that the content of the current voice request has a specified audio range; It is determined that the content of the current round of voice requests does not have a specified audio region, and the audio region of the sound source of the current round of voice requests is the same as the specified audio region in the content of the first round of voice requests.
2. The voice interaction method according to claim 1, characterized in that, The process of issuing control commands to the vehicle includes: If it is determined that the completed current voice request is executable, a first control command is issued to the vehicle so that the vehicle can perform the operation corresponding to the completed current voice request and interact with the user via voice according to the first control command. If it is determined that the completed voice request is not executable, a second control command is issued to the vehicle so that the vehicle can interact with the user via voice according to the second control command.
3. The voice interaction method according to any one of claims 1-2, characterized in that, After supplementing the omitted control names and / or vocal range information in the current voice request by combining the functional operation content of the first round of voice request, the method includes: When the content of the current round of voice request does not specify a sound region and the specified sound region in the content of the first round of voice request is inconsistent with the sound source sound region of the current round of voice request, ambiguity arises. If it is determined that the current round of voice request after completion is ambiguous, a voice broadcast to indicate the situation is determined and the voice broadcast is issued. The determination of the voice broadcast used to indicate the situation includes: If the sound source region of the current voice request can perform the function operation of the first voice request, a first broadcast is obtained based on the sound source region of the current voice request and the function operation of the first voice request. The first broadcast is used to indicate an unambiguous voice request. If the sound source region of the current voice request cannot perform the function operation of the first voice request, a second broadcast is obtained based on the sound source region of the current voice request and the function operation of the first voice request. The second broadcast is used to indicate that the sound source region of the current voice request cannot perform the function operation of the first voice request.
4. A voice interaction device, applied to a vehicle, characterized in that, The vehicle includes multiple audio output channels, and the device includes: The receiving module is used to receive the current voice request from the user in the vehicle cabin; The processing module is used to, when it is determined that the number of rounds of the current voice request is less than a target value, and the time interval from the first voice request is less than a target duration, supplement the control names and / or audio region information omitted in the current voice request by combining the functional operation of the first voice request; wherein, the first voice request is used to indicate a preset function; the round counting rule is: in addition to the rounds successfully inherited, other rule instructions and ambiguous rounds are also included in the inheritance rounds, and idle chat is not included in the inheritance rounds; The sending module is used to issue control commands to the vehicle when it is determined that the completed current voice request is unambiguous, so that the vehicle can perform the operation corresponding to the completed current voice request and interact with the user via voice according to the control commands; determining that the completed current voice request is unambiguous includes any of the following situations: determining that the default audio region corresponding to the first voice request is all audio regions; when the current voice request does not specify an audio region, it is not considered ambiguous regardless of whether it is consistent with the audio region of the first voice request; determining that the content of the current voice request has a specified audio region; determining that the content of the current voice request does not have a specified audio region, and the audio region of the sound source of the current voice request is the same as the specified audio region in the content of the first voice request.
5. A server, characterized in that, The server includes a memory and a processor, the memory storing a computer program that, when executed by the processor, implements the method according to any one of claims 1-3.
6. A non-volatile computer-readable storage medium for a computer program, characterized in that, When the computer program is executed by one or more processors, it implements the method according to any one of claims 1-3.
Citation Information
Patent Citations
Response method in man-machine conversation, conversation system and storage medium
CN111737411A
Control method and device, vehicle-mounted terminal, vehicle and storage medium
CN113990318A