Voice processing method, voice interaction method, server, and storage medium
By dividing the vehicle cabin into sound zones and updating the rejection mode, and combining the label information of voice requests, the accuracy problem of rejection in multi-sound zone voice interaction is solved, achieving efficient voice request recognition and rejection, and improving the user experience.
Patent Information
- Application Number
- CN202211255729.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-13
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2042-10-13
AI Technical Summary
Existing technologies cannot effectively recognize and reject voice requests from multiple areas within the vehicle cabin, resulting in poor accuracy of voice interaction and impacting user experience.
The vehicle cabin is divided into multiple sound zones. The rejection mode is updated based on the information from the wake-up sound zone and the dialogue sound zone. The rejection process is carried out by combining the speaker's label and intent classification label of the voice request, so as to ensure accurate recognition and rejection of voice requests.
In multi-zone interaction scenarios, the accuracy of voice request rejection is improved, enhancing the user experience and ensuring the efficiency and accuracy of voice interaction.
Smart Images

Figure CN115503639B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice technology, and in particular to a voice processing method, a voice interaction method, a server, and a computer-readable storage medium. Background Technology
[0002] With the development of autonomous driving technology, vehicles can support voice control services, such as voice control for opening windows. In real-world driving scenarios, users may issue voice commands from multiple audio zones within the vehicle, and not all of these commands are requests to the in-vehicle system. This requires the in-vehicle voice processor to reject irrelevant information from all voice commands, extract the voice requests specific to the user, and respond accordingly.
[0003] In related technologies, the rejection processing of voice requests can usually only be applied to single-voice scenarios. By combining current text information, automatic speech recognition technology, and confidence-based speech features, irrelevant voice requests can be rejected in single-voice scenarios, which cannot meet the needs of multi-voice interaction in vehicles. Summary of the Invention
[0004] This invention provides a voice processing method, a voice interaction method, a server, and a computer-readable storage medium.
[0005] The speech processing method of the present invention includes:
[0006] The vehicle receives the wake-up sound zone information forwarded by the user to activate the vehicle's voice function within the vehicle's cabin.
[0007] The initial rejection mode of each of the multiple sound zones in the vehicle cabin is determined based on the wake-up sound zone information.
[0008] Receive the user's voice request forwarded by the vehicle after the vehicle's voice function is activated, as well as the dialogue zone information confirmed according to the user's voice request;
[0009] The rejection mode of the corresponding voice region is updated according to the user's voice request and the dialogue voice region information to determine the rejection mode of each voice region.
[0010] Thus, in this invention, the vehicle cabin is divided into multiple audio zones. Upon receiving a voice request, the rejection mode for each audio zone is determined based on the voice request and its context, thereby meeting the rejection requirements for multi-zone voice interaction within the vehicle cabin. Furthermore, as the voice interaction progresses, the rejection modes for each audio zone are updated, resulting in high accuracy in voice request rejection and a superior user experience in multi-zone interaction scenarios.
[0011] The step of determining the initial rejection mode for each of the multiple sound zones in the vehicle cabin based on the wake-up sound zone information includes:
[0012] Based on the wake-up sound zone information, the initial rejection mode of the wake-up sound zone in the vehicle cabin is determined to be the first rejection mode;
[0013] The initial rejection mode of each sound zone in the vehicle cabin, excluding the wake-up sound zone, is determined to be the second rejection mode. The second rejection mode has a higher rejection level for voice requests than the first rejection mode.
[0014] In this way, the initial rejection mode of each sound zone can be determined based on the wake-up sound zone information. Specifically, the initial rejection mode of the wake-up sound zone is the first rejection mode, and the initial rejection mode of the non-wake-up sound zone is the second rejection mode with a higher degree of rejection.
[0015] The step of updating the rejection mode of the corresponding voice region based on the user's voice request and the dialogue voice region information to determine the rejection mode of each voice region includes:
[0016] If, based on the dialogue voice area information, it is confirmed that the rejection mode of the dialogue voice area is the first rejection mode and the user's voice request is a non-vehicle interaction voice request, then the rejection mode of the dialogue voice area is updated to the second rejection mode.
[0017] Thus, if during the interaction, a certain dialogue voice zone is in the first rejection mode, and the voice request in that voice zone is not a vehicle interaction voice request, then it can be considered that the voice zone has no real interaction intention for the time being, and the rejection mode of the voice zone is updated to the second rejection mode.
[0018] The step of updating the rejection mode of the corresponding voice region based on the user's voice request and the dialogue voice region information to determine the rejection mode of each voice region includes:
[0019] If the vehicle cabin denial mode is the first denial mode and the corresponding audio region does not receive a valid voice request within a first preset time period, then the denial mode of the corresponding audio region will be updated to the second denial mode.
[0020] Thus, if during the interaction, a certain dialogue voice region is in the first rejection mode, but does not receive a valid voice request within a preset time, it can be assumed that the voice region has no real interaction intention for the time being, and the rejection mode of the voice region is updated to the second rejection mode.
[0021] The step of updating the rejection mode of the corresponding voice region based on the user's voice request and the dialogue voice region information to determine the rejection mode of each voice region includes:
[0022] If the rejection mode of the dialogue voice area is confirmed to be the second rejection mode based on the dialogue voice area information, and if it is determined based on the user's voice request that a valid voice request is executed within a second preset time period in the dialogue voice area, then the rejection mode of the dialogue voice area is updated to the first rejection mode.
[0023] Thus, if during the interaction, a certain dialogue voice region is in the second rejection mode, but receives a valid voice request within a preset time, then it can be considered that the voice region has a genuine interaction intention, and the rejection mode of the voice region can be updated to the first rejection mode, that is, a rejection mode with a lower degree of rejection.
[0024] The speech processing method includes:
[0025] If no user voice request is received within a third preset time period after the vehicle voice function is activated, the vehicle voice function will be deactivated.
[0026] In this way, if no user in the cabin makes any voice request within a preset time, the vehicle's voice function will be temporarily deactivated and will wait for the next activation.
[0027] The method further includes:
[0028] Process the user's voice request to determine the speaking target label and intent classification label of the user's voice request;
[0029] The voice request is processed based on the rejection mode of the dialogue voice region, the speaker's label, and the intent classification label to obtain the rejection result.
[0030] In this way, the user's voice request is labeled by the speaking object label and the intent classification label. Then, combined with the rejection pattern of the voice request in the voice region, the rejection result of the voice request is determined, that is, whether it is clear and recallable or used as noise filtering.
[0031] The process of processing the voice request based on the rejection pattern of the dialogue voice region, the speaker label, and the intent classification label to obtain the rejection result includes:
[0032] When the rejection mode in the dialogue voice area is the first rejection mode, if the speaking object label is a voice assistant type label and the intent level label is a first-level label or a second-level label, then the rejection result obtained by processing the user's voice request is a clear result.
[0033] If the speaking target label is a non-voice assistant label and the intent classification label is a third-level label, then processing the user's voice request to obtain the rejection result is a noise result. The intent classification label characterizes the validity of the user's voice request, wherein the first-level label is greater than the second-level label and the second-level label is greater than the third-level label.
[0034] Thus, in the first rejection mode, for voice requests where the speaker's label is a voice assistant and the intent classification label is a first-level label or a second-level label, the rejection result is confirmed as a clear result; for voice requests where the speaker's label is not a voice assistant and the intent classification label is a third-level label, the rejection result is confirmed as a noise result.
[0035] The step of processing the voice request based on the rejection pattern, the speaker label, and the intent classification label to obtain the rejection result includes:
[0036] When the rejection mode in the dialogue voice area is the second rejection mode, if the speaking object label is a voice assistant type label and the intent classification label is a first-level label, then the rejection result obtained by processing the user's voice request is a clear result.
[0037] If the label of the speaking object is a non-voice assistant label and the label of the intent classification is a second-level label or a third-level label, then the rejection result obtained by processing the user's voice request is a noise result.
[0038] Thus, in the second rejection mode, for voice requests where the speaker's label is a voice assistant and the intent classification label is a first-level label, the rejection result is confirmed as a clear result; for voice requests where the speaker's label is not a voice assistant and the intent classification label is a second- or third-level label, the rejection result is confirmed as a noisy result. Compared to the first rejection mode, the second rejection mode is more stringent in rejecting labels with an intent classification label of level two.
[0039] The voice interaction method of the present invention includes:
[0040] The vehicle receives the wake-up sound zone information forwarded by the user to activate the vehicle's voice function within the vehicle's cabin.
[0041] The initial rejection mode of each of the multiple sound zones in the vehicle cabin is determined based on the wake-up sound zone information.
[0042] Receive the user voice request forwarded by the vehicle after the vehicle's voice function is activated, and the dialogue zone information confirmed according to the user voice request;
[0043] The rejection mode of the corresponding voice region is updated according to the user's voice request and the dialogue voice region information to determine the rejection mode of each voice region;
[0044] After determining the rejection mode for each of the aforementioned voice regions, the user's voice request is processed to obtain the speaking object label and the intent classification label;
[0045] The voice request is processed according to the rejection mode, the speaker label, and the intent classification label to obtain the rejection result;
[0046] The rejection result is sent to the vehicle to complete the voice interaction.
[0047] In this way, the vehicle cabin is divided into multiple sound zones. Upon receiving a voice request, the rejection mode for each sound zone is determined based on the request and its context, thus meeting the rejection requirements for multi-sound zone voice interaction within the vehicle cabin. Furthermore, as the voice interaction progresses, the rejection modes for each sound zone are updated, resulting in high accuracy in voice request rejection and a superior user experience in multi-sound zone interaction scenarios.
[0048] The server of the present invention includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, it implements the above-described method.
[0049] The present invention provides a computer-readable storage medium storing a computer program that, when executed by one or more processors, implements the above-described method.
[0050] Additional aspects and advantages of embodiments of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of embodiments of the invention. Attached Figure Description
[0051] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:
[0052] Figure 1 This is one of the flowcharts illustrating the speech processing method of the present invention;
[0053] Figure 2 This is a schematic diagram of the vehicle cabin of the present invention;
[0054] Figure 3 This is one of the state diagrams of the speech processing method of the present invention;
[0055] Figure 4 This is the second schematic diagram of the state of the speech processing method of the present invention;
[0056] Figure 5 This is the third schematic diagram of the state of the speech processing method of the present invention;
[0057] Figure 6 This is the fourth schematic diagram of the state of the speech processing method of the present invention;
[0058] Figure 7 This is the fifth schematic diagram of the state of the speech processing method of the present invention;
[0059] Figure 8 This is the sixth schematic diagram of the state of the speech processing method of the present invention;
[0060] Figure 9 This is the second flowchart of the speech processing method of the present invention;
[0061] Figure 10 This is the seventh schematic diagram of the state of the speech processing method of the present invention;
[0062] Figure 11 This is the eighth schematic diagram of the state of the speech processing method of the present invention;
[0063] Figure 12 This is a flowchart illustrating the voice interaction method of the present invention. Detailed Implementation
[0064] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of the present invention, and should not be construed as limiting the embodiments of the present invention.
[0065] Please see Figure 1 This invention provides a speech processing method, comprising:
[0066] 01: Receive the wake-up sound zone information forwarded by the vehicle when the user activates the vehicle's voice function inside the vehicle cabin;
[0067] 02: Determine the initial rejection mode for each of the multiple sound zones in the vehicle cabin based on the wake-up sound zone information;
[0068] 03: Receive user voice requests forwarded by the vehicle after the vehicle's voice function is activated, as well as the dialogue zone information confirmed based on the user's voice request;
[0069] 04: Update the rejection mode of the corresponding voice region based on the user's voice request and dialogue voice region information to determine the rejection mode of each voice region.
[0070] The present invention also provides a server, which includes a memory and a processor. The voice processing method of the present invention can be implemented by the server of the present invention. Specifically, the memory stores a computer program, and the processor is used to receive wake-up sound zone information of a user waking up the vehicle's voice function in the vehicle cabin, forwarded by the vehicle; to determine the initial rejection mode of each sound zone among multiple sound zones in the vehicle cabin based on the wake-up sound zone information; to receive user voice requests forwarded by the vehicle after the vehicle's voice function is woken up, as well as dialogue sound zone information confirmed based on the user's voice requests; and to update the rejection mode of the corresponding sound zone based on the user's voice requests and dialogue sound zone information to determine the rejection mode of each sound zone.
[0071] Specifically, in-vehicle voice assistants offer numerous conveniences to users in the cabin, allowing them to control software or vehicle components through voice interaction. For ease of interaction, voice assistants support continuous dialogue; that is, after a single activation, the user and voice assistant can engage in multi-turn conversations similar to natural language communication until the dialogue ends, without needing to activate the assistant every time they interact. However, to ensure vehicle safety, some related technologies only grant voice interaction permissions to the driver. Only the driver can perform voice interactions within the cabin, and other users can only request these functions from the driver. This can potentially distract the driver and affect driving safety. If permissions are granted to all users in the cabin, allowing all users to converse after the voice assistant is activated, the voice assistant may receive conversations from different users and between different users, as the in-car space is a shared environment. How to accurately process the received voice requests without limiting the interaction environment, and determine which voice requests need to be responded to, in order to better serve the user, will determine the user experience of voice interaction.
[0072] It's understandable that in multi-voice-zone continuous dialogue scenarios—that is, scenarios where, after the voice assistant is activated, users in different locations within the cabin can engage in multiple rounds of dialogue with it—multiple users may interact with the same topic with a high degree of freedom. Some of these interactions may be with the voice assistant, while others may be between users, making it more complex than a single-voice-zone scenario.
[0073] Activating the vehicle's voice function means activating the vehicle's voice assistant. The activation voice request can be set by the manufacturer or a user-defined wake-up word. After the voice assistant is activated, the user in the cabin can engage in multiple rounds of dialogue with it. The dialogue ends when the set threshold of rounds is reached, or when no voice request is received from the user within a predetermined time.
[0074] The cockpit is divided into different sound zones based on the areas where users might make noise. Please refer to [link / reference]. Figure 2 Taking a five-seat vehicle 100 as an example, the vehicle cabin can be divided into five audio zones, including the driver's seat audio zone 101, the passenger seat audio zone 102, the left rear left audio zone 103, the middle rear audio zone 104, and the right rear audio zone 105. Multiple voice pickup devices can be installed in the cabin to determine the location of the user making the voice request based on the acquired voice request status information.
[0075] The wake-up voice zone is the voice zone location of the user who issued the wake-up voice request. For example, if the driver wakes up the voice assistant, then the wake-up voice zone is the driver's voice zone. Wake-up voice zone information is the voice zone location information corresponding to the wake-up voice zone.
[0076] The dialogue voice zone is the location of the user's voice during a voice interaction, as determined by the voice assistant. For example, in a scenario where the driver and passenger interact with the voice assistant sequentially after it's activated, their voice requests are captured by the voice assistant sequentially, and both their voice zones belong to the dialogue voice zone. The dialogue voice zone can be the same as or different from the activation voice zone.
[0077] Rejection processing is used to identify which user voice requests are addressed to the voice assistant during the interaction process, recall and execute them, and filter out which are not addressed to the voice assistant as noise.
[0078] This invention provides multiple rejection modes. Different rejection modes recall or reject voice requests based on their annotations. Under different rejection modes, the same voice request may have different rejection results. Details are elaborated below.
[0079] In this invention, a state machine is introduced to record the rejection patterns of each voice region during voice interaction. The state machine is continuously updated based on the received corresponding voice region information and the user's voice requests. In real-world driving scenarios, user voice requests have a degree of randomness. When the voice assistant is activated, the rejection patterns of each voice region need to be updated along with the voice interaction process to ensure that every voice request with a clear intention to interact with the voice assistant can be accurately recognized, while other interactions not involving the voice assistant can be accurately rejected.
[0080] In summary, this invention divides the vehicle cabin into multiple audio zones. Upon receiving a voice request, it determines the rejection mode for each audio zone based on the voice request and its context, thereby meeting the rejection requirements for multi-zone voice interaction within the vehicle cabin. Furthermore, as the voice interaction progresses, the rejection modes for each audio zone are updated, resulting in high accuracy in voice request rejection and a superior user experience in multi-zone interaction scenarios.
[0081] Please see Figure 3 and Figure 4 Step 02 includes:
[0082] 021: Based on the wake-up sound zone information, the initial rejection mode of the wake-up sound zone in the vehicle cabin is determined to be the first rejection mode;
[0083] 022: Determine that the initial rejection mode for all sound zones in the vehicle cabin, excluding the wake-up sound zone, is the second rejection mode.
[0084] The processor is used to determine the initial rejection mode of the wake-up sound zone in the vehicle cabin as a first rejection mode based on the wake-up sound zone information, and to determine the initial rejection mode of each sound zone in the vehicle cabin other than the wake-up sound zone as a second rejection mode.
[0085] Specifically, this invention provides two rejection modes with different levels of rejection: a first rejection mode and a second rejection mode. The second rejection mode has a higher rejection rate for voice requests than the first rejection mode. For the same voice request, the rejection result will differ depending on the rejection mode used. For example, the voice request "Will it rain tomorrow?" may have an unclear intent, be somewhat ambiguous, and be relatively poorly expressed. However, if the first rejection mode is used, it can be recalled to confirm the intent to inquire about the weather, while if the second rejection mode is used, it will be directly rejected.
[0086] During the interaction, after the voice assistant is activated, an initial rejection mode is configured for each audio zone within the cabin. Subsequent rejection mode updates are then based on this initial rejection mode. It's understandable that users activating the voice assistant typically have a strong intention to interact; therefore, the initial rejection mode for the activation audio zone is set to the first rejection mode, while the initial rejection modes for other audio zones are set to the second rejection mode, to avoid potential interference from other audio zones with the interaction in the first audio zone.
[0087] In one example, if the vehicle's voice assistant is activated by a user in the driver's side voice zone 101, then the driver's side voice zone 101 is confirmed as the activation voice zone, and its rejection mode will be set to the first rejection mode. The rejection modes of other voice zones in the cabin, such as the passenger side voice zone 102, left rear voice zone 103, center voice zone 104, and right rear voice zone 105 in the previous example, will be set to the second rejection mode.
[0088] In this way, the initial rejection mode of each sound zone can be determined based on the wake-up sound zone information. Specifically, the initial rejection mode of the wake-up sound zone is the first rejection mode, and the initial rejection mode of the non-wake-up sound zone is the second rejection mode with a higher degree of rejection.
[0089] Please see Figure 3 and Figure 5 Step 04 includes:
[0090] 041: If the dialogue voice area information confirms that the rejection mode of the dialogue voice area is the first rejection mode and the user's voice request is a non-vehicle interaction voice request, then the rejection mode of the dialogue voice area is updated to the second rejection mode.
[0091] The processor is used to update the rejection mode of the dialogue voice area to the second rejection mode if it is confirmed from the dialogue voice area information that the rejection mode of the dialogue voice area is the first rejection mode and the user's voice request is a non-vehicle interaction voice request.
[0092] Specifically, during the interaction, the rejection mode of the dialogue voice zone can be determined based on the dialogue voice zone information. For example, if the dialogue voice zone is a wake-up voice zone, then the rejection mode of the dialogue voice zone is determined to be the first rejection mode. However, if the user's voice request is a non-vehicle interaction voice request, for example, if the obtained voice request is "Hello, who is this?", it can be confirmed that the user is on the phone; or if the obtained user request is "I don't know", it can be confirmed that the user is currently having casual conversation. Such voice requests can be considered non-vehicle interaction voice requests. In this case, it can be assumed that the user in this voice zone does not have a real intention to interact, and the rejection mode of this voice zone can be updated to the second rejection mode for a higher degree of rejection.
[0093] In one example, the driver wakes up the vehicle's voice assistant. The driver's voice zone 101 is set to the first rejection mode. However, based on the voice request obtained from the driver's voice zone 101, it is confirmed that the voice request is not a vehicle interaction voice request. Therefore, the rejection mode of the driver's voice zone 101 is updated to the second rejection mode. That is, it is determined that the driver's voice zone 101 does not have a clear interaction intention for the time being, so the rejection level is increased to prevent voice requests with low interaction intention from being missed.
[0094] Thus, if during the interaction, a certain dialogue voice zone is in the first rejection mode, and the voice request in that voice zone is not a vehicle interaction voice request, then it can be considered that the voice zone has no real interaction intention for the time being, and the rejection mode of the voice zone is updated to the second rejection mode.
[0095] Please see Figure 3 and Figure 6 Step 04 includes:
[0096] 042: If the voice zone of the vehicle cabin denial mode does not receive a valid voice request within the first preset time period, the denial mode of the corresponding voice zone will be updated to the second denial mode.
[0097] The processor is used to update the rejection mode of the corresponding audio zone to the second rejection mode if no valid voice request is obtained within a first preset time period when the audio zone of the vehicle cabin rejection mode is the first rejection mode.
[0098] Specifically, during the interaction, the rejection mode of the dialogue voice zone can be determined based on the dialogue voice zone information. For example, if the dialogue voice zone is a wake-up voice zone, then the rejection mode of the dialogue voice zone is determined to be the first rejection mode. However, if the voice zone does not receive a valid voice request within a certain period of time, for example, if the rejection mode of a certain voice zone is the first rejection mode, but no valid voice request is received within 20 seconds, then it can be considered that the user in that voice zone temporarily has no real intention to interact. In this case, the rejection mode of the voice zone can be updated to the second rejection mode, resulting in a higher degree of rejection. The failure to receive a valid voice request could mean that no voice request was received at all, or that although a voice request was received, it was unrelated to vehicle interaction.
[0099] The first preset duration is a time limit for the interval between valid voice requests from users. It can be set to an appropriate value based on actual conditions, such as 20s, 30s, 50s, or 1 minute. It's understandable that a first preset duration that is too short will cause frequent switching of the voice region's rejection mode, while a setting that is too long may lead to a higher false recall rate for voice requests.
[0100] In one example, the first preset duration can be set to 20 seconds. When the driver wakes up the vehicle's voice assistant, the driver's voice zone 101 is set to the first rejection mode. If no valid voice request is received in the driver's voice zone 101 within the first preset duration, that is, no voice request or voice request related to vehicle interaction is received within 20 seconds, then the rejection mode of the driver's voice zone 101 is updated to the second rejection mode. That is, it is determined that the driver's voice zone 101 does not have a clear interaction intention for the time being, and the rejection level is increased to prevent voice requests with low interaction intention from being missed.
[0101] If a valid instruction is received within the first preset time period, the first rejection mode of that vocal range will continue.
[0102] Thus, if during the interaction, a certain dialogue voice region is in the first rejection mode, but does not receive a valid voice request within a preset time, it can be assumed that the voice region has no real interaction intention for the time being, and the rejection mode of the voice region is updated to the second rejection mode.
[0103] Please see Figure 3 and Figure 7 Step 04 includes:
[0104] 043: If the rejection mode of the dialogue voice area is confirmed to be the second rejection mode based on the dialogue voice area information, and if it is determined based on the user's voice request that a valid voice request is executed within the second preset time period, then the rejection mode of the dialogue voice area is updated to the first rejection mode.
[0105] The processor is configured to update the rejection mode of the dialogue voice region to the first rejection mode if, based on the user's voice request, it is determined that a valid voice request has been executed within a second preset time period after confirming that the rejection mode of the dialogue voice region is the second rejection mode.
[0106] Specifically, executing a valid voice request means acquiring a valid voice request and generating a corresponding vehicle execution command. During the interaction, the rejection mode of the dialogue voice zone can be determined based on the dialogue voice zone information. For example, if the dialogue voice zone is a non-wake-up voice zone, then the initial rejection mode of the dialogue voice zone can be determined to be the second rejection mode. If the voice zone receives a valid voice request within a certain period of time, or in other words, acquires a voice request related to vehicle interaction, for example, if the rejection mode of a certain voice zone is the second rejection mode, and a valid voice request "open the window" is acquired within a second predetermined time period, then it can be considered that the user in that voice zone has a genuine intention to interact, and the rejection mode of that voice zone can be updated to the first rejection mode, performing a lower level of rejection.
[0107] The second preset duration, similar to the first, limits the interval between valid voice requests from the user. It can be set to an appropriate value based on actual conditions, such as 20s, 30s, 50s, or 1 minute. It's understandable that a first preset duration that is too short will cause frequent switching of the voice region's rejection mode, while a setting that is too long may result in a high false recall rate for voice requests.
[0108] In one example, the second preset duration can be set to 20 seconds, the driver's voice zone 101 is the wake-up voice zone, and the left rear voice zone 103 is the non-wake-up voice zone. The initial rejection state is the second rejection mode. If the left rear voice zone 103 obtains a valid voice request and executes it within 20 seconds, the rejection mode of the left rear voice zone 103 is updated to the first rejection mode with a lower rejection level. That is, it is determined that the subsequent left rear voice zone 103 has a relatively clear interaction intention, reducing the rejection level and preventing voice requests from being mistakenly rejected.
[0109] Understandably, if a vocal range in the second rejection mode does not receive a valid instruction within the second preset duration, the second rejection mode of that vocal range will continue to be maintained.
[0110] Thus, if during the interaction, a certain dialogue voice region is in the second rejection mode, but receives a valid voice request within a preset time, then it can be considered that the voice region has a genuine interaction intention, and the rejection mode of the voice region can be updated to the first rejection mode, that is, a rejection mode with a lower degree of rejection.
[0111] Please see Figure 3 and Figure 8 The speech processing method of the present invention further includes:
[0112] 044: If no user voice request is received within the third preset time after the vehicle voice function is activated, exit the vehicle voice function.
[0113] The processor is used to exit the vehicle voice function if no user voice request is received within a third preset time after the vehicle voice function is activated.
[0114] Specifically, during the interaction, if the time since the last time the voice assistant received the user's voice request exceeds a third preset duration, each voice zone can be timed separately until the last voice zone fails to receive the user's voice request within the third preset duration, at which point the vehicle's voice function will exit and wait for the next wake-up.
[0115] The third preset duration is a limit on the time before the vehicle voice function exits. An appropriate value can be set according to actual conditions, such as 100s, 120s, or 150s. It's understandable that a third preset duration that is too short will cause the vehicle voice function to exit frequently, affecting the user experience, while a setting that is too long may result in a longer period of ineffective operation, increasing the processing load.
[0116] In one example, the third preset duration can be set to 120 seconds. After the vehicle voice function is activated, if no voice request from the user is obtained in any of the various voice zones within 120 seconds after multiple rounds of interaction, the vehicle voice function will exit and wait for the next activation.
[0117] In this way, if no user in the cabin makes any voice request within a preset time, the vehicle's voice function will be temporarily deactivated and will wait for the next activation.
[0118] Please see Figure 9 Speech processing methods also include:
[0119] 05: Process user voice requests and determine the speaking target label and intent classification label of the user's voice request;
[0120] 06: The voice request is processed based on the rejection mode of the dialogue voice region, the speaker label, and the intent classification label to obtain the rejection result.
[0121] The processor is used to process user voice requests, determine the speaker label and intent classification label of the user voice request; and to process the voice request according to the rejection mode of the dialogue voice region, the speaker label, and the intent classification label to obtain a rejection result.
[0122] Specifically, the speaking target label is used to identify whether the user's voice request is directed to the voice assistant, and may include voice assistant category labels and non-voice assistant category labels.
[0123] Intent hierarchy labels are used to characterize the effectiveness of a user's voice request and the vehicle's intention to interact. They can be divided into first-level labels, second-level labels, and third-level labels according to their effectiveness from high to low.
[0124] In this invention, each user's voice request can be tagged using these two tags, and further combined with the previously determined rejection mode of the corresponding voice region, the final rejection result, i.e., recall or rejection, can be obtained.
[0125] In this way, the user's voice request is labeled by the speaking object label and the intent classification label. Then, combined with the rejection pattern of the voice request in the voice region, the rejection result of the voice request is determined, that is, whether it is clear and recallable or used as noise filtering.
[0126] Step 06 includes:
[0127] 061: When the rejection mode in the dialogue voice area is the first rejection mode, if the speaking object label is a voice assistant type label and the intent level label is a first-level label or a second-level label, then the rejection result obtained by processing the user's voice request is a clear result.
[0128] 062: If the speaker's label is not a voice assistant type and the intent classification label is a level 3 label, then the user's voice request will result in a noise result after processing.
[0129] The processor is used to process the user's voice request to obtain a clear result when the rejection mode in the dialogue voice area is the first rejection mode, and the speaker's label is a voice assistant type label and the intent classification label is a first-level label or a second-level label; and to process the user's voice request to obtain a noisy result when the speaker's label is a non-voice assistant type label and the intent classification label is a third-level label.
[0130] Specifically, please refer to Figure 10 In this invention, the speaking target label is used to indicate whether the voice request issued by the user is directed to the voice assistant. For example, it may include: "clearly speaking to the voice assistant", "highly likely speaking to the voice assistant", "clearly not speaking to the voice assistant", "highly likely not speaking to the voice assistant", "cannot be determined", "no speaker", etc. The voice assistant category label includes "clearly speaking to the voice assistant" and "highly likely speaking to the voice assistant", while the non-voice assistant category label includes "clearly not speaking to the voice assistant", "highly likely not speaking to the voice assistant", "cannot be determined", and "no speaker".
[0131] For example, a voice request to "open the car window" can be considered as "most likely being spoken to a voice assistant", and the target of the speech can be identified as a voice assistant.
[0132] For example, if the voice request is "hahahaha", it can be assumed that the voice request is "most likely not spoken to the voice assistant", and the target of the speech can be identified as a non-voice assistant type tag.
[0133] The intent level label is used to characterize the validity of the user's voice request, and may include: "strongly valid", "weakly valid", "no intent" and "cannot be determined", etc. According to the validity of the user's voice request, the labels can be divided into: first-level label "strongly valid", second-level label "weakly valid" and third-level label "no intent or cannot be determined".
[0134] Strongly effective voice requests are typically characterized by clear intent, unambiguity, standardized sentence structure, and strong relevance to vehicle functions. Examples include: turn on the air conditioning, straighten the seat back, turn up the instrument panel lights, play a song, open the music interface, and turn up the volume.
[0135] Weak, effective voice requests are typically characterized by unclear intent, potential ambiguity, non-standard sentence structure, and weak relevance to vehicle functions. Examples include: "Will it rain tomorrow?", "Why is the battery dead?", "What song is this?", "Turn up the volume?", "Air conditioning?", etc.
[0136] Unintentional voice requests are usually characterized by unclear intent, potential ambiguity, casual phrasing, or weak or no relation to vehicle functions. Examples include: "Whatever," "Our family," "I'd like to buy this car and get a loan," "Hurry up and come out," "Open the window," and "Change gears."
[0137] It cannot be determined and can be considered as a supplement to the above situations.
[0138] For example, a voice request to "open the car window" can be considered "most likely addressed to the voice assistant," thus confirming its target audience as a voice assistant. Furthermore, since this is a strongly valid voice request, its intent classification can be identified as a Level 1 label. If the voice region falls under the first rejection mode, the rejection result will be a clear result.
[0139] For example, regarding the voice request "hahahaha," it can be assumed that this voice request is "most likely not intended for the voice assistant," thus confirming that the target of the speech is tagged as non-voice assistant. Furthermore, since this voice request is an unintentional one, its intent classification tag can be confirmed as level three. If this voice region falls under the first rejection mode, the rejection result is considered noise.
[0140] In practical applications, when the dialogue area is in the first rejection mode, if the speaker label is a voice assistant-type label, indicating that the speaker request is from a voice assistant or is highly likely to be from a voice assistant, and the intent classification label is a first-level or second-level label (i.e., a strongly valid or weakly valid voice request), then processing the user's voice request will result in a clear rejection result, meaning the voice request will be recalled. Conversely, if the speaker label is not a voice assistant-type label, and the intent classification label is a third-level label, then processing the user's voice request will result in a noisy rejection result, meaning the voice request will be rejected.
[0141] Thus, in the first rejection mode, for voice requests whose speaker is labeled as a voice assistant and whose intent is labeled as a first or second level, the rejection result is confirmed as clear; for voice requests whose speaker is not labeled as a voice assistant and whose intent is labeled as a third level, the rejection result is confirmed as noise.
[0142] Step 06 also includes:
[0143] 063: When the rejection mode in the dialogue voice area is the second rejection mode, if the speaker's label is a voice assistant type label and the intent level label is a first-level label, then the rejection result obtained by processing the user's voice request is a clear result.
[0144] 064: If the target of the speech is labeled as a non-voice assistant and the intent classification label is a second-level or third-level label, then the rejection result obtained by processing the user's voice request is a noise result.
[0145] The processor is used to process the user's voice request to obtain a clear result when the rejection mode in the dialogue voice area is the second rejection mode, and the speaker's label is a voice assistant type label and the intent classification label is a first-level label; and to process the user's voice request to obtain a noisy result when the speaker's label is a non-voice assistant type label and the intent classification label is a second-level label or a third-level label.
[0146] Please see Figure 11 In practical applications, when the dialogue area is in the second rejection mode, if the speaker's label is a voice assistant-type label, indicating that the speaker request is from a voice assistant or is highly likely to be from a voice assistant, and the intent classification label is a first-level label (i.e., a strongly valid voice request), then processing the user's voice request will result in a clear rejection result, meaning the voice request will be recalled. Conversely, if the speaker's label is not a voice assistant-type label, and the intent classification label is a second-level or third-level label, then processing the user's voice request will result in a noisy rejection result, meaning the voice request will be rejected.
[0147] For example, a voice request to "open the car window" can be considered "most likely addressed to the voice assistant," thus confirming its target audience as a voice assistant. Furthermore, since this is a strongly valid voice request, its intent classification can be identified as a Level 1 label. If the voice region falls under the Level 2 rejection mode, the rejection result will be a clear result.
[0148] For example, regarding the voice request "hahahaha," it can be assumed that this voice request is "most likely not intended for the voice assistant," thus confirming that the target of the speech is tagged as non-voice assistant. Furthermore, since this voice request is an unintentional one, its intent classification tag can be confirmed as level three. If this voice region falls under the second rejection mode, the rejection result is considered noise.
[0149] Thus, in the second rejection mode, for voice requests where the speaker's label is a voice assistant and the intent classification label is Level 1, the rejection result is confirmed as a clear result; for voice requests that are not voice assistants and the intent classification label is Level 2 or Level 3, the rejection result is confirmed as a noisy result. Compared to the first rejection mode, the second rejection mode is more stringent in rejecting tags with an intent classification label of Level 2.
[0150] The following three scenario examples illustrate how processing voice requests based on rejection patterns, speaker labels, and intent classification labels yields rejection results:
[0151] Example 1: Referring to Table 1, a user in the driver's side audio zone 101 activates the vehicle's voice function. The driver's side audio zone 101 is confirmed as the activation zone, with an initial rejection mode of the first rejection mode. Other audio zones are non-activation zones, with an initial rejection mode of the second rejection mode. The user in the driver's side audio zone 101 issues a voice request to "turn on the air conditioning." The target of this voice request is labeled as a voice assistant, and the intent classification label is the first level, resulting in a clear rejection. Further, the user in the driver's side audio zone 101 issues a voice request to "20 degrees, level 3 fan." The target of this voice request is labeled as a voice assistant, and the intent classification label is the first level, resulting in a clear rejection. Further, the user in the left rear audio zone 103 issues a voice request to "a bit low, right?" The target of this voice request is labeled as a non-voice assistant, and the intent classification label is the second level, resulting in a noise rejection. Furthermore, the user in the left rear voice zone 103 issued voice requests "the vehicle temperature is a little higher" and "a little higher." The target of the speech was labeled as a voice assistant, and the intent classification label was the first level label. Since a valid voice request was executed within the preset time, the rejection mode of the left rear voice zone 103 will be updated to the first rejection mode, and a clear rejection result will be obtained.
[0152]
[0153] Table 1
[0154] Example 2: Referring to Table 2, a user in the left rear audio zone 103 activates the vehicle's voice function. Left rear audio zone 103 is confirmed as the wake-up audio zone, with an initial rejection mode of the first rejection mode. Other audio zones are non-wake-up audio zones, with an initial rejection mode of the second rejection mode. The user in left rear audio zone 103 issues a voice request, "How's the weather today?". The target of this voice request is tagged as a voice assistant, and the intent classification is at level one, resulting in a clear rejection. Further, the user in left rear audio zone 103 issues a voice request, "What about tomorrow?". The target of this voice request is tagged as a voice assistant, and the intent classification is at level one, resulting in a clear rejection. Subsequently, the user in the left rear voice zone 103 and the user in the right rear voice zone began chatting. The user in the left rear voice zone 103 made a voice request, "The weather is nice, how about we go hiking tomorrow?" Since a valid command was executed in the left rear voice zone 103 within the preset time, the rejection mode for the left rear voice zone 103 remained at the first rejection mode. The speaker's label for this voice request was "non-voice assistant type," and the intent classification label was level three, resulting in a noise rejection result. The user in the right rear voice zone 105 made a voice request, "Sure." The speaker's label for this voice request was "non-voice assistant type," and the intent classification label was level three, resulting in a noise rejection result. The user in the left rear voice zone 103 made a voice request, "Go to the Badaling Great Wall?" The speaker's label for this voice request was "non-voice assistant type," and the intent classification label was level three, resulting in a noise rejection result. The user in the right rear voice zone 105 made a voice request, "Let's see how long it will take to get there." The speaker's label for this voice request was "non-voice assistant type," and the intent classification label was level three, resulting in a noise rejection result. Furthermore, after the casual conversation ends, the user in the left rear audio zone 103 issues a voice request, "Help me navigate to the Badaling Great Wall." Since there is a valid command executed in the left rear audio zone 103 within the preset time, the rejection mode of the left rear audio zone 103 remains in the first rejection mode. The intent classification label of the voice request is determined to be the first level label, resulting in a clear rejection result.
[0155]
[0156]
[0157] Table 2
[0158] Example 3: Please refer to Table 3. After the user in the driver's voice zone 101 activates the vehicle's voice function, the driver's voice zone 101 is confirmed as the wake-up voice zone, and the initial rejection mode is the first rejection mode. The other voice zones are non-wake-up voice zones, and the initial rejection mode is the second rejection mode. At this time, the user in the driver's voice zone 101 starts making a phone call, issuing voice requests such as "Hello, hello," "I'm going to work now," and "Still on my way, haven't arrived yet." The speaking target labels for these voice requests are all non-voice assistant types, and the intent classification label is determined to be a level three label, resulting in a noise rejection result. Further, the user in the passenger's voice zone 102 issues a voice request "Turn the volume down a bit." The rejection mode of the passenger's voice zone 102 is updated to the first rejection mode. The speaking target label for this voice request is a voice assistant type, and the intent classification label is determined to be a level one label, resulting in a clear rejection result. The left rear vocal zone 103 issued a voice request, "Turn off the music." The rejection mode of the left rear vocal zone 103 was updated to the first rejection mode. The target of the voice request was labeled as a voice assistant, and the intent classification label was determined to be the first level label, resulting in a clear rejection result.
[0159]
[0160] Table 3
[0161] Please see Figure 12 The present invention also provides a voice interaction method, comprising:
[0162] 01: Receive the wake-up sound zone information forwarded by the vehicle when the user activates the vehicle's voice function inside the vehicle cabin;
[0163] 02: Determine the initial rejection mode for each of the multiple sound zones in the vehicle cabin based on the wake-up sound zone information;
[0164] 03: Receive user voice requests forwarded by the vehicle after the vehicle's voice function is activated, as well as the dialogue zone information confirmed based on the user's voice request;
[0165] 04: Update the rejection mode of the corresponding voice region based on the user's voice request and dialogue voice region information to determine the rejection mode of each voice region;
[0166] 07: After determining the rejection mode for each vocal region, process the user's voice request to obtain the speaker's label and intent classification label;
[0167] 08: The voice request is processed based on the rejection mode, the speaker's label, and the intent classification label to obtain the rejection result;
[0168] 09: Send the rejection result to the vehicle to complete the voice interaction.
[0169] The voice interaction method of the present invention can be implemented by the server of the present invention, the server including a memory and a processor. Specifically, the memory stores a computer program, the processor is used to receive wake-up sound zone information of the user waking up the vehicle's voice function in the vehicle cabin forwarded by the vehicle, and to determine the initial rejection mode of each sound zone among multiple sound zones in the vehicle cabin according to the wake-up sound zone information, and to receive user voice requests forwarded by the vehicle after the vehicle's voice function is woken up, as well as dialogue sound zone information confirmed according to the user voice request, and to update the rejection mode of the corresponding sound zone according to the user voice request and dialogue sound zone information to determine the rejection mode of each sound zone, and to process the user voice request after determining the rejection mode of each sound zone to obtain a speaking object label and an intent classification label, and to process the voice request according to the rejection mode, speaking object label and intent classification label to obtain a rejection result, and to send the rejection result to the vehicle to complete the voice interaction.
[0170] Specifically, after confirming the rejection of the voice request, the rejection result is sent to the vehicle, which can then execute the control command generated by the voice request or remain unresponsive, thus completing the voice interaction.
[0171] For details on the rejection mode and the confirmation method of rejection result, please refer to the explanation of each implementation method in the above processing method, which will not be repeated here.
[0172] In this way, the vehicle cabin is divided into multiple sound zones. Upon receiving a voice request, the rejection mode for each sound zone is determined based on the request and its context, thus meeting the rejection requirements for multi-sound zone voice interaction within the vehicle cabin. Furthermore, as the voice interaction progresses, the rejection modes for each sound zone are updated, resulting in high accuracy in voice request rejection and a superior user experience in multi-sound zone interaction scenarios.
[0173] The present invention provides a computer-readable storage medium storing a computer program that, when executed by one or more processors, implements the above-described method.
[0174] In the description of this specification, references to terms such as "above," "specifically," etc., indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of the present invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0175] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of executable request code comprising one or more steps for implementing a particular logical function or process, and the scope of preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.
[0176] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A speech processing method, characterized in that, include: The vehicle receives the wake-up sound zone information forwarded by the user to activate the vehicle's voice function within the vehicle's cabin. The initial rejection mode of each of the multiple sound zones in the vehicle cabin is determined based on the wake-up sound zone information. Receive the user voice request forwarded by the vehicle after the vehicle's voice function is activated, and the dialogue zone information confirmed according to the user voice request; The rejection mode of the corresponding voice region is updated according to the user's voice request and the dialogue voice region information to determine the rejection mode of each voice region; Process the user's voice request to determine the speaking target label and intent classification label of the user's voice request; When the rejection mode in the dialogue voice area is the first rejection mode, if the speaking object label is a non-voice assistant label and the intent classification label is a third-level label, then the rejection result obtained by processing the user's voice request is a noise result, and the intent classification label represents the validity of the user's voice request.
2. The speech processing method according to claim 1, characterized in that, The step of determining the initial rejection mode for each of the multiple sound zones in the vehicle cabin based on the wake-up sound zone information includes: Based on the wake-up sound zone information, the initial rejection mode of the wake-up sound zone in the vehicle cabin is determined to be the first rejection mode; The initial rejection mode of each sound zone in the vehicle cabin, excluding the wake-up sound zone, is determined to be the second rejection mode. The second rejection mode has a higher rejection level for voice requests than the first rejection mode.
3. The speech processing method according to claim 2, characterized in that, The step of updating the rejection mode of the corresponding voice region based on the user's voice request and the dialogue voice region information to determine the rejection mode of each voice region includes: If, based on the dialogue voice area information, it is confirmed that the rejection mode of the dialogue voice area is the first rejection mode and the user's voice request is a non-vehicle interaction voice request, then the rejection mode of the dialogue voice area is updated to the second rejection mode.
4. The speech processing method according to claim 2, characterized in that, The step of updating the rejection mode of the corresponding voice region based on the user's voice request and the dialogue voice region information to determine the rejection mode of each voice region includes: If the vehicle cabin denial mode is the first denial mode and the corresponding audio region does not receive a valid voice request within a first preset time period, then the denial mode of the corresponding audio region will be updated to the second denial mode.
5. The speech processing method according to claim 2, characterized in that, The step of updating the rejection mode of the corresponding voice region based on the user's voice request and the dialogue voice region information to determine the rejection mode of each voice region includes: If the rejection mode of the dialogue voice area is confirmed to be the second rejection mode based on the dialogue voice area information, and if it is determined based on the user's voice request that a valid voice request is executed within a second preset time period in the dialogue voice area, then the rejection mode of the dialogue voice area is updated to the first rejection mode.
6. The speech processing method according to claim 1, characterized in that, The aforementioned speech processing method includes: If no user voice request is received within a third preset time period after the vehicle voice function is activated, the vehicle voice function will be deactivated.
7. The speech processing method according to claim 1, characterized in that, The speech processing method includes: When the rejection mode in the dialogue voice area is the first rejection mode, if the speaking object label is a voice assistant type label and the intent level label is a first-level label or a second-level label, then the user's voice request is processed to obtain a clear rejection result, wherein the first-level label is greater than the second-level label and the second-level label is greater than the third-level label.
8. The speech processing method according to claim 7, characterized in that, The speech processing method includes: When the rejection mode in the dialogue voice area is the second rejection mode, if the speaking object label is a voice assistant type label and the intent classification label is a first-level label, then the rejection result obtained by processing the user's voice request is a clear result. If the label of the speaking object is a non-voice assistant label and the label of the intent classification is a second-level label or a third-level label, then the rejection result obtained by processing the user's voice request is a noise result.
9. A voice interaction method, characterized in that, The voice interaction method includes: The vehicle receives the wake-up sound zone information forwarded by the user to activate the vehicle's voice function within the vehicle's cabin. The initial rejection mode of each of the multiple sound zones in the vehicle cabin is determined based on the wake-up sound zone information. Receive the user voice request forwarded by the vehicle after the vehicle's voice function is activated, and the dialogue zone information confirmed according to the user voice request; The rejection mode of the corresponding voice region is updated according to the user's voice request and the dialogue voice region information to determine the rejection mode of each voice region; After determining the rejection mode for each of the aforementioned voice regions, the user's voice request is processed to obtain the speaking object label and the intent classification label; The voice request is processed according to the rejection mode, the speaker label, and the intent classification label to obtain the rejection result; The rejection result is sent to the vehicle to complete the voice interaction; When the rejection mode in the dialogue area is the first rejection mode, if the speaking object label is a non-voice assistant label and the intent classification label is a third-level label, then the rejection result obtained by processing the user's voice request is a noise result, and the intent classification label characterizes the validity of the user's voice request.
10. A server, characterized in that, The server includes a memory and a processor, the memory storing a computer program that, when executed by the processor, implements the method according to any one of claims 1-9.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by one or more processors, implements the method as described in any one of claims 1-9.
Citation Information
Patent Citations
Method for positioning sounding location, and terminal device
CN107430524A
Multi-sound-zone voice interaction method for vehicle and electronic equipment
CN111816189A
Rejection method and device, equipment and storage medium
CN114155853A