Directional Voice Capture for Accurate Multi-User Interaction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In multi-user interaction scenarios, existing electronic devices with voice interaction functions often suppress the voice of potential interactors due to directional audio collection, failing to recognize and respond to their inputs.
Innovation Solution
The method involves collecting audio signals from both the current and potential interactors, determining a target interactor based on a specified waiting time period, and prioritizing responses to ensure accurate interaction with the user who speaks within this period.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the electronic device collects only the voice in the direction in which the current interactor is located, then the accuracy of voice collection is improved, but the voice of the potential interactor is suppressed
Solution Approach 1:
The audio collection is segmented into multiple directional channels, each corresponding to a different user's location. The system divides the spatial audio field into separate segments and processes each segment independently to identify and select the target speaker's voice signal.
Solution Approach 2:
The system dynamically adjusts the audio collection direction and target user identification based on real-time detection of voice signals from multiple users. The beamforming direction and active user selection change dynamically according to who is currently speaking, allowing the system to adapt between different interaction scenarios.
2Object-affected harmful factors
If the electronic device suppresses voice in directions other than the target direction, then environmental noise interference is reduced, but the potential interactor's voice cannot be sensed
Solution Approach 1:
The system performs preliminary audio collection in multiple directions before determining the target speaker. By pre-capturing audio signals from all potential user directions and then selecting the active speaker's signal, the system ensures no voice information is lost while still achieving noise suppression for the final output.
Solution Approach 2:
The system introduces an intermediary selection mechanism that chooses between multiple directional audio signals based on activity detection. This intermediary layer allows the system to maintain multiple audio channels for information preservation while selecting only the relevant one for processing, effectively mediating between noise suppression and information retention.
3Measurement precision
If the electronic device responds to the current interactor only, then the interaction accuracy with the active user is improved, but the response to the potential interactor is delayed or missed
Solution Approach 1:
The system continuously monitors audio signals from multiple directions and provides feedback on which user is currently speaking. This feedback mechanism allows the system to dynamically switch between different target users based on real-time voice activity detection, ensuring timely and accurate responses to whichever user is actively interacting.
Data Source
Figure 1~2
Figure 3~4
Figure 5~6
AI summary
This application provides a voice interaction method, an electronic device, and a computer-readable storage medium. The voice interaction method includes: performing voice interaction with a first user, and collecting, in a voice collection time period of the voice interaction, a first audio signal in an angle range in which the first user is located and a second audio signal in an angle range in which a second user is located, where the second user is a user who performs voice interaction with the electronic device in a specified historical time period; determining whether a start moment of a first voice signal in the first audio signal is within a first time period, and determining a target voice signal from the first voice signal and a second voice signal based on a determining result, where the second voice signal is a voice signal included in the second audio signal, and the first time period is a time period of first duration after a start moment of the voice collection time period; and responding to the target voice signal. In this application, a target interactor can be accurately determined in a multi-person interaction scenario.