Multi-Device Voice Response Selection via Direct Sound Energy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In a multi-device voice interaction system, multiple electronic devices with sound pickup functions can be inadvertently activated by a single wake-up word, leading to confusing results and reduced user experience due to simultaneous responses.
Innovation Solution
A method that determines a target sound source point and direct sound energy for each electronic device, allowing only devices with sufficient direct sound energy to respond, thereby selectively activating the appropriate devices based on the user's voice command.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If multiple electronic devices use the same wake-up word for voice interaction, then the system becomes more accessible and easy to use, but multiple devices may be inadvertently activated simultaneously causing confusing results
Solution Approach 1:
The patent applies local quality by making each electronic device have distinct acoustic characteristics through unique device identifiers embedded in their audio responses. When a wake-up word is detected, each device calculates its own direct sound energy and compares it with others, allowing the system to locally determine which device should respond based on acoustic proximity rather than universal activation. This resolves the contradiction by maintaining ease of operation through wake-up word recognition while improving reliability through device-specific response selection.
2Adaptability or versatility
If all electronic devices with sound pickup functions respond to the same wake-up word, then the system provides comprehensive voice interaction coverage, but the sound pickup results become confusing and user experience deteriorates
Solution Approach 1:
The patent applies inversion by reversing the traditional approach where all devices respond to a wake-up word. Instead, when a wake-up word is detected by multiple devices, the system inverts the selection logic by having each device calculate its direct sound energy and select only the device with the highest value to respond. This maintains comprehensive voice interaction coverage through wake-up word detection while improving sound pickup accuracy by selecting a single appropriate device based on inverted selection criteria.
3Measurement precision
If electronic devices calculate direct sound energy to determine which device should respond, then the response accuracy improves, but the computational complexity and processing time increase
Solution Approach 1:
The patent applies partial action by having devices perform direct sound energy calculation only when necessary - specifically when multiple devices detect the same wake-up word. In normal operation, devices use simple wake-up word detection. The complex direct sound energy calculation is performed partially only for devices that need to determine response priority, reducing overall processing complexity while maintaining high measurement precision for sound source localization when needed.
Data Source
Figure 1~2
Figure 3~4
Figure 5
AI summary
The present disclosure relates to a method, an electronic device, a medium and a system for responding to a voice signal, the method including: receiving (S11) a voice signal by the plurality of electronic devices; determining (S12), for each of the electronic devices, a target sound source point associated with the voice signal; determining (S13), for each of the electronic devices, a direct sound energy associated with the voice signal according to the target sound source point; and selecting (S14) from the plurality of electronic devices at least one electronic device with a direct sound energy satisfying a predetermined condition, and responding to the voice signal by the at least one electronic device. It realizes a sound pickup decision system based on a plurality of electronic devices, which can make more accurate, more reasonable, and more targeted response to a user's voice, improving voice interaction experience.