Electronic Apparatus Sensor Prioritization for Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In electronic systems with multiple sensor devices for speech recognition, processing duplication and resource wastage occur due to simultaneous input from multiple devices, leading to inefficiencies in network transmission and calculation.
Innovation Solution
An electronic apparatus prioritizes sensor devices by identifying the effective sensor device based on audio signal similarity and quality characteristics using a mode-specific audio model, acquiring predicted audio components, and controlling operation states to minimize redundant processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple sensor devices receive and process user speech simultaneously, then speech recognition coverage is improved, but processing duplication and resource wastage occur
Solution Approach 1:
The system implements feedback by having sensor devices transmit audio signals to the electronic apparatus, which then provides feedback information indicating which sensor device should perform speech recognition. This closed-loop control prevents duplicate processing while maintaining multi-device coverage, as each device receives feedback about its processing role based on audio signal quality and similarity metrics.
Solution Approach 2:
The system changes parameters by evaluating audio signal quality characteristics and similarity metrics to dynamically determine which sensor device should process the speech. By comparing audio signals from multiple devices and assessing their quality parameters, the system optimizes processing assignment based on real-time signal conditions rather than static device selection.
2Measurement precision
If multiple sensor devices transmit audio signals to the electronic apparatus, then speech recognition accuracy is improved, but network transmission resources are wasted
Solution Approach 1:
The system extracts only the necessary audio signal from sensor devices by having them transmit audio signals to the electronic apparatus for quality assessment. Instead of all devices processing and transmitting full speech recognition results, only the audio signals from selected devices are transmitted for further processing, reducing redundant network traffic while preserving recognition accuracy.
Solution Approach 2:
The system applies partial action by having sensor devices transmit audio signals for quality comparison rather than performing complete speech recognition processing. This partial processing approach allows the system to evaluate signal quality from multiple devices without the full computational overhead of multiple complete recognition processes, optimizing the balance between accuracy and resource usage.
3Reliability
If the electronic apparatus processes audio signals from multiple sensor devices, then speech recognition reliability is improved, but calculation resources are consumed
Solution Approach 1:
The system performs preliminary action by having sensor devices transmit audio signals for quality assessment and comparison before initiating full speech recognition processing. This preliminary evaluation of audio signal quality and similarity allows the system to pre-select the most suitable device for processing, ensuring reliability through quality verification while avoiding the computational cost of processing signals from all devices.
Solution Approach 2:
The electronic apparatus acts as an intermediary by receiving audio signals from multiple sensor devices, evaluating their quality characteristics, and determining which device should perform the actual speech recognition. This intermediary role allows the system to benefit from multi-device input for reliability while concentrating calculation resources on a single selected device, rather than distributing processing across all devices.
Data Source
Figure 1~2a
Figure 2b~2c
Figure 2d~2e
AI summary
An electronic device is provided. The electronic apparatus includes a communication interface, and at least one processor configured to receive a first audio signal and a second audio signal from a first sensor device, and a second sensor device located away from the first sensor device, respectively, through the communication interface, acquire similarity between the first audio signal and the second audio signal, acquire a first predicted audio component from the first audio signal based on an operation state of an electronic apparatus located adjacent to the first sensor device, and a second predicted audio component from the second audio signal based on an operation state of an electronic apparatus located adjacent to the second sensor device in the case where the similarity is equal to or higher than a threshold value, identify one of the first sensor device or the second sensor device as an effective sensor device based on the first predicted audio component and the second predicted audio component, and perform speech recognition with respect to an additional audio signal received from the effective sensor device.