Multi-AI Speech Recognition via Environmental Weighting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing artificial intelligence systems fail to accurately recognize a user's utterance voice when multiple AI apparatuses are present in the same space, as each apparatus performs individual speech recognition without proper coordination, leading to reduced accuracy.
Innovation Solution
Determine weights between AI apparatuses based on environmental factors at the time of utterance, such as noise type and positional relation, to generate a final speech recognition result by combining individual recognition results, thereby improving accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple AI apparatuses perform individual speech recognition independently, then each apparatus can operate autonomously, but the speech recognition accuracy deteriorates due to lack of coordination and environmental consideration
Solution Approach 1:
The patent combines speech recognition results from multiple AI apparatuses by collecting recognition results from each apparatus and generating a final speech recognition result through integration. This merging approach leverages the collective recognition capabilities of multiple devices to improve overall accuracy while maintaining autonomous operation of each individual apparatus.
Solution Approach 2:
The patent determines weights for each AI apparatus based on environmental variables such as noise type and positional relationship with the user. By dynamically adjusting these weight parameters according to environmental conditions, the system optimizes the contribution of each apparatus to the final recognition result, thereby improving accuracy without requiring complex real-time coordination protocols.
2Measurement precision
If speech recognition results are collected from multiple AI apparatuses with weight determination based on environmental variables, then speech recognition accuracy is improved, but the processing complexity and time increase
Solution Approach 1:
The patent determines weights for each AI apparatus in advance based on environmental variables detected at the time of utterance. By pre-calculating these weights before combining recognition results, the system avoids complex real-time computations during the speech recognition process, thereby reducing processing time while maintaining high precision through environmentally-adapted weight selection.
3Reliability
If weights are determined based on noise type and environmental factors, then recognition accuracy in noisy environments is improved, but the system complexity for environmental analysis increases
Solution Approach 1:
The patent applies different weight values to different AI apparatuses based on their specific environmental conditions, such as noise type and positional relationship with the user. This localized quality adjustment allows each apparatus to be evaluated according to its own operational context rather than using a uniform weighting scheme, improving reliability in noisy environments while keeping the analysis complexity manageable through focused environmental parameter detection.
Data Source
AI summary
Embodiments provide an artificial intelligence apparatus for recognizing an utterance voice of a user. The artificial intelligence apparatus includes: a communication unit configured to communicate with at least one external artificial intelligence apparatus which obtains first sound data including the utterance voice of the user to generate a first speech recognition result from the first sound data; a microphone configured to obtain second sound data including the utterance voice; and a processor configured to receive first speech recognition results from each of the at least one external artificial intelligence apparatus, generate a second speech recognition result from the second sound data, generate a final speech recognition result for the utterance voice by using the first speech recognition results and the second speech recognition result, and perform a control corresponding to the final speech recognition result.


