Multi-View Display Voice Separation for Accurate Speaker Commands
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice recognition technologies in display devices struggle with accurately distinguishing multiple speakers and handling interference sounds, leading to repeated voice recognition processes and reduced user convenience.
Innovation Solution
The display device employs a speaker recognition system that identifies multiple speakers within a voice signal, separates their voice data, and processes these data differently based on the device's playback mode, allowing simultaneous or sequential execution of commands.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional voice recognition technology is used to recognize user input, then the display device can respond to user commands, but it fails to accurately distinguish multiple speakers and is easily interfered with by nearby sound sources, leading to repeated recognition processes
Solution Approach 1:
The patent segments the voice recognition process by first identifying individual speakers through speaker recognition technology, then separating their voice data, and finally performing voice recognition on each speaker's data independently. This segmentation resolves the contradiction by enabling accurate distinction between multiple speakers while maintaining reliable recognition even in the presence of interference sounds.
Solution Approach 2:
The patent introduces speaker recognition as an intermediary step between voice input and voice recognition processing. This intermediary component identifies and separates multiple speakers before the actual voice recognition occurs, thereby improving both measurement precision of speaker identity and reliability of the overall recognition system in multi-speaker environments.
2Productivity
If the display device processes voice input from multiple speakers using conventional methods, then it can handle basic commands, but it performs repeated voice recognition processes when interference sounds are present, reducing operational efficiency
Solution Approach 1:
The patent applies preliminary action by performing speaker recognition and voice data separation before the main voice recognition process. By pre-identifying and isolating each speaker's voice data, the system avoids repeated recognition attempts caused by interference sounds, thereby improving processing efficiency and reducing time loss.
3Adaptability or versatility
If the display device uses basic voice recognition without speaker identification, then the system remains simple, but it cannot provide personalized services or accurately match user intentions in multi-user scenarios
Solution Approach 1:
The patent implements multi-functionality by integrating speaker recognition, voice data separation, and voice recognition capabilities into a unified system. This allows the display device to handle both single-user and multi-user scenarios, providing personalized services while maintaining a cohesive system architecture that manages complexity effectively.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Disclosed are a display device and an operating method therefor. According to an aspect of the present disclosure, a method for operating a display device includes receiving voice data; separating the received voice data into pieces of voice data for speakers; and performing control such that pieces of content according to the pieces of voice data, which have been separated for the speakers are respectively output on multi-view screen areas, when a current playback mode is a multi-view mode.