Teleconference Terminal Voice-Adaptive Image Signal Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional teleconference systems experience excessive network traffic due to the transmission of motion images from all mobile units, leading to deteriorated image quality and increased data traffic, even when not all units are in use.
Innovation Solution
A teleconference terminal apparatus that adjusts the data amount of image signals based on voice signal levels by using an image output processing control section to change the compression ratio or stop image signal transmission, allowing for high-quality image display of active units while reducing network traffic.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If all mobile units transmit motion image signals to the parent station, then the image quality of transmitted content can be maintained, but the network traffic expands excessively
Solution Approach 1:
The patent changes the parameter of image signal transmission by switching between motion images and still images based on voice activity detection. When voice activity is detected, motion images are transmitted; when no voice activity is present, still images are transmitted instead. This parameter change reduces network traffic while maintaining image quality for important content.
Solution Approach 2:
The system dynamically adjusts the type of image signal transmitted (motion or still) based on real-time voice activity detection. This dynamic adaptation allows the system to optimize network traffic by transmitting only necessary motion images when speakers are active, while using still images during silent periods.
2Manufacturing precision
If all mobile units transmit motion image signals continuously, then the image quality can be maintained, but the data amount increases excessively
Solution Approach 1:
The patent applies parameter changes by switching the image signal type between motion and still formats based on voice activity. This reduces the data amount transmitted while preserving image quality for relevant content, as still images consume significantly less data than continuous motion images.
Solution Approach 2:
The system extracts and transmits only the essential image information needed for communication. By using still images during non-speech periods, the system extracts only the necessary visual content rather than transmitting continuous motion data, thereby reducing overall data transmission while maintaining quality when needed.
3Loss of information
If motion images are transmitted from all units, then complete visual information is provided, but network bandwidth is overwhelmed
Solution Approach 1:
The system changes the transmission parameter from continuous motion images to conditional still images based on voice activity detection. This maintains essential visual information for speakers while improving network efficiency by reducing unnecessary data transmission during silent periods.
Solution Approach 2:
The system dynamically adapts its transmission strategy by detecting voice activity and switching between motion and still image modes. This dynamic approach ensures that visual information is preserved when important (during speech) while maximizing network efficiency during non-speech periods.
Data Source
AI summary
The present invention provides a teleconference terminal apparatus for carrying out a teleconference by transmitting/receiving image and voice signals via a communications network. The teleconference terminal apparatus comprises: an image-capturing device which generates an image signal by capturing an image; an image processing section which converts the image signal to a signal mode corresponding to the communications network and outputs the converted signal; a microphone which performs detection of a voice along with the image capturing, and thereby generates a voice signal corresponding to the level of the voice; a voice processing section which converts the voice signal to a signal mode corresponding to the communications network and outputs the converted signal; and a computing section which increases or reduces the amount of data of the image signal outputted from the image processing section, on the basis of the level of the voice signal.


