Speech Keyword Screen Switching in Video Conferences
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conference systems require complex operations for participants to display images from one conference room to another, disrupting the smooth progression of meetings and limiting accessibility to only skilled users.
Innovation Solution
An image transmitting apparatus that automatically switches screens based on recognized keywords from speech, allowing for the generation and transmission of display screens to multiple users, including images and object data, to facilitate understanding and simplify the display process across multiple conference rooms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual operations are required to switch screens in video conference systems, then image display control is precise, but operation complexity increases and accessibility decreases
Solution Approach 1:
The system automatically detects keywords from speech and switches screens without requiring manual user input. The speech recognition unit processes spoken words, and the screen switching unit automatically selects and displays appropriate screens based on detected keywords, enabling the system to serve itself rather than requiring user intervention for each screen change.
Solution Approach 2:
The patent replaces manual mechanical operations (button presses, menu selections) with automated speech recognition and processing. The speech recognition unit converts spoken language into digital signals, which are then processed by the determination unit to trigger automatic screen switching, substituting physical user actions with automated electronic processing.
2Productivity
If multiple manual operations are required for image switching, then display control precision is maintained, but conference flow is interrupted
Solution Approach 1:
The system maintains continuous operation by automatically switching screens in response to speech keywords without interrupting the conference flow. The speech recognition and screen switching occur continuously and automatically, ensuring that the useful action of displaying appropriate images continues without pauses for manual intervention.
Solution Approach 2:
The determination unit is pre-configured with keyword associations to specific screens before the conference begins. When keywords are detected during speech, the corresponding screens are already prepared and can be switched to immediately, avoiding delays associated with manual selection and pre-positioning the system for rapid response.
3Adaptability or versatility
If skilled operations are required for image display, then display accuracy is improved, but user accessibility is limited
Solution Approach 1:
The system provides universal functionality where any participant can initiate screen switching through speech without requiring specialized skills. The speech recognition unit and automated determination mechanism work for all users uniformly, making the advanced screen switching capability accessible to everyone regardless of their technical expertise.
Solution Approach 2:
The speech recognition unit acts as an intermediary between the user's spoken words and the screen switching function. This intermediary layer translates natural speech into automated control signals, bridging the gap between simple user input and complex display control, thereby making the system accessible to users without technical expertise.
Data Source
AI summary
A multifunction peripheral (MFP) includes an accepting portion to accept a picked up image and a speech, a speech recognition portion to recognize the accepted speech, a display screen generating portion to generate a display screen in accordance with an output setting, and a transmission control portion to transmit the generated display screen. The display screen is generated in response to a recognized keyword included in a predetermined output setting. The display screen includes at least one of the picked up image and an image of object data stored in association with the keyword in advance and independently from the picked up image. The display screen is automatically switched including the image of object data stored in association with the recognized keyword in accordance with the output setting when a keyword is recognized by the speech recognition portion.


