Smart Speaker Visual Output via External Display Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Smart speakers without display screens face inefficiencies in conveying information, as they rely solely on synthesized voice responses, which are time-consuming and require users to manually transcribe visual information, while those with screens have higher hardware costs and complexities.
Innovation Solution
A smart speaker system that includes a network device, processor, sound playing component, and sound receiving device, which preloads linked settings for voiceprint data, user information, and authority settings, allowing the processor to convert voices into text, recognize users, and transmit relevant information to a cloud server, then determines whether to send visual responses to connected display devices based on privacy and content ratings.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If a display screen is added to the smart speaker, then visual information can be displayed to the user, but hardware costs and device complexity increase significantly
Solution Approach 1:
The patent extracts the display function from the smart speaker itself and relocates it to the user's existing display devices (smartphones, tablets, computers, televisions). The smart speaker only retains the audio output function and serves as a voice interface, while visual information is transmitted to external devices through network communication. This resolves the contradiction by eliminating the need for expensive display hardware in the speaker while still providing visual information delivery.
Solution Approach 2:
The patent leverages the universality of user-owned display devices to provide visual output. Instead of requiring each smart speaker to have its own display, the system can utilize any display device the user already possesses, making the solution applicable across different user scenarios without increasing speaker hardware complexity.
2Loss of information
If a display screen is added to the smart speaker, then visual information can be displayed, but manufacturing costs increase due to additional components
Solution Approach 1:
The display component is extracted from the smart speaker product definition and relocated to external user devices. The bill of materials for the smart speaker no longer includes display panels, touch cells, display drivers, or graphic processors, significantly reducing manufacturing costs while still enabling visual information delivery through networked devices.
3Loss of information
If the smart speaker generates and plays synthesized voice responses, then all relevant information can be conveyed, but the response time is extended
Solution Approach 1:
The patent adds a visual dimension to the response channel by transmitting information directly as text or images to display devices, bypassing the time-consuming synthesized voice generation process for certain information types. This parallel communication channel delivers information more quickly while maintaining completeness.
Data Source
AI summary
The present disclosure provides an operation method of a smart speaker. The method includes steps as follows. The linked settings among voiceprint registration data, user information and a cast setting of a user device are preloaded by the smart speaker. Wake-up words are received to set an operation mode of the smart speaker and to generate a voiceprint recognition result. In the operation mode, after receiving voice, the voice is converted into voice text and the voiceprint recognition result is compared to voiceprint registration data. When the voiceprint recognition result matches the voiceprint registration data, the user information and the voice text are transmitted to a cloud server, so that the cloud server returns the response message to the smart speaker. According to the cast setting, the response message is sent to the user device.


