Multimodal Advertising Personality via Server-Side Speech Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current multimodal applications face challenges in delivering effective voice recognition and synthesis on small devices due to limited interaction modalities and vocabularies, which hampers the presentation of advertising content in a manner appealing to users.
Innovation Solution
Establishing a multimodal advertising personality by associating vocal and visual demeanors with sponsors, allowing for the presentation of speech and visual portions of multimodal applications using these characteristics, thereby enhancing user interaction and advertising effectiveness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech recognition and synthesis are implemented on small devices, then user interaction capability is improved, but device complexity and processing requirements increase
Solution Approach 1:
The patent segments the speech processing functionality by introducing a server component that handles complex speech recognition and synthesis tasks, while the small device only needs to communicate with the server. This divides the system into client (device) and server components, reducing the complexity burden on the small device while maintaining enhanced interaction capabilities.
2Measurement precision
If limited vocabulary is used for speech recognition, then recognition accuracy is improved, but advertising content presentation flexibility deteriorates
Solution Approach 1:
The patent introduces a server as an intermediary between the user's speech input and the application logic. The server handles the speech recognition with its larger vocabulary capabilities, while the small device maintains a limited vocabulary for local processing. This intermediary approach allows the system to present flexible advertising content without requiring the small device to have extensive speech recognition vocabulary.
Data Source
AI summary
Establishing a multimodal advertising personality for a sponsor of a multimodal application, including associating one or more vocal demeanors with a sponsor of a multimodal application and presenting a speech portion of the multimodal application for the sponsor using at least one of the vocal demeanors associated with the sponsor.


