Live Video Interaction System Using Voice-Triggered Special Effects
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current live video streaming interactions are limited by a simple presentation mode that fails to reflect the interactive state of viewing users, leading to a poor experience and inadequate user participation.
Innovation Solution
A method and apparatus that capture and display streamer-end and user-end video data in real-time, monitor for preset voice instructions, and display video special effects based on recognized audio and video associations, enhancing user interaction and participation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If only streamer screen is displayed in real time in public screen region, then the system complexity is low, but the user participation sense is poor
Solution Approach 1:
The screen display is segmented into multiple regions: a public screen region displaying the streamer, and user screen regions displaying individual user videos. This segmentation allows the system to maintain low complexity in the public region while enhancing user participation through personalized display regions that show user-specific content and interactions.
Solution Approach 2:
The system transitions from a single-dimension public display to a multi-dimensional display architecture where each user has their own display dimension. User videos are displayed in dimensions specific to each user, allowing simultaneous display of streamer content and user-specific interactions without increasing overall system complexity.
2Adaptability or versatility
If user-end video data is captured and displayed in real time, then the user participation sense is improved, but the processing complexity increases
Solution Approach 1:
The system introduces an intermediary processing layer that receives video data from multiple users, performs selective processing based on interaction detection, and routes content to appropriate display regions. This intermediary layer manages processing complexity by not processing all user videos uniformly, but only those with detected interactions.
Solution Approach 2:
The system performs preliminary detection of interactions between users and streamer before processing and displaying user-end video data. By detecting interactions first (through audio/video analysis), the system only processes and displays user videos when interactions are present, reducing overall processing complexity while maintaining enhanced user participation.
3Adaptability or versatility
If video special effects are displayed based on voice instruction recognition, then the interaction presentation is enriched, but the recognition accuracy requirement increases
Solution Approach 1:
The system implements partial recognition by focusing on detecting specific preset voice instructions and interaction-related content rather than attempting to recognize all possible speech. This partial approach enriches interaction presentation for detected instructions while reducing the overall recognition accuracy burden compared to comprehensive speech recognition.
Solution Approach 2:
The system uses feedback mechanisms where detected interactions and voice instructions trigger specific video special effects. This feedback loop allows the system to confirm recognition accuracy through the appropriateness of triggered effects, and adjust recognition parameters based on effect outcomes, effectively managing the accuracy requirement.
Data Source
AI summary
The present application discloses techniques for interaction during live video streaming. The techniques comprise obtaining and playing streamer-end video data, and user-end video data captured by a user terminal in real time; monitoring and recognizing whether the streamer-end video data comprise a preset voice instruction; determining whether the user-end video data comprises a target audio or a target video when the streamer-end video data comprises the preset voice instruction; and displaying a video special effect corresponding to the preset voice instruction in a user video when the user-end video data comprise the target audio or the target video. By means of the present application, a video special effect can be played for a user video according to a result of interaction between a streamer and a user, which enriches the way of interaction presentation and enhances the sense of participation in interaction.


