Live Video Interaction System Using Voice-Triggered Special Effects

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current live video streaming interactions are limited by a simple presentation mode that fails to reflect the interactive state of viewing users, leading to a poor experience and inadequate user participation.

Innovation Solution

A method and apparatus that capture and display streamer-end and user-end video data in real-time, monitor for preset voice instructions, and display video special effects based on recognized audio and video associations, enhancing user interaction and participation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If only streamer screen is displayed in real time in public screen region, then the system complexity is low, but the user participation sense is poor

Engineering Contradiction:
Improvesystem complexityVSAvoiduser participation sense
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The screen display is segmented into multiple regions: a public screen region displaying the streamer, and user screen regions displaying individual user videos. This segmentation allows the system to maintain low complexity in the public region while enhancing user participation through personalized display regions that show user-specific content and interactions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from a single-dimension public display to a multi-dimensional display architecture where each user has their own display dimension. User videos are displayed in dimensions specific to each user, allowing simultaneous display of streamer content and user-specific interactions without increasing overall system complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If user-end video data is captured and displayed in real time, then the user participation sense is improved, but the processing complexity increases

Engineering Contradiction:
Improveuser participation senseVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system introduces an intermediary processing layer that receives video data from multiple users, performs selective processing based on interaction detection, and routes content to appropriate display regions. This intermediary layer manages processing complexity by not processing all user videos uniformly, but only those with detected interactions.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary detection of interactions between users and streamer before processing and displaying user-end video data. By detecting interactions first (through audio/video analysis), the system only processes and displays user videos when interactions are present, reducing overall processing complexity while maintaining enhanced user participation.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If video special effects are displayed based on voice instruction recognition, then the interaction presentation is enriched, but the recognition accuracy requirement increases

Engineering Contradiction:
Improveinteraction presentationVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system implements partial recognition by focusing on detecting specific preset voice instructions and interaction-related content rather than attempting to recognize all possible speech. This partial approach enriches interaction presentation for detected instructions while reducing the overall recognition accuracy burden compared to comprehensive speech recognition.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system uses feedback mechanisms where detected interactions and voice instructions trigger specific video special effects. This feedback loop allows the system to confirm recognition accuracy through the appropriateness of triggered effects, and adjust recognition parameters based on effect outcomes, effectively managing the accuracy requirement.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11889127B2Live video interaction method and apparatus, and computer device
Publication Date: 2024.01.30 SHANGHAI HODE INFORMATION TECH CO LTD
  • US11889127B2 patent drawing
  • US11889127B2 patent drawing
  • US11889127B2 patent drawing

AI summary

The present application discloses techniques for interaction during live video streaming. The techniques comprise obtaining and playing streamer-end video data, and user-end video data captured by a user terminal in real time; monitoring and recognizing whether the streamer-end video data comprise a preset voice instruction; determining whether the user-end video data comprises a target audio or a target video when the streamer-end video data comprises the preset voice instruction; and displaying a video special effect corresponding to the preset voice instruction in a user video when the user-end video data comprise the target audio or the target video. By means of the present application, a video special effect can be played for a user video according to a result of interaction between a streamer and a user, which enriches the way of interaction presentation and enhances the sense of participation in interaction.