Stream Merging for Speech Co-Hosting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In live stream co-hosting scenarios, there is a challenge in reducing the high uplink bandwidth pressure and CPU/GPU resource consumption at the co-hosting guest terminal while ensuring the audience can perceive the presence of the co-hosting guest.

Innovation Solution

A method and device for stream merging in speech co-hosting, which involves obtaining speech streams and images from both the co-hosting user and the live streamer, merging these streams, obtaining an image of the co-hosting user, and encoding it with the merged streaming data to create a final merged streaming data that includes both users' images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If the co-hosting terminal transmits its own video stream and audio stream separately, then the audience can perceive the presence of the co-hosting guest, but the uplink bandwidth pressure and CPU/GPU resource consumption at the co-hosting terminal increase significantly

Engineering Contradiction:
Improveaudience perception of co-hosting guest presenceVSAvoiduplink bandwidth pressure and CPU/GPU resource consumption
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The patent extracts the video processing and transmission task from the co-hosting terminal, leaving only audio transmission. The co-hosting terminal sends only its audio stream to the anchor terminal, which then synthesizes a virtual avatar video representing the co-hosting guest. This extraction reduces the co-hosting terminal's bandwidth and computational burden while maintaining audience perception of the guest's presence through the synthesized avatar.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The anchor terminal acts as an intermediary that receives the co-hosting guest's audio, generates a virtual avatar video, and synthesizes it with the anchor's own video stream. This intermediary processing allows the co-hosting guest to be represented visually without requiring the guest's terminal to perform heavy video encoding or transmission, thus resolving the contradiction between visual presence and resource consumption.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If the co-hosting terminal generates and transmits its own video stream, then the audience can see the co-hosting guest, but the device complexity and resource requirements at the co-hosting terminal increase

Engineering Contradiction:
Improvevisual presence of co-hosting guestVSAvoidvideo processing and transmission complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent removes the video generation and transmission function from the co-hosting terminal, transferring this complexity to the anchor terminal. The co-hosting terminal only needs to capture and transmit audio, which is computationally simple. The anchor terminal then creates the visual representation through avatar synthesis, eliminating the need for the co-hosting terminal to have sophisticated video processing capabilities.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of transmitting the actual co-hosting guest's video feed, the system creates a synthetic copy in the form of a virtual avatar at the anchor terminal. This avatar is generated from the co-hosting guest's audio input and is then combined with the anchor's video stream, providing visual representation without requiring the original terminal to perform complex video operations.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250184378A1Method and device of stream merging for speech co-hosting
Publication Date: 2025.06.05 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20250184378A1 patent drawing
  • US20250184378A1 patent drawing
  • US20250184378A1 patent drawing

AI summary

The disclosure provides a method, device, electronic device, computer-readable medium, computer program product, and computer program for stream merging for speech co-hosting. In the method, a device of stream merging for speech co-hosting obtains a first speech stream comprising speech information of a co-hosting user corresponding to a co-hosting terminal; obtains a second speech stream and a first image, the second speech stream comprising speech information of a live streamer user corresponding to a live streamer terminal, the first image comprising image information of the live streamer user corresponding to the live streamer terminal; merges the first speech stream, the second speech stream, and the first image to obtain first merged streaming data; obtains a second image indicating image information of the co-hosting user; and encodes the second image and the first merged streaming data to obtain second merged streaming data.