Live Streaming Script Generation With Multimodal Character Coordination

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing digital character live streaming systems struggle with expressiveness, interactive engagement, and sales capabilities in high-interaction and emotion-driven scenarios, lacking flexibility, multimodal coordination, and customization, leading to unnatural presentations and reduced viewer engagement.

Innovation Solution

A method for generating a live streaming script that integrates object description sub-segments with speech content, using large language models to synchronize tone, actions, and facial expressions, enabling multimodal coordination and real-time adaptability, with deep knowledge integration and persona consistency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional live streaming systems use real hosts, then expressiveness and interaction are improved, but operational costs increase and uninterrupted streaming cannot be achieved

Engineering Contradiction:
Improveuninterrupted streaming capabilityVSAvoidoperational cost
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent creates a virtual digital character as a copy or simulation of a real host, using AI-generated scripts and synthesized speech to replicate human-like live streaming behavior. This allows uninterrupted 24/7 streaming without the costs and limitations of real human hosts while maintaining the core function of content delivery and viewer interaction.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system enables the virtual character to autonomously generate and deliver live streaming content through automated script generation, speech synthesis, and real-time interaction capabilities. The virtual host serves itself by continuously generating relevant content based on product information and viewer feedback without requiring human intervention for each streaming session.

Inventive Principle:
Principle #25Self-service

2Productivity

If virtual digital characters are used for live streaming, then operational costs are reduced and uninterrupted streaming is achieved, but expressiveness, interactive engagement, and sales capabilities deteriorate

Engineering Contradiction:
Improveoperational cost efficiencyVSAvoidexpressiveness and interaction quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements dynamic script generation that adapts in real-time based on product information, viewer feedback, and interaction context. The virtual character's speech content, tone, and interaction responses are dynamically adjusted to maintain expressiveness and engagement, rather than using static pre-recorded content. This allows the system to respond naturally to different streaming scenarios and viewer reactions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes multiple parameters including speech tone, interaction style, content depth, and response timing to optimize expressiveness and viewer engagement. By adjusting these parameters based on real-time feedback and context, the virtual character can simulate human-like variations in delivery that maintain audience interest and interaction quality despite being AI-generated.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If simple script generation is used for virtual characters, then system complexity is reduced, but naturalness and viewer engagement deteriorate

Engineering Contradiction:
Improvescript generation system complexityVSAvoidpresentation naturalness
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent divides the script generation process into distinct segments: product information processing, speech content generation, object description generation, and synthesis integration. Each segment handles specific tasks independently, allowing the system to manage complexity through modular design while maintaining natural-sounding output through coordinated integration of all segments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer that translates product information and interaction context into natural speech content and object descriptions. This intermediary layer includes modules for generating speech content text, object description sub-segments, and coordinating them with actions and facial expressions to produce natural-sounding virtual character presentations.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Reliability

If detailed object description sub-segments are integrated with speech content, then multimodal coordination is improved, but processing time and system complexity increase

Engineering Contradiction:
Improvemultimodal coordination qualityVSAvoidscript generation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary generation of speech content and object description sub-segments during the script creation phase, before actual live streaming begins. By pre-processing and organizing these elements with their associated actions and facial expressions, the system reduces real-time processing requirements during streaming while maintaining high multimodal coordination quality through carefully structured pre-generated content.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20260012690A1Method for generating living streaming script, electronic device, and storage medium
Publication Date: 2026.01.08 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US20260012690A1 patent drawing
  • US20260012690A1 patent drawing
  • US20260012690A1 patent drawing

AI summary

A method for generating a live streaming script, an electronic device and a storage medium are provided, which relate to the field of artificial intelligence technologies, in particular to the fields of natural language processing, large models, and virtual digital characters. The method for generating a live streaming script includes: generating at least one first script segment according to an initial input information, where the first script segment includes a speech content text and an object description sub-segment for a live streaming object, the object description sub-segment includes an object description text for describing at least one of an action presented by the live streaming object or a presentation mode for the speech content text; and determining the live streaming script according to the at least one first script segment.