Live broadcast method for AI digital human multi-modal cloning and real-time interaction
By integrating digital human cloning, interaction, and output technologies, AI digital humans with sales-driving capabilities are generated, solving the problems of insufficient sales performance and compliance in live-streaming e-commerce scenarios, and achieving simplified operation and continuous optimization from material input to live-streaming launch.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-03-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing AI digital human technology lacks a closed loop from cloning to promotion in live-streaming e-commerce scenarios, making it unable to deeply participate in the complete e-commerce conversion chain from traffic acquisition to order completion. Furthermore, it lacks compliance and risk control, resulting in insufficient sales performance.
By integrating digital human cloning, interaction, and output technologies into a continuous and automated process, deep learning models are used to extract live streaming business characteristics, generate digital humans with a sales-oriented feel, and combine product knowledge and compliance review in real time to achieve a one-click live streaming process.
It simplifies the process from material input to live broadcast launch, improves the conversion efficiency of the live broadcast room, reduces the risk of violations, and allows for continuous optimization of the live broadcast effect.
Smart Images

Figure CN121711501A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a live streaming method for multimodal cloning and real-time interaction of AI digital humans. Background Technology
[0002] With the deep integration of artificial intelligence and multimodal technologies, digital human technology has evolved from an early tool for visual representation into an "intelligent agent" with perception, decision-making, and execution capabilities, becoming one of the most representative application interfaces in the "AI+" era. In particular, driven by large-scale pre-trained model (LLM) technology, digital humans have achieved significant breakthroughs in core capabilities such as speech synthesis, facial expression-driven communication, and natural language interaction, enabling large-scale applications in customer service, content delivery, education and training, and other fields, effectively helping enterprises reduce costs and increase efficiency.
[0003] Existing industry reports indicate that digital humans have become one of the most mature and widely applicable AI terminals. In the live-streaming e-commerce sector, some manufacturers have launched highly persuasive digital humans capable of dynamic decision-making and real-time interaction. In some cases, the sales data even rivals that of top live-streaming hosts, demonstrating enormous commercial potential.
[0004] Against the backdrop of rapid technological development and application exploration, the industry has undertaken numerous technological innovations regarding the interactive capabilities and evaluation systems of digital humans. These existing technological solutions have laid the foundation for development in this field, but there remains a significant gap between their original design intentions and technological focus and the specific needs of live-streaming e-commerce, a scenario characterized by high concurrency, strong marketing, and strict compliance.
[0005] For example, Chinese patent document with publication number CN120387712A provides an AI digital human interaction evaluation method and device; the core of this solution is to define a set of multi-dimensional evaluation indicators including task completion ability, task complexity, and independent ability, which aims to discover problems in the digital human interaction process through systematic evaluation, and then optimize its general interaction experience.
[0006] In addition, another Chinese patent document with publication number CN120045069A describes an interactive system that includes modules such as multimodal information acquisition, emotion recognition, dialogue management, and reinforcement learning. Its goal is to make the interactive behavior of digital humans more natural and vivid through emotion recognition and strategy learning.
[0007] In summary, existing technological solutions primarily focus on enhancing the general capabilities of digital humans in "understanding user intent and emotions" and "providing human-like responses." However, the essence of live-streaming e-commerce is a real-time marketing arena with "achieving sales conversion" as its core objective. Existing solutions lack key business intelligence modules such as structured integration of product knowledge, reasoning of promotional logic, and real-time capture and guidance of sales opportunities. As a result, while the generated interactive content may have a "human touch," its "sales-driving power" is insufficient, and it cannot deeply participate in the complete e-commerce conversion chain from traffic acquisition to order completion. Summary of the Invention
[0008] Based on this, it is necessary to provide a live streaming method for multimodal cloning and real-time interaction of AI digital human technology in live e-commerce scenarios, which addresses the technical problems of existing AI digital human technology, such as deviation in interaction goals (focusing on general companionship rather than sales conversion), functional process breaks (lacking a closed loop from cloning to streaming), and lack of compliance and risk control.
[0009] Specifically, the present invention aims to: 1. Integrate the separate digital human cloning, interaction, and output technologies into a continuous, automated process, allowing users to drive a complete live stream simply by providing basic materials.
[0010] 2. The cloning process not only replicates the image and voice, but also captures and learns business characteristics related to live-streaming e-commerce, such as promotional tone and guiding gestures, to generate digital humans with a sales-oriented feel.
[0011] 3. Upgrade real-time interaction from simple Q&A to a data-driven intelligent marketing process aimed at sales conversion, deeply integrating product knowledge.
[0012] 4. Ensure that the entire process of generating and interacting with live content undergoes a compliance review by the system and can automatically adapt to the streaming requirements of the live streaming platform to achieve stable, low-risk, unattended live streaming.
[0013] To achieve the above objectives, this application provides a live streaming method for multimodal cloning and real-time interaction of AI digital humans. This method consists of a series of sequentially executed and parallelly processed steps, forming a complete operational flow from preparation to execution and optimization. This method can be implemented through a system composed of software and / or hardware, and is independent of any specific hardware deployment.
[0014] Specifically, a live streaming method for multimodal cloning and real-time interaction of AI digital humans includes the following steps: Step S100: Multimodal digital human cloning based on business characteristics.
[0015] This step provides the digital foundation for all subsequent interactions. Its core lies not only in cloning the image and voice, but also in extracting and integrating the characteristics of the live streaming business.
[0016] S101: Input and Preprocessing. Receives a seed video uploaded by the user, containing a person's image, voice, and product explanations or promotional actions. The system automatically performs frame segmentation, face detection and alignment, and separates the audio stream.
[0017] S102: Multimodal Feature Joint Extraction and Business Analysis. Parallel processing using deep learning models: Visual analysis: Facial features and expression sequences are extracted using convolutional neural networks, and specific gesture fragments related to explanation, demonstration, and guidance, such as pointing and picking up products, are specifically identified and labeled.
[0018] Audio analysis: Extracting acoustic features of speech, such as Mel spectrum and fundamental frequency, and identifying segments of speech with business-related tones such as "promotion," "emphasis," and "call to action" through speech recognition and text analysis.
[0019] S103: Business-oriented joint-driven model training. Based on the features extracted in step S102, a multi-task driven model is trained. This model takes text and driving parameters as input, and can not only generate realistic lip movements and facial expressions synchronized with the text, but also autonomously match and generate corresponding business gestures and tone patterns learned in S102 based on the semantic tags of the text, such as whether it is a promotional script, thereby outputting a digital human driven model with basic sales performance capabilities.
[0020] Step S200: Real-time interaction that integrates bullet screen-driven and product knowledge.
[0021] This step processes user input during the live stream and generates targeted responses. At its core, the interaction aims to facilitate conversion and is deeply integrated with the product.
[0022] S201: Real-time bullet screen capture and intent recognition. The bullet screen stream is acquired in real-time from the live streaming platform interface and then cleaned and denoised. A natural language understanding model finely tuned for e-commerce live streaming scenarios is used to classify the intent of each bullet screen, such as: asking about product prices, inquiring about product materials, requesting detailed display, asking for discounts, and general interaction.
[0023] S202: Product Entity Linking and Knowledge Retrieval. The system maintains a real-time updated live-stream product knowledge base. It associates the intent identified in S201 with the currently explained product or a list of historical products. Based on the "intent-product" pair, it quickly retrieves structured standard answers, selling point information, or related promotional activities from the knowledge base.
[0024] S203: Personalized Marketing Response Generation and Digital Human-Driven Approach. The retrieved information is populated into a response template matching the digital human's style to form the final response text. If the comment contains a valid username, "@username" is embedded in the response, triggering the digital human to generate a subtle personalized feedback action, such as a nod. Finally, this text is input into the digital human-driven model generated in step S100 to synthesize a voice response with corresponding facial expressions, lip movements, and related business gestures.
[0025] Step S300: Virtual synthesis and compliance real-time streaming output.
[0026] This step transforms the digital human content generated in the previous steps into a signal that can be received by the live streaming platform. Its core purpose is to achieve high-quality, stable video stream output that meets the platform's requirements.
[0027] S301: Real-time compositing of multi-layered images. In the rendering engine, the digital human model driven by step S200 is combined with preset virtual backgrounds, such as those supporting green screen keying and real-time image and text layers, such as product labels and scrolling bullet comments.
[0028] S302: Virtual-Real Fusion and Virtual Camera Output, optional but recommended. To improve the realism of the live stream and circumvent platform detection, a unique step can be employed: capture footage from a real camera, adjust its transparency to an extremely low value that is barely perceptible to the naked eye, such as α=0.05, and use this as the bottom layer to perform alpha blending with the digital human image synthesized by S301. Finally, through the virtual camera driver interface, output the blended video frame sequence as a virtual video device signal.
[0029] S303: Simultaneous Compliance Review and Execution. Before the response text is generated in step S200 and the final broadcast occurs in step S300, the text content must pass through a parallel compliance filtering engine. This engine performs at least two levels of filtering: a first-level matching filter based on a sensitive word database; and a second-level semantic risk assessment based on a context model. Content deemed to be in violation will be replaced in real time with preset safety wording or silently skipped to ensure broadcast safety.
[0030] Step S400: Automated field control and feedback-based strategy iteration.
[0031] This step manages the live streaming process and optimizes its effectiveness. Its core purpose is to achieve live streaming operations and self-improvement without human intervention.
[0032] S401: Automated execution of preset processes. Based on the script prepared before the live stream, the system automatically performs operations such as scheduled start / end of the live stream, switching between products to be explained, and triggering marketing activities, such as distributing lucky bags, pop-up product links, and broadcasting the live stream control script.
[0033] S402: Real-time monitoring and protection, like a "live stream bodyguard." During the live stream, it monitors the video stream in parallel, preventing static and audio streams, as well as abnormal and interactive data. Once an anomaly is detected, it automatically triggers emergency plans, such as switching to backup scripts or issuing alarms.
[0034] S403: Data Collection and Reinforcement Learning Iteration. Collect multi-dimensional data throughout the live stream, such as user interaction behavior, product exposure and click data, and conversion data. Periodically use this data to train a reinforcement learning model. This model takes the live stream status as input, adjusts live stream strategies (such as changing the style of speech and interaction frequency) as actions, and aims to improve business metrics such as product conversion rate and extend viewing time as reward objectives, thereby continuously optimizing the interaction and session control strategies in steps S200 and S400.
[0035] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Closed-loop process, one-click start: The scattered technical links are integrated into an end-to-end automated method, which greatly reduces the user's operational complexity and technical threshold, and realizes a one-click process from "material input" to "live broadcast start".
[0036] 2. Cloning for Commercial Use: By injecting business feature learning during the cloning stage, the generated digital human is "naturally" capable of driving sales, solving the problem of the disconnect between general digital humans and commercial live streaming scenarios.
[0037] 3. Interaction equals conversion: Real-time interaction is deeply integrated with product knowledge and marketing goals, turning every user interaction into a potential sales opportunity and significantly improving the conversion efficiency of the live streaming room.
[0038] 4. Secure Output: The embedded parallel compliance review and virtual-real fusion output technology ensure content security while improving the platform compatibility and stability of the live stream, reducing the risk of violations and the probability of operational interruption.
[0039] 5. Evolution through Operation: Through a closed-loop data collection and strategy iteration mechanism, the entire live streaming method can continuously optimize itself during use, and the live streaming effect, such as the conversion rate, can continuously improve over time, thus possessing long-term value. Attached Figure Description
[0040] Figure 1 This is an integrated flowchart of the live streaming method for multimodal cloning and real-time interaction of AI digital human according to the present invention. Figure 2 This is a flowchart of the training process for the multimodal cloning module of a live streaming method for multimodal cloning and real-time interaction of AI digital humans, as described in this invention. Figure 3This is a flowchart of the real-time interaction engine for a live streaming method of AI digital human multimodal cloning and real-time interaction according to the present invention. Figure 4 This is a virtual synthesis and streaming system architecture diagram of a live streaming method for multimodal cloning and real-time interaction of AI digital humans according to the present invention. Figure 5 This is a schematic diagram illustrating the intelligent field control and strategy iteration principle of a live streaming method for multimodal cloning and real-time interaction of AI digital humans according to the present invention. Figure 6 This is an overview diagram of the system deployment and data interaction of a live streaming method for multimodal cloning and real-time interaction of AI digital humans according to the present invention. Detailed Implementation
[0041] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings and specific examples. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0042] Specifically, the live streaming method for multimodal cloning and real-time interaction of AI digital humans provided by this invention can be implemented on a computing device configured with a high-performance GPU, such as a computer or mobile phone, and equipped with necessary deep learning frameworks, such as PyTorch or TensorFlow, and streaming media service components. Figure 1 As shown, the implementation of this invention involves four core processing flows, which are coordinated through a central task scheduler: 1. Offline training process: Execute multimodal digital human cloning based on business characteristics, such as... Figure 2 As shown.
[0043] 2. Online Interaction Process: Execute real-time interaction that integrates bullet-screen comments with product knowledge, such as... Figure 3 As shown.
[0044] 3. Rendering and Streaming Process: Perform virtual compositing and compliant real-time streaming output, such as... Figure 4 As shown.
[0045] 4. Management and Optimization Process: Implement automated site control and feedback-based strategy iteration, such as... Figure 5 As shown.
[0046] Furthermore, in a specific implementation, the multimodal digital human cloning based on business characteristics, corresponding to step S100, is implemented as follows: 1.1 Input preprocessing and feature extraction, corresponding to steps S101-S102 During implementation, the system provides standardized video capture guidelines. The recommended seed video is an MP4 file that is 5-10 minutes long, 1080p resolution, with a clean background, showing a frontal half-body view of the subject and clearly demonstrating product demonstration actions. The preprocessing module uses OpenCV and Dlib libraries to automatically perform the following operations: Video framing and alignment: Frames are extracted at 30fps, and face detection and key point localization are performed using MTCNN or RetinaFace models to complete face alignment based on affine transformation, ensuring that all facial images are normalized.
[0047] Audio separation and processing: The Librosa library is used to separate the audio, and noise reduction and normalization are performed, with the sampling rate uniformly set to 16kHz.
[0048] Business Feature Annotation: The system calls a pre-trained "sales action recognition model" to analyze video frames. This model is a classification network trained on a manually labeled dataset containing categories such as "pointing to the screen," "holding a product for demonstration," "contrast gestures," and "gestures," such as a ResNet-based temporal model. Simultaneously, the audio text, after ASR conversion, is fed into a BERT-based fine-tuned "business tone classification model," which annotates segments such as "neutral statements," "emphasis on selling points," "call to action," and "creating a sense of urgency."
[0049] 1.2 Business-oriented joint-driven model training, corresponding to step S103 This step implements a multi-task learning architecture: Backbone network: Employs a non-autoregressive end-to-end speech-driven model, such as an improved version of MakeItTalk or SadTalker. This model takes acoustic features, including phoneme sequences and fundamental frequencies, as input.
[0050] Multi-task learning head: a. Facial Expression and Lip Shape Generation Head: This is the main task, outputting the vertex displacements or keypoint coordinates of the face mesh.
[0051] b. Business Gesture Generation Head: This is an auxiliary task. During training, the model receives not only audio features but also business action label vectors aligned with the audio timeline, generated in step S102. This auxiliary head is responsible for generating parameters controlling the upper body skeleton of the digital human, such as wrist and elbow movements, at the corresponding time points, driving predefined gesture animation primitives corresponding to the labels, such as Gesture_Point_At_Screen and Gesture_Hold_Product. Through joint training, the model learns the mapping relationship that "when the voice text is classified as 'selling point emphasis,' there is a high probability of triggering the 'pointing to screen' gesture."
[0052] c. Intonation style embedding: The business tone classification label obtained in step S102 is used as a style embedding vector, which is concatenated with the phoneme features and then input into the model to make the prosody of the synthesized speech, such as pitch and speech rate, adaptively adjusted according to different tones such as "promotional appeal" or "creating a sense of urgency".
[0053] Training and Output: The dataset is trained using audio-video annotations with business feature annotations. The loss function is a weighted sum of face reconstruction loss, gesture classification loss, and speech feature reconstruction loss. After training, a unified model file is exported, which can synchronously drive the digital human's lip movements, facial expressions, head poses, and business gestures based on the input text and optional action style labels.
[0054] Furthermore, in one specific implementation, the real-time interaction based on the fusion of bullet screen drive and product knowledge, corresponding to step S200, is implemented as follows: 2.1 Real-time bullet screen processing and intent recognition, i.e., step S201 The system connects to the bullet comment stream via the live streaming platform's official Webhook or WebSocket connection. During implementation, an asynchronous message queue, such as RabbitMQ or Kafka, is established to handle high-concurrency bullet comments.
[0055] Cleaning module: Filters out pure emojis, characters shorter than 2 characters, and high-frequency repetitive messages, such as three identical messages sent by the same user within 5 seconds.
[0056] Intent Recognition Module: Deploys a lightweight e-commerce intent classification model. This model is based on the ALBERT or RoBERTa-tiny architecture and fine-tuned on a self-built dataset containing millions of live stream comments and their manually annotated intents. The intent labeling system is specifically designed for sales promotion, for example: { "ask_price":"Inquire about the price", "ask_material":"Inquire about material / composition", "ask_spec":"Inquire about specifications / size", "ask_promotion":"Inquire about discounts / coupons", "ask_effect":"Inquire about the effect / function", "request_demo":"Request to show / try on", "positive_feedback":"positive feedback (looks good / want it)", "interactive":"Interactive (press 1, here it comes)", "other": "other" } The model outputs one or more intent labels and confidence scores for each bullet comment.
[0057] 2.2 Product entity linking and knowledge retrieval, i.e., step S202 This system maintains an in-memory database, such as a real-time product knowledge graph in the form of Redis.
[0058] Graph structure: Each product is a node, with attributes including product ID, name, current price, list of core selling points, detailed parameters in JSON format, FAQ list with answers to frequently asked questions, and associated marketing campaigns.
[0059] Entity Links: Using a semantic similarity model based on SimBERT or Sentence-BERT, the bullet screen text is matched with the product list name and alias in the current live stream. For example, the bullet screen text "How much is this little brown bottle?" can correctly link to the product "XX brand repair essence third generation".
[0060] Knowledge Retrieval: Perform efficient key-value queries based on "intent-product" pairs. For example, for the intent ask_price and product A, directly return product A.current_price and additional discount information; for the intent ask_effect, return the description of the effect from product A.core_selling_points.
[0061] 2.3 Personalized response generation and driving, i.e., step S203 Reply Templates and Fillers: The system provides multiple stylized reply templates for each common "intent-product" combination. For example, for ask_price + product A, the template might be: "@[Username] has great taste! This [Product A Name] is currently available at a live stream exclusive price of [Price]! Click the yellow shopping cart link below (number 1) to receive an extra [gift] when you order today!" Personalized trigger: If the comment contains a username, the system will verify and extract it from the user list and fill it into the [username] placeholder in the template. At the same time, a control instruction tag is generated, such as {"action":"greet","target":"user"}, which is passed to the digital human-driven model to trigger a short nodding or smiling animation facing the virtual questioner.
[0062] Content security pre-screening: Before the generated response text is synthesized into speech, it will be sent to a fast compliance filtering submodule for the first round of risk screening.
[0063] Furthermore, in a specific implementation, the real-time streaming output based on virtual synthesis and compliance is implemented as follows, corresponding to step S300: 3.1 Real-time image compositing, i.e., step S301 In practice, a real-time rendering engine is used, such as Unity's real-time rendering pipeline or OBS's plugin system.
[0064] Digital Human Rendering Layer: Loads the model trained by S100, receives text and audio streams from S203, and renders digital human video streams with facial expressions, lip movements, and business gestures in real time.
[0065] Background and Image / Text Layers: Supports importing images and videos as backgrounds, and applies ChromaKey (chroma keying) and other technologies to remove green screens. Through the engine's UI system, semi-transparent product images, dynamic price tags, and real-time scrolling featured comments can be overlaid at specified locations on the screen.
[0066] Synthesis: Alpha blending of the digital human layer, background layer, and image / text layer according to a predefined Z-order.
[0067] 3.2 Virtual-real fusion and virtual camera output, i.e., step S302 This is a key point in improving the "authenticity" of live streams to meet platform review requirements.
[0068] Real-world video capture: Capture video from a single USB camera at extremely low resolutions, such as 640x480, and frame rates, such as 5fps, via OpenCV or DirectShow interfaces. This low-load setup aims to minimize resource consumption.
[0069] Transparency blending: The captured real-world image is used as the bottom layer, and the alpha value (transparency) of each pixel is set to a very low constant α, such as α=0.05. Its typical range is 0-1, with 1 being completely opaque. This means that the real-world image contributes only 5%, which is almost indistinguishable to the naked eye, but leaves a weak, random environmental noise in the pixel data of the video stream.
[0070] Virtual Camera Creation and Streaming: Using the OBS-VirtualCam SDK or ManyCam API, create a virtual camera device named AI_Anchor_Cam in the operating system. The final mixed video frames, typically 1920x1080@30fps, are continuously written to the virtual device's buffer via callback functions provided by the SDK. Users can then simply select AI_Anchor_Cam as the video source in software such as Douyin Live Companion or Kuaishou Live Companion to start streaming.
[0071] 3.3 Simultaneous compliance review, i.e., step S303 The compliance engine runs as a separate service, reviewing all text and images to be output.
[0072] First-level filtering, or keyword matching: Based on the AC automaton algorithm, it matches a local thesaurus containing prohibited words on the platform and absolute terms banned by the advertising law. If a match is found, it directly triggers replacement or blocking.
[0073] Secondary filtering, or contextual semantic review: For content that passes the primary filter, a lightweight risk semantic classification model is invoked. This model can identify contexts such as "exaggerated claims," "implies of medical effects," and "unfair competition." For example, it identifies "This face cream can reverse aging" as high-risk.
[0074] Execution strategy: Based on the risk level, the system performs different actions: low risk allows passage; medium risk replaces keywords, such as replacing "most*" with "very"; high risk immediately interrupts the current script and switches to a preset safe transition script, such as "Let's take a look at another feature of the product...", and alerts are triggered through logs and notifications.
[0075] Furthermore, in a specific implementation, the implementation of step S400 based on automated field control and feedback-based strategy iteration is as follows: 4.1 Pre-set process automation, i.e., step S401 The system provides a visual task orchestrator. Operators can drag and drop components to create live streaming workflows as shown below: timeline: -time:"19:00:00" action: "start_streaming" -time:"19:05:00" action:"switch_product" params:{product_id:101} -time:"19:15:00" action: "trigger_lottery" # Triggers a lucky draw params:{duration_sec:300} -time:"19:20:00" action:"popup_product_card"#Popup product card params:{product_id:102} -time:"21:00:00" action: "end_streaming"
[0076] The orchestrator converts these tasks into precise timing instructions, which are then sent to the corresponding execution modules via message queues.
[0077] 4.2 Real-time monitoring and protection, i.e., step S402 "Live Stream Guardian" runs as a daemon process and performs the following monitoring: Still image detection: Calculate the Structural Similarity Index (SSIM) between consecutive video frames. If the SSIM value is higher than the threshold of 0.98 for 10 consecutive seconds, the image is determined to be still, triggering an alarm and automatically switching to a fault prompt screen.
[0078] Audio anomaly detection: Monitors the root mean square energy (RMS) of the audio stream. If the RMS remains close to zero for 30 seconds (i.e., silence), or if a momentary sharp pop occurs, an alarm is triggered.
[0079] Action taken: After the alarm is triggered, the system can automatically execute a predefined script, such as playing background music or using a digital human to announce a message like "The host is going to get a sample and will be right back," thus buying time for human intervention.
[0080] 4.3 Data collection and reinforcement learning iteration, i.e., step S403 This is the core of achieving self-optimization of the method provided by this invention.
[0081] Data collection: The system records all live streaming data, such as bullet comments, user behavior, product exposure and clicks, order conversions, etc., in a time-series database, such as InfluxDB.
[0082] State definition and policy modeling: The live streaming process is modeled as a Markov decision process.
[0083] State, S: The state vector at time t, can be represented as: S_t = [number of online users, like rate, click rate of product A, current speech style code, broadcast duration, ...].
[0084] Action, A: The operation that the system can perform, such as A={switch to product B, change the wording style to 'promotional', initiate a lottery, ...}.
[0085] Reward, R: Design a multi-objective reward function, for example: R_t = +0.5*Δ(conversion rate) + 0.3*Δ(average viewing time) - 0.2*(user churn rate) Where Δ represents the change compared to the previous time window. The reward function directly transforms the business objective into the model's optimization objective.
[0086] Offline Training and Online Updates: Regularly, such as every morning at midnight, a Deep Q-Network (DQN) or Proximal Policy Optimization (PPO) agent is trained using the collected (S_t, A_t, R_t, S_{t+1}) sequence as an experience pool. The trained policy model is deployed online as a "policy suggestion module." In subsequent live streams, this module analyzes the current state S_t in real time and provides action suggestions A_t. Initially, these suggestions may only serve as a reference for operations personnel; once the confidence level is high enough, they can be configured to be executed automatically within specific safety boundaries, thereby achieving continuous autonomous evolution of the live stream policy.
[0087] More specifically, the present invention also provides a complete embodiment, as follows: Taking a live-streaming sales event for apparel as an example, this method illustrates how it works collaboratively: 1. Preparation stage, i.e., before the broadcast: The operators uploaded a seed video featuring a real model explaining the T-shirt, along with product information for the T-shirt.
[0088] The system executes S100 to train a digital human that can mimic the model and naturally make gestures such as "showing the print" and "gesturing the pattern".
[0089] The system performs some of the functions of the S300 and configures virtual backgrounds, such as fashion street scenes and product image and text layers.
[0090] In the S400's editor, set up a process where each garment is explained for 5 minutes and a prize draw is held on the hour.
[0091] 2. Live Stream Execution Phase: Automatic broadcast starts at the designated time (S401). The digital human begins its presentation according to the generated script.
[0092] User comment: "Model, won't this outfit make you look fat?"
[0093] The bullet screen was captured (S201), and the intent was identified as ask_effect, which means asking for the effect.
[0094] The entity is linked to the T-shirt (S202) currently being discussed.
[0095] The system retrieves selling points such as "slim fit" and "draping fabric" from the knowledge base (S202).
[0096] Generate a reply: "@Xiaomei, this is a slim fit style with a great drape, it hides any extra weight!", and trigger the digital human to make a side-view gesture (S203).
[0097] The response text passes compliance review (S303), the voice and animation are synthesized (S301), and finally broadcast through a virtual camera (S302).
[0098] The "Live Stream Bodyguard" (S402) operated stably throughout the entire process.
[0099] 3. Optimization phase, i.e., after the live stream: System analysis of the data revealed that when the digital human made a "comparison of clothing" gesture, the click-through rate of the product increased by an average of 15%.
[0100] The reinforcement learning model (S403) learns this pattern and incorporates it into the strategy. In future live streams, when explaining similar products, the system will more frequently suggest or directly trigger the "comparison wearing" gesture. Through this closed loop, this method can continuously improve the interactivity and conversion efficiency of live streams.
[0101] Furthermore, in a specific implementation method, the specific implementation method of virtual-real fusion and virtual camera output, corresponding to the specific implementation method of step S302, aims to solve the long-standing technical pain point in the industry: pure virtual digital human live streams are easily judged as violations by live streaming platforms' "real person detection" or "non-real person content recognition" algorithms, leading to traffic restrictions, demotion, or interruption of streaming. Existing technologies typically only focus on generating and outputting virtual images, failing to fundamentally address this platform-level risk. This invention introduces an original low-opacity real image Alpha mixing technology to fundamentally modify the signal attributes of the output video stream, thereby significantly improving the platform approval rate and stability of the live stream; the specific method flow is as follows: 1. System architecture and parallel processing flow like Figure 4 As shown, this system adds a lightweight real signal acquisition and processing channel in parallel with the core virtual rendering pipeline, and finally synthesizes the signals through a dedicated mixer. The specific implementation is as follows: Virtual Rendering Main Channel: Receives the digital human driving instructions and audio output from the previous stage, and generates a high-resolution virtual main screen layer, such as 1920x1080, with a high frame rate, such as 30fps, by the digital human rendering engine.
[0102] Real signal acquisition channel: Signal source: The system automatically detects and calls up an available physical camera on a computer or mobile device, such as a USB webcam, a built-in camera on a laptop, or a built-in camera on a mobile phone.
[0103] Low-load acquisition: To avoid overburdening the system, this channel acquires data with significantly reduced parameters. Typical implementation parameters are: resolution set to 640x480, frame rate set to 5fps. The purpose of this setting is not to obtain clear visual content, but to continuously capture a video stream containing unstructured, random physical signals such as changes in real-world ambient light and sensor noise with minimal computational and I / O overhead.
[0104] Dedicated Alpha Mixer: This is a separate image processing unit responsible for receiving and processing the two signals mentioned above.
[0105] 2. Core processing steps: Transparency blending and signal synthesis The mixer's processing flow includes the following core steps, the key of which lies in the precise control of the transparency of the real image: Step 1: Real-world image preprocessing. The captured low-resolution real-world image is quickly scaled to the exact same size as the virtual main image layer, such as 1920x1080, using a bilinear interpolation algorithm. This step is only to ensure pixel alignment and does not prioritize image quality.
[0106] Step Two: Transparency, or Alpha value setting. A uniformly low and fixed transparency constant α is assigned to each pixel in the processed real-world image. Through extensive comparative experiments and effect verification, this invention has determined that the optimal value range for this parameter is 0.02 to 0.05, including the endpoints. In a preferred embodiment, α is set to 0.03.
[0107] Technical Principles and Parameter Basis: When the α value is greater than 0.1, interference from real-world images may be perceptible to the viewer, impairing the virtual streamer's viewing experience. When the α value is less than or equal to 0.01, the introduced signal is too weak to effectively influence the platform's detection algorithm. The range of 0.02 to 0.05 precisely allows for the injection of continuous and random real-world environmental noise into the video stream's pixel data, ensuring that the impact on the final image's subjective visual quality is negligible (virtually imperceptible to the human eye). This weak noise provides a continuous and effective feature for platform detection algorithms that rely on identifying patterns in "purely synthetic signals," indicating that the signal source contains "real-world components."
[0108] Step 3: Pixel-level Alpha blending. The mixer performs blending calculations for each corresponding pixel in each frame according to the following formula: Output pixel RGB value = Virtual image pixel RGB value × (1-α) + Real image pixel RGB value × α Because the α value is extremely small, the synthesized result is visually very close to the original virtual image, but at the signal level it has changed from a "purely synthesized signal" to a "mixed signal dominated by the synthesized signal and mixed with a small amount of real physical noise".
[0109] 3. Virtual camera output The synthesized final video frame sequence is continuously written to a virtual video device in the operating system by calling the virtual camera driver interface, such as using OBS-VirtualCamSDK or a similar development kit, to create a virtual camera named "AI_Anchor_Cam". Subsequently, users only need to select this virtual camera as the video source in any standard live streaming software, such as OBSStudio or Douyin Live Companion, to push the processed live stream to the live streaming platform.
[0110] Based on the implementation of this step, the present invention has achieved unexpected technical effects: Proactively mitigate risks: Instead of passively accepting platform detection, proactively preprocess the output signals to ensure compliance, fundamentally reducing the technical risk of being misjudged as "non-live broadcast" or "recorded broadcast".
[0111] Guaranteed visual quality: While achieving the above effects, the high-quality visual performance of the original digital human live stream is fully preserved, and the user experience is unaffected.
[0112] Low implementation cost: The added parallel processing channels have extremely low load, require no expensive hardware, and can run smoothly on ordinary commercial computers, greatly improving the practicality and scalability of the invention.
[0113] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0114] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. A live streaming method for multimodal cloning and real-time interaction of AI digital humans, characterized in that, It includes the following steps: Step S100: Multimodal digital human cloning based on business characteristics: Train and generate a digital human-driven model with sales performance based on seed videos containing product explanation actions and voices. Step S200: Real-time interaction driven by bullet comments and product knowledge integration, capturing and analyzing bullet comments in the live broadcast room in real time, generating personalized marketing responses by combining product knowledge graphs, and driving the digital human to broadcast; Step S300: Virtual synthesis and compliance-compliant real-time streaming output, the digital human image is synthesized with the virtual background and output to the live streaming platform through a virtual camera; Step S400: Automated live streaming control and feedback-based strategy iteration, execute the preset live streaming process, monitor the live streaming status, and optimize the interaction and live streaming control strategies based on the live streaming data.
2. The live streaming method for multimodal cloning and real-time interaction of AI digital humans according to claim 1, characterized in that, Step S100 includes: S101: Perform frame segmentation, face alignment, and audio separation on the seed video; S102: Perform business gesture recognition and annotation on video frames, and classify business tone in audio. S103: Based on the gesture and tone labels, train a multi-task driven model to synchronously generate the lip movements, facial expressions, and gestures that match the business semantics of the digital human.
3. The live streaming method for multimodal cloning and real-time interaction of AI digital humans according to claim 2, characterized in that, Step S200 includes: S201: Clean up real-time bullet comments and classify them using an e-commerce scenario intent recognition model; S202: Link the identified intent with the live-streamed product as an entity, and retrieve the corresponding structured information from the product knowledge graph; S203: Generate a reply text based on the search information and reply template, and after compliance pre-review, drive the digital human to synthesize a voice reply with personalized actions.
4. The live streaming method for multimodal cloning and real-time interaction of AI digital humans according to claim 3, characterized in that, Step S300 includes: S301: Perform real-time image synthesis of the digital human rendering layer, virtual background layer, and graphic overlay layer; S302: Acquire a low-resolution real camera image and blend it with the synthesized image using a preset extremely low transparency α; S303: Before the response is generated and the final broadcast is conducted, a compliance review is carried out through a multi-level filtering engine that includes keyword matching and semantic risk judgment.
5. The live streaming method for multimodal cloning and real-time interaction of AI digital humans according to claim 4, characterized in that: In step S302, the value of the extremely low transparency α ranges from 0.02 to 0.
05.
6. The live streaming method for multimodal cloning and real-time interaction of AI digital humans according to claim 5, characterized in that, Step S400 includes: S401: Based on the visually orchestrated script, automatically execute scheduled broadcasts, product switching, and marketing campaign triggering; S402: Monitor the abnormal status of video and audio streams in real time and automatically trigger the contingency plan when an abnormality is detected; S403: Collect live streaming data to train a reinforcement learning model, which takes the live streaming room status as input and actions to improve business metrics as output, and is used to optimize live streaming strategies.
7. A live streaming method for multimodal cloning and real-time interaction of AI digital humans according to claim 6, characterized in that, In step S102, the video frames are classified using a pre-trained sales action recognition model to identify at least one business gesture among pointing, displaying, and comparing; and at least one business tone among promotion, emphasis, and appeal is identified using a business tone classification model.
8. A live streaming method for multimodal cloning and real-time interaction of AI digital humans according to claim 7, characterized in that, In step S202, the product knowledge graph is maintained in the form of an in-memory database, and the node attributes include product ID, price, list of selling points, and answers to frequently asked questions.
9. A live streaming method for multimodal cloning and real-time interaction of AI digital humans according to claim 8, characterized in that, In step S402, static scenes are detected by calculating the structural similarity index of consecutive video frames, and audio anomalies are detected by calculating the root mean square energy of the audio stream.
10. A live streaming method for multimodal cloning and real-time interaction of AI digital humans according to claim 9, characterized in that, In step S403, the reward function of the reinforcement learning model is constructed based on the changes in conversion rate, average viewing time, and user churn rate.
Citation Information
Patent Citations
AI digital human interaction system
CN120045069A
Interaction evaluation method and device for AI (artificial intelligence) digital human
CN120387712A