A performance virtual digital human generation method and system

By collecting and analyzing audio and video streams of real actors in real time, and combining semantic recognition and reinforcement learning frameworks, the system generates lines and action instructions that match the style of the target actors. This solves the problem of improvisation and synchronous interaction of virtual digital humans in crosstalk performances, achieving highly realistic virtual digital human performances and enhancing the artistic effect and audience experience.

CN122289478APending Publication Date: 2026-06-26北京琰东艺术团
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
北京琰东艺术团
Filing Date
2026-03-25
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing technologies cannot enable improvisation and live interaction in crosstalk performances, resulting in a lack of artistic effect and immersive experience for the audience in human-machine collaborative performances. Virtual digital humans cannot adapt to the unique styles and immediacy requirements of different actors, and cannot achieve real-time multimodal collaborative generation of voice, facial expressions, and body movements.

Method used

The system captures real-time audio and video streams of live actors using audio and video acquisition devices, generates multimodal data packets, extracts text semantics and facial expressions using semantic recognition and sentiment analysis models, and combines character replication models and reinforcement learning frameworks to generate lines and action instructions that match the style of the target actors, enabling synchronous interaction between virtual digital humans and live actors.

Benefits of technology

It enables virtual digital humans to improvise and perform with high realism in crosstalk performances, with end-to-end latency controlled within 50ms, ensuring synchronization and artistic effect with real actors, and enhancing the audience's emotional resonance and artistic appeal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122289478A_ABST
    Figure CN122289478A_ABST
Patent Text Reader

Abstract

This invention provides a method and system for generating virtual digital humans for performances. The method includes: real-time acquisition of raw audio and video streams from live actors' performances, and generating multimodal data packets that are synchronized with the live actors' performance states in real time through data preprocessing; identifying the textual semantic information corresponding to the voice data and labeling it with corresponding emotional tags to form input data packets; inputting the input data packets into a character replication model to analyze the performance style characteristics under the current performance context; generating subsequent target dialogue text that conforms to the target actor's performance style based on the target actor's historical dialogue data; inputting the subsequent target dialogue text into a response generation model to obtain virtual digital human driving instruction packets to generate the virtual digital human's audio stream, facial expressions, and body movement animations, and rendering them on a display terminal. This enables highly realistic virtual digital human performances with creative capabilities, promoting the digital innovation and industrialization of traditional stage arts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of virtual reality performance technology, specifically to a method and system for generating virtual digital humans for performances. Background Technology

[0002] The technological development in the current crosstalk performance field is diversified, but the mainstream performance format still mainly relies on live performances by live actors. Audiences experience a purely live artistic presentation. With the development of digital media technology, some performances have begun to introduce pre-recorded video clips or pre-made animated characters to interact with live actors. However, these interactions are mostly one-way outputs; the timing and content of the virtual content are determined before the performance and cannot be adjusted according to the actual situation on site. In the field of human-machine collaborative performances, existing virtual digital human technology is mostly applied to scenarios such as news broadcasting, customer service Q&A, or simple talent displays. Its interaction logic is mostly based on preset scripts or limited keyword triggers, lacking the ability to understand complex contexts and improvise. In recent years, although the field of artificial intelligence has made significant progress in speech recognition and natural language processing, integrating these technologies into crosstalk performance scenarios that require high immediacy, artistry, and emotional resonance still faces many technical bottlenecks and has not yet formed a mature technical solution.

[0003] When applied to the specific art form of crosstalk performance, existing technologies have the following limitations that restrict the artistic effect of human-computer collaborative performance and the audience's immersive experience: Insufficient real-time performance: The essence of crosstalk performance lies in the seamless rhythmic coordination between the "supporting" and "teasing" roles, typically requiring response times controlled at the millisecond level. However, existing human-computer interaction systems often experience delays of hundreds of milliseconds or even longer when completing the entire chain from speech recognition, semantic understanding, sentiment analysis, content generation to speech synthesis. This delay completely disrupts the continuity of the crosstalk performance and the timeliness of the jokes, significantly diminishing the comedic effect. Poor role adaptability: Crosstalk is an art form with a highly individual style. Actors from different schools and different pairings have drastically different supporting and teasing logics and language habits. Even if existing virtual digital humans can generate language responses, their content often lacks distinct character characteristics. They cannot replicate a particular actor's unique way of delivering jokes, signature use of interjections, and the tacit understanding developed through decades of stage practice, resulting in stiff and mechanical interaction that fails to evoke emotional resonance and artistic identification from the audience. The lack of multimodal collaboration means that existing systems can only achieve single voice or text interaction. However, crosstalk performance is a comprehensive art that integrates language, facial expressions, and body movements. An actor's eye glance, gesture, or micro-expression can contain rich comedic information. Existing technology cannot achieve real-time multimodal collaborative generation of voice, facial expressions, and body movements, making the performance of virtual digital humans appear stiff and thin, unable to restore the three-dimensionality and appeal that crosstalk performance should have. Summary of the Invention

[0004] Therefore, this invention provides a method and system for generating virtual digital humans for performances, aiming to solve the technical problem that existing technologies cannot adapt to the large amount of improvisation and on-the-spot jokes in crosstalk performances, which greatly limits the application scope and artistic expression of the technology.

[0005] To achieve the above objectives, the present invention adopts the following technical solution:

[0006] According to a first aspect of the present invention, the present invention provides a method for generating virtual digital humans for performances, the method comprising: The original audio and video streams of live actors' performances are captured in real time by an audio and video acquisition device, and the original audio and video streams are preprocessed to generate multimodal data packets that are synchronized with the performance status of the live actors in real time. The multimodal data packets include voice data, facial feature parameter sequences, and whole-body skeletal motion data. The semantic recognition model is used to identify the text semantic information corresponding to the speech data, and the sentiment analysis model is used to label the sentiment tags corresponding to the text semantic information, forming an input data package containing the text semantic information, the sentiment tags, the facial feature parameter sequence, and the whole body skeletal motion data; The input data packet is fed into the character replication model for the target actor to analyze the performance style characteristics in the current performance context; the character replication model is a neural network model pre-trained using the target actor's historical performance data. The performance style features and the target actor's historical dialogue data are input into a real-time response model based on a reinforcement learning framework to generate subsequent target dialogue text that matches the target actor's performance style. The subsequent target dialogue text is input into the response generation model to obtain a virtual digital human driving instruction package; the virtual digital human driving instruction package is used to drive the virtual digital human to perform the subsequent target dialogue text based on the performance style of the target actor at the optimal response time point. The virtual digital human's audio stream, facial expressions, and body movement animations are generated according to the virtual digital human driving instruction package and then rendered on the display terminal.

[0007] Furthermore, the step of acquiring the raw audio and video streams of the live actors' performance in real time through the audio and video acquisition device, and performing data preprocessing on the raw audio and video streams to generate multimodal data packets that are synchronized in real time with the performance status of the live actors includes: The system uses microphone arrays and depth cameras deployed in AI performance equipment to capture raw audio and video streams from live performances or videos of real actors. Extract the clean audio stream from the original audio and video streams, convert the clean audio stream into text transcription information, and use it as the speech data; and, The facial expression image sequence of the live actor is extracted from the original audio and video stream, and key point detection and optical flow analysis are performed to obtain expression feature vector parameters, which are used as the facial feature parameter sequence; and, The skeletal movements of the live actor in the original audio and video stream are captured and the movements are analyzed to obtain a sequence of limb movements containing gesture trajectories, torso postures and displacement information, which serves as the whole-body skeletal movement data. The speech data, facial feature parameter sequence, and whole-body skeletal motion data are synchronized and packaged at the frame level using a time axis alignment algorithm to obtain the multimodal data packet.

[0008] Furthermore, before inputting the data packet to be input into the character replication model for the target actor, the method further includes: Building a role replication model based on the Transformer architecture, specifically including: The historical performance data of the target actor is obtained, and the character replication model is trained using the historical performance data to obtain a character replication model that deeply encodes the target actor's language habits, comedic logic, and / or physical style in a high-dimensional feature space.

[0009] Furthermore, the step of inputting the data packet to be input into the character replication model for the target actor and analyzing the performance style characteristics in the current performance context includes: The character replication model uses a self-attention mechanism to perform cross-modal comparison and association between the composite query tensor corresponding to the input data packet and the historical performance data, and deconstructs in real time to obtain the matching relationship between the current performance situation and historical experience; Based on the matching relationship, the performance intention, expected type of supporting and teasing responses, and stylized language features that should be present in the current performance context are dynamically generated as the performance style features.

[0010] Furthermore, before inputting the performance style features and the target actor's historical dialogue data into the real-time response model based on a reinforcement learning framework, the method further includes: Construct the real-time response model that includes a laugh prediction sub-model, a coherence calculation sub-model, a consistency calculation sub-model, and a policy network; The joke prediction sub-model is trained based on semantic ambush theory and historical joke data, and is used to calculate the punchline resounding index of the candidate dialogue text. The coherence calculation sub-model determines the interaction coherence index of the candidate dialogue text by calculating the semantic vector similarity between the dialogue text and the historical dialogue. The consistency calculation sub-model determines the style consistency index of the candidate dialogue text by calculating the cosine similarity between the performance style features and the style embedding of the dialogue text. The strategy model is used to generate multiple candidate dialogue texts based on the performance style features and the historical dialogue features, and to calculate the comprehensive reward score corresponding to each candidate dialogue text according to the punchline responsiveness index, the interaction coherence index and the style consistency index, and to determine the subsequent target dialogue text based on the comprehensive score index.

[0011] Further, the step of inputting the performance style features and the target actor's historical dialogue data into a real-time response model based on a reinforcement learning framework to generate subsequent target dialogue text that conforms to the target actor's performance style includes: The real-time response model generates multiple candidate dialogue texts based on the performance style features and the historical dialogue features, and calculates the punchline responsiveness index, interaction coherence index, and style consistency index for each candidate dialogue text. The comprehensive reward score for each candidate dialogue text is calculated based on the punchline responsiveness index, the interaction coherence index, and the style consistency index, as expressed in the following formula:

[0012] in, This indicates the overall reward score; This indicates the brightness index of the package; Indicators representing interactive coherence; Indicates style consistency index; , , These represent the dynamic adjustment weight coefficients for the punchline impact index, interaction coherence index, and style consistency index, respectively. With the goal of maximizing the expected cumulative reward, a beam search optimization is performed based on contextual information, and the candidate dialogue text with the highest comprehensive reward score is selected as the subsequent target dialogue text.

[0013] Furthermore, before inputting the subsequent target dialogue text into the response generation model, the method further includes: Construct the response generation model that includes a speech synthesis network, a facial expression generation network, and a body motion generation network; The speech synthesis network is used to convert the subsequent target dialogue text into an audio stream that is similar to the speech features of the target actor; the speech features include timbre, intonation, and speech rate; The facial expression generation network is used to generate facial expression parameters for the virtual digital human based on the subsequent target dialogue text and the emotion tags; the facial animation parameters include facial muscle movement control parameters and signature micro-expression control parameters. The body motion generation network is used to generate a skeletal animation sequence synchronized with speech and facial expressions based on the semantics corresponding to the subsequent target dialogue text and the performance rhythm.

[0014] Further, the step of inputting the subsequent target dialogue text into the response generation model to obtain the virtual digital human driving instruction package includes: The response generation model is used to generate an audio stream, facial animation parameters, and skeletal animation sequence to drive the performance of the virtual digital human based on the subsequent target dialogue text, as an initial driving instruction package; The process involves analyzing the recent performance status data of the live actor using coordination and scheduling techniques, and calculating and calibrating the optimal response timing for the virtual digital human to enter the scene based on this data. Specifically, this includes: In the anticipated stage before the live actor speaks, a timing control algorithm is used in conjunction with historical interaction rhythms to calculate and calibrate the response timing benchmark corresponding to the most recent performance state data; the most recent performance state data includes micro-expression parameters and premonitory signals of breathing movements; During the reaction phase where semantic information is gradually completed, the response timing benchmark is continuously corrected using a sliding window to determine the optimal response timing point for the virtual digital human to enter. The optimal response time is added to the initial driving instruction package to obtain the virtual digital human driving instruction package.

[0015] Further, the step of generating the audio stream, facial expressions, and body movement animations of the virtual digital human according to the virtual digital human driving instruction package, and rendering them on the display terminal, includes: The virtual digital human driving instruction package is sent to the rendering pipeline; The facial expressions and body movements are rendered based on the facial animation parameters and skeletal animation sequence; and the audio stream, facial expressions and body movements are displayed based on the optimal response timing, so that the timing of the virtual digital human's entry and the ending of the live actor's voice are within the preset perception and prediction range.

[0016] According to a second aspect of the present invention, the present invention provides a virtual digital human generation system for performances, the system comprising: The raw data acquisition module is used to acquire raw audio and video streams of live actors' performances in real time through audio and video acquisition devices, and to preprocess the raw audio and video streams to generate multimodal data packets that are synchronized with the performance status of the live actors in real time; the multimodal data packets include voice data, facial feature parameter sequences, and whole-body skeletal motion data; The input data processing module is used to identify the text semantic information corresponding to the speech data using a semantic recognition model, and to label the sentiment tags corresponding to the text semantic information using a sentiment analysis model, thereby forming an input data package containing the text semantic information, the sentiment tags, the facial feature parameter sequence, and the whole body skeletal motion data; The style feature analysis module is used to input the data packet to be input into the character replication model for the target actor and analyze the performance style features in the current performance context; the character replication model is a neural network model pre-trained using the historical performance data of the target actor. The dialogue text generation module is used to input the performance style features and the historical dialogue data of the target actor into a real-time response model based on a reinforcement learning framework to generate subsequent target dialogue text that conforms to the performance style of the target actor. The response instruction generation module is used to input the subsequent target dialogue text into the response generation model to obtain a virtual digital human driving instruction package; the virtual digital human driving instruction package is used to drive the virtual digital human to perform the subsequent target dialogue text based on the performance style of the target actor at the optimal response time point. The digital rendering driver module is used to generate the audio stream, facial expressions, and body movement animations of the virtual digital human according to the virtual digital human driving instruction package, and to render them on the display terminal.

[0017] The present invention, by adopting the above technical solution, has at least the following beneficial effects: This invention provides a method for generating a virtual digital human for performances, comprising: acquiring raw audio and video streams of a live actor's performance in real time using an audio and video acquisition device, preprocessing the raw audio and video streams to generate a multimodal data packet that is synchronized in real time with the live actor's performance state; identifying textual semantic information corresponding to the speech data using a semantic recognition model, and labeling the emotional tags corresponding to the textual semantic information using a sentiment analysis model to form an input data packet containing the textual semantic information, the emotional tags, the facial feature parameter sequence, and the whole-body skeletal motion data; inputting the input data packet into a character replication model for a target actor to analyze the performance style features in the current performance context; inputting the performance style features and the target actor's historical dialogue data into a real-time response model based on a reinforcement learning framework to generate subsequent target dialogue text that conforms to the target actor's performance style; inputting the subsequent target dialogue text into a response generation model to obtain a virtual digital human driving instruction package; generating the virtual digital human's audio stream, facial expressions, and body motion animations according to the virtual digital human driving instruction package, and rendering them on a display terminal. This has enabled highly realistic and creative virtual digital human performances, promoting the digital innovation and industrialization of traditional stage arts.

[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A flowchart illustrating a method for generating virtual digital human performances according to an embodiment of the present invention is shown. Figure 2 A flowchart illustrating a method for generating virtual digital humans for performances according to another embodiment of the present invention is shown; Figure 3 This diagram illustrates the structure of a virtual digital human generation system for performances according to an embodiment of the present invention. Figure 4 A schematic diagram of the structure of a virtual digital human generation system for performances provided in another embodiment of the present invention is shown; Figure 5 A schematic diagram of the physical structure of a computer device provided in an embodiment of the present invention is shown. Detailed Implementation

[0021] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0022] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.

[0023] This invention provides a method for generating virtual digital humans for performances, such as... Figure 1 As shown, it may include at least the following steps S101~S106: Step S101: The original audio and video streams of the live actors' performances are acquired in real time through the audio and video acquisition device, and the original audio and video streams are preprocessed to generate multimodal data packets that are synchronized with the live actors' performance status in real time.

[0024] It should be noted that the embodiments of the present invention can be applied to AI performance equipment. In practical applications, the original audio and video streams from live performances or videos can be captured in real time using microphone arrays and depth cameras deployed on the AI ​​performance equipment. Then, the clean audio stream is extracted from the original audio and video stream and converted into text transcription information as speech data. Specifically, echo cancellation, noise suppression, and speech activity detection can be performed on the original audio stream to obtain a clean audio stream. Simultaneously, facial expression image sequences of live actors can be extracted from the original audio and video stream and subjected to keypoint detection and optical flow analysis to obtain expression feature vectors (covering micro-motion parameters of the facial features) reflecting micro-expression changes of joy, anger, sorrow, and happiness, serving as expression feature vector parameters, i.e., facial feature parameter sequences. Furthermore, the skeletal movements of live actors in the original audio and video stream are captured and analyzed, then filtered and normalized to obtain limb movement sequences containing gesture trajectories, trunk postures, and displacement information, serving as full-body skeletal movement data. A time-axis alignment algorithm is used to perform frame-level synchronization and packaging of the speech data, facial feature parameter sequences, and full-body skeletal movement data to obtain multimodal data packets.

[0025] Understandably, to ensure temporal consistency among speech data, facial feature parameter sequences, and full-body skeletal motion data, the limb motion sequences generated after skeletal motion capture parsing, which include gesture trajectories, torso posture, and displacement information, need to be uniformly encapsulated and accompanied by precise timestamps. For these three modalities—the preprocessed clean audio stream and its corresponding text transcription information, the facial feature parameter sequences extracted through facial keypoint detection, and the aforementioned processed limb motion sequences—a layered compression coding approach is employed. First, each modal of data undergoes adaptive quantization compression to remove redundant information. Then, a timeline alignment algorithm is used to synchronize these three modalities at the frame level. Finally, a real-time transmission protocol is used for packaging and encapsulation to ensure low latency during multimodal data packet transmission over the network.

[0026] Step S102: Use a semantic recognition model to identify the text semantic information corresponding to the speech data, and use a sentiment analysis model to label the sentiment tags corresponding to the text semantic information, forming an input data package containing text semantic information, sentiment tags, facial feature parameter sequences, and whole-body skeletal motion data.

[0027] The semantic recognition model in this embodiment of the invention can adopt the BERT pre-trained language model framework. Specifically, it can be pre-trained based on a multi-layer Transformer encoder and through tasks such as masked language models, directly adapting to semantic tasks such as intent recognition and slot extraction to identify the text semantic information corresponding to semantic data. The sentiment analysis model can adopt a pre-trained encoder based on BERT / RoBERTa / ERNIE, trained using existing text content and existing sentiment tags to achieve polarity classification of text semantic information.

[0028] Then, the text semantic information, sentiment tags, facial feature parameter sequences, and whole-body skeletal motion data are integrated into an input data package for subsequent model input.

[0029] Step S103: Input the data packet to be input into the character replication model for the target actor and analyze the performance style characteristics in the current performance context.

[0030] like Figure 2 As shown, before step S103, the method for generating a virtual digital human for a performance provided in this embodiment of the invention may further include the following step S107: Step S107: Construct a role replication model based on the Transformer architecture.

[0031] Specifically, by acquiring the historical performance data of the target actor, the character replication model is trained using the historical performance data to obtain a character replication model that deeply encodes the target actor's language habits, comedic logic, and / or physical style in a high-dimensional feature space.

[0032] In this embodiment of the invention, the target actor can be a well-known classic crosstalk performer or a deceased famous crosstalk performer. The character replication model is trained using a massive amount of historical performance data of the target actor, deeply encoding the target actor's language habits, comedic logic, classic jokes, and unique body language style, forming a character identity representation in a high-dimensional feature space.

[0033] Furthermore, the semantic information, sentiment tags, facial feature parameter sequences, and full-body skeletal motion data in the input text data package are integrated into a composite query tensor, which is then input into the character replication model. The character replication model uses a self-attention mechanism to perform cross-modal comparison and association between the currently input composite query tensor and the target actor's historical performance data, deconstructing the matching relationship between the current performance context and historical experience in real time. Based on the matching relationship, it dynamically generates the target actor's expected performance intention, anticipated comedic or supporting role response type, and stylized language features in the current performance context, serving as performance style features. Thus, the mapping from multimodal input to characterized performance intention is completed.

[0034] Step S104: Input the performance style features and the target actor's historical dialogue data into the real-time response model based on the reinforcement learning framework to generate subsequent target dialogue text that matches the target actor's performance style.

[0035] like Figure 2 As shown, prior to step S104, the method for generating a virtual digital human for a performance provided in this embodiment of the invention may further include the following step S108: Step S108: Construct a real-time response model that includes a laugh prediction sub-model, a coherence calculation sub-model, a consistency calculation sub-model, and a policy network.

[0036] The system comprises three sub-models: a joke prediction sub-model, trained on semantic ambush theory and historical joke data, used to calculate the punchline responsiveness index of candidate dialogue texts; a coherence calculation sub-model, which determines the interaction coherence index of candidate dialogue texts by calculating the semantic vector similarity between the dialogue text and historical dialogues; and a consistency calculation sub-model, which determines the style consistency index of candidate dialogue texts by calculating the cosine similarity between performance style features and dialogue style embeddings. The strategy model generates multiple candidate dialogue texts based on performance style features and historical dialogue features, calculates a comprehensive reward score for each candidate dialogue text based on the punchline responsiveness index, interaction coherence index, and style consistency index, and determines the subsequent target dialogue text based on the comprehensive score index. Specifically, the scores of the above three dimensions can be dynamically weighted to form a comprehensive reward score that guides model optimization.

[0037] The generation strategy for subsequent dialogue text (the next line of dialogue that matches the target actor's performance characteristics) involves using a policy network to analyze the current performance context, emotional state, and character intent, aiming to maximize the expected cumulative reward. During the decoding process, a bundle search optimization is performed to select the candidate dialogue text with the highest overall reward score as the subsequent target dialogue text. Specifically, the real-time response model described above generates multiple candidate dialogue texts based on performance style features and historical dialogue features, and calculates the punchline responsiveness, interaction coherence, and style consistency indices for each candidate dialogue text. The overall reward score for each candidate dialogue text is calculated based on these indices, as shown in the following formula:

[0038] in, This indicates the overall reward score; The "Laugh-out-loudness" index is used to evaluate the humor potential of each candidate line text based on the joke prediction sub-model. The higher the value, the stronger the expected laugh effect. The interaction coherence metric is derived by calculating the semantic vector similarity between the generated text and the historical dialogue, measuring the naturalness of the contextual flow. The style consistency index is determined by the cosine similarity between the style features output by the character replication model and the style embedding of the generated text, reflecting the degree of matching with the target actor's comedic style. , , These represent the dynamic adjustment weight coefficients for the punchline responsiveness index, the interaction coherence index, and the style consistency index, respectively, which can be adjusted in real time according to the performance scenario.

[0039] In this embodiment of the invention, the real-time response model is trained with the goal of maximizing the expected value of future cumulative rewards. Parameters are optimized using a policy gradient method to ensure that the generated subsequent dialogue text consistently achieves high overall scores during long-term interaction. Specifically, in the decoding and generation stage, the real-time response model performs a bundle search based on the current context, maintains multiple candidate text sequences in parallel, and scores and prunes the candidate dialogue text at each step according to the overall reward score. Finally, it outputs the subsequent dialogue text with the highest overall reward score. This subsequent dialogue text not only contains semantic content but also includes emotional intensity values ​​and expected comedic effect labels predicted synchronously by the real-time response model.

[0040] Step S105: Input the subsequent target dialogue text into the response generation model to obtain the virtual digital human driving instruction package.

[0041] like Figure 2 As shown, before step S105, the method for generating a virtual digital human for a performance provided in this embodiment of the invention may further include the following step S109: Step S108: Construct a response generation model that includes a speech synthesis network, a facial expression generation network, and a body motion generation network.

[0042] The speech synthesis network is used to convert the subsequent target dialogue text into an audio stream with speech features similar to those of the target actor; speech features include timbre, intonation, and speech rate; the facial expression generation network is used to generate facial expression parameters for the virtual digital human based on the subsequent target dialogue text and emotional tags; facial animation parameters include facial muscle movement control parameters and signature micro-expression control parameters; the body movement generation network is used to generate a skeletal animation sequence that is synchronized with the speech and facial expressions and conforms to the unique stage movements of the target actor, based on the semantics corresponding to the subsequent target dialogue text and the performance rhythm.

[0043] Then, the response generation model can be used to generate audio streams, facial animation parameters, and skeletal animation sequences to drive the performance of the virtual digital human based on the subsequent target dialogue text, as the initial driving instruction package.

[0044] To ensure that the timing of the virtual digital human's entrance (either as the straight man or the comedian) is synchronized with the ending of the live actor's voice, this embodiment of the invention can also utilize coordination and scheduling technology to analyze the live actor's recent performance state data. Based on the recent performance state data, the optimal response timing point for the virtual digital human's entrance is calculated and calibrated. Specifically: in the anticipation stage before the live actor begins to speak, a timing control algorithm combined with historical interaction rhythms is used to calculate and calibrate the response timing benchmark corresponding to the recent performance state data; in the reaction stage where semantic information gradually becomes complete, a sliding window is used to continuously correct the response timing benchmark to determine the optimal response timing point for the virtual digital human's entrance; the optimal response timing point is added to the initial driving instruction package to obtain the virtual digital human driving instruction package.

[0045] In practical applications, collaborative scheduling technology is used to analyze the latest performance status data of the real actors in real time (i.e., the performance status data corresponding to the latest sentence or several lines of dialogue before the real-time interaction of the virtual digital human). At the same time, the driving instruction packets of the virtual digital human are received. In the expectation stage, the timing control algorithm analyzes the micro-expressions and breathing movements of the real actors before they speak, and combines them with the historical interaction rhythm to calculate the preliminary response timing benchmark in advance. In the reaction stage, as the semantic information gradually becomes complete, the sliding window is used to continuously correct the response timing benchmark and dynamically compress the entry time of the virtual digital human.

[0046] The timing control algorithm proposed in this invention uses a two-stage anticipatory-response correction combined with closed-loop feedback to dynamically calculate the optimal response time of the virtual digital human. Its core calculation formula is as follows:

[0047] in, This indicates the optimal response point after the virtual digital human has been finally calibrated, i.e., the target moment when the driving command is issued to the rendering pipeline; This represents the initial response timing baseline calculated during the anticipation phase. As a dynamic correction measure in the reaction phase, the completeness of the semantic information of the live actor is analyzed in real time using a sliding window. As the semantics become clearer, the adjustment is continuously made so that the response timing approaches the ideal ending point. It is represented as the closed-loop feedback correction amount, calculated based on the actual response delay error of the previous round of interaction (i.e., the deviation between the actual response time and the ideal ending time), and is used to compensate for the inherent delay of the system and the fluctuation of the rendering pipeline to achieve accurate synchronization of long-term interaction.

[0048] Through the calculation and real-time adjustment of the above formula, the end-to-end delay between the virtual digital human's entry moment and the live actor's ending moment is strictly controlled within 50ms, so as to eliminate the lag in human-computer interaction and restore the unique tacit understanding and smoothness of crosstalk performance.

[0049] Step S106: Generate the audio stream, facial expressions, and body movement animations of the virtual digital human according to the virtual digital human driving instruction package, and render them on the display terminal.

[0050] Specifically, the virtual digital human driving instruction package is sent to the rendering pipeline; facial expressions and body movement animations are rendered based on facial animation parameters and skeletal animation sequences; and the audio stream, facial expressions, and body movement animations are displayed based on the optimal response timing, so that the timing of the virtual digital human's entry and the ending of the live actor's voice are within the preset perception and prediction range.

[0051] Understandably, after the above-mentioned response timing benchmark, the final calibrated response command can be sent to the rendering pipeline through a closed-loop feedback mechanism to ensure that the timing of the virtual digital human's interjection or comedic timing is synchronized with the ending of the live actor's voice. For example, the end-to-end latency is strictly locked within 50ms, which is within the threshold of human hearing, so that every time the virtual digital human responds or every change in expression is in sync with the performance rhythm of the live actor, eliminating the lag in human-computer interaction and restoring the unique tacit understanding and fluency of crosstalk.

[0052] This invention provides a method for generating virtual digital humans for performances. Compared with the prior art, this invention has at least the following beneficial effects: 1) By constructing a character replication model based on the Transformer architecture, deep learning was used to learn the target actor's language habits, comedic logic, and physical style. This enabled precise mapping from multimodal input to character performance intentions, allowing virtual digital humans to reproduce the essence of a specific artist's performance in a highly realistic way. Through the digital replication and real-time reproduction of the target actor's performance style, the essence of classic artists' performances can break through the limitations of time and space, continuously resonating with contemporary audiences emotionally, and providing a replicable technical path for the digital protection and innovative dissemination of traditional art. 2) By applying multimodal collaborative scheduling technology and timing control algorithm, the end-to-end delay is strictly locked within 50ms through the expectation-response two-stage correction, ensuring that the timing of the virtual digital human's comedic interjection is synchronized with the performance rhythm of the real actor. 3) Through the collaborative work of voice, facial expression, and motion generation networks, virtual digital humans possess three-dimensional performance capabilities; dynamic adaptation to multi-person performance modes is supported, comprehensively solving the shortcomings of existing technologies in multiple key dimensions of human-computer collaboration. 4) By introducing a comprehensive reward function based on the loudness of the joke, the coherence of the interaction, and the consistency of the role, the real-time response model can continuously optimize the generation strategy under the reinforcement learning framework, so that the virtual digital human not only has the ability to imitate, but also has a certain ability to improvise, and can generate new artistic sparks in the interaction with real actors. 5) The multimodal generative model, driven by anatomical facial muscle parameters and diffusion-based speech synthesis, enables virtual digital humans to achieve a high degree of realism in micro-expressions and vocal details, greatly enhancing the artistic appeal of performances and potentially promoting the digital innovation and industrialization of traditional stage arts.

[0053] Furthermore, as Figure 1 In specific implementation, embodiments of the present invention provide a virtual digital human generation system for performances, such as... Figure 3 As shown, the system may include: a raw data acquisition module 310, an input data processing module 320, a style feature analysis module 330, a dialogue text generation module 340, a response instruction generation module 350, and a digital rendering drive module 360.

[0054] The raw data acquisition module 310 can be used to acquire raw audio and video streams of live actors' performances in real time through audio and video acquisition devices, and to preprocess the raw audio and video streams to generate multimodal data packets that are synchronized with the performance status of live actors in real time; the multimodal data packets include voice data, facial feature parameter sequences, and whole-body skeletal motion data; The input data processing module 320 can be used to identify the text semantic information corresponding to the speech data using a semantic recognition model, and to label the sentiment tags corresponding to the text semantic information using a sentiment analysis model, forming an input data package containing text semantic information, sentiment tags, facial feature parameter sequences and whole-body skeletal motion data. The style feature parsing module 330 can be used to input the data packet to be input into the character replication model for the target actor and analyze the performance style features in the current performance context; the character replication model is a neural network model pre-trained using the target actor's historical performance data. The dialogue text generation module 340 can be used to input performance style features and historical dialogue data of the target actor into a real-time response model based on a reinforcement learning framework to generate subsequent target dialogue text that conforms to the performance style of the target actor. The response instruction generation module 350 can be used to input the subsequent target dialogue text into the response generation model to obtain the virtual digital human driving instruction package; the virtual digital human driving instruction package is used to drive the virtual digital human to perform the subsequent target dialogue text based on the performance style of the target actor at the optimal response time point. The digital rendering driver module 360 ​​can be used to generate audio streams, facial expressions, and body movement animations of virtual digital humans according to the virtual digital human driving instruction package, and then render them on the display terminal.

[0055] Optionally, such as Figure 4As shown, another embodiment of the present invention provides a performance virtual digital human generation system, which may further include: a first model building module 370, a second model building module 380 and a third model building module 390.

[0056] The first model building module 370 can be used to build role replication models based on the Transformer architecture, specifically including: Acquire historical performance data of the target actor, use the historical performance data to train a character replication model, and obtain a character replication model that deeply encodes the target actor's language habits, comedic logic and / or physical style in a high-dimensional feature space.

[0057] The second model building module 380 can be used to build a real-time response model that includes a laugh prediction sub-model, a coherence calculation sub-model, a consistency calculation sub-model, and a policy network. Among them, the joke prediction sub-model is trained based on semantic ambush theory and historical joke data, and is used to calculate the punchline resounding index of the candidate dialogue text; The coherence calculation sub-model determines the interaction coherence index of candidate dialogue texts by calculating the semantic vector similarity between the dialogue text and the historical dialogue. The consistency calculation sub-model determines the style consistency index of the candidate dialogue text by calculating the cosine similarity between the performance style features and the style embedding of the dialogue text. The strategy model is used to generate multiple candidate dialogue texts based on performance style features and historical dialogue features. It calculates the comprehensive reward score corresponding to each candidate dialogue text based on the punchline responsiveness index, interaction coherence index, and style consistency index, and determines the subsequent target dialogue text based on the comprehensive score index.

[0058] The third model building module 390 can be used to build a response generation model that includes a speech synthesis network, a facial expression generation network, and a body motion generation network. Speech synthesis networks are used to convert subsequent target dialogue text into an audio stream that is similar to the speech features of the target actor; speech features include timbre, intonation, and speech rate; The facial expression generation network is used to generate facial expression parameters for virtual digital humans based on subsequent target dialogue text and emotion tags; facial animation parameters include facial muscle movement control parameters and signature micro-expression control parameters. The body motion generation network is used to generate skeletal animation sequences that are synchronized with speech and facial expressions, based on the semantics of the subsequent target dialogue text and the performance rhythm.

[0059] It should be noted that other corresponding descriptions of the functional modules involved in the performance virtual digital human generation system provided in this embodiment of the invention can be found in the following references. Figure 1The corresponding description of the method shown will not be repeated here.

[0060] Based on the above, Figure 1 Accordingly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the performance virtual digital human generation method of any of the above embodiments.

[0061] Based on the above, Figure 1 The method shown and as Figure 4 The embodiment of the system shown in the invention also provides a physical structure diagram of a computer device, such as... Figure 5 As shown, the computer device may include a communication bus, a processor, a memory, and a communication interface. It may also include input / output interfaces and a display device. The various functional units can communicate with each other via the bus. The memory stores a computer program, and the processor executes the program stored in the memory to perform the steps of the virtual digital human generation method described in the above embodiments.

[0062] Those skilled in the art will clearly understand that the specific working process of the systems, devices, modules and units described above can be referred to the corresponding process in the foregoing method embodiments. For the sake of brevity, it will not be repeated here.

[0063] Furthermore, the functional units in the various embodiments of the present invention can be physically independent of each other, or two or more functional units can be integrated together, or all functional units can be integrated into one processing unit. The integrated functional units described above can be implemented in hardware, or in software or firmware.

[0064] Those skilled in the art will understand that if the integrated functional unit is implemented in software and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or all or part of it, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computing device (e.g., a personal computer, server, or network device) to execute all or part of the steps of the methods described in the embodiments of the present invention when running the instructions. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0065] Alternatively, all or part of the steps of the foregoing method embodiments can be implemented by hardware (such as a computing device, personal computer, server, or network device) related to program instructions. The program instructions can be stored in a computer-readable storage medium. When the program instructions are executed by the processor of the computing device, the computing device executes all or part of the steps of the methods described in the various embodiments of the present invention.

[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that within the spirit and principles of the present invention, modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the corresponding technical solutions to depart from the protection scope of the present invention.

Claims

1. A method for generating virtual digital humans for performances, characterized in that, The method includes: The original audio and video streams of live actors' performances are captured in real time by an audio and video acquisition device, and the original audio and video streams are preprocessed to generate multimodal data packets that are synchronized with the performance status of the live actors in real time. The multimodal data packets include voice data, facial feature parameter sequences, and whole-body skeletal motion data. The semantic recognition model is used to identify the text semantic information corresponding to the speech data, and the sentiment analysis model is used to label the sentiment tags corresponding to the text semantic information, forming an input data package containing the text semantic information, the sentiment tags, the facial feature parameter sequence, and the whole body skeletal motion data; The input data packet is fed into the character replication model for the target actor to analyze the performance style characteristics in the current performance context; the character replication model is a neural network model pre-trained using the target actor's historical performance data. The performance style features and the target actor's historical dialogue data are input into a real-time response model based on a reinforcement learning framework to generate subsequent target dialogue text that matches the target actor's performance style. The subsequent target dialogue text is input into the response generation model to obtain a virtual digital human driving instruction package; the virtual digital human driving instruction package is used to drive the virtual digital human to perform the subsequent target dialogue text based on the performance style of the target actor at the optimal response time point. The virtual digital human's audio stream, facial expressions, and body movement animations are generated according to the virtual digital human driving instruction package and then rendered on the display terminal.

2. The method according to claim 1, characterized in that, The process of acquiring raw audio and video streams of live actors' performances in real time using an audio and video acquisition device, and preprocessing the raw audio and video streams to generate multimodal data packets that are synchronized in real time with the performance status of the live actors, includes: The system uses microphone arrays and depth cameras deployed in AI performance equipment to capture raw audio and video streams from live performances or videos of real actors. Extract the clean audio stream from the original audio and video streams, convert the clean audio stream into text transcription information, and use it as the speech data; and, The facial expression image sequence of the live actor is extracted from the original audio and video stream, and key point detection and optical flow analysis are performed to obtain expression feature vector parameters, which are used as the facial feature parameter sequence; and, The skeletal movements of the live actor in the original audio and video stream are captured and the movements are analyzed to obtain a sequence of limb movements containing gesture trajectories, torso postures and displacement information, which serves as the whole-body skeletal movement data. The speech data, facial feature parameter sequence, and whole-body skeletal motion data are synchronized and packaged at the frame level using a time axis alignment algorithm to obtain the multimodal data packet.

3. The method according to claim 1, characterized in that, Before inputting the data packet to be input into the character replica model for the target actor, the method further includes: Building a role replication model based on the Transformer architecture, specifically including: The historical performance data of the target actor is obtained, and the character replication model is trained using the historical performance data to obtain a character replication model that deeply encodes the target actor's language habits, comedic logic, and / or physical style in a high-dimensional feature space.

4. The method according to claim 1, characterized in that, The step of inputting the data packet to the target actor's character replication model and analyzing the performance style characteristics in the current performance context includes: The character replication model uses a self-attention mechanism to perform cross-modal comparison and association between the composite query tensor corresponding to the input data packet and the historical performance data, and deconstructs in real time to obtain the matching relationship between the current performance situation and historical experience; Based on the matching relationship, the performance intention, expected type of supporting and teasing responses, and stylized language features that should be present in the current performance context are dynamically generated as the performance style features.

5. The method according to claim 1, characterized in that, Before inputting the performance style features and the target actor's historical dialogue data into the real-time response model based on a reinforcement learning framework, the method further includes: Construct the real-time response model that includes a laugh prediction sub-model, a coherence calculation sub-model, a consistency calculation sub-model, and a policy network; The joke prediction sub-model is trained based on semantic ambush theory and historical joke data, and is used to calculate the punchline resounding index of the candidate dialogue text. The coherence calculation sub-model determines the interaction coherence index of the candidate dialogue text by calculating the semantic vector similarity between the dialogue text and the historical dialogue. The consistency calculation sub-model determines the style consistency index of the candidate dialogue text by calculating the cosine similarity between the performance style features and the style embedding of the dialogue text. The strategy model is used to generate multiple candidate dialogue texts based on the performance style features and the historical dialogue features, and to calculate the comprehensive reward score corresponding to each candidate dialogue text according to the punchline responsiveness index, the interaction coherence index and the style consistency index, and to determine the subsequent target dialogue text based on the comprehensive score index.

6. The method according to claim 5, characterized in that, The step of inputting the performance style features and the target actor's historical dialogue data into a real-time response model based on a reinforcement learning framework to generate subsequent target dialogue text that conforms to the target actor's performance style includes: The real-time response model generates multiple candidate dialogue texts based on the performance style features and the historical dialogue features, and calculates the punchline responsiveness index, interaction coherence index, and style consistency index for each candidate dialogue text. The comprehensive reward score for each candidate dialogue text is calculated based on the punchline responsiveness index, the interaction coherence index, and the style consistency index, as expressed in the following formula: in, This indicates the overall reward score; Indicates the brightness index of the package; Indicators representing interactive coherence; Indicates style consistency index; , , These represent the dynamic adjustment weight coefficients for the punchline impact index, interaction coherence index, and style consistency index, respectively. With the goal of maximizing the expected cumulative reward, a beam search optimization is performed based on contextual information, and the candidate dialogue text with the highest comprehensive reward score is selected as the subsequent target dialogue text.

7. The method according to claim 1, characterized in that, Before inputting the subsequent target dialogue text into the response generation model, the method further includes: Construct the response generation model that includes a speech synthesis network, a facial expression generation network, and a body motion generation network; The speech synthesis network is used to convert the subsequent target dialogue text into an audio stream that is similar to the speech features of the target actor; the speech features include timbre, intonation, and speech rate; The facial expression generation network is used to generate facial expression parameters for the virtual digital human based on the subsequent target dialogue text and the emotion tags; the facial animation parameters include facial muscle movement control parameters and signature micro-expression control parameters. The body motion generation network is used to generate a skeletal animation sequence synchronized with speech and facial expressions based on the semantics corresponding to the subsequent target dialogue text and the performance rhythm.

8. The method according to claim 7, characterized in that, The step of inputting the subsequent target dialogue text into the response generation model to obtain the virtual digital human driving instruction package includes: The response generation model is used to generate an audio stream, facial animation parameters, and skeletal animation sequence to drive the performance of the virtual digital human based on the subsequent target dialogue text, as an initial driving instruction package; The process involves analyzing the recent performance status data of the live actor using coordination and scheduling techniques, and calculating and calibrating the optimal response timing for the virtual digital human to enter the scene based on this data. Specifically, this includes: In the anticipated stage before the live actor speaks, a timing control algorithm is used in conjunction with historical interaction rhythms to calculate and calibrate the response timing benchmark corresponding to the most recent performance state data; the most recent performance state data includes micro-expression parameters and premonitory signals of breathing movements; During the reaction phase where semantic information is gradually completed, the response timing benchmark is continuously corrected using a sliding window to determine the optimal response timing point for the virtual digital human to enter. The optimal response timing is added to the initial drive instruction package to obtain the virtual digital human drive instruction package.

9. The method according to claim 8, characterized in that, The step of generating the audio stream, facial expressions, and body movement animations of the virtual digital human according to the virtual digital human driving instruction package, and rendering them on the display terminal, includes: The virtual digital human driving instruction package is sent to the rendering pipeline; The facial expressions and body movements are rendered based on the facial animation parameters and skeletal animation sequence; and the audio stream, facial expressions and body movements are displayed based on the optimal response timing, so that the timing of the virtual digital human's entry and the ending of the live actor's voice are within the preset perception and prediction range.

10. A virtual digital human generation system for performances, characterized in that, The system includes: The raw data acquisition module is used to acquire raw audio and video streams of live actors' performances in real time through audio and video acquisition devices, and to preprocess the raw audio and video streams to generate multimodal data packets that are synchronized with the performance status of the live actors in real time; the multimodal data packets include voice data, facial feature parameter sequences, and whole-body skeletal motion data; The input data processing module is used to identify the text semantic information corresponding to the speech data using a semantic recognition model, and to label the sentiment tags corresponding to the text semantic information using a sentiment analysis model, thereby forming an input data package containing the text semantic information, the sentiment tags, the facial feature parameter sequence, and the whole body skeletal motion data; The style feature analysis module is used to input the data packet to be input into the character replication model for the target actor and analyze the performance style features in the current performance context; the character replication model is a neural network model pre-trained using the historical performance data of the target actor. The dialogue text generation module is used to input the performance style features and the historical dialogue data of the target actor into a real-time response model based on a reinforcement learning framework to generate subsequent target dialogue text that conforms to the performance style of the target actor. The response instruction generation module is used to input the subsequent target dialogue text into the response generation model to obtain a virtual digital human driving instruction package; the virtual digital human driving instruction package is used to drive the virtual digital human to perform the subsequent target dialogue text based on the performance style of the target actor at the optimal response time point. The digital rendering driver module is used to generate the audio stream, facial expressions, and body movement animations of the virtual digital human according to the virtual digital human driving instruction package, and to render them on the display terminal.