Digital human video generation method and device based on dynamic mirror operation, medium and product

By generating speech feature data, lip-syncing animation, and limb movement timing alignment, combined with dynamic camera movement commands, the problem of independent optimization of digital human lip-syncing and movement in traditional technologies has been solved, thereby improving the vividness and realism of digital human videos.

CN121509780APending Publication Date: 2026-02-10ZHEJIANG BAORONG MEDIA TECH (ZHEJIANG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511848924.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Traditional video generation technologies optimize digital lip movements and actions independently, resulting in low video generation efficiency and stiff digital human performances, failing to achieve dynamic coordination between digital lip movements, actions, and camera movements.

Method used

By acquiring the text to be broadcast, the broadcast style, and parameters, speech feature data is generated. This data is then used to align the timing of lip-sync animation and body movements, generate dynamic camera movement instructions, and combine the body movements and speech feature data for animation rendering, thereby achieving dynamic coordination between digital lip movements, actions, and camera movements.

Benefits of technology

It enhances the vividness and realism of digital human videos, solves the problem of separation between digital human mouth movements and camera angles, and achieves natural and smooth performances in digital human videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509780A_ABST
    Figure CN121509780A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a digital human video generation method and device based on dynamic mirror operation, a medium and a product. The method comprises the following steps: acquiring a to-be-broadcasted text, a broadcasting style and broadcasting parameters, and generating voice feature data corresponding to the broadcasting text according to the to-be-broadcasted text, the broadcasting style and the broadcasting parameters; generating a corresponding mouth shape animation according to the voice feature data, and performing time sequence alignment on the voice feature data and the mouth shape animation; according to the to-be-broadcasted text, the voice feature data and the broadcasting style, generating a limb action corresponding to the voice feature data; according to the body movement, the voice feature data, the to-be-broadcasted text and the broadcasting style, generating a dynamic mirror moving instruction; under the dynamic mirror moving instruction, animation rendering is carried out according to the body movement, the voice feature data and the mouth shape animation, and a digital human video is obtained. Through alignment of voice and mouth shape and dynamic mirror movement of a digital human video, dynamic collaboration of the digital human shape, the action and the mirror movement lens can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology and can be applied to the field of computer video generation technology. In particular, it relates to a method, device, medium, and product for generating digital human videos based on dynamic camera movement. Background Technology

[0002] With the development of artificial intelligence technology, more and more scenarios are using artificial intelligence technology to generate videos in order to achieve information interaction.

[0003] In traditional video generation techniques, digitized lip movements and actions are typically optimized independently. Lip movements rely on speech feature mapping, while action generation depends on a predefined action database. Furthermore, traditional video generation techniques often employ fixed and singular camera movements, resulting in low video generation efficiency and stiff digitized human performances, failing to achieve dynamic coordination between digitized lip movements, actions, and camera movements.

[0004] Therefore, there is an urgent need to provide a dynamic camera movement method in digital human video generation to achieve dynamic coordination of digital human lip movements, actions, and camera movements, thereby enhancing the vividness of digital human videos. Summary of the Invention

[0005] This invention provides a method, device, medium, and product for generating digital human videos based on dynamic camera movement, so as to achieve dynamic coordination of digital human lip movements, actions, and camera movements, thereby enhancing the vividness of digital human videos.

[0006] According to one aspect of the present invention, a method for generating digital human videos based on dynamic camera movement is provided, the method comprising:

[0007] The system obtains the text to be played, the playing style, and the playing parameters, and generates speech feature data corresponding to the text to be played based on the text to be played, the playing style, and the playing parameters.

[0008] Generate a corresponding lip-sync animation based on the speech feature data, and align the speech feature data and the lip-sync animation in time sequence.

[0009] Based on the text to be broadcast, the voice feature data, and the broadcast style, generate body movements corresponding to the voice feature data;

[0010] Based on the body movements, the voice feature data, the text to be played, and the playback style, dynamic camera movement instructions are generated.

[0011] Under the dynamic camera movement command, animation rendering is performed based on the body movements, the voice feature data, and the lip-sync animation to obtain a digital human video.

[0012] According to another aspect of the present invention, a digital human video generation apparatus based on dynamic camera movement is provided, the apparatus comprising:

[0013] The speech feature data generation module is used to obtain the text to be played, the playback style, and the playback parameters, and to generate speech feature data corresponding to the text to be played based on the text to be played, the playback style, and the playback parameters.

[0014] The timing alignment module is used to generate corresponding lip-sync animations based on the speech feature data, and to perform timing alignment between the speech feature data and the lip-sync animations.

[0015] The body movement generation module is used to generate body movements corresponding to the voice feature data based on the text to be broadcast, the voice feature data, and the broadcast style.

[0016] The dynamic camera movement instruction generation module is used to generate dynamic camera movement instructions based on the body movements, the voice feature data, the text to be broadcast, and the broadcast style.

[0017] The digital human video generation module is used to perform animation rendering based on the body movements, the voice feature data, and the lip-sync animation under the dynamic camera movement command to obtain a digital human video.

[0018] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0019] At least one processor; and

[0020] A memory communicatively connected to the at least one processor; wherein,

[0021] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the digital human video generation method based on dynamic camera movement according to any embodiment of the present invention.

[0022] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the digital human video generation method based on dynamic camera movement as described in any embodiment of the present invention.

[0023] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the digital human video generation method based on dynamic camera movement as described in any embodiment of the present invention.

[0024] The technical solution of this invention obtains the text to be played, the playing style, and the playing parameters, and generates speech feature data corresponding to the text to be played based on the text to be played, the playing style, and the playing parameters; generates corresponding lip-sync animation based on the speech feature data, and aligns the speech feature data and the lip-sync animation in sequence; generates body movements corresponding to the speech feature data based on the text to be played, the speech feature data, and the playing style; generates dynamic camera movement instructions based on the body movements, speech feature data, the text to be played, and the playing style; and performs animation rendering based on the body movements, speech feature data, and lip-sync animation under the dynamic camera movement instructions to obtain a digital human video. This solves the problem of separation between lip movements, actions, and camera movements in digital human videos. By aligning speech and lip movements and performing dynamic camera movements on digital human videos, dynamic coordination between lip movements, actions, and camera movements can be achieved, improving the vividness of digital human videos.

[0025] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a flowchart of a digital human video generation method based on dynamic camera movement according to Embodiment 1 of the present invention;

[0028] Figure 2 This is a flowchart of a speech feature data generation process according to Embodiment 1 of the present invention;

[0029] Figure 3 This is a flowchart of a digital human video generation method based on dynamic camera movement according to Embodiment 2 of the present invention;

[0030] Figure 4 This is a flowchart illustrating a digital human video generation method based on dynamic camera movement according to Embodiment 2 of the present invention;

[0031] Figure 5 This is a flowchart of a digital human video generation method based on dynamic camera movement according to Embodiment 3 of the present invention;

[0032] Figure 6This is a deployment flowchart of a digital human video generation system based on dynamic camera movement according to Embodiment 3 of the present invention;

[0033] Figure 7 This is a schematic diagram of the structure of a digital human video generation device based on dynamic camera movement according to Embodiment 4 of the present invention;

[0034] Figure 8 This is a schematic diagram of the structure of an electronic device that implements the digital human video generation method based on dynamic camera movement according to an embodiment of the present invention. Detailed Implementation

[0035] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0036] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0037] Example 1

[0038] Figure 1 This is a flowchart of a digital human video generation method based on dynamic camera movement according to Embodiment 1 of the present invention. This embodiment is applicable to scenarios such as virtual anchors, news broadcasting, or online education for automated video content generation and playback. The method can be executed by a digital human video generation device based on dynamic camera movement, which can be implemented in hardware and / or software. This device can be configured in an electronic device, such as a mobile device like a mobile phone, tablet computer (PAD), or wearable device, or a personal computer (PC). Figure 1 As shown, the method includes:

[0039] Step 110: Obtain the text to be played, the playback style, and the playback parameters, and generate speech feature data corresponding to the text to be played based on the text to be played, the playback style, and the playback parameters.

[0040] The text to be played can be the text content that the user plans to play in different scenarios. The playback style can be different playback types, such as news, entertainment, or education. Playback parameters can include playback speed, tone, or timbre.

[0041] For example, a user interface can be provided, which can acquire user input or selection of the text to be played, the playback style, and playback parameters. Text-to-speech technology can then be used to generate speech feature data corresponding to the text to be played, based on the playback style and parameters.

[0042] When generating speech feature data, it can support multiple speech styles and emotional expressions, providing digital humans with rich speech expressiveness. Figure 2 This is a flowchart of a speech feature data generation process provided according to Embodiment 1 of the present invention. Figure 2 As shown, the process of generating speech feature data can be as follows: collecting speech samples of one or more styles and preprocessing the speech samples. The preprocessing can include at least one of the following: noise reduction (such as noise reduction by spectral subtraction or Wiener filtering), endpoint detection (such as the dual threshold method based on energy and zero-crossing rate), and frame windowing (such as Hamming window with a frame length of 25 milliseconds and a frame shift of 10 milliseconds).

[0043] like Figure 2 As shown, after preprocessing the speech samples, various features can be extracted. For example, speech sample feature extraction can include at least one of the following: time-domain feature extraction, frequency-domain feature extraction, pre-modulation feature extraction, and prosodic feature extraction. Time-domain feature extraction can include short-time energy calculation to reflect speech intensity and zero-crossing rate analysis to distinguish between voiced and unvoiced sounds. Frequency-domain feature extraction can include Fast Fourier Transform to obtain the spectrum, Mel filter banks to simulate human ear characteristics, and logarithmic discrete cosine transform of Mel frequency cepstral coefficients. Pitch feature extraction can include fundamental frequency detection (e.g., autocorrelation method or YIN algorithm), fundamental frequency trajectory smoothing (e.g., using median filtering), and pitch pattern analysis (e.g., rising, falling, or stationary). Prosodic feature extraction can include speech rate calculation (e.g., syllable count or duration determination), pause detection (silence segment analysis), and stress localization (e.g., by jointly determining energy and fundamental frequency).

[0044] like Figure 2As shown, the results of feature extraction from multiple aspects can be fused to form a sequence of feature vectors, which can then be used as input to a deep learning model for training, resulting in a speech feature data generation model. The text to be played, the playback style, and the playback parameters are input into the final speech feature data generation model to obtain speech feature data corresponding to the played text.

[0045] For example, a deep learning model can be an end-to-end speech synthesis model to improve the naturalness and realism of speech synthesis.

[0046] Step 120: Generate corresponding lip-sync animations based on speech feature data, and align the speech feature data and lip-sync animations in time sequence.

[0047] Speech feature data can contain timestamp information. Deep learning models can be used to generate lip-sync animations that are time-aligned with the speech feature data. In the time-alignment of speech feature data and lip-sync animation, a dynamic time warping algorithm and a deep learning model incorporating a temporal attention mechanism can be used to fuse and model the semantic features of the speech feature data with those of the text to be played. This forces the model to learn the strong correlation between speech with the same timestamp and lip movements, accurately generating lip-sync action frames corresponding to the speech feature data.

[0048] By aligning the timing of speech and lip movements, precise synchronization between lip-syncing animation and speech can be ensured in digital human videos, eliminating the problem of misalignment between lip movements and speech.

[0049] To further eliminate the misalignment between lip movements and speech, timing deviation detection and compensation can be performed after the speech and lip movements are time-aligned. Optionally, the speech feature data and lip-sync animation are time-aligned, including: binding the speech feature data with the lip movement data in the lip-sync animation, and performing initial timing alignment using an alignment algorithm; during animation rendering, detecting the timing deviation between the initially time-aligned speech feature data and the lip-sync animation; performing dynamic timing compensation on the speech feature data and lip-sync animation based on the timing deviation, and adjusting the alignment algorithm based on the dynamic timing compensation result.

[0050] The initial temporal alignment refers to the aforementioned alignment of speech and lip movements. After initial temporal alignment, further dynamic temporal compensation can be performed. During the rendering of digital human videos, the timestamps of speech feature data and lip-sync animation can be compared in real time. When there is a discrepancy in the timestamps, dynamic temporal compensation can be performed, for example, by adjusting the frame rendering rhythm to correct the timing. After temporal compensation, the compensation value can be fed back to the alignment algorithm, that is, fed back to a deep learning model that includes a temporal attention mechanism, to improve the accuracy of temporal alignment in subsequent applications and achieve precise synchronization between speech pronunciation and lip movements.

[0051] Step 130: Generate body movements corresponding to the voice feature data based on the text to be broadcast, voice feature data, and broadcast style.

[0052] Among these, body movements can be achieved through generative artificial intelligence technology. That is, generative artificial intelligence technology can generate corresponding body movements based on the text to be broadcast, voice feature data, and broadcast style.

[0053] To enhance the naturalness and expressiveness of body movements, body movements generated using generative artificial intelligence can be used as the initial result, and the movement sequence can be optimized based on this initial result. Optionally, body movements corresponding to the speech feature data can be generated based on the text to be broadcast, speech feature data, and broadcast style. This includes: generating initial body movements using a body movement generation tool based on the text to be broadcast, speech feature data, and broadcast style; acquiring audience feedback data on body movements during digital human video playback, and forming initial body movement reward / penalty factors based on the audience feedback data; and adjusting the initial body movements based on the reward / penalty factors to obtain the optimized body movement result.

[0054] Viewer feedback data can include comments or bullet screen messages from viewers during the digital human recognition playback. For example, specific feedback data could include issues like misalignment between movement and speech, pauses in movement transitions, or alignment between movement and speech. Based on this feedback data, the correctness and error of body movements can be identified. For instance, if the feedback data indicates misalignment between movement and speech, an error in the body movement can be determined. Based on the correctness and error of the identified body movements, rewards or penalties can be determined. The specific values ​​of the reward / penalty factors can be determined based on the degree of consistency when the body movements are correct and the degree of abnormality when they are incorrect. For example, if a large number of viewers consistently report misalignment between movement and speech, a higher penalty score can be assigned to the body movement as a reward / penalty factor.

[0055] For example, when the audience feedback data shows that the action and voice are aligned, the reward / penalty factor for the physical action is determined to be +10 points; when the audience feedback data shows that the action and voice are misaligned, the reward / penalty factor for the physical action is determined to be -5 points; and when the audience feedback data shows that the action is stuttered, the reward / penalty factor for the physical action is determined to be -3 points.

[0056] The correctness and incorrectness of corresponding body movements can be determined based on reward and punishment factors. When a body movement is incorrect, the priority for optimization can be determined according to the corresponding reward and punishment factors. Based on the priority, the incorrect body movements are optimized to improve the naturalness and expressiveness of the movements.

[0057] Step 140: Generate dynamic camera movement instructions based on body movements, voice feature data, text to be broadcast, and broadcast style.

[0058] Dynamic camera movement commands can be instructions used to control the camera movement in a digital human video, such as pushing, pulling, panning, and tilting. When generating dynamic camera movement commands, body movements, speech feature data, the text to be read, and the reading style can be combined. For example, the trigger conditions for camera movement commands can be determined from body movements, and corresponding commands can be matched based on these trigger conditions. Another example is determining the digital human's speech rate and emotions from speech feature data, and matching corresponding camera movement commands based on these speech rates and emotions. Yet another example is identifying the details and key points of the digital human video's explanation from the read text, and matching corresponding camera movement commands based on these details and key points. Yet another example is determining the corresponding shot type from the reading style, and then determining the corresponding camera movement command based on the shot type. Alternatively, one or more of the following can be combined to jointly determine the dynamic camera movement commands: body movements, speech feature data, the text to be read, and the reading style.

[0059] Dynamic camera movement commands enable automatic camera movement and free adjustment of perspective, avoiding the static and unnatural effects of fixed shots, such as missing key details. For example, when a digital human is emotionally agitated, the camera can zoom in from a medium or wide shot to enhance the emotional expression through a close-up of the face. Or, when a digital human suddenly moves, the camera can follow it in real time.

[0060] Step 150: Under the dynamic camera movement command, perform animation rendering based on body movements, voice feature data, and lip-sync animation to obtain a digital human video.

[0061] Animation rendering can be achieved using a 3D graphics rendering engine. Under dynamic camera movement commands, body movements, speech feature data, and lip-sync animation are applied to the digital human model and virtual scene, resulting in digital human video through a 3D graphics rendering engine.

[0062] The technical solution of this embodiment obtains the text to be played, the playing style, and the playing parameters, and generates speech feature data corresponding to the text to be played based on the text to be played, the playing style, and the playing parameters; generates corresponding lip-sync animation based on the speech feature data, and aligns the speech feature data and lip-sync animation in sequence; generates body movements corresponding to the speech feature data based on the text to be played, the speech feature data, and the playing style; generates dynamic camera movement instructions based on the body movements, speech feature data, the text to be played, and the playing style; and performs animation rendering based on the body movements, speech feature data, and lip-sync animation under the dynamic camera movement instructions to obtain a digital human video. This solves the problem of separation between lip movements, actions, and camera movements in digital human videos. By aligning speech and lip movements, optimizing body movements, and performing dynamic camera movements on digital human videos, dynamic coordination between lip movements, actions, and camera movements can be achieved, improving the vividness and realism of digital human videos.

[0063] Example 2

[0064] Figure 3 This is a flowchart of a digital human video generation method based on dynamic camera movement according to Embodiment 2 of the present invention. This embodiment is a further refinement of the above technical solution. The technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments.

[0065] Optionally, dynamic camera movement instructions are generated based on body movements, speech feature data, text to be broadcast, and broadcast style. This includes: collecting camera movement videos under different broadcast styles and identifying camera movement-related data in the video; training a model based on the video and the corresponding camera movement-related data to obtain a dynamic camera movement decision model; and inputting body movements, speech feature data, text to be broadcast, and broadcast style into the dynamic camera movement decision model to obtain the corresponding dynamic camera movement instructions.

[0066] like Figure 3 As shown, the method includes:

[0067] Step 310: Obtain the text to be played, the playback style, and the playback parameters, and generate speech feature data corresponding to the text to be played based on the text to be played, the playback style, and the playback parameters.

[0068] Step 320: Generate corresponding lip-sync animations based on speech feature data, and align the speech feature data and lip-sync animations in time sequence.

[0069] Step 330: Generate body movements corresponding to the voice feature data based on the text to be broadcast, voice feature data, and broadcast style.

[0070] Step 340: Collect camera movement videos under different broadcast styles and identify camera movement-related data in the video movement videos.

[0071] The broadcast style can mirror that of a virtual entertainment anchor, news broadcast, or online education. The camera movement video can be a video referenced by a high-quality lesson. Camera movement-related data includes at least one of the following: shot type, camera movement triggering conditions, camera movement parameters, digital human data associated with the camera movement, and scene layout information. Shot type can be camera switching types such as push, pull, pan, and tilt. Camera movement triggering conditions can be the emotion and speech rate of the digital human's voice characteristics, scene changes in the camera movement video, or changes in the digital human's movement state. Camera movement parameters can be data such as camera movement speed, switching angle, and focal length. The digital human data associated with the camera movement can be the digital human's lip-sync sequence, body movement trajectory, and speech semantic tags (such as excitement, detailed explanation). Scene layout information can be information such as the virtual stage area, whiteboard location, and digital human location.

[0072] Step 350: Train the model based on the camera movement video and the corresponding camera movement related data to obtain the dynamic camera movement decision model.

[0073] By training a model based on video footage and corresponding camera movement data, a dynamic camera movement decision model can be obtained. This model can learn the relationship between camera movement methods and related data within the video footage. Therefore, when applying the dynamic camera movement decision model, data in digital human videos can be analyzed, and corresponding dynamic camera movement instructions can be derived based on the related camera movement data within the digital human video.

[0074] Step 360: Input the body movements, voice feature data, text to be broadcast, and broadcast style into the dynamic camera movement decision model to obtain the corresponding dynamic camera movement instructions.

[0075] When applying the dynamic camera movement decision model, body movements, speech feature data, the text to be read, and the reading style can be input into the dynamic camera movement decision model to obtain dynamic camera movement instructions. These instructions, along with body movements, speech feature data, and lip-sync animation, are then combined to generate a digital human video.

[0076] Alternatively, when applying a dynamic camera movement decision model, an initial digital human video can be obtained by first rendering animations based on body movements, speech feature data, and lip-syncing. This initial digital human video is then input into the dynamic camera movement decision model to obtain dynamic camera movement instructions. Finally, the initial digital human video is optimized using these instructions to obtain the final digital human video.

[0077] Step 370: Under the dynamic camera movement command, perform animation rendering based on body movements, voice feature data, and lip-sync animation to obtain a digital human video.

[0078] By combining dynamic camera movement commands to generate digital human videos, the problem of digital human videos lacking dynamic depth due to the broadcast perspective relying entirely on the original fixed camera position can be solved, thus improving the visual effect of digital human videos. For example, through dynamic camera movement commands, when a digital human teacher is explaining the blackboard writing, the camera can automatically switch to a close-up of the blackboard writing without requiring the audience to manually drag the screen to view the details.

[0079] Based on the above implementation method, in order to further improve the picture effect of digital human video, the method may optionally include: during the generation of digital human video, obtaining at least one of the following camera movement adjustment related parameters: the movement trajectory of digital human, audience interaction parameters, and scene layout change parameters; and dynamically adjusting at least one of the following camera movement parameters in digital human video according to the camera movement adjustment related parameters: camera focus, movement speed, and composition ratio.

[0080] The movement trajectory of the digital human can be the initial and current coordinates of the head or torso. For example, when the current coordinates of the digital human deviate from the initial coordinates by a preset distance, the movement trajectory of the digital human can be determined to be walking, and the camera movement mode can be adjusted to follow, so that the movement speed of the camera is synchronized with the movement of the digital human.

[0081] Audience interaction parameters can include audience comments or bullet comments. Adjusting the camera movement of digital human videos using these parameters can be achieved by triggering the digital human character's turn based on audience comments or bullet comments. For example, if the audience interaction parameter is "The whiteboard is unclear," the camera focus can be adjusted to point the camera at the whiteboard, improving its clarity.

[0082] Scene layout changes can include virtual stage background switching, stage raising and lowering, or switching between multiple stages. For example, the aspect ratio of the digital human video can be adjusted when the scene layout changes. For instance, when the scene layout changes to a close-up, the aspect ratio of the digital human video can be adjusted to enlarge a specific area; when the scene layout changes to a distant view, the aspect ratio can be adjusted to display more of the scene. For example, when multiple people enter at a press conference, adjusting the aspect ratio can cover the entire scene and prevent compositional imbalance.

[0083] By adjusting the camera movement parameters of digital human videos based on the camera movement correlation parameters, the naturalness and adaptability of digital human video footage can be improved, breaking through the bottleneck of static fixed camera movement.

[0084] Figure 4 This is a flowchart illustrating a digital human video generation method based on dynamic camera movement, according to Embodiment 2 of the present invention. Figure 4 As shown, when generating a digital human video, the text to be broadcast can be obtained and preprocessed. Preprocessing can include word segmentation and part-of-speech tagging. After preprocessing, the text can be converted into speech feature data. When generating speech feature data, broadcast style and parameters can be combined. Speech feature data can include frequency cepstral coefficients, fundamental frequency F0, and energy spectrum. Lip animation can be generated based on the speech feature data. For example, lip-shape keypoint mapping and deep learning models can be used to generate lip animation. Based on the text to be broadcast, speech feature data, and broadcast style, body movements can be generated. For example, initial body movements can be generated using a body movement generation tool, and then optimized using reinforcement learning, such as audience feedback data. Animation rendering based on body movements, speech feature data, and lip animation yields the initial digital human video. Adjusting the initial digital human video using dynamic camera movement commands generates the final digital human video. This process can be further refined using methods such as... Figure 4The processing flow shown in this invention, and the technical solution of this embodiment, solve the problem of separation between lip movements, actions, and camera movements in digital human videos. By optimizing body movements, the naturalness of the body can be improved. By generating digital human videos under dynamic camera movement commands, the realism of digital human videos can be improved, the loss of key frames can be avoided, and the emotional expression of digital humans can be enhanced through camera changes, avoiding stiff performances.

[0085] Example 3

[0086] Figure 5 This is a flowchart of a digital human video generation method based on dynamic camera movement according to Embodiment 3 of the present invention. This embodiment is a further refinement of the above technical solution. The technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments.

[0087] Optionally, dynamic camera movement instructions are generated based on body movements, speech feature data, text to be played, and playback style. This includes: collecting camera movement videos under different playback styles and identifying camera movement-related data in the videos; generating a generalized camera movement decision framework based on the video movements under different playback styles and the corresponding camera movement-related data; and searching for the corresponding dynamic camera movement instructions in the generalized camera movement decision framework based on body movements, speech feature data, text to be played, and playback style.

[0088] Step 510: Obtain the text to be played, the playback style, and the playback parameters, and generate speech feature data corresponding to the text to be played based on the text to be played, the playback style, and the playback parameters.

[0089] Step 520: Generate corresponding lip-sync animations based on speech feature data, and align the speech feature data and lip-sync animations in time sequence.

[0090] Step 530: Generate body movements corresponding to the voice feature data based on the text to be broadcast, voice feature data, and broadcast style.

[0091] Step 540: Collect camera movement videos under different broadcast styles and identify camera movement-related data in the video movement videos.

[0092] Step 550: Generate a generalized camera movement decision framework based on the camera movement videos under different broadcast styles and the corresponding camera movement related data.

[0093] A generalized camera movement decision framework can be a set of camera movement schemes to be used in different situations. For example, it could include camera movement schemes to be used in news broadcasts when different digital human emotions, keywords, and positions are present; or camera movement schemes to be used in speeches when different digital human emotions and gestures are present. A generalized camera movement decision framework can also be a framework table that maps camera movement-related data to camera movement instructions.

[0094] Step 560: Based on body movements, voice feature data, text to be broadcast, and broadcast style, find the corresponding dynamic camera movement instructions in the generalized camera movement decision framework.

[0095] When applying a generalized camera movement decision framework, camera movement schemes can be matched within the framework using camera movement correlation data in the current digital human video. The target camera movement scheme with the highest matching degree can be found and used as the dynamic camera movement command. Therefore, the digital human video is generated using the camera movement command corresponding to the target camera movement scheme in the generalized camera movement decision framework.

[0096] Step 570: Under the dynamic camera movement command, perform animation rendering based on body movements, voice feature data, and lip-sync animation to obtain a digital human video.

[0097] Based on the above implementation method, optionally, under dynamic camera movement commands, animation rendering is performed based on body movements, speech feature data, and lip-sync animation to obtain a digital human video, including: under dynamic camera movement commands, animation rendering is performed based on body movements, speech feature data, and lip-sync animation to obtain an initial rendering result; based on the position of the digital human in the digital human video and the lighting direction of the virtual stage, the shadow of the digital human on the virtual stage ground is formed; the movement trajectory of the digital human in the digital human video is detected, and the shadow changes are controlled according to the movement trajectory to obtain the light and shadow tracking result in the digital human video.

[0098] For example, when the light source is far from the stage, the shadow of the digital human on the virtual stage can be shrunk; when the light source is close to the stage, the shadow can be enlarged. As another example, when the digital human turns or raises its hand, skeletal animation tracking can be used to determine the relative positions of different parts of the digital human's body to the light source, and the shadow can be tracked and adjusted accordingly. By changing the shadow based on the digital human's movement trajectory, real-time light and shadow tracking effects can be achieved, enhancing the immersiveness and realism of video broadcasts.

[0099] The technical solution of this invention involves acquiring the text to be played, the playback style, and playback parameters; generating speech feature data corresponding to the text to be played based on the text to be played, the playback style, and the playback parameters; generating corresponding lip-sync animation based on the speech feature data; aligning the speech feature data and lip-sync animation in sequence; generating body movements corresponding to the speech feature data based on the text to be played, the speech feature data, and the playback style; acquiring camera movement videos under different playback styles and identifying camera movement correlation data in the video movements; generating a generalized camera movement decision framework based on the camera movement videos under different playback styles and the corresponding camera movement correlation data; and generating a generalized camera movement decision framework based on body movements, Voice feature data, text to be read, and reading style are used to find corresponding dynamic camera movement instructions within a generalized camera movement decision framework. Under the dynamic camera movement instructions, animation rendering is performed based on body movements, voice feature data, and lip-sync animation to obtain digital human videos. This solves the problem of separation between lip movements, actions, and camera movements in digital human videos. Optimization of body movements can improve the naturalness of the movements. Generating digital human videos under dynamic camera movement instructions can improve the realism of digital human videos and avoid the loss of keyframes. The emotional expression of digital humans can be enhanced through camera changes, avoiding stiff performances. Real-time light and shadow tracking effects can improve the immersion and realism of the reading.

[0100] Figure 6 This is a deployment flowchart of a digital human video generation system based on dynamic camera movement, provided in Embodiment 3 of the present invention. Figure 6 As shown, a digital human video generation system based on dynamic camera movement can include: a digital human video production management subsystem, a rendering control subsystem, and an output processing subsystem. The dynamic camera movement-based digital human video generation system can perform initialization, load configuration files (such as loading rendering parameters, a generalized camera movement decision framework, and digital human video output formats), and preheat resources for digital human video generation (such as scene models, voice libraries, and preloading special effects materials). The digital human video production management subsystem can manage video broadcast production. During video broadcast production management, process control can be implemented, such as status monitoring, task scheduling, and exception handling. Status monitoring can include monitoring rendering progress, resource usage, and error logs. Task scheduling can be priority-based module calls. Exception handling can include timeout retries, degradation schemes, and user notifications. The digital human video production management subsystem can also include a user interaction layer. The user interaction layer can include text input, real-time preview (such as low-resolution draft rendering), and style selection (such as news, entertainment, or technology).

[0101] like Figure 6As shown, the rendering control subsystem can perform rendering and camera movement control of the virtual scene. Virtual scene rendering and camera movement control can include dynamic camera movement, a 3D rendering engine, and a scene bus. Dynamic camera movement can extract emotion tags from text, map shot language based on emotion, perform dynamic path planning (such as avoiding clipping and maintaining compositional balance), and execute camera movement commands. During the execution of camera movement commands, keyframes can be smoothly transitioned using Bezier curves, and rendering results can be fine-tuned in real time for correction. The 3D rendering engine can include performance optimization (such as LOD grading, occlusion culling, and GPU instantiation) and engine configuration (such as PBR materials, global illumination, and post-processing stack). Dynamic data can be injected through the scene bus, such as real-time text, voice duration, and facial expression parameters.

[0102] like Figure 6 As shown, the output processing subsystem can process digital human videos. For example, video output processing can include a compositing pipeline, multi-platform adaptation, and a real-time interactive layer. The compositing pipeline can include multi-track mixing (such as voice, background music, and spatial positioning of sound effects) and dynamic bitrate control (such as adjusting quality based on network conditions). Multi-platform adaptation can include format encapsulation (such as MP4, HLS, DASH, WebM formats), resolution adaptation (such as 480p-8K adaptive), and encryption processing (such as DRM encryption for content copyright protection). The real-time interactive layer can include a streaming media server, interactive interfaces (such as bullet comments, likes, and real-time subtitle switching), and CDN distribution (such as edge node caching strategies).

[0103] like Figure 6 As shown, data analysis can also be performed in the digital human video generation system based on dynamic camera movement. Data analysis can be used to optimize the model (such as iterating camera movement rules and adjusting rendering parameters), perform playback statistics (such as completion rate, bounce rate, and interactive hotspots), and conduct quality assessment (such as user ratings).

[0104] Through such Figure 6The system integration and output shown enable the specific deployment of a digital human video generation system based on dynamic camera movement, applying the method for generating digital human videos based on dynamic camera movement. The digital human video generation method based on dynamic camera movement provided in this invention tightly integrates digital human performance with the dynamic camera movement control of a virtual scene. Through dynamic camera movement, the lip movements, actions, and emotional states of the digital human can be analyzed in real time, dynamically generating camera movement commands such as push-pull shots, close-up / far-close switching, and full-body / half-body perspective transitions. This ensures a high degree of coordination between camera movement and the digital human performance content. Camera movement commands are generated in real time based on the digital human's performance content (such as emotions and speech rate), achieving intelligent camera scheduling and free adjustment of perspective. This solves the problems of unschedulable cameras and stiff performances in traditional technologies, making digital human performances more natural, fluid, and dynamically layered, enhancing the audience's immersion. For example, in online education scenarios, when a teacher explains the blackboard, the camera can automatically switch to a close-up of the blackboard, eliminating the need for students to manually drag the screen to view details, thus improving user experience and interactivity.

[0105] The digital human video generation method based on dynamic camera movement achieves end-to-end intelligent collaboration between lip movements, actions, and camera movement, eliminating the timeline misalignment problem caused by independent control of each module in traditional technologies. This improves the naturalness and smoothness of the performance, increasing naturalness by more than 50%. Furthermore, it supports real-time rendering and playback, making the generation process more efficient and faster. Generating body movements through a body movement generation tool significantly reduces data acquisition costs and time cycles, shortening the single-scene video generation cycle by more than 90%, and solving the problems of poor material reusability and high costs in traditional technologies.

[0106] By using multimodal scene perception to detect the movement trajectory of the digital human character, audience interaction events, and changes in scene layout in real time, the system dynamically adjusts the camera focus, movement speed, and composition ratio. This allows the system to perform exceptionally well in complex and dynamic scenes, avoiding the loss of key shots or compositional imbalances caused by rigid preset parameters in traditional rule-based solutions. Simultaneously, it supports seamless integration with existing AI generation engines, automatically adapting to the camera movement requirements of different styles and platforms, thus improving AI scene compatibility.

[0107] Example 4

[0108] Figure 7 This is a schematic diagram of a digital human video generation device based on dynamic camera movement, according to Embodiment 4 of the present invention. Figure 7 As shown, the device includes: a voice feature data generation module 710, a timing alignment module 720, a body motion generation module 730, a dynamic camera movement command generation module 740, and a digital human video generation module 750. Among them:

[0109] The speech feature data generation module 710 is used to acquire the text to be played, the playback style and the playback parameters, and generate speech feature data corresponding to the text to be played based on the text to be played, the playback style and the playback parameters.

[0110] The timing alignment module 720 is used to generate corresponding lip-sync animations based on speech feature data and to perform timing alignment between the speech feature data and the lip-sync animations.

[0111] The body movement generation module 730 is used to generate body movements corresponding to the voice feature data based on the text to be broadcast, voice feature data, and broadcast style.

[0112] The dynamic camera movement instruction generation module 740 is used to generate dynamic camera movement instructions based on body movements, voice feature data, text to be broadcast, and broadcast style.

[0113] The digital human video generation module 750 is used to generate digital human videos by rendering animations based on body movements, voice feature data, and lip-syncing under dynamic camera movement commands.

[0114] Optionally, the dynamic camera movement command generation module 740 includes:

[0115] The camera movement data acquisition unit is used to collect camera movement videos under different broadcast styles and identify camera movement data in the video.

[0116] The dynamic camera movement decision model generation unit is used to train the model based on the camera movement video and the corresponding camera movement related data to obtain the dynamic camera movement decision model.

[0117] The first dynamic camera movement instruction generation unit is used to input body movements, speech feature data, text to be broadcast, and broadcast style into the dynamic camera movement decision model to obtain the corresponding dynamic camera movement instructions.

[0118] Optionally, the dynamic camera movement command generation module 740 includes:

[0119] The camera movement association data unit is used to collect camera movement videos under different broadcast styles and identify camera movement association data in the video.

[0120] The generalized camera movement decision framework unit is used to generate a generalized camera movement decision framework based on camera movement videos under different broadcast styles and corresponding camera movement related data.

[0121] The second dynamic camera movement instruction generation unit is used to find the corresponding dynamic camera movement instruction in the generalized camera movement decision framework based on body movements, speech feature data, text to be broadcast, and broadcast style.

[0122] Optionally, the device may also include:

[0123] The camera movement adjustment related parameter acquisition module is used to acquire at least one of the following camera movement adjustment related parameters during the digital human video generation process: the digital human's movement trajectory, audience interaction parameters, and scene layout change parameters;

[0124] The camera movement parameter adjustment module is used to dynamically adjust at least one of the following camera movement parameters in the digital human video based on the camera movement adjustment related parameters: camera focus, movement speed, and composition ratio.

[0125] Optionally, the limb motion generation module 730 includes:

[0126] The initial body movement generation unit is used to generate initial body movements based on the text to be broadcast, speech feature data, and broadcast style using a body movement generation tool.

[0127] The body movement reward and punishment factor determination unit is used to obtain audience feedback data on body movements during the playback of digital human videos, and to form initial body movement reward and punishment factors based on the audience feedback data.

[0128] The body movement optimization unit is used to adjust the initial body movements according to reward and punishment factors to obtain the body movement optimization results; and to update the digital human video according to the body movement optimization results.

[0129] Optionally, the timing alignment module 720 includes:

[0130] The initial temporal alignment unit is used to bind speech feature data with lip movement data in lip animation and perform initial temporal alignment using an alignment algorithm;

[0131] The timing deviation detection unit is used to detect the timing deviation between the initially time-aligned speech feature data and the lip-sync animation during the animation rendering process.

[0132] The temporal dynamic compensation unit is used to perform temporal dynamic compensation on speech feature data and lip-sync animation based on temporal deviations, and adjust the alignment algorithm based on the temporal dynamic compensation results.

[0133] Optional, the digital human video generation module 750 includes:

[0134] The initial rendering unit is used to perform animation rendering based on body movements, speech feature data, and lip-sync animation under dynamic camera movement commands to obtain the initial rendering result.

[0135] The shadow forming unit is used to form the shadow of the digital human on the virtual stage ground based on the position of the digital human in the digital human video and the direction of lighting on the virtual stage;

[0136] The light and shadow tracking unit is used to detect the movement trajectory of the digital human in the digital human video and control the shadow change according to the movement trajectory to obtain the light and shadow tracking results in the digital human video.

[0137] The digital human video generation device based on dynamic camera movement provided in the embodiments of the present invention can execute the digital human video generation method based on dynamic camera movement provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0138] In the technical solutions of this invention, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information (such as voice, body movements, audience feedback data, etc.) all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0139] The information collected is information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding operation portals are provided for users to choose to authorize or refuse.

[0140] Provide users with corresponding operation entry points, allowing them to choose to agree to or reject the automated decision results; if the user chooses to reject, the process will proceed to the expert decision-making process.

[0141] Example 5

[0142] Figure 8 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0143] like Figure 8As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0144] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0145] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a digital human video generation method based on dynamic camera movement.

[0146] In some embodiments, the dynamic camera-based digital human video generation method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the dynamic camera-based digital human video generation method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the dynamic camera-based digital human video generation method by any other suitable means (e.g., by means of firmware).

[0147] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0148] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0149] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0150] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0151] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0152] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0153] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0154] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for generating digital human videos based on dynamic camera movement, characterized in that, include: The system obtains the text to be played, the playing style, and the playing parameters, and generates speech feature data corresponding to the text to be played based on the text to be played, the playing style, and the playing parameters. Generate a corresponding lip-sync animation based on the speech feature data, and align the speech feature data and the lip-sync animation in time sequence. Based on the text to be broadcast, the voice feature data, and the broadcast style, generate body movements corresponding to the voice feature data; Based on the body movements, the voice feature data, the text to be played, and the playback style, dynamic camera movement instructions are generated. Under the dynamic camera movement command, animation rendering is performed based on the body movements, the voice feature data, and the lip-sync animation to obtain a digital human video.

2. The method according to claim 1, characterized in that, Based on the body movements, the speech feature data, the text to be broadcast, and the broadcast style, dynamic camera movement instructions are generated, including: Collect video footage of camera movements under different broadcasting styles, and identify camera movement-related data in the video footage; The model is trained based on the camera movement video and the corresponding camera movement data to obtain a dynamic camera movement decision model; The body movements, the speech feature data, the text to be broadcast, and the broadcast style are input into the dynamic camera movement decision model to obtain the corresponding dynamic camera movement instructions.

3. The method according to claim 1, characterized in that, Based on the body movements, the speech feature data, the text to be broadcast, and the broadcast style, dynamic camera movement instructions are generated, including: Collect video footage of camera movements under different broadcasting styles, and identify camera movement-related data in the video footage; Based on the camera movement videos under different broadcasting styles and the corresponding camera movement related data, a generalized camera movement decision framework is generated. Based on the body movements, the voice feature data, the text to be broadcast, and the broadcast style, the corresponding dynamic camera movement instruction is searched in the generalized camera movement decision framework.

4. The method according to claim 1, characterized in that, Also includes: During the digital human video generation process, at least one of the following camera movement adjustment parameters is obtained: the digital human's movement trajectory, audience interaction parameters, and scene layout change parameters; Based on the aforementioned camera movement adjustment parameters, at least one of the following camera movement parameters in the digital human video is dynamically adjusted: camera focus, movement speed, and composition ratio.

5. The method according to claim 1, characterized in that, Based on the text to be broadcast, the speech feature data, and the broadcast style, generate body movements corresponding to the speech feature data, including: Based on the text to be broadcast, the voice feature data, and the broadcast style, an initial body movement is generated using a body movement generation tool. During the playback of the digital human video, audience feedback data on body movements is acquired, and the initial body movement reward and punishment factor is formed based on the audience feedback data. The initial body movements are adjusted according to the reward and punishment factors to obtain optimized body movement results; and the digital human video is updated according to the optimized body movement results.

6. The method according to claim 1, characterized in that, The speech feature data and the lip-sync animation are time-series aligned, including: The speech feature data is bound to the lip movement data in the lip-sync animation, and initial temporal alignment is performed using an alignment algorithm; During animation rendering, the timing deviation between the initial timing-aligned speech feature data and the lip-sync animation is detected. The speech feature data and the lip-sync animation are dynamically compensated for according to the time-series deviation, and the alignment algorithm is adjusted according to the result of the dynamic time-series compensation.

7. The method according to claim 1, characterized in that, Under the dynamic camera movement command, animation rendering is performed based on the body movements, the speech feature data, and the lip-sync animation to obtain a digital human video, including: Under the dynamic camera movement command, animation rendering is performed based on the body movements, the speech feature data, and the lip-sync animation to obtain the initial rendering result; Based on the position of the digital human in the digital human video and the direction of the virtual stage lighting, the shadow of the digital human on the virtual stage ground is formed; The movement trajectory of the digital human in the digital human video is detected, and the shadow changes are controlled according to the movement trajectory to obtain the light and shadow tracking results in the digital human video.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the digital human video generation method based on dynamic camera movement as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the digital human video generation method based on dynamic camera movement as described in any one of claims 1-7.

10. A computer program product comprising a computer program that, when executed by a processor, implements the digital human video generation method based on dynamic camera movement according to any one of claims 1-7.