Short video data real-time inference method and system

CN122840233APending Publication Date: 2026-09-29HONGDONG FUTURE (JILIN) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610923452.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0006]本发明解决了如何在保持可接受的生成质量的前提下,将数字人模型的训练数据需求从数小时降至分钟级、训练时间从数天缩短至小时级以及推理硬件门槛从高端GPU降低至消费级GPU水平的问题

Benefits of technology

[0016]本发明解决了如何在保持可接受的生成质量的前提下,将数字人模型的训练数据需求从数小时降至分钟级、训练时间从数天缩短至小时级以及推理硬件门槛从高端GPU降低至消费级GPU水平的问题。具体有益效果包括:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840233A_ABST
    Figure CN122840233A_ABST
Patent Text Reader

Abstract

The application discloses a short video data real-time inference method and system, relates to the technical field of data processing, and solves the problem of how to reduce the training data requirement of a digital human model from several hours to minutes, shorten the training time from several days to hours, and reduce the inference hardware threshold from a high-end GPU to a consumer-grade GPU level while keeping acceptable generation quality. The method comprises the following steps: S1, collecting short video data and performing preprocessing; S2, performing synchronous feature extraction on the preprocessed short video data, and outputting 512-dimensional audio features; S3, constructing a light-weight U-Net model; and S4, inputting the 512-dimensional audio features into the light-weight U-Net model, and performing real-time inference on the short video data on the basis of the preprocessed short video data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, specifically to a method and system for real-time reasoning of short video data. Background Technology

[0002] Audio-driven talking face generation is a hot research topic in computer vision and multimedia. This technology aims to generate natural-looking videos of human faces speaking, synchronized with the audio content, based on audio input. From a technological evolution perspective, this field has roughly gone through three research and development phases: In the first phase of research and development, GAN-based lip editing methods were represented by Wav2Lip. This method employs a generative adversarial network architecture to map audio features to the lip region, replacing the lips in the original video. Its core consists of a pre-trained lip-sync expert and a generator network. However, such methods typically require hours of video data for training and rely on high-performance GPUs (usually 24GB or more of VRAM). The generated results are limited to local modifications in the lip region, making it difficult to fully generate dynamic facial expressions including head pose and facial expressions.

[0003] In the second research and development phase, parameterization methods based on 3D models / motion coefficients were represented by SadTalker. This method introduces 3D facial reconstruction technology to predict 3D motion coefficients (including head rotation, expression coefficients, etc.) from audio, and then drives a static reference image to generate animation. Its main problem is the need to pre-collect a large amount of facial data from different angles for 3D model fitting, and because it relies on the 3D prior of a single reference image, it is prone to distortion when facing non-frontal poses or significant facial expression changes. Furthermore, the inference speed on consumer-grade GPUs is approximately 100-200ms / frame, which is insufficient to meet real-time standards.

[0004] In the third research and development phase, facial rendering methods based on Neural Radiation Fields (NeRF) were represented by GeneFace++ and RAD-NeRF. These methods utilize neural radiation fields to reconstruct the head in 3D and combine this with audio features to drive rendering, offering advantages in image quality and multi-view consistency. However, their computational overhead is extremely high. GeneFace++'s inference speed on an A100 GPU is approximately 200-500ms / frame, and training requires hours to days of video data and a GPU with at least 24GB of video memory, making it unsuitable for real-time applications and widespread adoption.

[0005] In recent years, real-time streaming methods based on U-Net and implicit keypoint-driven methods have improved inference speed to some extent, but they still face common bottlenecks: training a usable digital human model requires at least tens of minutes to several hours of high-quality training data, and inference still needs to run on high-end GPUs (such as RTX 3090 and above), making it impossible to achieve smooth real-time interaction on consumer-grade GPUs. Summary of the Invention

[0006] This invention solves the problem of how to reduce the training data requirements of digital human models from hours to minutes, the training time from days to hours, and the inference hardware threshold from high-end GPUs to consumer-grade GPUs while maintaining acceptable generation quality.

[0007] The real-time inference method for short video data described in this invention includes the following steps: Step S1: Collect short video data and perform preprocessing; Step S2: Extract synchronous features from the preprocessed short video data and output 512-dimensional audio features; Step S3: Construct a lightweight U-Net model; Step S4: Input the 512-dimensional audio features into the lightweight U-Net model, and perform real-time inference on the short video data based on the preprocessed short video data.

[0008] Furthermore, in one embodiment of the present invention, step S1, which involves collecting and preprocessing short video data, includes the following steps: Step S101: Capture short video data, which includes silent segments and audio segments; Step S102: Based on short video data, perform face detection and key feature extraction to obtain a standardized face image sequence; Step S103: Extract audio features based on the audio segments of the short video data; Step S104: Extract the silent reference frame based on the silent segment of the short video data.

[0009] Furthermore, in one embodiment of the present invention, step S2, which involves extracting synchronous features from the preprocessed short video data and outputting 512-dimensional audio features, specifically includes: Based on the SyncNet model, 512-dimensional audio features and 512-dimensional normalized face image sequences are extracted frame by frame, and then joint feature extraction is performed to output 512-dimensional audio features.

[0010] Furthermore, in one embodiment of the present invention, the lightweight U-Net model in step S3 specifically refers to: The image input encoder outputs a 12×12 feature image. This feature image passes through a bottleneck layer to achieve cross-modal feature fusion before being input into the decoder, which outputs a 96×96 feature image.

[0011] Furthermore, in one embodiment of the present invention, the image input encoder outputs a feature image with a size of 12×12, specifically: The image input is a 3×3 convolutional layer with a stride of 2, and the output is an image with a size of 48×48 and 64 channels. This image is then continuously input into two levels of inverse residual modules, and the output is a feature image with a size of 12×12. The step size of the inverted residual module is 2, and the expansion factor is 6.

[0012] Furthermore, in one embodiment of the present invention, the feature map is passed through a bottleneck layer to achieve cross-modal feature fusion before being input into the decoder, specifically as follows: The feature map is processed by the inverse residual module, and the output channel is a 512 feature map. After feature concatenation with the audio features, it is input into the inverse residual module again, and the output channel is a 512 feature fusion map.

[0013] Furthermore, in one embodiment of the present invention, the input to the decoder outputs a feature image with a size of 96×96, specifically as follows: The 512-channel feature fusion image is input to the upsampling layer, then downsampled to 256 channels using a 3×3 convolutional layer. After being concatenated with the output of the inverse residual module in the encoder, it is input to the upsampling layer and downsampled to 128 channels using a 3×3 convolutional layer. This image is then concatenated with the features output from a 3×3 convolutional layer with a stride of 2 in the encoder, resulting in 192 channels. Finally, a 1×1 convolutional layer outputs a feature image with a size of 96×96.

[0014] Furthermore, in one embodiment of the present invention, the reverse residual module specifically comprises: The number of input channels is increased to a factor of 1×1 point convolution, and the output is after batch normalization and ReLU activation. The feature map after dimensionality increase is subjected to a 3×3 depthwise separable convolution, where the number of convolution groups is equal to the number of channels after dimensionality increase, and the convolution stride is set to 1 or 2 according to the module configuration. After batch normalization and ReLU activation, the output is processed. The feature map channels are projected to the output channels using 1×1 point convolution, and then output after batch normalization. When the number of input channels and the number of output channels of the inverse residual module are the same and the step size is 1, the output of the inverse residual module is the element-wise addition of the output of the dimensionality reduction projection operation and the input of the inverse residual module; otherwise, the output of the inverse residual module is the output of the dimensionality reduction projection operation.

[0015] The present invention discloses a real-time inference system for short video data, which is implemented based on the real-time inference method for short video data described above, and includes the following modules: Module S1 collects short video data and performs preprocessing; Module S2 performs synchronous feature extraction on the preprocessed short video data and outputs 512-dimensional audio features; Module S3, constructs a lightweight U-Net model; Module S4 inputs 512-dimensional audio features into a lightweight U-Net model and performs real-time inference on the preprocessed short video data.

[0016] This invention solves the problem of reducing the training data requirements for digital human models from hours to minutes, shortening training time from days to hours, and lowering the inference hardware threshold from high-end GPUs to consumer-grade GPUs while maintaining acceptable generation quality. Specific beneficial effects include: The present invention discloses a real-time inference method for short video data. In order to overcome the technical difficulties existing in the prior art, the present invention performs synchronous feature pre-extraction on the short video data after collecting and preprocessing the short video data, and finally realizes real-time inference of short videos through the construction of a lightweight U-Net model. The training data requirement is reduced from several hours to minutes (3-5 minutes), and the training time is shortened from several days to hours (3-6 hours). Moreover, the model inference can be run in real time (80-125fps) on consumer-grade GPUs (such as NVIDIA RTX 2080 / 3060). Attached Figure Description

[0017] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a diagram of the lightweight U-Net model structure described in Implementation Method 1; Figure 2 This is a schematic diagram of the internal structure of the reverse residual module described in Implementation Method 1. Detailed Implementation

[0018] Various embodiments of the present invention will now be clearly and completely described with reference to the accompanying drawings. The embodiments described with reference to the drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0019] Implementation Method 1: This implementation method aims to solve the technical problems of how to reduce the training data requirements of digital human models from hours to minutes, the training time from days to hours, and the inference hardware threshold from high-end GPUs to consumer-grade GPUs while maintaining acceptable generation quality.

[0020] The essence of this technical problem is a balancing act between three mutually constraining technical indicators, primarily because: 1) The contradiction between data volume and training stability: Reducing the amount of training data will lead to underfitting of the model. Existing technologies lack effective cross-modal pre-training to compensate for insufficient information in small sample scenarios. 2) The contradiction between lightweight model and generation quality: Compressing the network parameter size will inevitably result in a loss of feature representation ability. How to maintain acceptable lip-sync accuracy and visual realism when the number of parameters is greatly reduced? 3) The contradiction between computational load and real-time performance: Consumer-grade GPUs have limited computing power. How can we complete the entire inference pipeline of audio feature extraction, visual feature processing, feature fusion and image generation within a limited FLOPs budget?

[0021] Therefore, in order to overcome the above-mentioned technical difficulties, this embodiment proposes a real-time reasoning method for short video data, taking into account the interaction between the various steps, including the following steps: Step S1: Short video data acquisition and preprocessing Step S101, the collected video segment data includes two stages: Seconds 1-20: The photographer keeps their mouth shut and maintains a natural expression (silent segment).

[0022] From the 21st second to the end: The photographer reads the pre-selected text (audio segment, it is recommended to cover different vowel and consonant combinations).

[0023] The total duration should be controlled between 3 and 5 minutes. This implementation method has been verified through extensive experiments. Training data shorter than 3 minutes results in insufficient coverage of pronunciation mouth shapes by the model; while training data longer than 5 minutes leads to diminishing marginal returns and violates the design goal of lowering the data threshold.

[0024] Step S102, Face detection and key point extraction: Lightweight detectors such as MTCNN or RetinaFace are used to extract face regions and crop them to a uniform size of 96×96 pixels. Validation shows that 96×96 resolution achieves a good balance between model size and generation quality: lower resolutions (e.g., 64×64) result in blurred lip details, while higher resolutions (e.g., 128×128) significantly increase model parameters and inference time.

[0025] Step S103, Audio Feature Extraction: Audio feature extraction is performed by switching between Wenet and Hubert encoders. The reason for retaining both options instead of fixing one is based on a trade-off between different deployment scenarios: the Wenet encoder has fewer parameters and faster inference speed, making it suitable for mobile devices / low-computing-power scenarios; the Hubert encoder has stronger feature representation capabilities and better performance, making it suitable for server-side deployment. Switching between the two can be achieved through configuration options without affecting the overall workflow.

[0026] Step S104, Extraction of the silent reference frame: The best-quality frame from the first 20 frames of the silent segment was selected as the baseline frame for silence. The reason for selecting multiple frames instead of fixing the first frame is that there are often subtle facial expression adjustments at the moment of shooting, and directly taking the first frame may contain unnatural muscle tension.

[0027] Step S2, Synchronous Feature Pre-extraction Step S201: Load the pre-trained SyncNet model.

[0028] The SyncNet model is an audio-video synchronization discrimination model that includes an audio encoder branch and a video encoder branch. It maps audio and video pairs to a shared feature space through contrastive learning.

[0029] In this implementation, the SyncNet model is used as a "pre-trained feature extractor." Before training the main model, the SyncNet model extracts joint audio-visual features from the training data, which serve as the initialization conditions for the main model's training. Its purpose is to accelerate the convergence speed of the main model, enabling good alignment results even with limited training data (3-5 minutes).

[0030] In step S202, 512-dimensional audio feature vectors are extracted frame by frame from the training data in step S103, and 512-dimensional visual feature vectors are extracted frame by frame from the training data in step S102.

[0031] Step S203: The 512-dimensional audio feature vector and 512-dimensional visual feature vector extracted frame by frame are cached synchronously.

[0032] Existing technologies require hours of video training because the model needs to learn the high-dimensional prior of "natural facial motion distribution." Therefore, this implementation reduces data requirements in steps S1 and S2 through two methods: First, it utilizes cross-modal alignment initialization provided by the pre-trained SyncNet model to reduce the main model's dependence on the amount of training data; second, it collects silent segments from the video to extract baseline facial features in a "non-speaking state," which are then reused during the inference phase.

[0033] Step S3: Construct a lightweight U-Net model Existing U-Net models typically employ encoder channel number increments from 128 to 1024 (as seen in the original U-Net model and most variants), resulting in tens of millions of parameters. Therefore, directly applying the existing U-Net model to this implementation lacks systematic simplification of the training process, still requiring over 30 minutes of data, and fails to achieve the desired reduction in training time. Therefore, as... Figure 1 As shown, this implementation proposes a lightweight U-Net model, specifically: The specific parameters for each layer of the encoder are as follows: The first layer is a 3×3 standard convolution with a stride of 2, which takes a single frame image from the standardized face image sequence obtained in step S102 as input. The image is a 96×96 pixel RGB three-channel image (obtained from video frames after face detection and cropping). The first layer is a standard 3×3 convolution with a stride of 2, which converts the RGB three-channel input to a 64-channel output, reducing the image size from 96×96 to 48×48. The second layer is an inverse residual module with 64 inputs and 128 outputs, a stride of 2, and a spread factor of 6, reducing the size to 24×24. The third layer also uses an inverse residual module with 128 inputs and 256 outputs, with the same parameters as above, reducing the size to 12×12.

[0034] like Figure 2 As shown, the input-output relationship of each layer of the inverted residual module is as follows: First layer: 1×1 point convolution (dimensionality increase layer), with C_in as the number of input channels and C_in×6 as the number of output channels (expansion factor is 6). The feature map size remains unchanged; it is output after batch normalization and ReLU activation.

[0035] The second layer is a 3×3 depthwise separable convolution (feature extraction layer) with C_in×6 input channels and C_in×6 output channels. The number of convolutional groups is equal to the number of input channels; when the stride is 2, the feature map size is reduced to 1 / 2 of its original size, and remains unchanged when the stride is 1. Output after batch standardization and ReLU activation.

[0036] The third layer: 1×1 point convolution (dimensionality reduction projection layer), with C_in×6 input channels and C_out output channels; The feature map size remains unchanged; it is output after batch standardization.

[0037] Residual connection: When the number of input channels C_in equals the number of output channels C_out and the step size is 1, the final output of the module is the element-wise addition of the third-layer output and the module input; otherwise, the output of the third layer is directly output.

[0038] Taking the second layer of the encoder as an example: the input is a 64-channel, 48×48 feature map, and the output after passing through the inverse residual module is a 128-channel, 24×24 feature map. The number of up-dimensional channels is 384 (64×6), and the down-dimensional output is 128 channels.

[0039] Compared to ordinary convolution, the inverse residual module described in this embodiment can extract richer local features with the same number of parameters; compared to the standard residual module, the expand mechanism compensates for the information loss caused by channel reduction, and is especially suitable for low-channel-count scenarios.

[0040] The encoder only performs downsampling twice, which is 3-4 times less than the U-Net model in the existing technology. This is because experiments have shown that for a 96×96 input, the spatial resolution is still 12×12 after two downsamplings, which is sufficient to preserve the key spatial information of the lip region, while avoiding information loss and increased computation caused by over-downsampling.

[0041] The bottleneck layer is the connection channel between the encoder and decoder, and also the key location for receiving audio features from the SyncNet model. It is designed as two cascaded inverse residual modules with 512 channels. Specifically: The 12×12 image is fed into the first inverse residual module. Through feature concatenation, the 512-dimensional audio feature vector output by the SyncNet model is concatenated with the visual feature vector output by the encoder along the channel dimension.

[0042] The reason for choosing concatenation instead of weighted summation or attention mechanisms is due to the constraints of lightweight design: concatenation does not introduce additional learnable parameters, while summation or attention mechanisms introduce fully connected layers or attention weight matrices, significantly increasing the number of model parameters and inference time. This parameter-free concatenation method has proven to be an efficient and effective cross-modal feature fusion strategy in lightweight models.

[0043] Finally, after passing through the second inverted residual module, the output channel is a 512 feature fusion map.

[0044] In the process of injecting 512-dimensional audio features, this implementation initially injected audio features into the input layer and each subsequent layer separately, but the results were unsatisfactory: injection into the input layer interfered with the texture feature extraction of lower-level networks, resulting in the generated lip shape being in the correct position but with blurred texture; injection into each subsequent layer increased the number of parameters and computational cost exponentially, violating the lightweight design principle. Ultimately, this implementation chose to inject at the bottleneck layer, since the encoder had already completed the extraction of visual features from low to high levels by this point, and the bottleneck layer contains the highest-level visual feature representation. Injecting audio features (512-dimensional) here, followed by the decoder gradually restoring the spatial resolution, resulted in almost no increase in the number of parameters, while the synchronization accuracy was most significantly improved. This injection strategy maximized the guiding role of audio features in lip shape generation while ensuring the model's lightweight nature.

[0045] The decoder consists of two upsampling stages, both implemented using bilinear interpolation. The first layer receives the 512-channel input from the bottleneck layer, upsamples it to 24×24, then uses a 3×3 convolution to reduce it to 256 channels, and then concatenates it with the 128-channel output of encoder Down1 (skip connection), resulting in 384 channels. The second layer upsamples it to 48×48, then uses a 3×3 convolution to reduce it to 128 channels, and then concatenates it with the 64-channel output of encoder Conv1 (skip connection), resulting in 192 channels. Finally, a 1×1 convolution is used to output a three-channel 96×96 RGB image.

[0046] The lightweight U-Net model is trained using a composite loss function with three components, designed to balance the visual quality, perceptual realism, and audio-video synchronization accuracy of the generated images. L1 Reconstruction Loss: Calculates the L1 distance between the generated frame and the ground truth frame pixel by pixel. L1 loss is chosen over L2 loss because it is more robust to outliers and effectively prevents blurring in the generated image. Experiments show that L1 loss provides approximately 0.3-0.5 dB higher PSNR than L2 loss for the same number of training epochs.

[0047] Perceptual loss: High-level features are extracted using a pre-trained VGG16 network, and the distance between the generated frame and the real frame in the feature space is calculated. This component is used to maintain the naturalness of the generated image at the visual perception level, avoid the over-smoothing problem caused by pure L1 loss, and make the generated result more consistent with human visual perception.

[0048] Synchronization Loss: Utilizing the discriminative power of the SyncNet model, a synchronization score is calculated between the generated frame and the corresponding audio. The purpose of this component is to explicitly optimize the audio-video synchronization accuracy during training, ensuring a high degree of match between the generated lip movements and the audio.

[0049] The L1 reconstruction loss weight is set to 1.0, the perceptual loss weight to 0.1, and the synchronization loss weight to 0.5. This configuration has been validated through multiple ablation experiments. Excessive weighting of the synchronization loss leads to the model prioritizing synchronization accuracy at the expense of image quality, while insufficient weighting results in inadequate synchronization accuracy. The current configuration achieves the optimal balance between generation quality and synchronization accuracy.

[0050] The above describes the lightweight U-Net model constructed in this embodiment. Why not directly train the U-Net model end-to-end during the design process? That is, why not directly input the preprocessed data into the lightweight U-Net model? Because this embodiment found that when training end-to-end on 3-5 minutes of data, the lip shape generated by the U-Net model, while visually close to the target, has poor temporal synchronization accuracy (LSE-D score approximately 6.0-6.5). By introducing the synchronization features pre-extracted from the SyncNet model, the U-Net model can directly obtain prior information on audio-visual alignment during training, and the LSE-D score can be improved to approximately 4.2-4.8 with the same amount of data. This indicates that the cross-modal prior knowledge provided by the SyncNet model is crucial for training with small sample sizes.

[0051] Step S4, Real-time Inference Step S401: Receive the real-time audio stream and extract audio features using a Wenet or HuBERT encoder (frame rate set to 20fps or 25fps, depending on the encoder type).

[0052] Step S402: Load the silent reference frame pre-extracted in step S104 into video memory.

[0053] Step S403: Inject audio features into the bottleneck layer of the lightweight U-Net module, using the silent reference frame as the initial visual state, and generate the corresponding lip shape image frame by frame.

[0054] Step S404: The output 96×96 lip area is spliced ​​onto the original video frame (aligned by facial key points).

[0055] Existing methods require real-time processing of camera input to obtain the current facial state. This implementation replaces real-time video input with pre-extracted silent reference frames. Its feasibility is based on the fact that during the inference phase, the U-Net model only needs to know the "current silent state" as a visual baseline, while audio features provide dynamic information driving lip movements. The pre-extracted silent reference frames provide this baseline. This design offers a key advantage in practical deployment: the camera does not need to be activated during inference; only microphone audio input is required to drive the digital human, significantly reducing input device requirements and system complexity. Of course, this also means that the digital human's head posture and expressions are fixed (limited to the silent reference frame state). This implementation is positioned for "lightweight and readily usable" scenarios, rather than high-degree-of-freedom, fully dynamic scenarios.

[0056] To illustrate the technical effects of the real-time inference method for short video data described in this embodiment, as shown in Table 1, the following data was used for comparison and verification: Table 1 Data Comparison Table

[0057] It should be noted that the above-mentioned technical effects in lowering the training threshold achieved by this implementation method do not rely on a single breakthrough in a module, but are the result of the synergistic effect of three levels, specifically manifested as follows: At the data level: The synchronization features pre-extracted by the SyncNet model provide powerful cross-modal prior knowledge for small sample training, effectively alleviating the underfitting problem caused by insufficient training data.

[0058] At the model level: By designing a lightweight U-Net architecture and a fusion strategy that splices audio features at the bottleneck layer, the number of model parameters is compressed to approximately 1.9 million while maintaining acceptable generation quality.

[0059] At the process level: the adoption of silent reference frame reuse and short video acquisition standards greatly reduces the threshold for input data preparation and inference deployment, enabling non-professional users to get started quickly.

[0060] These three levels are indispensable. This implementation method, through the organic integration and optimization of these technical means, achieves a total effect that surpasses the sum of the effects of each individual technology, resulting in significant technical effects.

[0061] To better illustrate the technical effects of the real-time inference method for short video data described in this embodiment, the following examples provide a detailed description: Example 1 Application scenario: Individual users can quickly create their own digital avatars using mobile phones and consumer-grade laptops in a typical home environment for video calls, live streaming, or social media content creation. The specific steps include: Step S1: The user uses the front-facing camera of a regular smartphone (such as an iPhone 12) to record a 3-minute reading video in natural indoor light. The user keeps their mouth shut for the first 20 seconds and reads a passage of text covering various lip movements (such as tongue twisters like "Eight hundred soldiers rushed to the north slope").

[0062] Step S2: Upload the video file to the training application running on the local laptop (equipped with an NVIDIA RTX 3060 / RTX2080 GPU).

[0063] Step S3: After face detection and cropping (96×96 pixels), audio feature extraction (Wenet encoder) begins, followed by feature extraction through the SyncNet model (loading pre-trained weights). Finally, the lightweight U-Net model is trained (composite loss function, approximately 200 epochs, taking 3-4 hours).

[0064] Step S4: After training is complete, the system automatically exports the model in ONNX format.

[0065] In step S5, the user starts the reasoning application and can drive the digital human in real time by speaking through the microphone. The actual test of single-frame reasoning is about 10 milliseconds, achieving a smooth real-time interactive experience.

[0066] Example 2 Application scenario: E-commerce merchants train their own digital human anchors to replace or assist human anchors in achieving 24 / 7 uninterrupted live streaming. This includes the following steps: Step S1: The host records a 5-minute training video in a normal live streaming environment. The host remains silent for the first 20 seconds, and then reads the product description aloud for the last 4 minutes and 40 seconds in a live streaming style. In this embodiment, the background light in the video recording environment is soft and uniform, and the host faces the camera directly.

[0067] Step S2: Use a server-side GPU (RTX 2080 or higher) to complete model training. During the inference phase, you can choose between Wenet or Hubert encoders depending on the device. For server-side deployment, Hubert encoders are recommended for better results.

[0068] Step S3: The inference module is integrated into the live streaming software to receive audio streams from the microphone or pre-recorded audio files and drive the digital human anchor in real time.

[0069] Step S4: After testing, the end-to-end latency (from audio input to digital human generation) is approximately 15-25 milliseconds, which meets the latency requirements for live streaming.

[0070] Implementation Method 2: A real-time inference system for short video data as described in this implementation method is based on a real-time inference method for short video data as described in Implementation Method 1, and includes the following modules: Module S1 collects short video data and performs preprocessing; Module S2 performs synchronous feature extraction on the preprocessed short video data and outputs 512-dimensional audio features; Module S3, constructs a lightweight U-Net model; Module S4 inputs 512-dimensional audio features into a lightweight U-Net model and performs real-time inference on the preprocessed short video data.

[0071] The above provides a detailed description of the real-time reasoning method and system for short video data proposed in this invention. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

Claims

1. A real-time inference method for short video data, characterized in that, Includes the following steps: Step S1: Collect short video data and perform preprocessing; Step S2: Extract synchronous features from the preprocessed short video data and output 512-dimensional audio features; Step S3: Construct a lightweight U-Net model; Step S4: Input the 512-dimensional audio features into the lightweight U-Net model, and perform real-time inference on the short video data based on the preprocessed short video data.

2. The real-time inference method for short video data according to claim 1, characterized in that, In step S1, the collection and preprocessing of short video data includes the following steps: Step S101: Capture short video data, which includes silent segments and audio segments; Step S102: Based on short video data, perform face detection and key feature extraction to obtain a standardized face image sequence; Step S103: Extract audio features based on the audio segments of the short video data; Step S104: Extract the silent reference frame based on the silent segment of the short video data.

3. The real-time inference method for short video data according to claim 2, characterized in that, In step S2, the synchronous feature extraction of the preprocessed short video data and the output of 512-dimensional audio features are specifically as follows: Based on the SyncNet model, 512-dimensional audio features and 512-dimensional normalized face image sequences are extracted frame by frame, and then joint feature extraction is performed to output 512-dimensional audio features.

4. The real-time inference method for short video data according to claim 3, characterized in that, In step S3, the lightweight U-Net model specifically refers to: The image input encoder outputs a 12×12 feature image. This feature image passes through a bottleneck layer to achieve cross-modal feature fusion before being input into the decoder, which outputs a 96×96 feature image.

5. The real-time inference method for short video data according to claim 4, characterized in that, The image input encoder outputs a 12×12 feature image, specifically: The image input is a 3×3 convolutional layer with a stride of 2, and the output is an image with a size of 48×48 and 64 channels. This image is then continuously input into two levels of inverse residual modules, and the output is a feature image with a size of 12×12. The step size of the inverted residual module is 2, and the expansion factor is 6.

6. The real-time inference method for short video data according to claim 4, characterized in that, The feature map, after passing through a bottleneck layer to achieve cross-modal feature fusion, is then input into the decoder, specifically as follows: The feature map is processed by the inverse residual module, and the output channel is a 512 feature map. After feature concatenation with the audio features, it is input into the inverse residual module again, and the output channel is a 512 feature fusion map.

7. The real-time inference method for short video data according to claim 4, characterized in that, The input to the decoder outputs a feature image with a size of 96×96, specifically: The 512-channel feature fusion image is input to the upsampling layer, then downsampled to 256 channels using a 3×3 convolutional layer. After being concatenated with the output of the inverse residual module in the encoder, it is input to the upsampling layer and downsampled to 128 channels using a 3×3 convolutional layer. This image is then concatenated with the features output from a 3×3 convolutional layer with a stride of 2 in the encoder, resulting in 192 channels. Finally, a 1×1 convolutional layer outputs a feature image with a size of 96×96.

8. A real-time inference method for short video data according to any one of claims 5-7, characterized in that, The aforementioned residual module is specifically as follows: The number of input channels is increased to a factor of 1×1 point convolution, and the output is after batch normalization and ReLU activation. The feature map after dimensionality increase is subjected to a 3×3 depthwise separable convolution, where the number of convolution groups is equal to the number of channels after dimensionality increase, and the convolution stride is set to 1 or 2 according to the module configuration. After batch normalization and ReLU activation, the output is processed. The feature map channels are projected to the output channels using 1×1 point convolution, and then output after batch normalization. When the number of input channels and the number of output channels of the inverse residual module are the same and the step size is 1, the output of the inverse residual module is the element-wise addition of the output of the dimensionality reduction projection operation and the input of the inverse residual module; otherwise, the output of the inverse residual module is the output of the dimensionality reduction projection operation.

9. A real-time inference system for short video data, wherein the system is implemented based on the real-time inference method for short video data as described in claim 1, characterized in that, Includes the following modules: Module S1 collects short video data and performs preprocessing; Module S2 performs synchronous feature extraction on the preprocessed short video data and outputs 512-dimensional audio features; Module S3, constructs a lightweight U-Net model; Module S4 inputs 512-dimensional audio features into a lightweight U-Net model and performs real-time inference on the preprocessed short video data.