An AIGC Video Creation Method for Cultural Relic Explanation

By combining interactive segmentation and tracking models, video object elimination models, 3D modeling technology, speech separation and extraction models and text generation speech models, the problems of insufficient segmentation accuracy of complex scenes, poor naturalness of enhancement and synthesis, speech segmentation accuracy and cloning realistic in cultural relics explanation video creation are solved, and high-quality AIGC video creation is achieved, improving the audience's immersive experience and the display effect of cultural relics.

CN119516059BActive Publication Date: 2025-07-01UNIV OF ELECTRONIC SCI & TECH OF CHINA CHENGDU COLLEGE

Patent Information

Application Number
CN202411622092.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2025-07-01
Estimated Expiration
2044-11-14

AI Technical Summary

Technical Problem

The existing technology has problems in the creation of cultural relics explanation videos with insufficient accuracy of complex scene segmentation, poor naturalness of enhancement and synthesis, speech segmentation accuracy and cloning realistic problems, which affect the quality of the video and the immersive experience of the audience.

Method used

Interactive segmentation and tracking model, video object elimination model, human body construction model, 3D modeling technology based on pure computer vision, speech separation and extraction model and text generation speech model are used to process video and audio, and synchronously and fusion to generate high-quality AIGC video.

Benefits of technology

It improves the accuracy and efficiency of video processing, enhances the naturalness and consistency of the video, improves the authenticity of voice and the immersive experience of the audience, and at the same time realizes the multi-angle display and interactivity of cultural relics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119516059B_ABST
    Figure CN119516059B_ABST
Patent Text Reader

Abstract

The present invention provides an AIGC video creation method for cultural relics explanation, which belongs to the field of AIGC video creation technology. The method includes: S1. virtual character video processing; S2. cultural relics video processing; S3. audio data processing; S4. fusion synthesis processing. The present invention can increase the high-precision video processing effect and audio authenticity processing effect, and at the same time, add 3D personnel to increase the fun. The 3D restoration of cultural relics can understand the cultural relics from multiple angles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of AIGC video creation, and more particularly to an AIGC (Artificial Intelligence Generated Content) video creation method for cultural relic explanation. Background Art

[0002] As an intuitive and vivid communication method, cultural relic explanation videos can showcase the charm and connotation of cultural relics to the general public. Currently, the existing cultural relic explanation video creation solutions have the following problems:

[0003] 1. In terms of video processing:

[0004] Accuracy of complex scene segmentation: In complex backgrounds, segmentation models such as SAM may not be able to accurately distinguish cultural relics from the background, resulting in unsatisfactory video generation effects.

[0005] Naturalness of enhancement and synthesis: When ProPainter performs image restoration and enhancement, problems such as over - restoration or inconsistency with the original style may occur, affecting the naturalness and consistency of the video.

[0006] 2. In terms of audio processing:

[0007] Accuracy of speech segmentation: Existing speech segmentation technologies may lack accuracy and efficiency when dealing with long - duration, multi - person audio files, affecting subsequent processing effects.

[0008] Realism of voice cloning: When voice cloning technology generates the voices of specific historical figures, problems such as unnatural pronunciation and insufficient emotional expression may occur, affecting the immersive experience of the audience.

[0009] 3. In terms of cultural relic reconstruction:

[0010] Reconstruction accuracy and efficiency: When the NeRF model reconstructs highly complex or texture - detailed cultural relics, it may require a large amount of computing resources and time, affecting efficiency and practicality.

[0011] Fidelity of lighting and material: During the cultural relic reconstruction process, the NeRF model may not be able to perfectly restore the true material and reflection characteristics of cultural relics under different lighting conditions, affecting the visual effect. Summary of the Invention

[0012] The present invention provides an AIGC video creation method for cultural relic explanation to solve one or more of the above problems.

[0013] An embodiment of this specification discloses an AIGC video creation method for cultural relic explanation, including:

[0014] S1. Virtual character video processing:

[0015] Obtain a live-action video, and use an interactive segmentation and tracking model, a video object elimination model, and a human body construction model to process the video to obtain a character video;

[0016] S2. Cultural relic video processing:

[0017] Obtain a real-shot video of cultural relics, and use 3D modeling technology based on pure computer vision to reconstruct the cultural relics to obtain a cultural relic video;

[0018] S3. Audio data processing:

[0019] Based on the live-action video, use the human voice track segmentation technology to remove the ambient sound of the video and separately extract the voice tracks of different characters, then use speech text recognition to identify the human voice text, and input the text content of the human voice text into the text-to-speech model to obtain audio data;

[0020] S4. Fusion and synthesis processing:

[0021] Synchronously fuse the character video, cultural relic video, and audio data to obtain a cultural relic explanation video based on AIGC.

[0022] In some embodiments of this specification, S1 includes:

[0023] S11. Image segmentation and target tracking:

[0024] Use the Segment and Track Anything model to perform image segmentation and target tracking on each frame of the live-action video, and generate the mask and image frame of the characters in the video;

[0025] S12. Video object elimination:

[0026] Input the generated image frame and mask into the ProPainter model to perform video object elimination and background completion, and generate a complete video frame after eliminating the characters;

[0027] S13. Generate character point cloud data and mesh:

[0028] Use the SMPLer-X model to process the image frame containing human body data, and generate the point cloud data and mesh file of each frame of the image to obtain 3D data;

[0029] Use the PYrender library of Python to render the 3D data generated by the SMPLer-X model, add textures to the rendered 3D data for off-screen rendering, and perform pixel superposition with the original video to obtain a virtual character rendered video, which is the character video.

[0030] In some embodiments of the present specification, in S2, a video of the explanation of cultural relics recorded on site and a video of cultural relics with surround camera angles are used as video sources, and the video source is a real-shot video of cultural relics. The real-shot video data of the cultural relics is separated, and after the image frames are extracted, they are input into the NeRF model for 3D reconstruction, and the 3D model of the cultural relics is derived to obtain a video of the cultural relics.

[0031] In some embodiments of this specification, S3 includes:

[0032] S31. Speech separation and extraction:

[0033] The MossFormer2 model is used to separate and extract the audio from the on-site recorded cultural relic explanation video to obtain the voice data of different characters;

[0034] S32. Speech to text:

[0035] Use speech-to-text technology to convert speech data into text data that can be recognized by the TTS model;

[0036] S33.Text to Speech Synthesis:

[0037] The recognized text data is input into the VITS2 model, the target timbre is selected, and high-quality speech data with different timbre characteristics is generated, that is, audio data is obtained.

[0038] In some embodiments of the present specification, the Segment and Track Anything model uses a fusion tracking algorithm mode, which includes an interactive tracking panel and an automatic tracking panel. The interactive tracking panel SAM is used for the first frame of the video, allowing the user to accurately mark the object of interest by clicking, drawing or text input; the automatic tracking panel is used to track new objects that appear in subsequent frames in the video;

[0039] The algorithm structure of the interactive tracking section is:

[0040] The user interacts with SAM by clicking the mouse to segment the target object in the keyframe;

[0041] SAM uses interactive prompts from the user to generate masks of objects;

[0042] Use Grounding-DINO for open-set object detection based on textual cues to enhance semantic understanding and assist SAM in generating more accurate masks;

[0043] Receive the SAM-generated mask as a reference frame annotation through Deep Automatic Object Tracking and initialize tracking through the Gated Propagation Module;

[0044] The algorithm structure of the automatic tracking section is:

[0045] Full paragraph segmentation: SAM generates masks for each object in the key reference frame, and Deep Automatic Object Tracking merges annotations to track new and old objects;

[0046] Object of Interest Segmentation: Grounding-DINO is used to detect new objects based on predefined textual cues, SAM generates masks, and Deep Automatic Object Tracking handles tracking;

[0047] Among them, the calculation method of the mask of the new object involves comparing the segmentation result of SAM and the current tracking result of Deep Automatic Object Tracking, and determining the new object based on the difference between the two. The formula is as follows:

[0048] N = T0*S;

[0049] N is the new object mask, T0 is the background of the Deep Automatic Object Tracking tracking result, and S is the SAM annotation result. Only where there is a difference between the SAM annotation and the Deep Automatic Object Tracking tracking result will it be marked in the new object mask.

[0050] The criterion for determining new objects is based on the ratio between each object's annotation result S in SAM and the new object mask N. When the ratio is greater than a preset minimum threshold t, the object corresponding to the ratio is defined as a new object. The formula is as follows:

[0051]

[0052] Among them, x s represents the size of the object in S, x n represents the size of the object in N, t represents the minimum threshold for defining a new object, and when x s and x n If the ratio is greater than t, the object will be defined as a new object.

[0053] In some embodiments of this specification, the algorithm flow of the ProPainter model is:

[0054] 1. Cyclic optical flow completion: ProPainter uses a cyclic network to complete the optical flow field of the masked area, which helps to perform pixel-level propagation in subsequent steps;

[0055] Use the RAFT algorithm to extract the forward and backward optical flow of the video;

[0056] The damaged optical flow field is completed through an efficient recurrent network to generate a complete optical flow map;

[0057] The network adopts deformable alignment and propagates optical flow information bidirectionally based on deformable convolution to complete the optical flow of the masked area.

[0058] 2. Dual-domain propagation: The model performs global and local propagation in both the image and feature domains, using the completed optical flow information to fill in the missing areas;

[0059] Image Propagation: Perform global propagation in the image domain, using optical flow-guided deformation operations to propagate pixels from unmasked areas to masked areas;

[0060] Feature Propagation: Perform local propagation in the feature domain and use the deformable alignment module to guide feature propagation based on the completed optical flow map to improve the robustness to occlusion and inaccurate optical flow completion;

[0061] 3. Mask-guided Sparse Video Transformer: Use a mask-guided sparse Transformer module to refine the propagated features and enhance the spatial and temporal coherence of the infill content through a self-attention mechanism;

[0062] Use the self-attention mechanism in the Transformer architecture to capture long-range dependencies in the video;

[0063] Introducing a mask-guided sparsity strategy to reduce computational complexity and memory consumption by selectively applying the attention mechanism;

[0064] Sparse Query Space: Attention is applied only to query windows that intersect the masked area, thus reducing unnecessary computations.

[0065] Sparse Key / Value Space: Reduces the size of the key / value space by selecting key / value frames at intervals of 2 time steps, thereby reducing computational and memory costs;

[0066] The refined features extracted by the self-attention mechanism are used in subsequent modules to generate the final restored video sequence.

[0067] In some embodiments of this specification, the framework of the SMPLer-X model:

[0068] 1. Backbone network:

[0069] Vision Transformer: As an image feature extractor, it can process large-scale data and extract rich visual features;

[0070] Embedding: The input image is segmented into patches of fixed size and converted into feature vectors through the embedding layer;

[0071] Positional Encoding: Add positional encoding to maintain spatial information in the image;

[0072] 2. Neck network:

[0073] BoxNet: predicts bounding boxes of hands and faces using features extracted from a backbone network;

[0074] Region of Interest Module: Crops the region of interest from the feature map based on the predicted bounding box;

[0075] 3. Head network:

[0076] Hand Head and Face Heads: For the ROIs of the hand and face, a deformable convolutional network is used for feature alignment, and then 3D key points and shape parameters are obtained through regression through a fully connected layer;

[0077] Body Head: The body parts use additional task tokens combined with the features of the backbone network, and the posture and shape parameters of the body are obtained through full connection layer regression;

[0078] The algorithm flow of the SMPLer-X model:

[0079] 1. Input video;

[0080] 2. Feature extraction: The input image is processed through the backbone network to extract deep visual features;

[0081] 3. Keypoint and bounding box prediction: Use the neck network to predict the bounding boxes of the hands and face, and crop the ROI from the features of the spine network;

[0082] 4. Feature alignment and parameter regression:

[0083] For the ROIs of hands and faces, a deformable convolutional network is used for feature alignment to adapt to different human postures;

[0084] The features after head network alignment are regressed to obtain 3D key points and shape parameters;

[0085] 5. Virtual human rendering and synthesis:

[0086] Rendering a 3D human body mesh using the SMPL model and predicted parameters;

[0087] The rendered 3D mesh is added with a texture for off-screen rendering and pixel-wise superimposed with the original video to obtain a virtual character rendering video.

[0088] In some embodiments of this specification, the NeRF model includes:

[0089] Scene representation: NeRF represents the scene as a continuous 5D function whose input is the coordinates (x, y, z) in 3D space and the 2D viewing direction (θ, φ), and the output is the RGB color and volume density (σ, c) at that location;

[0090] Neural Network: NeRF uses a multi-layer perceptron as a neural network to approximate 5D functions. The input is 5D coordinates (x, y, z, θ, φ), and the output is volume density and RGB color.

[0091] Volume rendering: NeRF samples 3D points along each camera ray and uses an MLP network to predict the color and density of the 3D points. It then accumulates the color and density of the 3D points using volume rendering techniques to generate a 2D image.

[0092] The formula for volume rendering is expressed as:

[0093]

[0094] Among them, r(t)=o+td is the ray starting from the camera origin o, with direction d and parameter t;

[0095] Optimization process: NeRF optimizes model parameters by minimizing the error between the observed image and the corresponding view rendered from NeRF. This process is differentiable and uses gradient descent method for optimization.

[0096] Positional encoding: To help MLP represent high-frequency functions, NeRF uses positional encoding to map the input coordinates to a higher-dimensional space as follows:

[0097] g(p)=(sin(2 0 πp),cos(2 0 πp),…,sin(2 L πp),cos(2 L πp));

[0098] The function g(p) is applied to the 3D coordinates (x, y, z) and the unit vector d of the viewing direction respectively;

[0099] Hierarchical Sampling: To improve efficiency, NeRF uses hierarchical sampling, first using a coarse network for sampling, and then performing more targeted sampling on the fine network based on the output of the coarse network.

[0100] In some embodiments of this specification, the MossFormer2 model includes:

[0101] 1. Encoder and Decoder

[0102] Encoder:

[0103] Input: Mixed speech waveform x∈R 1×T ;

[0104] Method: Use a one-dimensional convolutional layer (Cony1D) combined with a rectified linear unit (ReLU) to encode the input.

[0105] Output: Encoded sequence X∈R N×S , where N is the embedding dimension and S is the encoding sequence length. The encoding sequence length S is calculated as Where K1 is the convolution kernel size;

[0106] Decoder:

[0107] Input: masked encoded output;

[0108] Method: Use transposed 1D convolutional layers with the same kernel size and stride as the encoder;

[0109] Output: reconstructed original waveform;

[0110] 2. Mixing MossFormer and Cycle Modules

[0111] MossFormer Module:

[0112] Core: Apply self-attention on the entire sequence;

[0113] Strategy: Joint local-global self-attention strategy, performing full computational self-attention on non-overlapping local segments and using a linearized self-attention mechanism on the entire sequence;

[0114] Goal: Capture long-range, coarse-scale dependencies;

[0115] Loop Module:

[0116] Goal: Model complex temporal dependencies in speech signals and capture local cyclic patterns;

[0117] Method: Perform cyclic learning for each embedding dimension;

[0118] 3. No RNN loop module

[0119] Module composition:

[0120] Bottleneck layer: Use a 1×11×1 convolutional layer to reduce the embedding dimension;

[0121] GCU layer: built on the gated convolutional unit (GLU) and the expanded FSMN;

[0122] Output layer: LayerNorm followed by a 1×11×1 convolutional layer;

[0123] Extended FSMN blocks:

[0124] Structure: It includes a feed-forward layer and a storage layer. The storage layer uses stacked two-dimensional dilated convolution blocks.

[0125] Goal: Cover a wider receptive field and enhance information flow and gradient propagation;

[0126] Conv-U Block:

[0127] Structure: Contains LayerNorm layer, linear layer, SiLU activation and one-dimensional deep convolution layer;

[0128] Goal: Assist the GCU layer to capture the local pattern of position;

[0129] 4. Masking network: maps the encoded output to a set of masks that are used to separate different speech sources;

[0130] The workflow of the MossFormer2 model:

[0131] Input mixed speech waveform;

[0132] Encoding: Convert the waveform to a coded sequence through Conv1 D and ReLU;

[0133] MossFormer module: applies a joint local-global self-attention strategy to capture long-range dependencies;

[0134] Loop module: uses GCU layers and dilated FSMN modules to capture complex temporal dependencies and local loop patterns;

[0135] Masking network: maps the encoded output to a set of masks;

[0136] Decoding: Reconstruct the original waveform using a transposed 1D convolutional layer.

[0137] In some embodiments of this specification, the VITS2 model includes:

[0138] 1. Random duration predictor:

[0139] Input: Hidden representation of text h textand Gaussian noise z d ;

[0140] Generator G: Use h text and z d As input, generate the predicted duration

[0141] Discriminator D: The discriminator receives h text and the logarithmic duration d or predicted duration obtained from Monotonic Alignment Search As input;

[0142] Training strategy: Using adversarial learning, the discriminator classifies according to the actual and predicted durations;

[0143] The loss function is as follows:

[0144] Adversarial loss function:

[0145]

[0146] Mean square error loss function:

[0147] L mse =MSE(G(z d ,h text ),d);

[0148] 2. Monotonic alignment search: Find the alignment with the highest probability between text and audio among all possible monotonic alignments, and add Gaussian noise during initial training to increase the diversity of alignment search;

[0149] algorithm:

[0150]

[0151] where ∈ is the product of the noise sampled from the standard normal distribution and the standard deviation of P;

[0152] 3. Normalization flow: Use convolutional blocks to capture adjacent data and add a small Transformer block with residual connections to capture long-range dependencies;

[0153] 4. Speaker-Conditioned Text Encoder: Generates speech with different characteristics based on speaker conditions in a multi-speaker model;

[0154] Method: Introduce speaker vector in the third Transformer block of the text encoder to learn and express the unique speech characteristics of each speaker;

[0155] Workflow diagram of the VITS2 model:

[0156] Input text text;

[0157] Text encoder: Encodes the input text into a hidden representation h text ;

[0158] Duration prediction: Generate predicted duration through random duration predictor

[0159] MAS alignment: Use MAS to align text and audio, and add Gaussian noise to enhance the diversity of alignment;

[0160] Normalization flow: Use a combination of convolutional blocks and Transformer blocks to normalize the data and capture long-range and local dependencies;

[0161] Speaker conditioning: adding speaker vectors to the text encoder to enhance the generation capabilities of multi-speaker models;

[0162] Adversarial Learning: Use the generator G and the discriminator D for adversarial training to improve the naturalness of duration prediction and speech synthesis.

[0163] The embodiments of this specification can at least achieve the following beneficial effects:

[0164] Accuracy and efficiency: Traditional video processing technologies are inaccurate or inefficient when it comes to object recognition, segmentation, and tracking. STA can automatically identify dynamic objects in videos and perform precise tracking and segmentation.

[0165] Reduce costs: Manual video editing and processing usually requires a lot of manpower and time investment, but the use of new technology ProPainter can reduce the need for manual operations and reduce the cost of editing and processing.

[0166] Increased creativity: Traditional videos have limited creativity, but by adding ProPainter technology and SMPler-X for three-dimensional creation, more realistic dynamic expression and analysis can be provided.

[0167] Flexible perspective generation: Traditional 3D modeling technology may be limited in perspective generation, while NeRF technology can generate images from nearly infinite perspectives, making it possible to display artifacts in all directions. Users can observe the model from any angle, increasing interactivity; it also overcomes occlusion problems.

[0168] The present invention can increase the high-precision video processing effect and the audio authenticity processing effect. At the same time, adding 3D personnel can increase the fun. The 3D restoration of cultural relics can understand the cultural relics from multiple angles. BRIEF DESCRIPTION OF THE DRAWINGS

[0169] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0170] Figure 1 It is a schematic diagram of the AIGC video creation method for cultural relics explanation involved in some embodiments of the present invention.

[0171] Figure 2 It is a flowchart of the AIGC video creation method for cultural relics explanation involved in some embodiments of the present invention.

[0172] Figure 3 It is a schematic diagram of the SMPLer-X model framework involved in some embodiments of the present invention.

[0173] Figure 4 It is a flowchart of the MossFormer2 model involved in some embodiments of the present invention.

[0174] Figure 5 It is a flowchart of the VITS2 model involved in some embodiments of the present invention. Detailed implementation manners

[0175] In the following text, only some exemplary embodiments are briefly described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the spirit or scope of the embodiments of the present invention. Therefore, the drawings and the description are considered to be exemplary in nature rather than restrictive.

[0176] The following disclosure provides many different implementation manners or examples for implementing different structures of the embodiments of the present invention. To simplify the disclosure of the embodiments of the present invention, the components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the embodiments of the present invention. In addition, the embodiments of the present invention may repeat reference numerals and / or reference letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various implementation manners and / or settings discussed.

[0177] The following will detail the embodiments of the present invention with reference to the drawings.

[0178] As Figure 1 and Figure 2 shown, the embodiments of this specification disclose an AIGC video creation method for cultural relics explanation, including:

[0179] S1. Virtual character video processing:

[0180] Obtain real-person shot videos, and use the Segment and Track Anything model, the ProPainter model, and the SMPLer-X model to process the videos to obtain person videos;

[0181] S2. Cultural relic video processing:

[0182] Obtain real-shot videos of cultural relics, and use the 3D modeling technology (NeRF) based on pure computer vision to reconstruct the cultural relics to obtain cultural relic videos;

[0183] S3. Audio data processing:

[0184] Based on the real-person shot video, use the MossFormer2 technology to remove the ambient sound of the video and separate and extract the audio tracks of different people, then use speech text recognition to identify the voice text, and input the text content of the voice text into the VITS2 model to obtain audio data of the intended tone color audio (such as the voice of Lazy Yangyang);

[0185] S4. Fusion and synthesis processing:

[0186] Synchronously fuse the person video, the cultural relic video, and the audio data to obtain a cultural relic explanation video based on AIGC.

[0187] Through the above steps, the present invention can generate high-quality and natural virtual person videos, cultural relic reconstruction videos, and sound videos, and realize video creation in multiple scenarios.

[0188] In some embodiments of this specification, S1 includes:

[0189] S11. Image segmentation and object tracking:

[0190] Use the Segment and Track Anything model to perform image segmentation and object tracking on each frame of the real-person shot video, and generate the mask and image frame of the person in the video;

[0191] S12. Video object elimination:

[0192] Input the generated image frame and mask into the ProPainter model to perform video object elimination and background completion, and generate a complete video frame after eliminating the person;

[0193] S13. Generate person point cloud data and mesh:

[0194] Process the image frames containing human body data using the SMPLer-X model to generate point cloud data and mesh files for each frame of the image, obtaining 3D data;

[0195] S14. Use the PYrender library in Python to render the 3D data generated by the SMPLer-X model, add textures to the rendered 3D data for off-screen rendering, and perform pixel superposition with the original video to obtain a virtual character rendered video, which is the said character video.

[0196] In some embodiments of this specification, in S2, use the on-site recorded cultural relics explanation video and the cultural relics video with a surrounding camera view as the video source. The said video source is the actual video of the cultural relics. After separating the actual video data of the cultural relics and extracting the image frames, input them into the NeRF model for 3D reconstruction, and export to generate the 3D model of the cultural relics, thus obtaining the cultural relics video.

[0197] In some embodiments of this specification, S3 includes:

[0198] S31. Voice separation and extraction:

[0199] Use the MossFormer2 model to separate and extract the audio in the on-site recorded cultural relics explanation video to obtain the voice data of different people;

[0200] S32. Voice to text conversion:

[0201] Use voice-to-text technology to convert the voice data into text data that can be recognized by the TTS (text-to-speech) model;

[0202] S33. Text to speech synthesis:

[0203] Input the recognized text data into the VITS2 model, select the target voice (such as the voice of Lazy Goat), and generate high-quality voice data with different voice characteristics, thus obtaining the audio data.

[0204] In some embodiments of this specification, the Segment and Track Anything model uses a fusion tracking algorithm mode, which includes an interactive tracking section SAM (Segment Anything Model) and an automatic tracking section (Interactive tracking mode). The interactive tracking section is used for the first frame of the video, allowing users to accurately mark the objects of interest through clicking, drawing, or text input; the automatic tracking section is used to track new objects appearing in subsequent frames of the video;

[0205] The algorithm structure of the interactive tracking section is:

[0206] The user interacts with SAM by clicking the mouse (such as clicking on the body of the person to be segmented) to segment the target object in the key frame;

[0207] SAM uses the user's interactive prompts (such as clicking on the object contour) to generate a mask for the object;

[0208] Utilize Grounding-DINO for open-set object detection based on text prompts to enhance semantic understanding and assist SAM in generating more accurate masks;

[0209] Receive the mask generated by SAM as a reference frame annotation through Deep Automatic Object Tracking and initialize the tracking through the Gated Propagation Module;

[0210] The algorithm structure of the automatic tracking section is as follows:

[0211] Full Paragraph Segmentation: SAM generates masks for each object in the key reference frame, and Deep Automatic Object Tracking combines the annotations to track new and old objects;

[0212] Object of Interest Segmentation: Use Grounding-DINO to detect new objects based on predefined text prompts, SAM generates masks, and Deep Automatic Object Tracking processes the tracking;

[0213] Among them, the calculation method of the mask of the new object involves comparing the segmentation result of SAM and the current tracking result of Deep Automatic Object Tracking, and determining the new object based on the difference between the two. The formula is as follows:

[0214] N = T0 * S;

[0215] N represents the new object mask, T0 is the background of the tracking result of Deep Automatic Object Tracking, and S is the annotation result of SAM; only where there is a difference between the annotation of SAM and the tracking result of Deep Automatic Object Tracking will it be marked in the new object mask;

[0216] The criterion for determining a new object is based on the ratio between the annotation result S of each object in SAM and the new object mask N. When this ratio is greater than a preset minimum threshold t, the object corresponding to this ratio is defined as a new object. The formula is expressed as follows:

[0217]

[0218] Among them, CMR(x) is a discriminant function for determining a new object, where x s represents the object size in S, and x n represents the object size in N, and t represents the minimum threshold for defining a new object. When the ratio of x s and x n is greater than t, then the object will be defined as a new object.

[0219] In some embodiments of this specification, the algorithm process of the ProPainter model is as follows:

[0220] 1. Recurrent Flow Completion (RFC): ProPainter uses a recurrent network to complete the optical flow field in the masked area, which helps to perform pixel-level propagation in subsequent steps;

[0221] Use the RAFT algorithm to extract the forward and backward optical flows of the video;

[0222] Complete the damaged optical flow field through an efficient recurrent network to generate a complete optical flow map;

[0223] This network adopts Deformable Alignment and is based on deformable convolutional networks (DCN) to bidirectionally propagate optical flow information, thereby completing the optical flow in the masked area;

[0224] 2. Dual-Domain Propagation (DDP): The model performs global and local propagation simultaneously in the image and feature domains, and uses the completed optical flow information to fill in the missing areas;

[0225] Image Propagation: Perform global propagation in the image domain, and use the optical flow-guided deformation operation to propagate the pixels in the unmasked area to the masked area;

[0226] Feature Propagation: Perform local propagation in the feature domain, and use the deformable alignment module to guide feature propagation according to the completed optical flow map, improving the robustness to occlusion and inaccurate optical flow completion;

[0227] Through this dual-domain propagation, the model can effectively utilize global and local information to improve the visual consistency and coherence of the filled content;

[0228] 3. Mask-Guided Sparse Video Transformer (MSVT): It uses a mask-guided sparse Transformer module to refine the propagated features and enhances the spatial and temporal coherence of the filled content through the self-attention mechanism;

[0229] It uses the self-attention mechanism in the Transformer architecture to capture long-range dependencies in the video;

[0230] It introduces a mask-guided sparse strategy to reduce computational complexity and memory consumption by selectively applying the attention mechanism;

[0231] Sparse Query Space: It applies the attention mechanism only to the query windows that intersect with the masked regions, thus reducing unnecessary computations;

[0232] Sparse Key / Value Space: It selects key / value frames at an interval of 2 time steps to reduce the size of the key / value space, thereby reducing computational and memory costs;

[0233] The refined features extracted through the self-attention mechanism are used in subsequent modules to generate the final inpainted video sequence.

[0234] In some embodiments of this specification, as Figure 3 shown, the framework of the SMPLer-X model:

[0235] 1. Backbone Network:

[0236] Vision Transformer (ViT): As an image feature extractor, it can process large-scale data and extract rich visual features;

[0237] Embedding: The input image is divided into patches of a fixed size and converted into feature vectors through an embedding layer;

[0238] Positional Encoding: Positional encoding is added to preserve the spatial information in the image;

[0239] 2. Neck Network:

[0240] BoxNet: It uses the features extracted by the backbone network to predict the bounding boxes of the hands and faces;

[0241] Region of Interest (ROI) Module: It crops the regions of interest from the feature map according to the predicted bounding boxes;

[0242] 3. Head Networks:

[0243] Hand Head and Face Heads: For the ROIs of the hand and face, the Deformable Convolutional Networks (DCN) are used for feature alignment, and then the 3D key points and shape parameters are obtained by regression through the fully connected layer;

[0244] Body Head: For the body part, the features of the backbone network are combined with additional task tokens, and the body pose and shape parameters are obtained by regression through the fully connected layer;

[0245] The algorithm process of the SMPLer-X model:

[0246] 1. Input view;

[0247] 2. Feature extraction: The input image is processed by the backbone network to extract deep visual features;

[0248] 3. Key point and bounding box prediction: The neck network is used to predict the bounding boxes of the hand and face, and the ROIs are cropped from the features of the backbone network;

[0249] 4. Feature alignment and parameter regression:

[0250] For the ROIs of the hand and face, the Deformable Convolutional Networks are used for feature alignment to adapt to different human postures;

[0251] Regression is performed on the features aligned by the head network to obtain the 3D key points and shape parameters;

[0252] 5. Virtual human rendering and synthesis:

[0253] The SMPL model and the predicted parameters are used to render the 3D human mesh;

[0254] The rendered 3D mesh is added with textures for off-screen rendering, and pixel superposition is performed with the original video to obtain the virtual human rendering video.

[0255] In some embodiments of this specification, NeRF (Neural Radiance Fields) is a deep learning model for 3D scene representation and rendering. It can learn the continuous 3D representation of the scene from a set of sparse 2D images and can render new views.

[0256] The NeRF model includes:

[0257] Scene Representation: NeRF represents a scene as a continuous 5D function. The input of this function is the coordinates (x, y, z) in 3D space and the 2D viewing direction (θ, φ), and the output is the RGB color and volume density (σ, c) at that position; θ is the azimuth angle, φ is the elevation angle, c is the RGB color, and σ is the volume density.

[0258] Neural Network: NeRF uses a multi-layer perceptron (MLP) as a neural network to approximate the 5D function. The input is the 5D coordinates (x, y, z, θ, φ), and the output is the volume density and RGB color.

[0259] Volume Rendering: NeRF samples 3D points along each camera ray and uses the MLP network to predict the color and density of the 3D points. The volume rendering technique is used to accumulate the color and density of the 3D points to generate a 2D image.

[0260] The formula for volume rendering is expressed as:

[0261]

[0262] Where: C(r) represents the accumulated color value along ray r, σ(r(t)) is the volume density function, representing the density at a certain point on ray r. c(r(t), d) is the color function, representing the color at a certain point on ray r. r(t) = O + t·d is the ray starting from the camera origin O and along direction d, and t is the parameter.

[0263] Optimization Process: NeRF optimizes the model parameters by minimizing the error between the observed image and the corresponding view rendered from NeRF. This process is differentiable and uses the gradient descent method for optimization.

[0264] Position Encoding: To help the MLP represent high-frequency functions, NeRF uses position encoding to map the input coordinates to a higher-dimensional space as follows:

[0265] g(p) = (sin(2 0 πp), cos(2 0 πp), …, sin(2 L-1 πp), cos(2 L-1 πp));

[0266] The function g(p) is applied to the 3D coordinates (x, y, z) and the unit vector d of the viewing direction respectively; 2 0 is the sine and cosine terms with the lowest frequency in the position encoding, p is the component of the input coordinate or direction vector, 2 L-1 is the sine and cosine terms with the highest frequency in the position encoding, and L is the number of frequencies of the position encoding.

[0267] Stratified Sampling: To improve efficiency, NeRF uses stratified sampling. First, it samples using a coarse network, and then, based on the output of the coarse network, it performs more targeted sampling on the fine network.

[0268] In some embodiments of this specification, MossFormer2 is a deep learning model for speech separation and extraction, and its main algorithm principle can be divided into several key steps: encoding, hybrid self-attention and recurrent modules, masking, and decoding.

[0269] The MossFormer2 model includes:

[0270] 1. Encoder and Decoder

[0271] Encoder:

[0272] Input: Mixed speech waveform x ∈ R 1×T ;

[0273] Method: Use a one-dimensional convolutional layer (Cony1D) combined with a rectified linear unit (ReLU) to encode the input;

[0274] Output: Encoded sequence X ∈ R N×S , where N is the embedding dimension, S is the length of the encoded sequence, and the length of the encoded sequence S is calculated as where K1 is the kernel size; R is the number of repetitions, and T is the length of the input sequence.

[0275] Decoder:

[0276] Input: Masked encoded output;

[0277] Method: Use a transposed one-dimensional convolutional layer with the same kernel size and stride as the encoder;

[0278] Output: Reconstructed original waveform;

[0279] 2. Hybrid MossFormer and Recurrent Modules

[0280] MossFormer Module:

[0281] Core: Apply self-attention over the entire sequence;

[0282] Strategy: Combine local-global self-attention strategy, perform full computational self-attention on non-overlapping local segments, and adopt a linearized self-attention mechanism over the entire sequence;

[0283] Objective: Capture long-range, coarse-scale dependencies;

[0284] Recurrent Module:

[0285] Objective: Simulate complex temporal dependencies in speech signals and capture local recurrent patterns;

[0286] Method: Perform recurrent learning on each embedding dimension;

[0287] 3. Without RNN recurrent module

[0288] Module composition:

[0289] Bottleneck layer: Use a 1×11×1 convolutional layer to reduce the embedding dimension;

[0290] GCU layer: Constructed based on the gated convolutional unit (GLU) and dilated FSMN;

[0291] Output layer: Followed by LayerNorm and then a 1×11×1 convolutional layer;

[0292] Extended FSMN block:

[0293] Structure: Includes a feed-forward layer (FFN) and a storage layer, and the storage layer uses stacked two-dimensional dilated convolutional blocks;

[0294] Objective: Cover a wider receptive field, enhance information flow and gradient propagation;

[0295] Conv-U block:

[0296] Structure: Contains a LayerNorm layer, a linear layer, SiLU activation, and a one-dimensional depth convolutional layer;

[0297] Objective: Assist the GCU layer in capturing local positional patterns;

[0298] 4. Masking network: Map the encoded output to a set of masks, which are used to separate different speech sources;

[0299] As Figure 4 shown, the working process of the MossFormer2 model:

[0300] Input the mixed speech waveform x;

[0301] Encoding: Convert the waveform into an encoded sequence X through Conv1D and ReLU;

[0302] MossFormer module: Apply the joint local-global self-attention strategy to capture long-range dependencies;

[0303] Recurrent module: Use the GCU layer and the dilated FSMN module to capture complex temporal dependencies and local recurrent patterns;

[0304] Masking network: Map the encoded output to a set of masks;

[0305] Decoding: Use a transposed 1D convolutional layer to reconstruct the original waveform.

[0306] Through these steps, the MossFormer2 model can effectively separate and extract different speech sources, especially suitable for processing mixed speech containing multiple human voices.

[0307] In some embodiments of this specification, VITS2 (Variational Inference Text-to-Speech) is an improved text-to-speech synthesis model that combines variational inference, normalizing flow, and self-attention mechanisms to improve the quality and efficiency of speech generation.

[0308] The VITS2 model includes:

[0309] 1. Random duration predictor:

[0310] Input: The hidden representation h of the text text and Gaussian noise z d ;

[0311] Generator G: Use h text and z d as inputs to generate the predicted duration

[0312] Discriminator D: The discriminator receives h text and the logarithmic duration d obtained from Monotonic Alignment Search (MAS) or the predicted duration as inputs;

[0313] Training strategy: Use adversarial learning, and the discriminator classifies according to the actual and predicted durations;

[0314] Loss function:

[0315] Adversarial loss function (least-squares loss):

[0316]

[0317] L adv (D) is the adversarial loss of the discriminator, and L adv (G) is the adversarial loss of the generator.

[0318] Mean squared error loss function (MSE loss):

[0319] L mse = MSE(G(z d , h text ), d);

[0320] 2. Monotonic Alignment Search: Find the alignment with the highest probability between text and audio among all possible monotonic alignments, and add Gaussian noise during initial training to increase the diversity of alignment search;

[0321] Algorithm:

[0322] where ∈ is the product of the noise sampled from the standard normal distribution and the standard deviation of P; P i,j is the Gaussian probability logarithm value calculated between positions i and j. is the normal distribution (also known as the Gaussian distribution), z j is the latent variable transformed from the normal flow. μ i is the mean of the normal distribution, σ i is the standard deviation of the normal distribution, Q i,j is the log maximum likelihood value of the monotonic alignment search. Q i-1,j-1 and Q i,j-1 are the previously calculated alignment probability values.

[0323] 3. Normalizing Flow: Use convolutional blocks to capture adjacent data, and add a small Transformer block with residual connections to capture long-range dependencies; Effect: The Transformer block can collect information from different positions when transforming the distribution.

[0324] 4. Speaker-Conditioned Text Encoder: In a multi-speaker model, generate speech with different features according to the speaker condition;

[0325] Method: Introduce the speaker vector into the third Transformer block of the text encoder to learn and express the unique speech features of each speaker;

[0326] As Figure 5 shown, the workflow diagram of the VITS2 model:

[0327] Input text text;

[0328] Text Encoder: Encode the input text into a hidden representation h text ;

[0329] Duration Prediction: Generate the predicted duration through a random duration predictor

[0330] MAS Alignment: Use MAS for the alignment of text and audio, and add Gaussian noise to enhance the diversity of alignment;

[0331] Normalizing Flow: Use a combination of convolutional blocks and Transformer blocks to normalize the data and capture long-range and local dependencies;

[0332] Speaker condition: Add a speaker vector to the text encoder to enhance the generation ability of the multi-speaker model;

[0333] Adversarial learning: Use the generator G and discriminator D for adversarial training to improve the naturalness of duration prediction and speech synthesis.

[0334] Through these steps, the VITS2 model can efficiently generate natural and high-quality speech in a multi-speaker environment.

[0335] The above-described embodiments are used to illustrate the present invention, not to limit the present invention. Therefore, changes in the example values or replacement of equivalent elements should still fall within the scope of the present invention.

[0336] From the above detailed description, those of ordinary skill in the art can clearly understand that the present invention can indeed achieve the aforementioned objectives and has actually met the requirements of the patent law.

[0337] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention. The above description is only the preferred embodiments of the present invention and is not intended to limit the present invention. It should be noted that any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

[0338] It should be noted that the above description of the process is only for illustration and explanation, and does not limit the scope of application of this specification. For those skilled in the art, various corrections and changes can be made to the process under the guidance of this specification. However, these corrections and changes are still within the scope of this specification.

[0339] The basic concept has been described above. Obviously, for those of ordinary skill in the art after reading this application, the above invention disclosure is only an example and does not constitute a limitation to this application. Although not explicitly stated here, those of ordinary skill in the art may make various modifications, improvements, and corrections to this application. Such modifications, improvements, and corrections are proposed in this application, so such modifications, improvements, and corrections still belong to the spirit and scope of the exemplary embodiments of this application.

[0340] Meanwhile, this application uses specific terms to describe the embodiments of this application. For example, "one embodiment", "an embodiment", and / or "some embodiments" mean a certain feature, structure, or characteristic related to at least one embodiment of this application. Therefore, it should be emphasized and noted that "an embodiment" or "one embodiment" or "an alternative embodiment" mentioned twice or more at different positions in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of this application can be appropriately combined.

[0341] In addition, those of ordinary skill in the art can understand that various aspects of this application can be described and illustrated by several patentable types or situations, including any new and useful process, machine, product, or composition of matter, or any new and useful improvement thereof. Therefore, various aspects of this application can be implemented entirely by hardware, can be implemented entirely by software (including firmware, resident software, microcode, etc.), or can be implemented by a combination of hardware and software. The above hardware or software can be referred to as "unit", "module", or "system". In addition, various aspects of this application can take the form of a computer program product embodied in one or more computer-readable media, in which computer-readable program code is included.

[0342] The computer program code required for the operation of each part of this application can be written in any one or more programming languages, including object-oriented programming languages such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB.NET, Python, etc., conventional procedural programming languages such as C programming language, VisualBasic, Fortran2103, Perl, COBOL2102, PHP, ABAP, dynamic programming languages such as Python, Ruby, and Groovy, or other programming languages. The program code can run entirely on the user's computer, or run as an independent software package on the user's computer, or run partially on the user's computer and partially on a remote computer, or run entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any network form, such as a local area network (LAN) or a wide area network (WAN), or connected to an external computer (for example, through the Internet), or in a cloud computing environment, or used as a service such as software as a service (SaaS).

[0343] In addition, unless otherwise specified in the claims, the order of the processing elements and sequences, the use of numerical letters, or the use of other names described in this application are not used to limit the order of the processes and methods of this application. Although some currently useful embodiments of the invention are discussed through various examples in the above disclosure, it should be understood that such details are only for illustrative purposes, and the appended claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that conform to the essence and scope of the embodiments of this application. For example, although the implementation of the various components described above can be embodied in a hardware device, it can also be implemented as a pure software solution, for example, by installing it on an existing server or mobile device.

[0344] Similarly, it should be noted that, in order to simplify the description of the disclosure of this application and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of this application, sometimes multiple features are grouped into one embodiment, drawing, or description thereof. However, this method of this application should not be construed as reflecting an intention that the claimed subject matter requires more features than are expressly recited in each claim. On the contrary, the subject matter of the invention should have fewer features than the above single embodiment.

Claims

1. A method for creating AIGC videos of cultural relics explanations, characterized in that: include: S1. Virtual character video processing: Obtain real-life video, and use the interactive segmentation and tracking model, the video object elimination model, and the human body construction model to process the video to obtain a person video; S2. Cultural Relics Video Processing: Obtain real-life video of cultural relics, and reconstruct the cultural relics using 3D modeling technology based on pure computer vision to obtain a video of the cultural relics; S3. Audio data processing: Based on real-life video, the human voice track segmentation technology is used to remove the ambient sound of the video and separate the audio tracks of different characters. Then, the human voice text is recognized using voice text, and the text content of the human voice text is passed into the text generation voice model to obtain audio data. S4. Fusion synthesis processing: Synchronously fusing the character video, cultural relic video and audio data to obtain a cultural relic explanation video based on AIGC; S1 includes: S11. Image segmentation and target tracking: Use the Segment and Track Anything model to perform image segmentation and target tracking on each frame of the real-life video to generate masks and image frames of the characters in the video; S12. Video object elimination: The generated image frame and mask are input into the ProPainter model to remove video objects and complete the background, thus generating a complete video frame after removing the person. S13. Generate character point cloud data and mesh: Use the SMPLer-X model to process the image frames containing human body data, generate point cloud data and mesh files for each frame of image, and obtain 3D data; S14. Use Python's PYrender library to render the 3D data generated by the SMPLer-X model, add a texture to the rendered 3D data for off-screen rendering, and overlay the pixels with the original video to obtain a virtual character rendering video, which is the character video; In S2, the on-site recorded cultural relic explanation video and the cultural relic video with surround camera angles are used as video sources. The video source is the real-shot video of the cultural relic. The real-shot video data of the cultural relic is separated, and after the image frames are extracted, they are input into the NeRF model for 3D reconstruction, and the 3D model of the cultural relic is derived to obtain the cultural relic video.

2. The AIGC video creation method for cultural relics explanation according to claim 1 is characterized in that: S3 includes: S31. Speech separation and extraction: The MossFormer2 model is used to separate and extract the audio from the on-site recorded cultural relic explanation video to obtain the voice data of different characters; S32. Speech to text: Use speech-to-text technology to convert speech data into text data that can be recognized by the TTS model; S33.Text to Speech Synthesis: The recognized text data is input into the VITS2 model, the target timbre is selected, and high-quality speech data with different timbre characteristics is generated, that is, audio data is obtained.

3. The AIGC video creation method for cultural relics explanation according to claim 1 is characterized in that: The Segment and Track Anything model uses a fusion tracking algorithm mode, which includes an interactive tracking section and an automatic tracking section. The interactive tracking section is used for the first frame of the video, allowing the user to accurately mark the object of interest by clicking, drawing or text input; the automatic tracking section is used to track new objects that appear in subsequent frames in the video; The algorithm structure of the interactive tracking section is: The user interacts with the interactive tracking panel by clicking the mouse to segment the target object in the keyframe; The interactive tracking section uses the user's interactive prompts to generate masks of objects; Use Grounding-DINO for open-set object detection based on textual cues to enhance semantic understanding and assist the interactive tracking module in generating more accurate masks; Receive the mask generated by the interactive tracking block as a reference frame annotation through Deep Automatic Object Tracking and initialize tracking through the Gated Propagation Module; The algorithm structure of the automatic tracking section is: Full-paragraph segmentation: The interactive tracking module generates masks for each object in the key reference frame, and Deep Automatic Object Tracking merges annotations to track new and old objects; Object of Interest Segmentation: Grounding-DINO is used to detect new objects based on predefined textual cues, interactive tracking panels generate masks, and Deep Automatic Object Tracking handles tracking; The calculation method of the mask of the new object involves comparing the segmentation result of the interactive tracking plate and the current tracking result of DeepAutomatic Object Tracking, and determining the new object based on the difference between the two. The formula is as follows: N = T0*S; N is the new object mask, T0 is the background of the Deep Automatic Object Tracking tracking result, and S is the interactive tracking plate annotation result. Only where there is a difference between the annotation of the interactive tracking plate and the tracking result of Deep Automatic Object Tracking will it be marked in the new object mask. The criterion for determining new objects is based on the ratio between the annotation result S of each object in the interactive tracking plate and the new object mask N. When the ratio is greater than a preset minimum threshold t, the object corresponding to the ratio is defined as a new object. The formula is as follows: Among them, x s represents the size of the object in S, x n represents the size of the object in N, t represents the minimum threshold for defining a new object, when x s and x n If the ratio of is greater than t, the object will be defined as a new object.

4. The AIGC video creation method for cultural relics explanation according to claim 1 is characterized in that: The algorithm flow of the ProPainter model is: Circular optical flow completion: ProPainter uses a cyclic network to complete the optical flow field of the masked area, which helps with pixel-level propagation in subsequent steps; Use the RAFT algorithm to extract the forward and backward optical flow of the video; The damaged optical flow field is completed through an efficient recurrent network to generate a complete optical flow map; The network adopts deformable alignment and propagates optical flow information bidirectionally based on deformable convolution to complete the optical flow of the masked area. Dual-domain propagation: The model performs global and local propagation in both the image and feature domains, using the completed optical flow information to fill in the missing areas; Image Propagation: Perform global propagation in the image domain, using optical flow-guided deformation operations to propagate pixels from unmasked areas to masked areas; Feature Propagation: Perform local propagation in the feature domain and use the deformable alignment module to guide feature propagation based on the completed optical flow map to improve the robustness to occlusion and inaccurate optical flow completion; Mask-guided Sparse Video Transformer: Uses a mask-guided sparse Transformer module to refine propagated features and enforces spatial and temporal coherence of infill content through a self-attention mechanism; Use the self-attention mechanism in the Transformer architecture to capture long-range dependencies in videos; introduce a mask-guided sparsity strategy to reduce computational complexity and memory consumption by selectively applying the attention mechanism; Sparse Query Space: Attention is applied only to query windows that intersect the masked area, thus reducing unnecessary computations. Sparse Key / Value Space: Reduces the size of the key / value space by selecting key / value frames at intervals of 2 time steps, thereby reducing computational and memory costs; The refined features extracted by the self-attention mechanism are used in subsequent modules to generate the final restored video sequence.

5. The AIGC video creation method for cultural relics explanation according to claim 1 is characterized in that: The framework of the SMPLer-X model: Backbone Network: Vision Transformer: As an image feature extractor, it can process large-scale data and extract rich visual features; Embedding: The input image is segmented into patches of fixed size and converted into feature vectors through the embedding layer; Positional Encoding: Add positional encoding to maintain spatial information in the image; Neck Network: BoxNet: predicts bounding boxes of hands and faces using features extracted from a backbone network; Region of Interest Module: Crops the region of interest from the feature map based on the predicted bounding box; Head network: Hand Head and Face Heads: For the ROIs of the hand and face, a deformable convolutional network is used for feature alignment, and then 3D key points and shape parameters are obtained through regression through a fully connected layer; Body Head: The body parts use additional task tokens combined with the features of the backbone network, and the posture and shape parameters of the body are obtained through full connection layer regression; The algorithm flow of the SMPLer-X model is as follows: Input view; Feature extraction: The input image is processed through the backbone network to extract deep visual features; Keypoint and bounding box prediction: Use the neck network to predict the bounding boxes of the hands and face, and crop the ROIs from the features of the spine network; Feature Alignment and Parameter Regression: For the ROIs of hands and faces, a deformable convolutional network is used for feature alignment to adapt to different human postures; The features after head network alignment are regressed to obtain 3D key points and shape parameters; Virtual Human Rendering and Synthesis: Rendering a 3D human body mesh using the SMPL model and predicted parameters; The rendered 3D mesh is added with a texture for off-screen rendering and pixel-wise superimposed with the original video to obtain a virtual character rendering video.

6. The AIGC video creation method for cultural relics explanation according to claim 1 is characterized in that: The NeRF model includes: Scene representation: NeRF represents the scene as a continuous 5D function whose input is the coordinates (x, y, z) in 3D space and the 2D viewing direction (θ, φ), and the output is the RGB color and volume density (σ, c) at that location; Neural Network: NeRF uses a multi-layer perceptron as a neural network to approximate 5D functions. The input is 5D coordinates (x, y, z, θ, φ), and the output is volume density and RGB color. Volume rendering: NeRF samples 3D points along each camera ray and uses an MLP network to predict the color and density of the 3D points. It then accumulates the color and density of the 3D points using volume rendering techniques to generate a 2D image. The formula for volume rendering is expressed as: Among them, r(t)=o+td is the ray starting from the camera origin o, with direction d and parameter t; Optimization process: NeRF optimizes model parameters by minimizing the error between the observed image and the corresponding view rendered from NeRF. This process is differentiable and uses gradient descent method for optimization. Positional encoding: To help MLP represent high-frequency functions, NeRF uses positional encoding to map the input coordinates to a higher-dimensional space as follows: g(p)=(sin(2 0 πp),cos(2 0 πp),…,sin(2 L πp),cos(2 L πp)); The function g(p) is applied to the 3D coordinates (x, y, z) and the unit vector d of the viewing direction respectively; p is the component of the input coordinate or direction vector, and L is the number of frequencies of the position encoding; Hierarchical Sampling: To improve efficiency, NeRF uses hierarchical sampling, first using a coarse network for sampling, and then performing more targeted sampling on the fine network based on the output of the coarse network.

7. The AIGC video creation method for cultural relics explanation according to claim 2 is characterized in that: The MossFormer2 model includes: Encoder: Input: Mixed speech waveform x∈R 1×T ; Method: Use one-dimensional convolutional layer combined with trimmed linear unit to encode the input line; Output: Encoded sequence X∈R N×S , where N is the embedding dimension and S is the encoding sequence length. The encoding sequence length S is calculated as Where K1 is the convolution kernel size; R is the number of repetitions, and T is the input sequence length; Decoder: Input: masked encoded output; Method: Use transposed 1D convolutional layers with the same kernel size and stride as the encoder; Output: reconstructed original waveform; MossFormer Module: Core: Apply self-attention on the entire sequence; Strategy: Joint local-global self-attention strategy, performing full computational self-attention on non-overlapping local segments and using a linearized self-attention mechanism on the entire sequence; Goal: Capture long-range, coarse-scale dependencies; Loop Module: Goal: Model complex temporal dependencies in speech signals and capture local cyclic patterns; Method: Perform cyclic learning for each embedding dimension; No RNN cycle module: Bottleneck layer: Use a 1×11×1 convolutional layer to reduce the embedding dimension; GCU layer: built based on gated convolutional units and dilated FSMN; Output layer: LayerNorm followed by a 1×11×1 convolutional layer; Extended FSMN blocks: Structure: It includes a feed-forward layer and a storage layer. The storage layer uses stacked two-dimensional dilated convolution blocks. Goal: Cover a wider receptive field and enhance information flow and gradient propagation; Conv-U Block: Structure: Contains LayerNorm layer, linear layer, SiLU activation and one-dimensional deep convolution layer; Goal: Assist the GCU layer to capture the local pattern of position; Masking network: maps the encoded output to a set of masks that are used to separate different speech sources; The workflow of the MossFormer2 model: Input mixed speech waveform; Encoding: Convert the waveform into a coding sequence through Conv1D and ReLU; MossFormer module: applies a joint local-global self-attention strategy to capture long-range dependencies; Recurrent module: uses GCU layers and dilated FSMN modules to capture complex temporal dependencies and local recurrent patterns; Masking network: maps the encoded output to a set of masks; Decoding: Reconstruct the original waveform using a transposed 1D convolutional layer.

8. The AIGC video creation method for cultural relics explanation according to claim 2 is characterized in that: The VITS2 model includes: Random duration predictor: Input: Hidden representation of text h text and Gaussian noise z d ; Generator G: Use h text and z d As input, generate the predicted duration Discriminator D: The discriminator receives h text and the logarithmic duration d or predicted duration obtained from Monotonic Alignment Search As input; Training strategy: Using adversarial learning, the discriminator classifies according to the actual and predicted durations; Adversarial loss function: Mean square error loss function: L mse =MSE(G(z d ,h text ),d); Monotonic alignment search: Find the alignment with the highest probability between text and audio among all possible monotonic alignments, and add Gaussian noise during initial training to increase the diversity of alignment search; algorithm: where ∈ is the product of the noise sampled from the standard normal distribution and the standard deviation of P; is a normal distribution, z j is the latent variable after transformation from normal flow; μ i is the mean of the normal distribution, σ i is the standard deviation of the normal distribution; Normalization flow: Use convolutional blocks to capture adjacent data and add a small Transformer block with residual connections to capture long-range dependencies; Speaker-Conditioned Text Encoder: Generates speech with different characteristics based on speaker conditions in a multi-speaker model; Method: Introduce speaker vector in the third Transformer block of the text encoder to learn and express the unique speech characteristics of each speaker; Workflow diagram of the VITS2 model: Input text text; Text encoder: Encodes the input text into a hidden representation h text ; Duration prediction: Generate predicted duration through random duration predictor Monotone Alignment Search: Use monotone alignment search to align text and audio, adding Gaussian noise to enhance the diversity of alignment; Normalization flow: Use a combination of convolutional blocks and Transformer blocks to normalize the data and capture long-range and local dependencies; Speaker conditioning: adding speaker vectors to the text encoder to enhance the generation capabilities of multi-speaker models; Adversarial Learning: Use the generator G and the discriminator D for adversarial training to improve the naturalness of duration prediction and speech synthesis.

Citation Information

Patent Citations

  • Video local stylization method, computer system, storage medium and program product

    CN118524255A

  • Multiview neural human prediction using implicit differentiable renderer for facial expression, body pose shape and clothes performance capture

    US20220319055A1

Cited By

  • AIGC-based sports explanation generation method

    CN122153804A