AI-based children's story video generation method and system

By constructing an AI model that integrates emotion modeling and an image generation model, the problems of singular emotional expression and incoherence between frames in children's story videos have been solved, enabling more engaging video generation and user-friendly content adjustment capabilities.

CN120812370BActive Publication Date: 2025-11-21KUAISHANGYUN (SHANGHAI) NETWORK TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511289500.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2025-11-21
Estimated Expiration
2045-09-10

AI Technical Summary

Technical Problem

Existing technologies lack accurate modeling for specific emotional expressions in children's story video generation, resulting in one-dimensional emotional delivery in the generated video content, frame incoherence or detail distortion during image generation, and a lack of effective joint optimization mechanisms.

Method used

An AI model integrating emotion modeling capabilities is constructed, combining StyleGAN3-T and LDM models to generate contour and color images. The Lookahead optimizer is used for iterative solving, and frame alignment is performed using Sobel edge detection and block matching algorithms. An emotion modulation function is defined as a conditional audio generation function for the WaveNet model, and FFmpeg is used to synthesize videos. A visual interface is built for users to adjust.

Benefits of technology

The generated children's story videos are more emotionally engaging, with improved frame continuity and consistency in visual style. Users can adjust the content according to their needs to adapt to diverse creative requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120812370B_ABST
    Figure CN120812370B_ABST
Patent Text Reader

Abstract

The application discloses an AI-based children's story video generation method and system, relates to the field of artificial intelligence and multimedia cross technology, and comprises the following steps: constructing an AI model integrating emotion modeling capability to generate a script from original text input by a user; constructing an image generation combined model, defining a joint loss function, using a Sobel edge detection algorithm to calculate an edge intensity map of a contour image, using a block matching algorithm to calculate an optical flow field of frame changes, and performing dynamic frame alignment of a color image; using a fine-tuning WaveNet model to generate audio; constructing an image generation combined model, combining a StyleGAN3-T model and an LDM model, defining a joint loss function, using a Sobel edge detection algorithm and a block matching algorithm to calculate an edge intensity map and an optical flow field, realizing dynamic frame alignment of a color image, and improving interframe continuity of a generated video.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence and multimedia, in particular to a children's story video generation method and system based on AI. BACKGROUND

[0002] With the rapid development of artificial intelligence technology, AI-based content generation technology has been widely used in the multimedia field. In the field of children's story video generation, related technologies mainly involve natural language processing, generative adversarial networks, diffusion models, and speech synthesis technology, which provide new possibilities for children's education and entertainment.

[0003] The existing technology still has deficiencies. In terms of emotion modeling, the existing technology is mostly based on general language models or speech synthesis models, lacking precise modeling for specific emotional expression of children's stories, resulting in single emotional transmission of generated video content, which is difficult to effectively attract the attention of children viewers. In the image generation process, the existing method usually generates contour and color images independently, lacking effective joint optimization mechanism, resulting in problems such as frame discontinuity or detail distortion in the generated images in dynamic scenes. SUMMARY

[0004] The present application aims to provide a children's story video generation method and system based on AI to solve the above problems.

[0005] The present application is achieved by the following technical solutions:

[0006] A children's story video generation method based on AI, comprising:

[0007] An AI model with emotion modeling capability is constructed to generate a script from the user input original text;

[0008] An image generation combined model is constructed, including a StyleGAN3-T model to generate contour images, an LDM model to generate color images, a joint loss function is defined, a Lookahead optimizer is used for iterative solution, a Sobel edge detection algorithm is used to calculate the edge intensity map of the contour image, a block matching algorithm is used to calculate the optical flow field of frame changes, and dynamic frame alignment of color images is performed;

[0009] A sentiment modulation function is defined as a condition for fine-tuning the WaveNet model, and a fine-tuned WaveNet model is used to generate audio;

[0010] FFmpeg tools are used to generate children's story videos from video and audio, and a visual interface is constructed for user adjustment and modification.

[0011] As a preferred scheme of the AI-based children's story video generation method, the AI model with fusion emotional modeling capability is used to generate a script from the original text input by the user, including:

[0012] The AI model includes a semantic encoder, an emotional semantic embedding unit, a prompt engineering unit, a logical consistency checking unit, and a content-emotion disentanglement unit.

[0013] The original text input by the user is converted into a semantic vector using the semantic encoder, and the original text input by the user is mapped to an emotional space using the emotional semantic embedding unit to generate an emotional vector. The semantic vector and the emotional vector are classified using a BERT fine-tuned age-appropriateness classification model to filter out inappropriate content for children, obtaining an age-appropriate semantic vector and an emotional vector.

[0014] The age-appropriate semantic vector and the emotional vector are respectively retrieved in the children's story template library using the Faiss index, the age-appropriate semantic vector, the age-appropriate emotional vector, and the similar template are integrated into a structured prompt text using the prompt engineering unit, and the decoder is used to decode the text structure.

[0015] The text structure is mapped to a graph structure, a graph neural network GNN model is constructed, the graph structure of the children's story template is used for training, node embedding is learned to capture the global narrative structure, the graph structure is input, the narrative graph is output, the subgraph of the narrative graph is extracted, and the ChatGLM model is used to generate a text prompt corresponding to the subgraph.

[0016] Based on the input original text, the ChatGLM model is used to generate a sub-script from the text prompt, and the bidirectional attention mechanism is used to extract context information from the sub-script and the text prompt to generate an emotional label.

[0017] The content-emotion disentanglement unit is used to map the sub-script and the emotional label to a unified semantic space using a pre-trained Transformer model, output a joint embedding vector, and separate the content feature vector and the emotional feature vector in the joint embedding vector through a variational autoencoder.

[0018] The logical consistency checking unit is used to verify the logical coherence of the sub-script using cosine similarity and verify the age-appropriateness through the BERT fine-tuned age-appropriateness classification model. If one of the verification conditions is not met, the sub-script is regenerated, the content feature vector and the emotional feature vector of each sub-script are associated and integrated to generate a complete script.

[0019] As a preferred scheme of the AI-based children's story video generation method, wherein: the image generation combination model is constructed, a color image is generated, a joint loss function is defined, and a Lookahead optimizer is used for iterative solving, including:

[0020] Set Shared by the StyleGAN3-T model and the LDM model for the feature extractor;

[0021] The StyleGAN3-T model is constructed, including a contour image generator and a discriminator;

[0022] The contour image generator is configured to output a contour image based on a conditional input vector;

[0023] The contour image discriminator is configured to determine whether the generated contour conforms to the cartoon style;

[0024] After collecting children's story set data and using an AI model to generate a script, the emotional feature vector, the sub-script, and the random noise are spliced into a vector, and a pre-trained MLP model is used for dimension reduction to obtain a conditional input vector suitable for the StyleGAN3-T model;

[0025] Cartoon style loss and anti-aliasing loss are added, a StyleGAN3-T model loss function is defined, contour image features are extracted from the contour image generator of the StyleGAN3-T model using an attention mechanism, the contour image features and the emotional feature vector are spliced into a vector, and a pre-trained MLP model is used for dimension reduction on the spliced vector to obtain a joint conditional vector;

[0026] The LDM model is constructed, including a color image encoder, a diffusion model, and a decoder;

[0027] The color image encoder is configured to encode the joint conditional vector into a latent representation;

[0028] The color image diffusion model is configured to generate a color image in the latent space based on the emotional feature vector;

[0029] The color image decoder is configured to decode the latent representation into a high-resolution color image;

[0030] The LDM model diffusion loss function is defined, and the StyleGAN3-T model loss function is connected in parallel to construct a diffusion loss function and a joint loss function, with the goal of minimizing the joint loss function, using a Lookahead optimizer for optimization, while updating the parameters of the StyleGAN3-T model and the LDM model, stopping iteration after reaching the maximum number of iterations, outputting a color latent representation, and using up-sampling and deconvolution by the color image decoder to restore the color latent representation to a color image.

[0031] Set a structural consistency score threshold and calculate the structural consistency score. If the structural consistency score is If the score is below the structural consistency score threshold, the weight of the cross-modal consistency loss is increased to re-optimize the model; otherwise, a validated color image is output.

[0032] As a preferred embodiment of the AI-based children's story video generation method of the present invention, the step of using the Sobel edge detection algorithm to calculate the edge intensity map of the contour image, using the block matching algorithm to calculate the optical flow field of frame changes, and performing dynamic frame alignment of the color image includes:

[0033] The contour image features, color image features, and sentiment feature vectors are concatenated to obtain a comprehensive feature.

[0034] The Sobel edge detection algorithm is used to calculate the edge intensity map of the contour image, and an edge intensity threshold is set. Edge intensity filtering is performed, a sentiment mask is set, a weighted summation method is used to calculate a weight map, and a weighted SAD is calculated based on the weight map. Optical flow estimation is performed using the block-matching algorithm to calculate the optical flow field of frame changes. An initial animation sequence is generated using a nonlinear interpolation method, and the initial animation sequence is smoothed using a bilateral filtering algorithm to obtain smoothed frames. Sentiment consistency constraints are defined, smoothed frames are adjusted, and histogram matching is used to match sentiment-adjusted smoothed frames, keyframes, and time steps to obtain an optimized animation sequence.

[0035] As a preferred embodiment of the AI-based children's story video generation method of the present invention, wherein: defining an emotion modulation function as a condition for fine-tuning the WaveNet model, and using the fine-tuned WaveNet model to generate audio, includes:

[0036] The WaveNet model was pre-trained using the LJSpeech and VCTK datasets, and then fine-tuned using a children's cartoon speech dataset.

[0037] Define an emotion modulation function and use it as the output condition of the fine-tuned WaveNet model. Input the content vector and emotion vector into the fine-tuned WaveNet model and output a conditional probability distribution. For each time step, sample from the conditional probability distribution using maximum likelihood estimation to generate a continuous waveform sequence.

[0038] Based on emotion vectors and speech waveforms, a pre-trained MusicGen model is used to generate emotion-matched music, and the speech and music are mixed using FFmpeg to obtain mixed audio.

[0039] The optimized animation sequence is merged into a video stream using the FFmpeg tool, the length of the video and the difference of the mixed audio are calculated, the speech rate of the mixed audio is adjusted based on the difference using the FFmpeg tool, and a simultaneous long audio is obtained.

[0040] As a preferred scheme of the AI-based children's story video generation method, the video and audio are generated into a children's story video using the FFmpeg tool, including:

[0041] Based on the complete script timestamp, the character lip movement is synchronized with the voiceover, the video stream and the simultaneous long audio are aligned using the FFmpeg tool, an initial children's story audio video is obtained, and the image resolution of the initial children's story audio video is improved using the Lanczos interpolation algorithm;

[0042] Based on the complete script scene, the initial children's story audio video is added with a fade-in and fade-out effect, and an HLS multi-bit rate video version is generated, and a final children's story video is obtained.

[0043] As a preferred scheme of the AI-based children's story video generation method, the visual interface is constructed for user adjustment and modification, including:

[0044] The visual interface is constructed using the React tool, the final children's story video is displayed, and the user can select video clips based on the timeline for adjustment and deletion;

[0045] The adjustment includes image and audio adjustment;

[0046] The deletion includes that the user can delete the timeline based on the content of the final children's story video, which is not suitable for children to watch.

[0047] In a second aspect, the application provides an AI-based children's story video generation system, including,

[0048] The fusion module is used to construct an AI model with fusion emotion modeling capability, and the original text input by the user is generated into a script;

[0049] The alignment module is used to construct an image generation combination model, including a StyleGAN3-T model for generating contour images, an LDM model for generating color images, a joint loss function is defined, a Lookahead optimizer is used for iterative solution, a Sobel edge detection algorithm is used to calculate the edge intensity map of the contour image, a block matching algorithm is used to calculate the optical flow field of frame changes, and the color image dynamic frame alignment is performed.

[0050] A definition generation module is configured to define an emotion modulation function as a condition for fine-tuning a WaveNet model, and generate audio using the fine-tuned WaveNet model;

[0051] A design interaction module is configured to generate a children's story video using FFmpeg tools to combine video and audio, and build a visual interface for users to adjust and modify.

[0052] Compared with the prior art, the present application has the following advantages and beneficial effects:

[0053] 1、The present application constructs an AI model with emotion modeling capability, generates a script with emotional color from the user's input original text, and fine-tunes a WaveNet model using an emotion modulation function as a condition to generate audio.

[0054] 2、The present application constructs an image generation combination model, generates contour images using StyleGAN3-T model, generates color images using LDM model, defines a joint loss function, uses Lookahead optimizer for iterative solution, and uses Sobel edge detection algorithm and block matching algorithm to calculate edge intensity map and optical flow field to realize dynamic frame alignment of color images, improve inter-frame continuity and picture style consistency of generated video, compared with the common inter-frame flicker or style inconsistency in the prior art, the video sequence generated by the present application is smoother in dynamic scene transition, the picture details are richer, and the overall visual effect is more smooth and artistic;

[0055] 3、The present application constructs a visual interface and integrates FFmpeg tools to allow users to make real-time adjustments and modifications to the generated children's story video, enhancing the interactivity and flexibility of the system. BRIEF DESCRIPTION OF DRAWINGS

[0056] The accompanying drawings described herein are used to provide further understanding of the embodiments of the present application, form a part of the present application, and do not constitute a limitation of the embodiments of the present application.

[0057] Fig. 1A flowchart of an AI-based children's story video generation method of the present application;

[0058] Fig. 2 A schematic diagram of an AI-based children's story video generation system of the present application. DETAILED DESCRIPTION

[0059] In order to make the objectives, technical solutions and advantages of the present application clearer, further detailed description will be given below in combination with embodiments and drawings, the illustrative embodiments of the present application and the description thereof are only used to explain the present application, and do not limit the present application. It should be noted that the present application has been in the actual research and development stage. Embodiment 1

[0060] As shown in the following table, the present embodiment includes. Figs. 1-2

[0061] S1, constructing an AI model with fusion emotion modeling capability, generating a script from the original text input by the user;

[0062] Preferably, based on the GLM framework of the ChatGLM model, the AI model with fusion emotion modeling capability is constructed to process the input of the user;

[0063] The AI model includes a semantic encoder, an emotion semantic embedding unit, a prompt engineering unit, a logical consistency checking unit and a content-emotion disentanglement unit (wherein the semantic encoder, the prompt engineering unit and the logical consistency checking unit belong to the GLM framework, and the emotion semantic embedding unit and the content-emotion disentanglement unit are added);

[0064] The original text input by the user is converted into a semantic vector using the semantic encoder;

[0065] The original text input by the user is mapped to an emotion space (based on an emotion dictionary) using the emotion semantic embedding unit to generate an emotion vector;

[0066] The semantic vector and the emotion vector are classified using a BERT fine-tuned age-appropriate classification model, and the content inappropriate for children is filtered to obtain an age-appropriate semantic vector and an emotion vector;

[0067] The age-appropriate semantic vector and the emotion vector are respectively retrieved in a children's story template library using a Faiss index (the template with the highest cosine similarity);

[0068] The age-appropriate semantic vector, the age-appropriate emotion vector and the similar template are integrated into a structured prompt text using the prompt engineering unit, and a decoder is used to decode the text structure;

[0069] ​Map the text structure to the graph structure, define the roles, scenes and plot events as nodes, and represent the relationship between the nodes as edges (for example, the little rabbit got lost in the forest, then met the little hedgehog, the little hedgehog helped the little rabbit find the way home, the little rabbit and the little hedgehog are roles, the forest is a scene, and getting lost, meeting and helping home are plot events, the edges are little rabbit→lost→forest, little rabbit→met→little turtle, little turtle→helped→little rabbit, little rabbit→found→way home);

[0070] Build a graph neural network (GNN) model, train it using the graph structure of the children's story template, learn node embeddings to capture the global narrative structure, input the graph structure, and output the narrative graph;

[0071] Extract subgraphs from the narrative graph, and use the ChatGLM model to generate text prompts corresponding to the subgraphs;

[0072] Based on the input original text, use the ChatGLM model to generate subscripts from the text prompts, and use the bidirectional attention mechanism to extract context information from the subscripts and text prompts to generate sentiment labels;

[0073] The content-emotion disentanglement unit is configured to use a pre-trained Transformer model to map the subscripts and sentiment labels to a unified semantic space, output joint embedding vectors, and separate the content feature vectors and emotion feature vectors in the joint embedding vectors through a variational autoencoder;

[0074] The content feature vector refers to a phoneme sequence;

[0075] The emotion feature vector includes emotion categories and emotion intensity scores (using a Sigmoid function to map the emotion labels to the range [0, 1] to generate emotion intensity scores);

[0076] The logical consistency checking unit is configured to use cosine similarity to verify the logical coherence of the subscripts, and use a BERT fine-tuned age-appropriateness classification model to verify the age-appropriateness, and if one of the conditions is not met, regenerate the subscripts;

[0077] Integrate the content feature vectors and emotion feature vectors of each subscript to generate a complete script.

[0078] By converting the original text into semantic vectors, the conversion from unstructured text to structured representation is realized, by the sentiment semantic embedding unit, the accurate extraction and quantification of text sentiment features are realized, by the accurate age filtering, the child-friendly of the generated content is significantly improved, by the efficient retrieval of Faiss index, the fast and accurate template matching is realized, by the structured prompt generation and decoding, the smooth transition from abstract vector to specific text is realized, the readability and narrative coherence of the generated text are improved, by the mapping of the graph structure, the structured and relational expression of the narrative elements is realized, by the training and optimization of GNN, the overall and logical of the story narrative is significantly improved, by the subgraph extraction and bidirectional attention mechanism, the accurate generation and sentiment labeling of narrative segments are realized, the accuracy of emotional expression and context coherence of story segments are improved, by the content-emotion disentanglement, the independent modeling and optimization of content and emotion are realized, by the strict logic and age verification, the quality and reliability of the generated script are significantly improved, the occurrence of incoherent or inappropriate content is reduced.

[0079] S2, construct an image generation combination model, including a StyleGAN3-T model to generate an outline image, an LDM model to generate a color image, define a joint loss function, use a Lookahead optimizer to iteratively solve, use a Sobel edge detection algorithm to calculate an edge intensity map of the outline image, use a block matching algorithm to calculate an optical flow field of frame changes, and perform dynamic frame alignment on the color image;

[0080] Preferably, set shared by the StyleGAN3-T model and the LDM model for the feature extractor;

[0081] Construct a StyleGAN3-T model, including an outline image generator and a discriminator;

[0082] The outline image generator is configured to output an outline image based on a conditional input vector;

[0083] The outline image discriminator is configured to determine whether the generated outline conforms to a cartoon style;

[0084] After collecting children's story set data, using an AI model to generate a script, the emotional feature vector, the sub-script and the random noise are spliced into a vector, and a pre-trained MLP model is used to reduce the dimension to obtain a conditional input vector suitable for the StyleGAN3-T model;

[0085] Add cartoon style loss and anti-aliasing loss to optimize the outline image generator, formula:

[0086] ,

[0087] wherein, is the total loss function of the StyleGAN3-T model, is the expectation of the real contour x and the conditional input , is the output probability of the discriminator for the real contour x and the conditional input , is the expectation of the conditional input , is the contour image output by the StyleGAN3-T model generator, is the output probability of the discriminator for the generated contour image , and are the weights of the cartoon style loss and the anti-aliasing loss respectively, which are set using an experimental method, is the cartoon style loss, is the contour image output by the StyleGAN3-T contour image generator, is the DINOv2 self-supervised feature extractor, is the square of the L2 norm, which calculates the feature difference, is the anti-aliasing loss, is the gradient of the contour image, which measures the spatial changes of the image;

[0088] The contour image features are extracted from the contour image generator of the StyleGAN3-T model using an attention mechanism, the contour image features and the emotion feature vector are vector spliced to obtain a spliced vector, and the spliced vector is dimensionally reduced using a pre-trained MLP model to obtain a joint condition vector;

[0089] An LDM model is constructed, including a color image encoder, a diffusion model and a decoder;

[0090] The color image encoder is configured to encode the joint condition vector into a latent representation;

[0091] The color image diffusion model is configured to generate a color image in the latent space based on the emotion feature vector;

[0092] The color image decoder is configured to decode the latent representation into a high-resolution image;

[0093] The LDM model diffusion loss function is as follows:

[0094] ,

[0095] wherein, is the LDM model diffusion loss function, is the contour image, is the real color image, is the joint condition vector, is a random noise, is a diffusion time, is a denoising network of the LDM model predicting noise conditioned on the joint condition vector, is a latent representation at time step t, generated by a forward diffusion process;

[0096] A joint loss function is defined, formula:

[0097] ,

[0098] wherein, is the joint loss function, is a weight of the cross-modal consistency loss, set using an experimental method, is the cross-modal consistency loss, is a color image generated by the LDM model;

[0099] To minimize the joint loss function, the Lookahead optimizer is used to optimize the parameters of the StyleGAN3-T model and the LDM model, and after reaching the maximum number of iterations, the iteration is stopped, and the color latent representation is output. The color image decoder uses upsampling and deconvolution to restore the color latent representation to a color image;

[0100] A structure consistency score threshold is set, and the structure consistency score is calculated, formula:

[0101] ,

[0102] If the structure consistency score is lower than the structure consistency score threshold, the value of the weight of the cross-modal consistency loss is increased, and the model is re-optimized, otherwise the color image that passes the verification is output.

[0103] By constructing a style-controllable contour image generation network, a cartoon-style contour image can be generated from conditional vectors. By fusing sentiment features, textual semantics, and randomness, and using a pre-trained MLP to reduce the dimensionality to a StyleGAN3-T acceptable format, structured conditional control is achieved. By introducing cartoon style loss (based on self-supervised deep feature extraction) and anti-aliasing loss (based on gradient smoothing constraints), multi-objective generation optimization is achieved. Contour image features are captured through an attention mechanism and fused with sentiment features, achieving a deep binding between image structure and emotional semantics. By establishing a matching mechanism between real and generated images in the latent space, the diffusion model is guided to converge toward the real image. By jointly training StyleGAN3-T and LDM, a cross-modal consistency loss is introduced to ensure semantic consistency between contour images and color images. By scoring structural consistency, dynamic feedback on the model training effect is achieved, and the loss weights are automatically adjusted accordingly.

[0104] Furthermore, the contour image features, color image features, and emotion feature vectors are concatenated to obtain comprehensive features;

[0105] The Sobel edge detection algorithm is used to calculate the edge intensity map of the contour image, and an edge intensity threshold is set. Perform edge strength filtering and set a sentiment mask. Formula:

[0106] ,

[0107] in, For pixel coordinates Emotional mask In pixel coordinates Edge intensity map;

[0108] Based on the sentiment mask and edge intensity map, a weight map is calculated using a weighted summation method. Then, based on the weight map, a weighted SAD is calculated using the following formula:

[0109] ,

[0110] in, for and The weighted sum of absolute differences of the matching blocks, and They belong to keyframes and The center position of the matching block, keyframe and For the start and end frames of a color image, based on a subscript, These are the pixel coordinates within the block. and Keyframes and In pixels and Pixel value at that location, and The candidate optical flow vector represents the block from the keyframe. arrive The horizontal and vertical displacements were generated using the Diamond Search algorithm. In pixel coordinates Weighted graph;

[0111] Based on weighted SAD, the Block-Matching algorithm is used for optical flow estimation to calculate the frame-varying optical flow field, as shown in the formula:

[0112] ,

[0113] in, For optical flow field, representing pixel From keyframes arrive displacement vector ;

[0114] Based on the optical flow field, an initial animation sequence is generated using a nonlinear interpolation method, with the following formula:

[0115] ,

[0116] ,

[0117] in, For the interpolated frame at time step t, For image deformation operations;

[0118] The initial animation sequence is smoothed using a bilateral filtering algorithm to obtain smoothed frames. Emotional consistency constraints are defined, and the smoothed frames are adjusted using the following formula:

[0119] ,

[0120] in, In pixel coordinates Emotional adjustment smooth frames, For time t, the pixel coordinates are Smooth frames, For the color map table, from the current, LUT stands for Lookup Table Mapping. For color lookup tables;

[0121] Histogram matching is used to match emotion adjustment smoothing frames, keyframes, and time steps to obtain an optimized animation sequence.

[0122] By combining edge strength and emotion mask, this step realizes accurate positioning of key regions of the image and preliminary quantification of emotional features. By introducing a weight map and weighted SAD, this step significantly improves the accuracy and robustness of optical flow estimation. By combining the weighted SAD block matching algorithm, the accuracy of optical flow estimation is significantly improved. By combining the optical flow field with a nonlinear interpolation method, complex motion trajectories can be accurately simulated, generating smooth and natural animation sequences, reducing motion distortion or blur that may be caused by traditional linear interpolation. By bilateral filtering and emotion consistency constraints, this step significantly improves the visual quality of the animation sequence, reducing noise interference. By matching the color distribution, color jumps or inconsistencies between frames are eliminated, making the final output animation sequence more natural and attractive in overall visual perception.

[0123] S3, define an emotion modulation function as a condition for fine-tuning the WaveNet model, and use the fine-tuned WaveNet model to generate audio;

[0124] Preferably, the LJSpeech and VCTK datasets are used to pre-train the WaveNet model, and the WaveNet model is fine-tuned using a children's cartoon voice dataset;

[0125] Define the emotion modulation function, formula:

[0126] ,

[0127] wherein, is the emotion modulation function, is the emotion vector, is the emotion intensity score, , and are the weights of pitch adjustment, speech rate adjustment and volume adjustment, respectively, which are set based on the sound characteristics of children's stories, , and are the pitch adjustment, speech rate adjustment and volume adjustment values based on , respectively, with initial values set based on the sound characteristics of children's stories;

[0128] The emotion modulation function is used as an output condition for the fine-tuned WaveNet model, and the content vector and emotion vector are input into the fine-tuned WaveNet model to output a conditional probability distribution. For each time step, a continuous waveform sequence is generated by sampling from the conditional probability distribution using maximum likelihood estimation;

[0129] Based on the emotion vector and the speech waveform, the pre-trained MusicGen model is used to generate emotion-matching music, the FFmpeg tool is used to mix the speech and music, and the mixed audio is obtained.

[0130] The optimized animation sequence is merged into a video stream using the FFmpeg tool, the difference between the video length and the mixed audio is calculated, the speech rate of the mixed audio is adjusted based on the difference using the FFmpeg tool, and the simultaneous length audio is obtained.

[0131] By using two high-quality speech data sets, LJSpeech and VCTK, to pre-train the WaveNet model, the basic speech generation capability is built, and then the model is fine-tuned by using child cartoon speech data to further adapt the voice features of the model in the context of children's stories. By defining an emotion modulation function containing an emotion vector and an intensity score, a weight adjustment mechanism in three dimensions of pitch, speech rate, and volume is introduced, and initial values are set for the unique sound style of children's stories. The emotion modeling and parameter control of the speech synthesis process are realized, and the emotion modulation function is input as a conditional input to the fine-tuned WaveNet model together with the content vector, realizing conditional speech generation based on semantic and emotional double information. The audio and animation sequence are integrated into a video stream by the FFmpeg tool, and the time length difference between the two is automatically analyzed, and the audio speech rate is dynamically adjusted, so that the final output audio and video are accurately aligned on the time axis, realizing natural audio-visual synchronization.

[0132] S4, use FFmpeg tool to generate children's story video from video and audio, and build a visual interface for users to adjust and modify;

[0133] Preferably, based on the complete script timestamp, the character's lip movement is synchronized with the voiceover, the video stream and the simultaneous length audio are aligned using the FFmpeg tool, and the initial children's story video is obtained. The Lanczos interpolation algorithm is used to improve the image resolution of the initial children's story audio and video (for adapting to high-definition devices);

[0134] Based on the complete script, fade-in and fade-out effects are added to the initial children's story audio and video, and HLS multi-rate video versions are generated, and the final children's story video is obtained.

[0135] Through the precise time marking of the script timestamp, the lip movement of the characters in the video is completely synchronized with the voiceover audio on the time axis. Through the high-quality image interpolation technology of the Lanczos algorithm, the resolution of the video is improved to adapt to high-definition devices (such as 4K displays, tablets), while maintaining the smoothness and detail clarity of the image edges. Through the fade-in and fade-out effect, smooth transition between scenes is realized, and at the same time, HLS technology generates multiple rate video streams to adapt to different network environments and device performance.

[0136] Further, a visual interface is built using React tools to display the final children's story video, allowing users to select video segments based on the timeline for adjustment and deletion;

[0137] The adjustment includes image and audio adjustment;

[0138] The deletion includes the user's feeling that the content of the final children's story video is not suitable for children to watch, and the deletion of the timeline.

[0139] Through the componentized development and efficient rendering of React, an intuitive and easy-to-use interactive interface is built, allowing users to accurately select video segments through the timeline, adjust images (such as brightness, contrast) or audio (such as volume, sound effects), and delete content that does not meet the children's viewing standards, suitable for scenarios that require user involvement in content customization.

[0140] The embodiment also provides an AI-based children's story video generation system, comprising:

[0141] A fusion module is built to build an AI model with fusion emotion modeling capabilities to generate scripts from user inputted raw text;

[0142] A calculation alignment module is used to build an image generation combination model, including a StyleGAN3-T model to generate outline images, an LDM model to generate color images, a joint loss function is defined, a Lookahead optimizer is used for iterative solution, a Sobel edge detection algorithm is used to calculate the edge intensity map of the outline image, a block matching algorithm is used to calculate the optical flow field of frame changes, and the color image dynamic frame alignment is performed;

[0143] A generation module is defined to define an emotion modulation function as a condition for fine-tuning a WaveNet model, and a fine-tuned WaveNet model is used to generate audio;

[0144] An interaction module is designed to use FFmpeg tools to generate children's story videos from video and audio, and build a visual interface for user adjustment and modification.

[0145] The above specific embodiments further detail the purpose, technical solutions and benefits of the present application. It should be understood that the above description is only a specific embodiment of the present application and does not limit the scope of protection of the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. An AI-based children's story video generation method, The application relates to a method for generating a children's story video, comprising the following steps: constructing an AI model with a fusion emotion modeling capability to generate a script from original text input by a user; constructing an image generation combination model, including a StyleGAN3-T model for generating an outline image, an LDM model for generating a color image, defining a joint loss function, using a Lookahead optimizer for iterative solving, using a Sobel edge detection algorithm to calculate an edge intensity map of the outline image, using a block matching algorithm to calculate an optical flow field of frame changes, and performing dynamic frame alignment of the color image; defining an emotion modulation function as a condition for fine-tuning a WaveNet model, and using the fine-tuned WaveNet model to generate audio; using an FFmpeg tool to generate a children's story video from the video and the audio, and constructing a visual interface for a user to adjust and modify; the AI model with the fusion emotion modeling capability includes a semantic encoder, an emotion semantic embedding unit, a prompt engineering unit, a logical consistency checking unit and a content-emotion disentanglement unit; the semantic encoder is used to encode the original text input by the user into a semantic vector, the emotion semantic embedding unit is used to map the original text input by the user to an emotion space to generate an emotion vector, and a BERT fine-tuned age-appropriateness classification model is used to classify the semantic vector and the emotion vector, filter the content unsuitable for children, and obtain age-appropriate semantic vectors and emotion vectors; the age-appropriate semantic vectors and the emotion vectors are respectively searched in a children's story template library using a Faiss index, the age-appropriate semantic vectors, the age-appropriate emotion vectors and the similar templates are integrated into a structured prompt text using the prompt engineering unit, and the text structure is decoded into a text structure using a decoder; the text structure is mapped into a graph structure, a graph neural network GNN model is constructed, the graph structure of the children's story template is used for training, node embedding is learned to capture a global narrative structure, the graph structure is input, a narrative graph is output, a subgraph of the narrative graph is extracted, and a ChatGLM model is used to generate a text prompt corresponding to the subgraph; based on the input original text, a text prompt is generated from the subgraph using the ChatGLM model, context information is extracted from the text prompt and the subgraph using a bidirectional attention mechanism, and an emotion label is generated; the content-emotion disentanglement unit and the logical consistency checking unit include: the content-emotion disentanglement unit is used to map the subgraph and the emotion label to a unified semantic space using a pre-trained Transformer model, output a joint embedding vector, and separate a content feature vector and an emotion feature vector in the joint embedding vector through a variational autoencoder; the logical consistency checking unit is used to verify the logical coherence of the subgraph using a cosine similarity, verify the age-appropriateness through a BERT fine-tuned age-appropriateness classification model, and regenerate the subgraph if one of the verification conditions is not met, and the content feature vector and the emotion feature vector of each subgraph are associated and integrated to generate a complete script. the StyleGAN3-T model for generating an outline image includes:

2. The AI-based children's story video generation method of claim 1, wherein: ​ Setting shared with the StyleGAN3-T model and the LDM model for the feature extractor; The StyleGAN3-T model is constructed, including a contour image generator and a contour image discriminator; The contour image generator is configured to output a contour image based on a conditional input vector; The contour image discriminator is configured to determine whether the generated contour conforms to a cartoon style; After collecting children's story set data and generating a script using an AI model, the emotional feature vector, sub-script, and random noise are spliced into a vector, and a pre-trained MLP model is used for dimension reduction to obtain a conditional input vector suitable for the StyleGAN3-T model; Cartoon style loss and anti-aliasing loss are added to define the StyleGAN3-T model loss function, and an attention mechanism is used to extract contour image features from the contour image generator of the StyleGAN3-T model. The contour image features and emotional feature vectors are spliced into a vector, and a pre-trained MLP model is used for dimension reduction to obtain a joint conditional vector.

3. The AI-based children's story video generation method of claim 2, wherein: The LDM model generates a color image, defines a joint loss function, and uses a Lookahead optimizer for iterative solution, including: The LDM model is constructed, including a color image encoder, a color image diffusion model, and a color image decoder; The color image encoder is configured to encode the joint conditional vector into a latent representation; The color image diffusion model is configured to generate a color image in the latent space based on the emotional feature vector; The color image decoder is configured to decode the latent representation into a high-resolution color image; The diffusion loss function of the LDM model is defined, and the StyleGAN3-T model loss function is connected in parallel to construct the diffusion loss function and the joint loss function. The Lookahead optimizer is used for optimization to update the parameters of the StyleGAN3-T model and the LDM model simultaneously. After reaching the maximum number of iterations, the iteration is stopped, and the color latent representation is output. The color image decoder uses upsampling and deconvolution to restore the color latent representation to a color image; setting a structure consistency score threshold, calculating the structure consistency score , if the structure consistency score is lower than the structure consistency score threshold, re-optimizing the model by increasing the value of the weight of the cross-modal consistency loss, otherwise outputting the color image that passes the verification.

4. The AI-based children's story video generation method of claim 1, wherein: The Sobel edge detection algorithm is used to calculate the edge intensity map of the contour image, the block matching algorithm is used to calculate the optical flow field of the frame change, and the color image dynamic frame alignment includes: The contour image features, color image features, and emotional feature vectors are spliced into a comprehensive feature; An edge intensity map of the contour image is calculated using a Sobel edge detection algorithm, an edge intensity threshold is set , edge intensity screening is performed, an emotion mask is set, a weight map is calculated using a weighted summation method, a weighted SAD is calculated based on the weight map, optical flow estimation is performed using a block-matching algorithm Block-Matching, an optical flow field of frame changes is calculated, an initial animation sequence is generated using a nonlinear interpolation method, the initial animation sequence is smoothed using a bilateral filtering algorithm to obtain a smoothed frame, an emotion consistency constraint is defined, the smoothed frame is adjusted, and the emotion-adjusted smoothed frame, the key frame and the time step are matched using a histogram matching method to obtain an optimized animation sequence.

5. The AI-based children's story video generation method of claim 1, wherein: The emotional modulation function is defined as the condition of the fine-tuned WaveNet model, and the fine-tuned WaveNet model is used to generate audio, including: The LJSpeech and VCTK datasets are used to pre-train the WaveNet model, and the WaveNet model is fine-tuned using the children's cartoon voice dataset; The emotional modulation function is defined as the output condition of the fine-tuned WaveNet model. The content vector and emotional vector are input into the fine-tuned WaveNet model, and the conditional probability distribution is output. For each time step, the maximum likelihood estimation is used to sample from the conditional probability distribution to generate a continuous waveform sequence; Based on the emotion vector and the voice waveform, the pre-trained MusicGen model is used to generate emotion matching music, the FFmpeg tool is used to mix the voice and the music, and the mixed audio is obtained; The optimized animation sequence is merged into a video stream using the FFmpeg tool, the difference between the video length and the mixed audio is calculated, the speech rate of the mixed audio is adjusted based on the difference using the FFmpeg tool, and the simultaneous length audio is obtained.

6. The AI-based children's story video generation method of claim 1, wherein: The video and audio are generated into a children's story video using the FFmpeg tool, including: Based on the complete script timestamp, the character lip movement is synchronized with the voiceover, the video stream and the simultaneous length audio are aligned using the FFmpeg tool, the initial children's story audio video is obtained, and the Lanczos interpolation algorithm is used to improve the image resolution of the initial children's story audio video; Based on the complete script scene, the initial children's story audio video is added with fade-in and fade-out effect, and the HLS multi-rate video version is generated, and the final children's story video is obtained.

7. The AI-based children's story video generation method of claim 1, wherein: The visual interface is constructed for user adjustment and modification, including: The visual interface is constructed using the React tool to display the final children's story video, which can be selected by the user based on the timeline to adjust and delete; The adjustment includes image and audio adjustment; The deletion includes that the user can feel that the content of the final children's story video is not suitable for children to watch, and the timeline is deleted.

8. An AI-based children's story video generation system based on the AI-based children's story video generation method of any one of claims 1 to 7. Including, A fusion module is constructed to construct an AI model with fusion emotion modeling capability, which generates a script from the user input original text; A calculation alignment module is constructed to construct an image generation combination model, including a StyleGAN3-T model to generate a contour image, an LDM model to generate a color image, a joint loss function is defined, a Lookahead optimizer is used for iterative solution, a Sobel edge detection algorithm is used to calculate the edge intensity map of the contour image, a block matching algorithm is used to calculate the optical flow field of frame change, and the color image dynamic frame alignment is performed; A generation module is defined to define an emotion modulation function as a condition for fine-tuning the WaveNet model, and the fine-tuned WaveNet model is used to generate audio; An interactive module is designed to generate a children's story video using the FFmpeg tool to combine video and audio, and a visual interface is constructed for user adjustment and modification.

Citation Information

Patent Citations

  • Fine adjustment method and device for generative model, equipment, medium and product

    CN119558372A

  • Audio book video generation method and device, electronic equipment, medium and program product

    CN120353967A