Cartoon dynamic and speech synthesis method and system based on multi-modal generation

By using multimodal frame animation generation and speech synthesis models, dynamic videos and character voices are automatically generated, solving the problem of low processing efficiency in animation production and achieving efficient animation production.

CN121053262APending Publication Date: 2025-12-02CHENGDU MEGAYOU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510931151.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

In the animation production process, the low efficiency of dynamic compositing and insufficient automation result in high costs and long production cycles.

Method used

A multimodal generation-based approach is adopted, using a frame animation generation model and a speech synthesis model. Through layered processing, motion trajectory simulation, and audio-visual synchronization, dynamic videos and character voices are automatically generated.

Benefits of technology

It improves the processing efficiency of animation production, reduces manual operation, shortens the production cycle, and enhances the degree of automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053262A_ABST
    Figure CN121053262A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a cartoon dynamic and voice synthesis method and system based on multi-modal generation, and the method can generate a dynamic video according to a static cartoon image sequence in animation generation data through employing a frame animation generation model after the animation generation data is obtained, and employs a voice synthesis model. And according to the cartoon image corresponding to the key frame in the dynamic video, binding a timestamp for the role voice so as to generate a mixed audio. Therefore, the mixed audio and the dynamic video are synthesized into an output animation. According to the method, layering processing, motion trail simulation based on a physical engine and middle frame adding processing can be performed on a cartoon image through a frame animation generation model, role voiceprint features and role identifiers are bound through a voice synthesis model, emotion parameter injection is performed on role voice based on semantic analysis of dubbing texts, and the character voiceprint features and the role identifiers are combined. The processing time for converting the cartoon image into the animation can be shortened, and the processing efficiency of the dynamic synthesis process of the cartoon is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video synthesis technology, and in particular to a method and system for animation and speech synthesis of comics based on multimodal generation. Background Technology

[0002] Video compositing is the process of combining multiple multimedia elements such as images, video clips, and audio clips to create a dynamic video work. Taking comic book video compositing as an example, animators need to draw multiple frames of comic book images and play them sequentially according to their action sequences to create a dynamic comic book video. Because animation production involves compositing static comic book images into dynamic animated videos, the comic book video compositing process is also known as comic book mobilization. After creating the dynamic comic book video, it also needs to be dubbed; that is, voice actors and sound effects companies record audio clips, which are then embedded into the comic book video to create a video animation containing sound.

[0003] Because animation production requires manually drawing multiple frames and special effects images, the cost is relatively high. Therefore, animation generation applications can be used to assist in video frame creation during the animation production process. This involves drawing multiple comic book images as keyframes for the dynamic video and using the animation generation application to create intermediate transition frames, thus compositing a dynamic comic book video.

[0004] However, because the audio recording for animated videos requires manual recording, it's difficult to guarantee the consistency of voiceprints for the characters. Furthermore, embedding audio clips into the animated video requires manual alignment of keyframes with the audio track to ensure audio-visual synchronization. Therefore, the animation compositing process for comics is inefficient, lacks automation, and prolongs the animation production cycle. Summary of the Invention

[0005] In view of this, embodiments of this application provide a method and system for comic animation and speech synthesis based on multimodal generation, in order to solve the problem of low processing efficiency in the comic animation synthesis process.

[0006] According to one aspect of this application, a method for comic animation and speech synthesis based on multimodal generation is provided, the method comprising:

[0007] Acquire animation generation data, which includes a static comic image sequence, voice-over text, and character voiceprint features; the static comic image sequence includes multiple comic images.

[0008] A frame animation generation model is used to generate dynamic videos from the static comic image sequence. The frame animation generation model is a neural network model based on a diffusion transformer architecture and a self-attention mechanism architecture. The frame animation generation model is configured to perform hierarchical processing on the comic images, motion trajectory simulation based on a physics engine, and add intermediate frames.

[0009] Using a speech synthesis model, character speech is generated based on the dubbing text and the character's voiceprint features. The speech synthesis model is a deep learning-based neural network model. The speech synthesis model is configured to bind the character's voiceprint features to the character's identifier and to perform emotional parameter injection on the character's speech based on semantic analysis of the dubbing text.

[0010] The character's voice is time-stamped based on the comic images corresponding to keyframes in the dynamic video to generate mixed audio;

[0011] The mixed audio and the dynamic video are combined to form an output animation.

[0012] In some embodiments, obtaining animation generation data includes:

[0013] Get multiple comic images input by the user;

[0014] A sequence number of multiple comic images is set according to the sequence association parameters of the comic images. The sequence association parameters include at least one of the following: image content order, user-specified display order, image generation order, and image input order.

[0015] The multiple comic images are combined into the static comic image sequence based on the sequence number.

[0016] In some embodiments, obtaining animation generation data includes:

[0017] Obtain text data input by the user based on the static comic image sequence, the text data including character dialogue text and psychological activity text;

[0018] Read the grouping parameters of the text data, wherein the grouping parameters include at least one of screen state parameters, input state parameters, and specified grouping parameters;

[0019] Query the associated image frames of the text data based on the grouping parameters;

[0020] The text data is grouped based on the associated image frames to generate the dubbing text.

[0021] In some embodiments, obtaining animation generation data includes:

[0022] Acquire sample audio data, which is voice-over audio recorded for the target character in the comic image;

[0023] Denoising processing is performed on the sample audio data;

[0024] Extract sample speech segments from the denoised sample audio data. The sample speech segments are audio segments whose audio energy value is greater than or equal to a preset energy value and whose duration is greater than or equal to a preset duration.

[0025] Extract the voiceprint features of the target character from the sample speech segments.

[0026] In some embodiments, a frame animation generation model is used to generate a dynamic video based on the static comic image sequence, including:

[0027] Invoke the frame animation generation model;

[0028] The static comic image sequence is input into the frame animation generation model to perform layering processing on the comic image through the frame animation generation model to obtain a layering result, which includes a foreground layer, a background layer, and a character layer.

[0029] Based on the physics engine, the target motion patterns in the foreground and character layers of the layered results are identified;

[0030] Simulate the motion trajectory based on the target motion mode;

[0031] Intermediate frames are generated according to the motion trajectory, and the target characters in the character layer in the multiple intermediate frames are arranged according to the motion trajectory.

[0032] In some embodiments, generating intermediate frames according to the motion trajectory includes:

[0033] The importance of multiple comic images in the static comic image sequence is evaluated to obtain importance information;

[0034] Based on the importance information, differential compression is performed on multiple comic images to obtain multiple reference images;

[0035] Based on the motion trajectory, the start image and the end image are extracted from the multiple reference images;

[0036] A bidirectional context sampling method is used to extract image features from the starting image and the ending image respectively to generate intermediate frame images;

[0037] Phase alignment is performed on the intermediate frame image to obtain dynamic video.

[0038] In some embodiments, a speech synthesis model is used to generate character speech based on the dubbed text and the character's voiceprint features, including:

[0039] The voiceprint feature parameters of the character's voiceprint feature are obtained, including Mel frequency cepstral coefficients and linear prediction cepstral coefficients;

[0040] Based on the voiceprint feature parameters, a character identifier is bound to the character's voiceprint feature;

[0041] Extract the character's voice text from the dubbing text based on the character identifier;

[0042] Invoke the speech synthesis model;

[0043] The character's voice text and voiceprint features are input into the speech synthesis model to generate the character's voice.

[0044] In some embodiments, the character's voice text and the character's voiceprint features are input into the speech synthesis model to generate the character's voice through the speech synthesis model, including:

[0045] The speech text of the character is encoded using the speech synthesis model to obtain word vectors.

[0046] Based on the word vectors, the semantic information of the dubbed text is identified;

[0047] Based on the semantic information, prosodic features and sentiment parameters are added to the word vectors to generate text features;

[0048] The text features are fused with the character's voiceprint features to obtain the character's voice.

[0049] In some embodiments, timestamps are attached to the character's voice based on the comic images corresponding to keyframes in the dynamic video to generate mixed audio, including:

[0050] Define an audio-visual synchronization function;

[0051] The dynamic video and the character's voice are obtained through the audio-visual synchronization function, and the character's voice includes multiple synthesized voice segments.

[0052] Identify keyframes in the dynamic video, wherein the keyframes include multiple cartoon images in the static cartoon image sequence;

[0053] Extract the playback time of the keyframe in the dynamic video, and set the timestamp of the synthesized speech segment based on the playback time;

[0054] Based on the timestamp, multiple synthesized speech segments are merged into a mixed audio.

[0055] According to another aspect of this application, a comic animation and speech synthesis system based on multimodal generation is provided, the system comprising:

[0056] A multimodal input module is used to acquire animation generation data, which includes a static comic image sequence, voice-over text, and character voiceprint features; the static comic image sequence includes multiple comic images.

[0057] The video generation module is used to generate dynamic videos from the static comic image sequence using a frame animation generation model. The frame animation generation model is a neural network model based on a diffusion transformer architecture and a self-attention mechanism architecture. The frame animation generation model is configured to perform hierarchical processing on the comic images, motion trajectory simulation based on a physics engine, and add intermediate frames.

[0058] The speech generation module is used to generate character speech based on the dubbing text and the character's voiceprint features using a speech synthesis model. The speech synthesis model is a neural network model based on deep learning. The speech synthesis model is configured to bind the character's voiceprint features to the character's identifier and to perform emotional parameter injection on the character's speech based on semantic analysis of the dubbing text.

[0059] The audio-visual synchronization module is used to bind timestamps to the character's voice based on the comic images corresponding to key frames in the dynamic video, so as to generate mixed audio;

[0060] An output module is used to combine the mixed audio and the dynamic video into an output animation.

[0061] According to another aspect of this application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described method for animation of comics and speech synthesis based on multimodal generation.

[0062] According to another aspect of this application, a storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the above-described method for animation of comics and speech synthesis based on multimodal generation.

[0063] By employing the above technical solutions, embodiments of this application provide a method and system for comic animation and speech synthesis based on multimodal generation. The method, after acquiring animation generation data, can use a frame animation generation model to generate dynamic video based on the static comic image sequence in the animation generation data, and then use a speech synthesis model. Furthermore, it binds timestamps to the character's voice based on the comic images corresponding to keyframes in the dynamic video to generate mixed audio. The mixed audio and dynamic video are then synthesized into an output animation. The method can perform layered processing on comic images through the frame animation generation model, simulate motion trajectories based on a physics engine, and add intermediate frames. It also binds character voiceprint features and character identifiers through the speech synthesis model, and injects emotional parameters into the character's voice based on semantic analysis of the dubbing text. This can shorten the processing time for converting comic images into animation, improve the processing efficiency of the comic animation synthesis process, and solve the problem of low processing efficiency in the comic animation synthesis process.

[0064] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0065] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0066] Figure 1 This is a schematic diagram of the video synthesis process provided in an embodiment of this application;

[0067] Figure 2 This is a schematic diagram of the method for comic animation and speech synthesis based on multimodal generation provided in an embodiment of this application;

[0068] Figure 3 This is a schematic diagram of the process for generating dynamic video provided in an embodiment of this application;

[0069] Figure 4 This is a schematic diagram of the character voice generation process provided in an embodiment of this application;

[0070] Figure 5 This is a schematic diagram of the audio synthesis and mixing process provided in the embodiments of this application;

[0071] Figure 6 This application provides a system architecture diagram for generating output animations.

[0072] Figure 7 This is a schematic diagram of the model training process provided in the embodiments of this application;

[0073] Figure 8 This is a schematic diagram of the structure of a comic animation and speech synthesis system based on multimodal generation provided in an embodiment of this application. Detailed Implementation

[0074] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.

[0075] In this embodiment, the animation of comics and speech synthesis based on multimodal generation is a video synthesis process. Wherein, as... Figure 1 As shown, video compositing is the process of combining multiple multimedia elements such as images, video clips, and audio clips to create a dynamic video work.

[0076] Taking the comic book video compositing process as an example, animators need to draw multiple frames of comic book images and play them sequentially according to the action sequence to create a dynamic comic book video. Since animation production involves compositing static comic book images into a dynamic animated video, the comic book video compositing process is also known as comic book mobilization. After creating the dynamic comic book video, it also needs to be dubbed; that is, voice actors and sound effects companies record audio clips, which are then embedded into the comic book video to create a video animation containing sound.

[0077] Because animation production requires manually drawing multiple frames and special effects images, the cost of animation production is relatively high. Therefore, in some embodiments, animation generation applications can be used to assist in video frame creation during the animation production process. That is, multiple comic images are drawn as keyframes for the dynamic video, and intermediate transition frames are drawn based on the animation generation application to synthesize a dynamic comic video.

[0078] However, because the audio recording for animated videos requires manual recording, it's difficult to guarantee the consistency of voiceprints for the characters. Furthermore, embedding audio clips into the animated video requires manual alignment of keyframes with the audio track to ensure audio-visual synchronization. Therefore, the animation compositing process for comics is inefficient, lacks automation, and prolongs the animation production cycle.

[0079] To address the low processing efficiency of comic animation synthesis, this application provides a method for comic animation and speech synthesis based on multimodal generation in some embodiments. This method can be applied to electronic devices with data processing capabilities. These electronic devices include, but are not limited to, computers, servers, mobile terminals, smart wearable devices, and industrial control machines. For ease of description, this application uses an electronic device as the execution subject of the method in its embodiments. It should be understood that the method can also be applied to other types of execution subjects, which are not illustrated in all embodiments of this application. Figure 2 As shown, the method includes:

[0080] S101. Obtain animation generation data.

[0081] To animate comics, electronic devices can first acquire animation generation data, which refers to a combination of multimodal data used to generate animated videos. Animation generation data can include static comic image sequences, voice-over text, and character voiceprint features.

[0082] A static comic image sequence can include multiple comic images. These images can be drawn by the user in real time; in this case, the electronic device can display a drawing interface to acquire them. The user can perform drawing operations based on the drawing interface to generate a comic image in image format. After drawing is complete, the electronic device can acquire the comic image.

[0083] In some embodiments, the comic image can also be image data uploaded by the user. That is, the electronic device can display the application interface of an animation generation application, which may include an image upload control. Users can click the image upload control to specify a file upload path, thereby uploading the comic image stored on a specific storage medium to the electronic device, enabling the electronic device to access the comic image.

[0084] In some embodiments, the comic image can also be obtained by taking a screenshot of a specific interface. That is, when an electronic device displays an interface containing comic content, a screenshot tool can be invoked in response to a user's input of a screenshot operation, and a screenshot of the currently displayed interface can be taken using the screenshot tool to obtain a comic image containing comic content.

[0085] It should be noted that, in addition to the methods for acquiring comic images described above, electronic devices can also acquire comic images through other means. For example, generating comic images using a text-based image model, acquiring comic images from a network resource database, or extracting frames from animated videos. Other methods for acquiring comic images that are conceived by those skilled in the art based on the above methods also fall within the scope of protection of this application.

[0086] Multiple comic images can be arranged in a specific order to form a static comic image sequence. Therefore, in some embodiments, when acquiring animation generation data, the electronic device can first acquire multiple comic images input by the user, then set sequence numbers for the multiple comic images according to the sequence association parameters of the comic images, and combine the multiple comic images into the static comic image sequence based on the sequence numbers.

[0087] The sequence association parameter refers to a parameter that determines the order in which multiple comic images are arranged in a static comic image sequence. Specifically, the sequence association parameter includes at least one of the following: image content order, user-specified display order, image generation order, and image input order. The image content order refers to the sequence association parameter determined by image recognition of the specific content within the multiple comic images.

[0088] For example, by performing image recognition on multiple comic book images and determining that the content corresponding to the images is a number displayed by a comic book character, the order of the images can be determined based on the numerical values ​​displayed in the images. If image A displays the number 1, image B displays the number 3, image C displays the number 2, and image D displays the number 4, then the order of the images can be determined as image A, image C, image B, image D.

[0089] User-specified display order means that when a user inputs multiple comic images, they can specify the display order for each comic individually. For example, when inputting image A, the display order can be specified as 1, and when inputting image B, the display order can be specified as 2, and so on.

[0090] Image generation order refers to the arrangement order determined by the user's real-time drawing process of comic images. For example, if the user draws image A first, and then draws image B, the electronic device can determine that image A takes precedence over image B. Similarly, image input order refers to the arrangement order determined by the upload order when the user uploads multiple comic images sequentially.

[0091] Therefore, after acquiring multiple comic images input by the user, the electronic device can sort the multiple comic images based on at least one sequence association parameter in the above example, that is, determine the sequence number of the multiple comic images, and then combine the multiple comic images according to the sequence number to obtain a static comic image sequence.

[0092] Dubbing text is text data input by a user during the input of a static comic image sequence, used to represent the dialogue and inner thoughts of a character. Dubbing text can be used to generate character voiceovers. Therefore, dubbing text can be grouped according to multiple comic images in the static comic image sequence. That is, in some embodiments, the electronic device can first acquire the text data input by the user based on the static comic image sequence when acquiring animation generation data. This text data includes character dialogue text and inner thoughts text. Then, grouping parameters of the text data are read, and associated image frames are queried according to the grouping parameters, thereby grouping the text data based on the associated image frames to generate dubbing text.

[0093] The grouping parameters refer to parameters that can affect the frame-by-frame grouping results of character dialogue text or psychological activity text. Grouping parameters can include at least one of screen state parameters, input state parameters, and specified grouping parameters. Screen state parameters refer to parameters that can represent the content of a comic book image, such as lip movements, emotions, and scene recognition results.

[0094] Input state parameters are parameters used to characterize the state when a user inputs text data. For example, input state parameters may include input nodes, input time, etc. That is, as the user draws a comic, each time the user draws a comic image, they can input the character's dialogue text or inner thoughts text through the text input control. The electronic device can then obtain the input state parameters by reading the relevant parameters of the text input process.

[0095] Specifying grouping parameters refers to the association relationship that the user specifies between comic images and text data. For example, the user can specify the text data corresponding to image A as TX1, and the electronic device can directly determine the specified grouping parameters based on the user-specified association relationship to facilitate grouping the text data by frame.

[0096] After acquiring text data input by the user based on the static comic image sequence, the electronic device can determine the comic image to which each text data or text fragment belongs based on at least one of the above grouping parameters, thereby realizing the grouping of text data by frame to obtain dubbing text.

[0097] Voiceprint features are characteristic parameters extracted from speech signals that can characterize the speaker's identity. Voiceprint features can include spectrum, cepstral, linear prediction coefficients, pitch, tone, formants, timbre, phonetics, idioms, etc.

[0098] Electronic devices can obtain voiceprint features of characters by extracting features from dubbed audio. Specifically, in some embodiments, when acquiring animation generation data, the electronic device can first acquire sample audio data, which is dubbing audio recorded for the target character in the comic image. Then, denoising processing is performed on the sample audio data, and sample speech segments are extracted from the denoised sample audio data. Finally, the voiceprint features of the target character are extracted from the sample speech segments.

[0099] The sample audio segments are defined as audio segments with an audio energy value greater than or equal to a preset energy value and a duration greater than or equal to a preset duration. For example, to extract a character's voiceprint features, sample audio data of the target character reading aloud according to preset text content can be obtained. Then, the sample audio data is denoised to improve data quality. Audio segments are then extracted from the sample audio data. If the audio energy value of a speech segment in the target audio data is greater than or equal to a preset energy value and the duration is greater than or equal to a preset duration, the speech segment is cut out. Based on the cut-out speech segment, the voiceprint features of the target character are extracted.

[0100] In some embodiments, when an electronic device extracts the voiceprint features of a target character from the sample speech segments, it can utilize the core processing unit in the circuit to run feature extraction algorithms, such as Mel-Frequency Cepstral Coefficients (MFCC), to extract feature parameters from the speech samples and normalize the extracted feature parameters to improve the accuracy of recognition.

[0101] In some embodiments, the electronic device can also extract voiceprint features from sample speech segments based on a voiceprint feature extraction model. That is, the electronic device can first train the voiceprint feature extraction model based on voiceprint training data. After obtaining sample speech segments, the sample speech segments can be input into the voiceprint feature extraction model. The voiceprint feature extraction model can perform preprocessing operations on the sample speech segments, such as pre-emphasis, framing, and windowing. Then, a Fast Fourier Transform is performed on the preprocessed audio frames to convert them from time-domain signals to frequency-domain signals. A Mel filter is then used to filter the audio frames, achieving smoothing and eliminating harmonics in multiple frequency domains, thus obtaining the Mel spectral features corresponding to the audio frames. By taking the logarithm of the Mel spectral features, the sample static features can be obtained.

[0102] S102. Using a frame animation generation model, generate a dynamic video based on the static comic image sequence.

[0103] After acquiring the animation generation data, electronic devices can perform dynamic processing of comics based on the animation generation data. That is, electronic devices can use frame animation generation models to generate dynamic videos based on the static comic image sequences in the animation generation data.

[0104] The frame animation generation model is a neural network model based on a diffusion transformer architecture and a self-attention mechanism. For example, the frame animation generation model can be a neural network model based on the FramePack algorithm. The FramePack algorithm is an efficient video generation algorithm that can generate videos based on progressive compression mechanisms and anti-drift sampling methods.

[0105] Frame animation generation models can utilize neural network models based on diffusion transformer (DiT) architectures. For example, the FramePack algorithm uses the HunyuanVideo model as its base model. HunyuanVideo is a model with 13 billion parameters that addresses forgetting and drift issues in video generation by compressing the context length of input frames. The diffusion transformer model structure allows frame animation generation models to efficiently process a large number of image frames while maintaining low computational complexity.

[0106] When generating dynamic video, the electronic device applies a frame animation generation model to perform layered processing, motion trajectory simulation based on a physics engine, and intermediate frame addition on the comic images. Therefore, in some embodiments, when the electronic device uses the frame animation generation model to generate dynamic video from the static comic image sequence, it can first call the frame animation generation model and then input the static comic image sequence into the frame animation generation model to perform layered processing on the comic images. Through layered processing, the frame animation generation model can separate a layered result from the comic image, including a foreground layer, a background layer, and a character layer.

[0107] After obtaining the layered results, the electronic device can also identify the target motion patterns of the foreground and character layers in the layered results based on the physics engine, simulate motion trajectories according to the target motion patterns, and then generate intermediate frames according to the motion trajectories. In these intermediate frames, the target characters in the character layer are arranged according to the motion trajectories.

[0108] like Figure 3As shown, in order to generate intermediate frames, in some embodiments, when generating intermediate frames according to the motion trajectory, the electronic device can first evaluate the importance of the comic images based on an intelligent importance assessment algorithm, that is, by evaluating the importance of multiple comic images in the static comic image sequence to obtain importance information. Then, based on the importance information, differential compression is performed on the multiple comic images to obtain multiple reference images.

[0109] To this end, frame animation generation models can include a frame compression processing module, which allows electronic devices to progressively compress multiple comic images in a static comic image sequence. For example, a frame animation generation model based on the FramePack algorithm can differentially compress frames according to their importance. For less important frames, a higher compression ratio is applied to reduce memory usage while preserving key visual information. For instance, the most recent frame might retain more tokens, while more distant frames are compressed to fewer tokens. This geometric compression method ensures that the total context length converges to a fixed upper limit.

[0110] During progressive compression, the frame animation generation model can also provide specialized processing for the tail frame. For example, for the last comic image in a static comic image sequence, since it serves as the tail frame of a dynamic video, the frame animation generation model can allow each tail frame to increase the context length by one potential pixel, or perform global average pooling on all tail frames and process them with the largest kernel.

[0111] The frame animation generation model may also include an anti-drift sampling processing module. After obtaining multiple reference images, the electronic device can invoke the anti-drift sampling processing module of the frame animation generation model so that the anti-drift sampling processing module can extract a start image and an end image from the multiple reference images based on the motion trajectory. The start image and end image can be determined based on a single motion step in the motion trajectory; that is, the first cartoon image of each motion step is used as the start image, and the last cartoon image of each motion step is used as the end image.

[0112] The electronic device then uses bidirectional context sampling to extract image features from the start and end images, respectively, to generate intermediate frame images. For example, a frame animation generation model can employ a bidirectional context anti-drift sampling method, generating both the start and end portions simultaneously in the first iteration, with subsequent iterations filling in the gaps. This bidirectional context anti-drift sampling method ensures the end frame is determined from the outset, causing subsequent generated frames to converge towards this target, thus preventing drift.

[0113] In some embodiments, the electronic device may also employ reverse drift-resistant sampling to generate intermediate frame images. For example, in the process of generating dynamic video, the frame animation generation model may use the reverse drift-resistant sampling method, taking the user input image as a high-quality first frame, and then generating subsequent frames in reverse time order, continuously optimizing the generated frames to approximate the user input first frame, thereby generating a high-quality video.

[0114] The frame animation generation model may also include an alignment module. After generating multiple intermediate frame images, the frame animation generation model can perform phase alignment on the intermediate frame images through the alignment module to obtain dynamic video. For example, the alignment module can perform Rotary Position Embedding Alignment (RoPE) processing on multiple intermediate frame images. Since the input context lengths of different compression kernels are different, the frame animation generation model needs to perform RoPE processing. RoPE processing can generate a complex number with real and imaginary parts for each marked position, called the "phase". In order to match the compressed RoPE encoding, the frame animation generation model can directly use average pooling to downsample the RoPE phase to match the compression kernel.

[0115] S103. Using a speech synthesis model, generate the character's voice based on the dubbing text and the character's voiceprint features.

[0116] Electronic devices can generate character voices simultaneously with dynamic videos. Specifically, they can use a speech synthesis model to generate character voices based on the dubbing text in the animation generation data and the character's voiceprint features.

[0117] The speech synthesis model is a deep learning-based neural network model. For example, the text-to-speech (TTS) model can utilize open-source models such as CosyVoice TTS to achieve high-quality, multilingual, and low-latency speech synthesis. The speech synthesis model can be based on diffusion models and the Transformer architecture for text processing and speech generation. Diffusion models can generate data by progressively removing noise, resulting in high-quality speech samples. The Transformer architecture effectively captures long-distance dependencies between text and speech, ensuring the coherence and consistency of the generated speech. The speech synthesis model can also employ flow matching techniques, generating new samples by learning the distribution of data when decoding text tokens into audio, thereby improving the quality and diversity of speech generation.

[0118] Electronic devices can bind character voiceprint features and character identifiers based on a speech synthesis model, and perform emotion parameter injection on the character's speech based on semantic analysis of the dubbing text. Therefore, in some embodiments, when an electronic device generates character speech using a speech synthesis model based on the dubbing text and the character's voiceprint features, it can first obtain the voiceprint feature parameters of the character's voiceprint features. These voiceprint feature parameters include Mel-frequency cepstral coefficients and linear predictive cepstral coefficients (LPCCs).

[0119] After acquiring voiceprint feature parameters, the electronic device can bind a character identifier to the character's voiceprint features based on these parameters, and extract the character's voice text from the dubbing text based on the character identifier. For example, the voiceprint feature parameters can be used to determine the character's voice type, where the character's voice type refers to voiceprint features with the same voiceprint feature parameters. Then, an association is established between this character's voice type and a specified character in a comic book image, thus binding the character identifier to the character's voiceprint features.

[0120] After binding the character identifier to the character's voiceprint features, the electronic device can also call the speech synthesis model and input the character's voice text and the character's voiceprint features into the speech synthesis model to generate the character's voice through the speech synthesis model.

[0121] like Figure 4 As shown, in some embodiments, to generate character voice, the electronic device, during the process of inputting the character voice text and the character voiceprint features into the speech synthesis model, can perform text encoding on the character voice text through the speech synthesis model to obtain word vectors. Then, based on the word vectors, the semantic information of the dubbing text is identified, and prosodic features and emotional parameters are added to the word vectors according to the semantic information to generate text features. The text features are then fused with the character voiceprint features to obtain the character voice.

[0122] For example, in the process of generating character voices, the speech synthesis model can perform text processing on the dubbing text and audio processing on the character's voiceprint features. For the text processing, after the electronic device calls the speech synthesis model, it can first perform text preprocessing on the dubbing text, that is, preprocess the input text, including text cleaning, word segmentation, part-of-speech tagging, and other operations.

[0123] For the pre-processed dubbing text, the speech synthesis model can perform text encoding, using techniques such as word embedding to convert the pre-processed text into a numerical representation that the model can understand. This maps each word to a fixed-dimensional vector space, allowing the semantic information in the text to be captured through word vectors. In other words, the speech synthesis model can utilize a pre-trained Language Model (LLM) to perform semantic understanding of the text. LLMs can deeply understand the meaning, contextual relationships, and sentiment of the text, thus providing more accurate semantic guidance for speech synthesis.

[0124] After semantic understanding, the speech synthesis model can also perform prosodic modeling and emotion parameter injection. Prosodic feature modeling can include intonation, speech rate, pauses, etc. That is, it models the prosody of speech through autoregression to make the generated speech more natural and fluent. Emotion parameter injection, on the other hand, guides the model to generate speech with specific emotional coloring through specific cues or parameters.

[0125] In audio processing, speech synthesis models can extract voiceprint features from reference audio using a single timbre encoder. The timbre encoder analyzes the spectral characteristics, formants, and other information of the reference audio, encoding them into a compact feature vector to represent the voice characteristics of a specific speaker. After extracting voiceprint features, the speech synthesis model can also fuse these features with text features to generate speech that matches the text content. During speech generation, the model also mimics the timbre and style of a reference speaker by using their voiceprint features.

[0126] S104. Based on the comic images corresponding to the key frames in the dynamic video, timestamps are bound to the character's voice to generate mixed audio.

[0127] After generating dynamic videos using a frame animation generation model and generating character voices using a speech synthesis model, electronic devices can synchronize the generated dynamic videos and character voices, that is, bind timestamps to the character voices based on the comic images corresponding to key frames in the dynamic videos to generate mixed audio.

[0128] like Figure 5 As shown, in some embodiments, in order to perform audio-visual synchronization, when the electronic device binds timestamps to the character's voice based on the cartoon images corresponding to keyframes in the dynamic video to generate mixed audio, it can first define an audio-visual synchronization function and obtain the dynamic video and the character's voice through the audio-visual synchronization function. The character's voice includes multiple synthesized voice segments.

[0129] For example, such as Figure 6As shown, electronic devices can synchronize audio clips with video using Python functions. Using the moviepy library for video editing, the moviepy.editor module can be imported using "import moviepy.editor as mpe" and renamed to mpe for later use. Then, a function named sync_audio_video for audio-video synchronization can be defined using "def sync_audio_video(video_path, audio_segments:"), which accepts two parameters: the path to the video file "video_path" and a list of audio clips "audio_segments".

[0130] After defining the audio-visual synchronization function, the electronic device can identify keyframes in the dynamic video. These keyframes include multiple cartoon images from the static cartoon image sequence. The playback time of each keyframe in the dynamic video is extracted, and a timestamp for the synthesized speech segment is set based on this playback time. Then, multiple synthesized speech segments are merged into a mixed audio based on the timestamps.

[0131] For example, an electronic device can load a video file using `mpe.VideoFileClip` by executing `video = mpe.VideoFileClip(video_path); audio_tracks = []`, and create an empty list `audio_tracks` to store audio segments. Then, it creates an `AudioClip` for each audio segment and binds a timestamp to it using `for seg in audio_segments:audio = mpe.AudioFileClip(seg.audio_path).set_start(seg.start_time); audio_tracks.append(audio)`. This is done by iterating through each audio segment (seg) in the `audio_segments` list. For each audio segment, an `AudioFileClip` object is created, and the `set_start` method is used to set the start time of the audio segment (seg.start_time). Finally, this audio segment is added to the `audio_tracks` list.

[0132] S105. Combine the mixed audio with the dynamic video to create an output animation.

[0133] By setting timestamps and combining the audio tracks, electronic devices can combine the mixed audio with dynamic video to create an output animation. For example, an electronic device can combine audio tracks and synthesize video by executing `final_audio = mpe.CompositeAudioClip(audio_tracks); final_video = video.set_audio(final_audio)`. Specifically, `mpe.CompositeAudioClip` merges all audio clips into a single audio track (`final_audio`). Then, the `set_audio` method is used to set this audio track as the audio for the video, generating the final video (`final_video`).

[0134] After mixing the audio tracks and compositing the video, the electronic device can output the composite animated video. For example, the electronic device can use the `write_videofile` method to write the final video to the file "output.mp4" by executing "return final_video.write_videofile("output.mp4",codec='libx264')" and specifying the use of the libx264 codec.

[0135] Therefore, electronic devices can synchronize multiple audio clips with a single video file and generate a new video file. After loading the video file, an AudioFileClip object is created for each audio clip and its start time is set. Then, all audio clips are merged into a single audio track, which is set as the audio for the video, to output a new video file.

[0136] By applying the technical solutions of the above embodiments, the comic animation and speech synthesis method based on multimodal generation provided in the above embodiments can perform layered processing on comic images through frame animation generation models, simulate motion trajectories based on physics engines, and add intermediate frame processing. It can also bind character voiceprint features and character identifiers through speech synthesis models, and inject emotional parameters into character speech based on semantic analysis of dubbing text. This can shorten the processing time for converting comic images into animation, improve the processing efficiency of comic animation synthesis, and solve the problem of low processing efficiency in comic animation synthesis.

[0137] In some embodiments, as a refinement and extension of the specific implementation of the above embodiments, and in order to fully illustrate the specific implementation process of this embodiment, some embodiments of this application also provide a method for comic animation and speech synthesis based on multimodal generation, such as... Figure 7 As shown, the method includes:

[0138] S201. Obtain training sample data, wherein the training sample data includes sample image data and sample speech data;

[0139] S202. Call the trained model, which includes a frame animation generation model and a speech synthesis model;

[0140] S203. Perform model training on the trained model using the training sample data, including: training the frame animation generation model using the sample image data, and training the speech synthesis model using the sample speech data;

[0141] S204. When the trained model satisfies the convergence condition, output the model parameters of the trained model.

[0142] Electronic devices can train frame animation generation models and speech synthesis models based on training sample data. During training, sample image data can be input into the frame animation generation model being trained, and sample speech data can be input into the speech synthesis model being trained, respectively, to obtain the model output data. The training loss is then calculated based on the output data. If the training loss exceeds a loss threshold, it indicates that the current trained model does not meet the convergence condition. Therefore, the model parameters can be modified based on the training loss, and the output data can be re-obtained and the training loss recalculated by re-inputting training sample data, thus achieving iterative training.

[0143] When the training loss is less than or equal to the loss threshold, or when the number of iterations reaches the preset number of iterations threshold, it means that the current trained model has been trained to convergence, that is, the trained model meets the convergence condition. At this time, the model parameters of the trained model can be output to obtain a frame animation generation model and a speech synthesis model with a certain output accuracy.

[0144] By applying the technical solutions of the above embodiments, the comic animation generation and speech synthesis method based on multimodal generation provided in the above embodiments can perform model training through electronic devices, so as to improve the frame animation generation model and speech synthesis model through continuous iterative training, thereby improving the output accuracy of the model and the accuracy of comic animation generation and speech synthesis.

[0145] In some embodiments, as a specific implementation of the multimodal generation-based comic animation and speech synthesis method described in the above embodiments, some embodiments of this application also provide a multimodal generation-based comic animation and speech synthesis system, such as... Figure 8 As shown, the system includes:

[0146] A multimodal input module is used to acquire animation generation data, which includes a static comic image sequence, voice-over text, and character voiceprint features; the static comic image sequence includes multiple comic images.

[0147] The video generation module is used to generate dynamic videos from the static comic image sequence using a frame animation generation model. The frame animation generation model is a neural network model based on a diffusion transformer architecture and a self-attention mechanism architecture. The frame animation generation model is configured to perform hierarchical processing on the comic images, motion trajectory simulation based on a physics engine, and add intermediate frames.

[0148] The speech generation module is used to generate character speech based on the dubbing text and the character's voiceprint features using a speech synthesis model. The speech synthesis model is a neural network model based on deep learning. The speech synthesis model is configured to bind the character's voiceprint features to the character's identifier and to perform emotional parameter injection on the character's speech based on semantic analysis of the dubbing text.

[0149] The audio-visual synchronization module is used to bind timestamps to the character's voice based on the comic images corresponding to key frames in the dynamic video, so as to generate mixed audio;

[0150] An output module is used to combine the mixed audio and the dynamic video into an output animation.

[0151] By applying the technical solutions of the above embodiments, this application provides a method and system for comic animation and speech synthesis based on multimodal generation. The method, after acquiring animation generation data, can use a frame animation generation model to generate dynamic video based on the static comic image sequence in the animation generation data, and use a speech synthesis model. Then, based on the comic images corresponding to keyframes in the dynamic video, timestamps are bound to the character's voice to generate mixed audio. The mixed audio and dynamic video are then synthesized into an output animation. This method can perform layered processing on comic images through the frame animation generation model, simulate motion trajectories based on a physics engine, and add intermediate frames. It can also bind character voiceprint features and character identifiers through the speech synthesis model, and inject emotional parameters into the character's voice based on semantic analysis of the dubbing text. This can shorten the processing time for converting comic images into animation, improve the processing efficiency of the comic animation synthesis process, and solve the problem of low processing efficiency in the comic animation synthesis process.

[0152] It should be noted that other corresponding descriptions of the functional units involved in the multimodal generation-based comic animation and speech synthesis system provided in this application embodiment can be found in the corresponding descriptions in the multimodal generation-based comic animation and speech synthesis method provided in the above embodiments, and will not be repeated here.

[0153] This application also provides a computer device, specifically a personal computer, server, network device, etc. The computer device includes a bus, processor, memory, and communication interface, and may also include input / output interfaces and a display device. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores location information. The network interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the various method embodiments.

[0154] Those skilled in the art will understand that the structure of the computer device described above is only a partial structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. A specific computer device may include more or fewer components, or combine certain components, or have different component arrangements.

[0155] In one embodiment, a computer-readable storage medium is also provided, which may be non-volatile or volatile, and a computer program is stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0156] In one embodiment, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0157] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0158] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods.

[0159] Any references to memory, database, or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc.

[0160] Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can take many forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0161] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchain. The processors involved in the embodiments provided in this application may be, but are not limited to, general-purpose processors, graphics processors, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc.

[0162] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0163] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for comic animation and speech synthesis based on multimodal generation, characterized in that, The method includes: Acquire animation generation data, which includes a static comic image sequence, voice-over text, and character voiceprint features; the static comic image sequence includes multiple comic images. A frame animation generation model is used to generate dynamic videos from the static comic image sequence. The frame animation generation model is a neural network model based on a diffusion transformer architecture and a self-attention mechanism architecture. The frame animation generation model is configured to perform hierarchical processing on the comic images, motion trajectory simulation based on a physics engine, and add intermediate frames. Using a speech synthesis model, character speech is generated based on the dubbing text and the character's voiceprint features. The speech synthesis model is a deep learning-based neural network model. The speech synthesis model is configured to bind the character's voiceprint features to the character's identifier and to perform emotional parameter injection on the character's speech based on semantic analysis of the dubbing text. The character's voice is time-stamped based on the comic images corresponding to keyframes in the dynamic video to generate mixed audio; The mixed audio and the dynamic video are combined to form an output animation.

2. The method according to claim 1, characterized in that, Obtain animation generation data, including: Get multiple comic images input by the user; A sequence number of multiple comic images is set according to the sequence association parameters of the comic images. The sequence association parameters include at least one of the following: image content order, user-specified display order, image generation order, and image input order. The multiple comic images are combined into the static comic image sequence based on the sequence number.

3. The method according to claim 1, characterized in that, Obtain animation generation data, including: Obtain text data input by the user based on the static comic image sequence, the text data including character dialogue text and psychological activity text; Read the grouping parameters of the text data, wherein the grouping parameters include at least one of screen state parameters, input state parameters, and specified grouping parameters; Query the associated image frames of the text data based on the grouping parameters; The text data is grouped based on the associated image frames to generate the dubbing text.

4. The method according to claim 1, characterized in that, Obtain animation generation data, including: Acquire sample audio data, which is voice-over audio recorded for the target character in the comic image; Denoising processing is performed on the sample audio data; Extract sample speech segments from the denoised sample audio data. The sample speech segments are audio segments whose audio energy value is greater than or equal to a preset energy value and whose duration is greater than or equal to a preset duration. Extract the voiceprint features of the target character from the sample speech segments.

5. The method according to claim 1, characterized in that, Using a frame animation generation model, a dynamic video is generated based on the static comic image sequence, including: Invoke the frame animation generation model; The static comic image sequence is input into the frame animation generation model to perform layering processing on the comic image through the frame animation generation model to obtain a layering result, which includes a foreground layer, a background layer, and a character layer. Based on the physics engine, the target motion patterns in the foreground and character layers of the layered results are identified; Simulate the motion trajectory based on the target motion mode; Intermediate frames are generated according to the motion trajectory, and the target characters in the character layer in the multiple intermediate frames are arranged according to the motion trajectory.

6. The method according to claim 5, characterized in that, Generating intermediate frames according to the motion trajectory includes: The importance of multiple comic images in the static comic image sequence is evaluated to obtain importance information; Based on the importance information, differential compression is performed on multiple comic images to obtain multiple reference images; Based on the motion trajectory, the start image and the end image are extracted from the multiple reference images; A bidirectional context sampling method is used to extract image features from the starting image and the ending image respectively to generate intermediate frame images; Phase alignment is performed on the intermediate frame image to obtain dynamic video.

7. The method according to claim 1, characterized in that, Using a speech synthesis model, the character's voice is generated based on the dubbed text and the character's voiceprint features, including: The voiceprint feature parameters of the character's voiceprint feature are obtained, including Mel frequency cepstral coefficients and linear prediction cepstral coefficients; Based on the voiceprint feature parameters, a character identifier is bound to the character's voiceprint feature; Extract the character's voice text from the dubbing text based on the character identifier; Invoke the speech synthesis model; The character's voice text and voiceprint features are input into the speech synthesis model to generate the character's voice.

8. The method according to claim 7, characterized in that, The character's voice text and voiceprint features are input into the speech synthesis model to generate the character's voice, including: The speech text of the character is encoded using the speech synthesis model to obtain word vectors. Based on the word vectors, the semantic information of the dubbed text is identified; Based on the semantic information, prosodic features and sentiment parameters are added to the word vectors to generate text features; The text features are fused with the character's voiceprint features to obtain the character's voice.

9. The method according to claim 1, characterized in that, The process involves binding timestamps to the character's voice based on the comic images corresponding to keyframes in the dynamic video to generate mixed audio, including: Define an audio-visual synchronization function; The dynamic video and the character's voice are obtained through the audio-visual synchronization function, and the character's voice includes multiple synthesized voice segments. Identify keyframes in the dynamic video, wherein the keyframes include multiple cartoon images in the static cartoon image sequence; Extract the playback time of the keyframe in the dynamic video, and set the timestamp of the synthesized speech segment based on the playback time; Based on the timestamp, multiple synthesized speech segments are merged into a mixed audio.

10. A comic animation and speech synthesis system based on multimodal generation, characterized in that, The system includes: A multimodal input module is used to acquire animation generation data, which includes a static comic image sequence, voice-over text, and character voiceprint features; the static comic image sequence includes multiple comic images. The video generation module is used to generate dynamic videos from the static comic image sequence using a frame animation generation model. The frame animation generation model is a neural network model based on a diffusion transformer architecture and a self-attention mechanism architecture. The frame animation generation model is configured to perform hierarchical processing on the comic images, motion trajectory simulation based on a physics engine, and add intermediate frames. The speech generation module is used to generate character speech based on the dubbing text and the character's voiceprint features using a speech synthesis model. The speech synthesis model is a neural network model based on deep learning. The speech synthesis model is configured to bind the character's voiceprint features to the character's identifier and to perform emotional parameter injection on the character's speech based on semantic analysis of the dubbing text. The audio-visual synchronization module is used to bind timestamps to the character's voice based on the comic images corresponding to key frames in the dynamic video, so as to generate mixed audio; An output module is used to combine the mixed audio and the dynamic video into an output animation.