Video generation method and device, equipment and storage medium

Through the combination of large language models and large video generation models, audio and video works are automatically generated, which solves the problems of complex and high cost of production of traditional audio and video works, and realizes simplified production and personalized generation.

CN120302124APending Publication Date: 2025-07-11QILIN HESHENG NETWORK TECH INC
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510410513.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The production methods of traditional audio and video works are cumbersome and costly, which is difficult to meet the users' simplified production needs.

Method used

By obtaining user audio and content text, using a large language model to generate storyboard description text, combining video to generate storyboard video, and synthesize it with audio to automatically create audio and video works.

Benefits of technology

It simplifies the production process of audio and video works, reduces production costs, and generates personalized audio and video works.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302124A_ABST
    Figure CN120302124A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a video generation method and device, electronic equipment and a storage medium, and relates to the technical field of computers. The video generation method comprises the following steps: acquiring an audio of a user and a content text of the audio; inputting the content text into a large language model, and enabling the large language model to output a plurality of split description texts of an associated video of the audio based on the content text; generating a plurality of split video generation cue words based on the plurality of split description texts, inputting the plurality of split video generation cue words into a video generation large model, and enabling the video generation large model to generate a plurality of split videos of the associated video based on the plurality of split video generation cue words; and synthesizing the plurality of split videos of the associated video with the audio to generate an audio and video work corresponding to the audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a video generation method, apparatus, device, and storage medium. Background Art

[0002] In today's digital content consumption era, simple audio alone can no longer meet users' needs for dissemination, sharing, and expression. Creating audio-visual works for audio can not only enhance expressiveness but also improve the dissemination effect. Although creating audio-visual works for audio brings many benefits, the traditional production methods of audio-visual works are often cumbersome, making the process of creating audio-visual works for users relatively complex and costly. Summary of the Invention

[0003] The purpose of the embodiments of this application is to provide a video generation method for simplifying the production process of audio-visual works and reducing the production cost of audio-visual works.

[0004] In a first aspect, the embodiments of this application provide a video generation method, which includes: Obtain the user's audio and the content text of the audio; Input the content text into a large language model, so that the large language model generates multiple storyboard description texts of the accompanying video of the audio based on the content text; Generate multiple storyboard video generation prompts based on the multiple storyboard description texts, input the multiple storyboard video generation prompts into a video generation large model, so that the video generation large model generates multiple storyboard videos of the accompanying video based on the multiple storyboard video generation prompts; Synthesize the multiple storyboard videos of the accompanying video with the audio to generate an audio-visual work corresponding to the audio.

[0005] In a second aspect, the embodiments of this application provide a video generation apparatus, which includes: An acquisition module, configured to obtain the user's audio and the content text of the audio; A storyboard description generation module, configured to input the content text into a large language model, so that the large language model generates multiple storyboard description texts of the accompanying video of the audio based on the content text; A storyboard video generation module, configured to generate multiple storyboard video generation prompts based on the multiple storyboard description texts, input the multiple storyboard video generation prompts into a video generation large model, so that the video generation large model generates multiple storyboard videos of the accompanying video based on the multiple storyboard video generation prompts; An audio-visual generation module, configured to synthesize the multiple storyboard videos of the accompanying video with the audio to generate an audio-visual work corresponding to the audio.

[0006] In a third aspect, an embodiment of the present application provides an electronic device, which includes: a processor; and a memory arranged to store computer-executable instructions, which when executed cause the processor to execute the above video generation method.

[0007] In a fourth aspect, an embodiment of the present application provides a readable storage medium for storing computer-executable instructions, which when executed by a processor implement the above video generation method.

[0008] The technical solution provided by the embodiment of the present application obtains the user's audio and the content text of the audio; inputs the content text into a large language model, so that the large language model generates multiple storyboard description texts of the accompanying video of the audio based on the content text; generates multiple storyboard video generation prompts based on the multiple storyboard description texts, inputs the multiple storyboard video generation prompts into a video generation large model, so that the video generation large model generates multiple storyboard videos of the accompanying video based on the multiple storyboard video generation prompts; synthesizes the multiple storyboard videos of the accompanying video with the audio to generate an audiovisual work corresponding to the audio, and can automatically generate an audiovisual work for the user based on the user's audio, simplify the production process of the audiovisual work, and reduce the production cost of the audiovisual work. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0010] Figure 1 It is a schematic flowchart of a video generation method provided by an embodiment of the present application; Figure 2 It is a schematic diagram of the modules of a video generation device provided by an embodiment of the present application; Figure 3 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0011] An embodiment of the present application provides a video generation method, device, equipment and storage medium.

[0012] To enable those skilled in the art to better understand the technical solutions in this application, the following will clearly and completely describe the technical solutions in the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0013] Please refer to Figure 1 , which is a schematic flowchart of a video generation method provided by an embodiment of this application. This video generation method can be applied to a content creation and distribution platform and is executed by the content creation and distribution platform. The content creation and distribution platform is, for example, a platform for users to create audio-visual works and distribute the created audio-visual works to various social media platforms. As Figure 1 shown, this video generation method includes the following processes: S102, obtaining the user's audio and the content text of this audio.

[0014] The user's audio and the content text of this audio are user-specific and can vary with different users. For example, the user's audio can be the user's song audio, and the content text of this audio can be the lyrics. Another example is that the user's audio can be the user's recitation audio, and the content text of this audio can be the recitation words.

[0015] In one implementation, obtaining the user's audio includes one of the following: collecting the audio recorded by the user according to the audio recording operation performed by the user on the audio recording interface; obtaining the audio uploaded by the user according to the audio upload operation performed by the user on the audio upload interface.

[0016] In one implementation, obtaining the content text of this audio includes one of the following: Inputting this audio into an audio content recognition model for audio content recognition to obtain the content text of this audio; obtaining the content text of this audio uploaded by the user according to the content text upload operation performed by the user on the text upload interface.

[0017] Specifically, optionally, the user's audio can be collected on-site. Obtaining the user's audio can specifically include: collecting the audio recorded by the user according to the audio recording operation performed by the user on the audio recording interface.

[0018] In implementation, the content creation and distribution platform can provide an audio recording interface, and the user can choose to perform an audio recording operation on the above audio recording interface. When the user records audio on the audio recording interface, the content creation and distribution platform can collect the audio recorded by the user on-site.

[0019] In one implementation, the audio recording interface can provide multiple functions to facilitate user recording and meet the needs of different users.

[0020] For example, the audio recording interface can provide a real-time volume monitoring and display function to help users adjust the microphone distance and input gain; the audio recording interface can provide a silent environment detection function to give users a prompt when the detected background noise is too loud; the audio recording interface can provide a countdown preparation function to give users a few seconds of preparation time; the audio recording interface can provide a visual recording progress bar and time display function to help users understand the recording progress.

[0021] Again, the audio recording interface can provide a recording model selection function for users to select a recording model, such as free mode or timed mode; the audio recording interface can provide functions of pausing, continuing, and re-recording during recording to facilitate user recording.

[0022] In one implementation, when the user records the song audio of an existing song, to facilitate user recording, a karaoke-like song following interface can be provided for the user to follow the song.

[0023] In one implementation, to improve the recording effect, noise suppression and dynamic range compression processing can be performed during the user's audio recording process; to avoid the loss of the recorded audio due to interruption other than recording, during the user's audio recording process, the audio recorded by the user can be collected in real time and stored in a buffer file in real time; after the user's audio recording is completed, the collected audio can be stored in a lossless format (such as WAV format), and at the same time, a preview file in a compressed format (such as MP3 format) can be generated for the user to quickly view and listen.

[0024] Optionally, the user's audio can be uploaded by the user. Obtaining the user's audio can specifically include: obtaining the audio uploaded by the user according to the audio upload operation performed by the user on the audio upload interface.

[0025] In implementation, the content creation and distribution platform can provide an audio upload interface, and the user can choose to upload audio on the above audio upload interface. After the user uploads the audio on the audio upload interface, the content creation and distribution platform can obtain the audio uploaded by the user.

[0026] Among them, the audio uploaded by the user can be an audio file in a specified format, and the specified format can be any one of the following formats: WAV, MP3, FLAC, AAC, etc.

[0027] Optionally, the content text of the above audio can be obtained by performing audio content recognition on the above audio. Obtaining the content text of the above audio can specifically include: inputting the above audio into an audio content recognition model for audio content recognition to obtain the content text of the above audio.

[0028] Among them, the audio content recognition model is used to recognize the content in the audio and output the content text of the audio. When the above audio is a song audio, the audio content recognition model can recognize the content of the song audio and output the lyrics. When the above audio is a recitation audio, the audio content recognition model can recognize the content of the recitation audio and output the recitation words.

[0029] In one implementation, to achieve accurate audio content recognition, inputting the above audio into the audio content recognition model for audio content recognition to obtain the content text of the above audio can be specifically: inputting the above audio into the audio content recognition model corresponding to the audio type of the above audio for audio content recognition to obtain the content text of the above audio.

[0030] Among them, the audio type of the above audio can be a song type or a recitation type. When the audio type of the above audio is a song type, the above audio can be input into the audio content recognition model corresponding to the song type for audio content recognition. When the audio type of the above audio is a recitation type, the above audio can be input into the audio content recognition model corresponding to the recitation type for audio content recognition.

[0031] The audio type of the above audio can be determined according to the audio type selected by the user, or can be automatically recognized by the audio type recognition model.

[0032] In one implementation, the above audio recording interface or the above audio uploading interface can provide audio type options. When the user records or uploads the above audio, the user can select the audio type of the above audio. The audio type selected by the user can be determined as the audio type of the above audio. When the user does not select an audio type, the above audio can be input into the audio type recognition model, and the audio type recognition model can recognize the audio type of the above audio.

[0033] In one implementation, the above audio recording interface or the above audio uploading interface may not provide audio type options either. After obtaining the above audio, the above audio can be input into the audio type recognition model, and the audio type recognition model can recognize the audio type of the above audio.

[0034] Among them, the audio type recognition model is used to recognize the audio type, and it can adopt an existing model.

[0035] In one implementation, to make the audio type recognition more accurate and improve the audio-visual effect of the subsequent generated audio-visual works. Before inputting the above audio into the audio type recognition model for audio type recognition, the above audio can be preprocessed for audio signals. The above audio signal preprocessing includes but is not limited to: noise reduction, silence removal, volume normalization, etc. The above step of inputting the above audio into the audio type recognition model for audio type recognition can be specifically: inputting the preprocessed audio into the audio type recognition model for audio type recognition. The above step of inputting the above audio into the audio content recognition model for audio content recognition can be specifically: inputting the preprocessed audio into the audio content recognition model corresponding to the audio type of the above audio for audio content recognition.

[0036] Optionally, the content text of the above audio can also be uploaded by the user. Obtaining the content text of the above audio can specifically include: obtaining the content text of the above audio uploaded by the user according to the content text upload operation performed by the user on the text upload interface.

[0037] In implementation, the content creation and distribution platform can provide a text upload interface. The user can perform a text upload operation on this text upload interface to upload the content text of the above audio.

[0038] S104, input the above content text into the large language model, so that the large language model generates multiple storyboard description texts of the accompanying video of the above audio based on the above content text.

[0039] The accompanying video of the above audio is an auxiliary video that matches the audio content of the above audio and is used to provide visual assistance or enhancement for the above audio. It can be a video without sound.

[0040] The above accompanying video can be composed of multiple storyboard videos. Each storyboard description text in the above multiple storyboard description texts is used to describe a storyboard video, corresponding to a text segment in the content text. The storyboard videos described by the multiple storyboard description texts are combined in the order of their corresponding text segments in the content text to form the accompanying video.

[0041] The specific number of the multiple storyboard description texts and the text segment corresponding to each storyboard description text can be planned by the large language model according to the above content text.

[0042] A large language model (LLM) is an artificial intelligence model trained on a vast amount of text data. It can understand, generate, and process natural language, analyze the input text, and give reasonable responses. The above text content can be input into the large language model, along with appropriate storyboard description prompts, so that the large language model can generate the storyboard description text of the accompanying video of the above audio and the text segments corresponding to each storyboard description text based on the above text content and the storyboard description prompts.

[0043] A prompt is an instruction or question input to the large language model, used to guide the large language model to generate an output that meets the requirements. In the embodiments of this application, the storyboard description prompt is used to guide the large language model to generate the storyboard description text of the accompanying video of the above audio based on the content text of the above audio.

[0044] For example, when the user's audio is song audio and the content text of the audio is the lyrics, the above storyboard description prompt can be: The following is a piece of lyrics. We hope to split this piece of lyrics into multiple semantically coherent scene units, generate a supporting accompanying video for this piece of lyrics, such as an MV. Please help me design the scene units, that is, the storyboards, of the accompanying video, give the lyrics segments corresponding to each storyboard and the storyboard description text. The requirements are: each storyboard description text should contain the following information: duration, scene description, main object, action requirements, emotional tone, and photography style.

[0045] In one implementation, to make the information in the storyboard description text more abundant and enrich the video performance, the embodiments of this application can also guide the large language model to perform creative expansion and enrichment on the content text through the storyboard description prompt to generate the storyboard description text. The above-mentioned inputting the above text content into the large language model to make the large language model generate multiple storyboard description texts of the accompanying video of the above audio can be specifically: inputting the above text content and the storyboard description prompt into the large language model, so that the large language model outputs multiple storyboard description texts of the accompanying video and the text segments corresponding to each storyboard description text based on the above text content and the above storyboard description prompt.

[0046] Among them, the above storyboard description prompt is used to guide the large language model to perform creative expansion and enrichment on the above text content, and output multiple storyboard description texts and the text segments corresponding to each storyboard description text. The large language model can perform creative expansion and enrichment under the guidance of this storyboard description prompt.

[0047] For example, the storyboard description prompt can require the large language model to make reasonable inferences and expansions on the implicit content in the content text, enrich the video performance, automatically supplement transition scenes according to the context, ensure the smoothness and coherence of the video, and generate possible visual representation schemes for the abstract concepts in the content text.

[0048] In one implementation, to make the description of the storyboard description text more in line with the user's needs, after the large language model outputs the storyboard description text, the storyboard description text can be shown to the user so that the user can modify or confirm the storyboard description text. Subsequently, the storyboard description text modified or confirmed by the user can be determined as the final storyboard description text.

[0049] S106. Generate multiple storyboard video generation prompts based on the above-mentioned multiple storyboard description texts, input the above-mentioned multiple storyboard description texts into the video generation large model, and enable the video generation large model to generate multiple storyboard videos of the associated video based on the above-mentioned multiple storyboard description texts.

[0050] The video generation large model is used to generate storyboard videos according to the storyboard video generation prompts. One storyboard description text can generate one storyboard video generation prompt, and one storyboard video generation prompt can generate one storyboard video. In the embodiments of the present application, there is a one-to-one correspondence between the text segment, the storyboard description text, the storyboard video generation prompt, and the storyboard video. One text segment corresponds to one storyboard description text corresponds to one storyboard video generation prompt corresponds to one storyboard video.

[0051] In implementation, each storyboard video generation prompt in the above-mentioned multiple storyboard video generation prompts can be input into the video generation large model, and the video generation large model generates one storyboard video based on each storyboard video generation prompt, thereby generating multiple storyboard videos.

[0052] In one implementation, to make the generated video more personalized for the user, before processing S106, target visual style information can be obtained. In processing S106, generating multiple storyboard video generation prompts based on the above-mentioned multiple storyboard description texts can specifically be: generating each storyboard video generation prompt based on each storyboard description text and the above-mentioned target visual style information.

[0053] Among them, the target visual style is, for example: art style, realistic style, animation style, documentary style, etc.

[0054] The target visual style can be selected by the user, or determined according to the visual style of the reference picture uploaded by the user, or determined according to the audio characteristics of the audio. The above-mentioned obtaining of the target visual style information includes one of the following: According to the visual style selection operation performed by the user on the visual style selection interface, determine the visual style selected by the user as the target visual style; Based on the reference picture uploading operation performed by the user on the visual style reference interface, determine the visual style in the reference picture uploaded by the user as the target visual style; Extract the audio features of the above audio, input the above audio features into the visual style analysis model, and let the visual style analysis model analyze the visual style of the associated video of the above audio based on the above audio features, and determine the visual style analyzed by the visual style analysis model as the target visual style.

[0055] Specifically, a visual style selection interface can be provided, and multiple visual styles can be provided in the visual style selection interface for the user to choose. The visual style selected by the user can be determined as the target visual style.

[0056] When the user does not select a visual style, based on the reference picture uploading operation performed by the user on the visual style reference interface, determine the visual style in the reference picture uploaded by the user as the target visual style.

[0057] When the user neither selects a visual style nor uploads a reference picture, extract the audio features of the above audio, input the above audio features into the visual style analysis model, and let the visual style analysis model analyze the visual style of the associated video of the audio. The visual style analyzed by the visual style analysis model can be determined as the target visual style.

[0058] Among them, the above audio features are, for example: Mel Frequency Cepstral Coefficients (MFCC), zero-crossing rate, spectral centroid, rhythm information (BPM), pitch and other features, which can be extracted by an audio feature extraction model.

[0059] The above visual style analysis model is used to analyze and output the most suitable visual style of the associated video of the above audio based on the input audio features of the above audio. It can be trained by using visual style analysis samples on a deep learning model, such as a Convolutional Neural Network (CNN) or a time series-based neural network (LSTM). Each visual style analysis sample can include: audio features and the visual style label corresponding to the audio features. Visual style labels are, for example: art style, realistic style, documentary style and other labels.

[0060] In implementation, audio data including audio files and their corresponding visual style tags can be collected to construct an audio dataset, where each audio in the audio dataset corresponds to a visual style tag. These style tags can include, but are not limited to: artistic style, realistic style, documentary style, electronic style, classical style, rock style. The visual style tag corresponding to each audio can be manually labeled or automatically obtained through an existing music style classification model. An audio feature extraction model can be used to extract the audio features of each audio. Audio features include, but are not limited to: Mel Frequency Cepstral Coefficients (MFCC), zero-crossing rate, spectral centroid, rhythm information (BPM), pitch. Based on the audio features of each audio and the corresponding visual style tags above, a dataset for training a visual style model can be constructed.

[0061] Then, a base model to be trained can be selected. The base model to be trained is a classification model, such as a Convolutional Neural Network (CNN) or a time-series based neural network (such as LSTM or Transformer). This model can include: an input layer for inputting audio features (such as MFCC, rhythm, pitch, etc.); a convolutional layer / fully connected layer for extracting deep features of the audio; an LSTM / Transformer layer for processing the time-series information in the audio and learning the changes in rhythm and emotion; an output layer for outputting the most suitable MV style tag (such as, artistic style, realistic style, documentary style, etc.).

[0062] After that, the above dataset can be divided into a training set, a validation set, and a test set, and the above base model can be supervised-trained to obtain a visual style analysis model.

[0063] Among them, during training, the loss function can be selected as the Cross-Entropy Loss function. This cross-loss function is, for example:

[0064] Among them, C is the number of visual style categories, y i is the i-th visual style category, and p i is the probability of the i-th visual style category predicted by the model.

[0065] During training, the Adam optimizer can be used to optimize the parameters of the model. For example, the learning rate can be set to 0.0001 and the momentum to 0.9. The parameters can be continuously adjusted according to the actual training situation to find the optimal values.

[0066] Further, to improve the generalization ability of the model, data augmentation can be performed on the audio data and video data. Audio augmentation can include volume change, time stretching, audio noise, etc. Video augmentation can include cropping, flipping, rotation, color enhancement, etc. After training, metrics such as accuracy, precision, recall, etc. can be used to evaluate the performance of the model, and a confusion matrix can be used to analyze the performance of the model on different categories.

[0067] In one implementation, to make the generated video meet the user's needs, before processing S106, video additional requirement information input by the user on the video additional requirement interface can be obtained. In processing S106, generating multiple shot video generation prompts based on the multiple shot description texts can specifically be: generating each shot video generation prompt based on each shot description text, the above-mentioned target visual style information, and the above-mentioned video additional requirement information.

[0068] Among them, the video additional requirement information is, for example: composition requirements, such as "golden ratio", "symmetric composition"; color tone and light and shadow effect requirements, such as "high contrast", "soft backlight", etc.; technical parameter requirements, such as "resolution", "aspect ratio", etc.; generation quality requirements, such as "high detail", "4K", "basic photography", etc.

[0069] In the embodiments of the present application, the finally generated audio-visual work can be published to the target platform. The target platform can be a social media platform specified by the user. In one implementation, the video generation large model can be the video generation large model corresponding to the target platform. In processing S106, inputting the above-mentioned multiple shot video generation prompts into the video generation large model, and enabling the video generation large model to generate multiple shot videos of the above-mentioned accompanying video based on the above-mentioned multiple shot video generation prompts can specifically be: inputting the above-mentioned multiple shot video generation prompts into the video generation large model corresponding to the target platform, and enabling the video generation large model corresponding to the target platform to generate multiple shot videos based on the above-mentioned multiple shot video generation prompts.

[0070] The video large model corresponding to the target platform is a video large model optimized for the target platform. By inputting the shot video generation prompts into the video generation large model corresponding to the target platform, and enabling the video generation large model corresponding to the target platform to generate shot videos based on the shot video generation prompts, the generated shot videos can be more adapted to the target platform.

[0071] The format of the storyboard description text required for the video generation large model corresponding to each platform may be different. Further, to match the video generation large model corresponding to the target platform, in process S106, the generation of multiple storyboard video generation prompts based on the multiple storyboard description texts and the target visual style information can be specifically: generating each target structure storyboard video generation prompt based on each storyboard description text and the above-mentioned target visual style information; the input of the multiple storyboard video generation prompts into the video generation large model to enable the video generation large model to generate the multiple storyboard videos of the above-mentioned associated video based on the multiple storyboard video generation prompts can be specifically: inputting each of the above-mentioned target structure storyboard video generation prompts into the video generation large model corresponding to the target platform, so that the video generation large model corresponding to the target platform generates each storyboard video based on each target structure storyboard video generation prompt.

[0072] Among them, the above-mentioned target structure is determined based on the prompt structure of the video generation large model corresponding to the target platform. By generating the video generation prompts of the target structure and inputting the video generation prompts of the target structure into the video generation large model corresponding to the target platform, the structure of the prompts can be better adapted to the video generation large model corresponding to the target platform, facilitating the generation of better-quality storyboard videos.

[0073] In implementation, a large language model can be used to generate video generation prompts. For example, the above-mentioned storyboard description text, the above-mentioned target style information, the above-mentioned video additional requirement information, and the target format information can be input into the large language model, and with appropriate prompts, the large language model generates video generation prompts in the target format based on the above-mentioned storyboard description text, the above-mentioned target style information, and the above-mentioned video additional requirement information.

[0074] In implementation, to improve the generation speed and quality of the storyboard videos, a queue management mechanism can be adopted when generating the storyboard videos. During the process of generating one storyboard video, the next storyboard video generation prompt can be generated.

[0075] After generating the storyboard videos, the visual understanding ability of the large model can be used to evaluate the quality of the generated storyboard videos, detect problems such as distortion and blurring. If any exist, the video generation large model can be made to regenerate the storyboard videos. Multiple alternative versions can be generated for each storyboard for the user to choose, and the version selected by the user is determined as the final version of the storyboard video.

[0076] S108, synthesize the multiple storyboard videos of the above-mentioned associated video with the above-mentioned audio to generate the audio-video work corresponding to the above-mentioned audio.

[0077] The audio-visual work corresponding to the above audio is a work that simultaneously includes the above audio and the accompanying video of the above audio, and it is a multimedia presentation mode of the audio content of the above audio. When the above audio is a song audio, the audio-visual work corresponding to the above audio is the audio-visual work of the song, such as the MV of the song. When the above audio is a recitation audio, the audio-visual work corresponding to the above audio is the audio-visual work of the recitation, such as an audio-visual work in the style of an MV of the recitation.

[0078] In the embodiments of the present application, one storyboard video corresponds to one storyboard description text, one storyboard description text corresponds to one text segment of the content text, and one text segment corresponds to one audio segment. There is a one-to-one correspondence among the storyboard video, the storyboard description text, the text segment, and the audio segment.

[0079] In one implementation, the storyboard video and the audio segment corresponding to the same text segment can be synthesized to obtain an audio-visual segment corresponding to each text segment; according to the position of the text segment corresponding to each audio-visual segment in the content text, multiple audio-visual segments are synthesized to generate an audio-visual work.

[0080] For example, the storyboard video and the audio segment corresponding to the first text segment can be synthesized to generate an audio-visual segment corresponding to the first text segment, the storyboard video and the audio segment corresponding to the second text segment can be synthesized to generate an audio-visual segment corresponding to the second text segment,..., the storyboard video and the audio segment corresponding to the Nth text segment can be synthesized to generate an audio-visual segment corresponding to the Nth text segment, and then, the N audio-visual segments corresponding to the above N text segments can be combined in the order of the N text segments in the content text to generate an audio-visual work.

[0081] In one implementation, in order to synthesize the storyboard video and the audio segment corresponding to the same text segment, the storyboard video and the audio segment corresponding to the same text segment can be time-aligned.

[0082] Among them, the storyboard video generated by the video generation large model is usually a video segment with a fixed length. When time-aligning the storyboard video and the audio segment corresponding to the same text segment, the storyboard video can be adjusted to make the storyboard video and the audio segment time-aligned. Among them, adjusting the storyboard video includes but is not limited to: intercepting the storyboard video according to the time range, and adjusting the playing speed of the storyboard video.

[0083] For example, if the total duration of a storyboard video is 5 seconds, and the duration of the audio clip corresponding to the same text clip as the storyboard video is 4 seconds, the playback speed of the storyboard video can be increased to 125% to align the storyboard video with the audio clip in time. Of course, the first 4 seconds of the storyboard video can also be intercepted to align the storyboard video with the audio clip in time.

[0084] Among them, in order to avoid the video playback appearing unnatural due to excessive acceleration or slowing down of the playback speed, the playback speed of the storyboard video can be adjusted within a preset speed range, and the preset speed range is, for example, 70% to 130%. That is, the playback speed of the storyboard video can be adjusted to 130% of the original speed at the highest, and the playback speed of the storyboard video can be adjusted to 70% of the original speed at the lowest. In order to avoid excessive video truncation resulting in the inability to fully express the video content, the storyboard video can be intercepted within a preset interception ratio range, and the preset interception ratio range is, for example, 70% to 100%, that is, the original time range of the original storyboard video can be intercepted at the lowest 70%, and the time range of the original storyboard video can be intercepted at the highest 100%. Of course, the two methods of adjusting the speed of the storyboard video and intercepting the storyboard video can be used at the same time. If the playback speed of the storyboard video is adjusted and the storyboard video is intercepted according to the time range at the same time, and the time of the storyboard video and the audio clip cannot be aligned, the user can be prompted to intervene, such as prompting the user to generate a new video clip, modify the video to generate text (modify the prompt word), tune the audio, etc.

[0085] Furthermore, in order to improve the overall listening experience of the audio-visual work and enhance the sound expression of the audio-visual work, optionally, before processing S108, the above-mentioned audio can be processed according to the audio processing operation selected by the user in the audio processing interface; in processing S108, the above-mentioned multiple storyboard videos of the accompanying video are synthesized with the audio to generate an audio-visual work corresponding to the audio, including: synthesizing the multiple storyboard videos of the accompanying video with the audio after audio processing to generate an audio-visual work corresponding to the audio.

[0086] Among them, the above-mentioned audio processing includes one of the following: original sound optimization processing, voice cloning processing, and accompaniment mixing processing.

[0087] Original sound optimization processing is to optimize the clarity, dynamic range and spatial sense of the sound through intelligent audio processing technology under the premise of completely retaining the user's original singing timbre. It may include at least one of the following processing: retaining 100% of the user's original sound and only performing audio enhancement; applying an adaptive equalizer to adjust the audio frequency characteristics to make the sound clearer and more natural; using an intelligent compressor to control the dynamic range to increase the sound density and presence; and using spatial effect processing to increase the environmental and three-dimensional sense of the sound.

[0088] Voice cloning processing clones the voiceprint features through a small number of user's original voice samples to generate highly realistic synthetic speech (TTS). The user's audio and the content text of the audio can be input into the voice cloning model. The voice cloning model clones the user's voice in the audio and regenerates the user's audio with the cloned voice. Among them, for song audio, the voice cloning model can clone the user's voice in the song audio, and the music generation model can re-sing the song in the song audio through AI to generate an AI song. Then, the cloned user's voice can be used to replace the voice in the AI song, so that the AI can re-sing the song with the user's voice. For the case of reading audio, the cloned user's voice and the recitation words can be used as the input of the TTS model, so that the TTS model can re-read the recitation words with the user's voice, so that the AI can re-recite with the user's voice. Among them, the speed, intonation, and tone of reading the manuscript can be controlled by the TTS model itself.

[0089] Accompaniment mixing processing mixes the user's audio with the accompaniment. When the user's audio is the song audio of an existing song, such as the audio of an existing song sung by the user, the accompaniment of the song can be obtained from the music library, and the user's audio can be mixed with the accompaniment. When the user's audio is the song audio of an original song, such as the audio of an original song sung by the user, the above audio can be input into the accompaniment generation model to generate an accompaniment for the above audio. Then, the above audio can be mixed with the accompaniment.

[0090] In the application, whether to perform audio processing on the user's audio and what kind of audio processing to perform can be determined according to the user's needs. An audio processing interface can be provided. The audio processing interface can include audio processing options such as original sound optimization processing, voice cloning processing, and accompaniment mixing processing. The user's audio can be processed accordingly according to the selected audio processing option. For example, when the user is originally out of tune and has poor singing skills, etc., the voice cloning processing can be selected so that the AI can replace the person to sing the whole song, while ensuring the pitch accuracy, pitch, and timbre, and maintaining the user's original voice, so that the generated song sounds like it is sung by the user himself.

[0091] Furthermore, to enhance the overall visual experience of the audio-visual work and strengthen the visual expression of the audio-visual work, further visual optimization can be performed on the generated audio-visual work.

[0092] For example, color correction can be performed on the entire audio-visual work to ensure the consistency of the overall visual style of the audio-visual work; intelligent sharpening and detail enhancement can be performed on the audio-visual work to improve the picture clarity of the audio-visual work; the video parameters of the audio-visual work can be adjusted accordingly according to the video parameters required by the audio-visual work on the target platform; the audio-visual work can be converted into the required video format and resolution to meet different usage scenarios.

[0093] The video generation method provided by the embodiments of the present application obtains the user's audio and the content text of the above audio, inputs the above content text into a large language model, enables the large language model to generate multiple storyboard description texts of the accompanying video of the above audio based on the above content text, inputs the above multiple storyboard description texts into a large video generation model, enables the large video generation model to generate multiple storyboard videos of the above accompanying video based on the above multiple storyboard description texts, and then synthesizes the multiple storyboard videos of the above accompanying video with the above audio to generate an audiovisual work corresponding to the above audio, which can automatically generate an audiovisual work for the user and reduce the cost of the user making an audiovisual work.

[0094] Moreover, in the embodiments of the present application, the storyboard description text, the storyboard video, etc. are all generated by a large model, which has a certain degree of randomness, and various needs of the user are considered during the generation process. Therefore, the audiovisual works generated for each user are different, which can improve the user personalization degree of the audiovisual works.

[0095] Furthermore, the video generation method provided by the embodiments of the present application can solve the user's intonation problem and improve the overall listening experience of the audiovisual work by optimizing the audio for the user.

[0096] In some cases, considering issues such as personal image and privacy, the user may not want to appear on camera and hopes to use a virtual model to appear on camera. For this reason, in some embodiments of the present application, the video generation method provided by the embodiments of the present application may further include: Obtaining the virtual image requirements input by the user on the virtual image requirements interface; The above process S104 may specifically be: inputting the above content text and the above virtual image requirements into a large language model, enabling the large language model to output multiple storyboard description texts of the accompanying video of the above audio based on the above content text and the above virtual image requirements. Thus, the generated storyboard description text can be a storyboard description text considering the above virtual image requirements, and the finally generated audiovisual work is an audiovisual work that meets the user's virtual image requirements. For example, the generated audiovisual work can be an audiovisual work in which a virtual image that meets the user's virtual image requirements is singing or reciting.

[0097] Among them, in order to make the generated audiovisual work an audiovisual work that meets the user's virtual image requirements, some storyboard videos may include a virtual image model that meets the user's virtual image requirements.

[0098] In the case of a storyboard video being a video of a virtual character model, for example, in the case of a storyboard video being a video of a virtual character model singing or reciting, when synthesizing the storyboard video with its corresponding audio clip, lip-sync processing can be performed on the storyboard video and its corresponding audio clip to make the generated video more natural and smooth.

[0099] By generating audio-visual works for users by considering the virtual character needs of users, audio-visual works that meet the virtual character needs of users can be generated for users without the need for users to appear on camera, solving the problem that users are reluctant to appear on camera.

[0100] The above is the video generation method provided by the embodiments of the present application. Based on the same idea, the embodiments of the present application also provide a video generation device. Figure 2 It is a module schematic diagram of a video generation device provided by an embodiment of the present application. As Figure 2 shown, the video generation device 200 includes: An acquisition module 210, configured to acquire the audio of the user and the content text of the audio; A storyboard description generation module 220, configured to input the content text into a large language model, so that the large language model generates multiple storyboard description texts of the accompanying video of the audio based on the content text; A storyboard video generation module 230, configured to generate multiple storyboard video generation prompts based on the multiple storyboard description texts, input the multiple storyboard video generation prompts into a video generation large model, so that the video generation large model generates multiple storyboard videos of the accompanying video based on the multiple storyboard video generation prompts; An audio-visual generation module 240, configured to synthesize the multiple storyboard videos of the accompanying video with the audio to generate an audio-visual work corresponding to the audio.

[0101] In one implementation, acquiring the audio of the user includes one of the following: Collecting the audio input by the user according to the audio input operation performed by the user on the audio input interface; Obtaining the audio uploaded by the user according to the audio upload operation performed by the user on the audio upload interface; Obtaining the content text of the audio includes one of the following: Inputting the audio into an audio content recognition model for audio content recognition to obtain the content text of the audio; Obtaining the content text of the audio uploaded by the user according to the content text upload operation performed by the user on the text upload interface.

[0102] In one implementation, the storyboard description generation module 220 is configured to: Input the content text and the storyboard description prompt into the large language model, so that the large language model generates multiple storyboard description texts for the companion video and the text segment corresponding to each storyboard description text based on the content text and the storyboard description prompt; Among them, the storyboard description prompt is used to guide the large language model to creatively expand and enrich the content text, and generate multiple storyboard description texts for the companion video and the text segment corresponding to each storyboard description text.

[0103] In one implementation, the audio-visual generation module 240 is used for: Synthesize the storyboard video and the audio segment corresponding to the same text segment to obtain the audio-visual segment corresponding to each text segment; According to the position of the text segment corresponding to each audio-visual segment in the content text, synthesize multiple audio-visual segments to generate an audio-visual work.

[0104] In one implementation, the video generation device 200 further includes: A visual style acquisition module, which is used to acquire target visual style information; The storyboard video generation module is used for: Generate a storyboard video generation prompt for each based on each storyboard description text and the target visual style information.

[0105] Among them, the acquisition of the target visual style information includes one of the following: According to the visual style selection operation performed by the user on the visual style selection interface, determine the visual style selected by the user as the target visual style; According to the reference picture uploading operation performed by the user on the visual style reference interface, determine the visual style in the reference picture uploaded by the user as the target visual style; Extract the audio features of the audio, input the audio features into a visual style analysis model, and the visual style analysis model analyzes the visual style of the companion video of the audio based on the audio features, and determine the visual style analyzed by the visual style analysis model as the target visual style.

[0106] In one implementation, the video generation large model includes: the video generation large model corresponding to the target platform; The storyboard video generation module is used for: Generate a target structure storyboard video generation prompt for each based on each storyboard description text and the target visual style information; where the target structure is determined based on the prompt structure of the video generation large model corresponding to the target platform. Input the prompting words for generating each target structured storyboard video into the video generation large model corresponding to the target platform, so that the video generation large model corresponding to the target platform generates each storyboard video based on the prompting words for generating each target structured storyboard video.

[0107] In one implementation, the video generation device 200 further includes: An audio processing module, configured to perform audio processing on the audio according to the audio processing operations selected by the user in the audio processing interface, where the audio processing includes one of the following: original sound optimization processing, voice cloning processing, and accompaniment mixing processing; The audio-video generation module is configured to: Synthesize the multiple storyboard videos of the accompanying video with the audio after audio processing to generate an audio-video work corresponding to the audio.

[0108] The above is the video generation device provided by the embodiments of the present application. Based on the same concept, the embodiments of the present application also provide an electronic device, as Figure 3 shown.

[0109] The electronic device 300 may be a terminal device that executes the above video generation method, such as a mobile phone, a computer, etc.

[0110] The electronic device 300 may vary greatly due to different configurations or performances, and may include one or more processors 301 and a memory 302. The memory 302 is arranged to store computer-executable instructions, and when the executable instructions are executed by the processor 301, the processor 301 executes the following process: Obtain the user's audio and the content text of the audio; Input the content text into a large language model, so that the large language model generates multiple storyboard description texts of the accompanying video of the audio based on the content text; Generate multiple prompting words for generating storyboard videos based on the multiple storyboard description texts, input the multiple prompting words for generating storyboard videos into a video generation large model, so that the video generation large model generates multiple storyboard videos of the accompanying video based on the multiple prompting words for generating storyboard videos; Synthesize the multiple storyboard videos of the accompanying video with the audio to generate an audio-video work corresponding to the audio.

[0111] In one implementation, when the executable instructions are executed by the processor 301, the processor 301 executes the following process: Collect the audio input by the user according to the audio input operation performed by the user in the audio input interface; or Obtain the audio uploaded by the user according to the audio upload operation performed by the user in the audio upload interface.

[0112] Input the audio into an audio content recognition model to perform audio content recognition and obtain the content text of the audio; or Obtain the content text of the audio uploaded by the user according to the content text upload operation performed by the user on the text upload interface.

[0113] In one implementation, when the executable instruction is executed by the processor 301, the processor 301 executes the following process: Input the content text and the storyboard description prompt words into a large language model, so that the large language model generates multiple storyboard description texts of the associated video and the text segments corresponding to each storyboard description text based on the content text and the storyboard description prompt words; Among them, the storyboard description prompt words are used to guide the large language model to perform creative expansion and enrichment on the content text, and generate multiple storyboard description texts of the associated video and the text segments corresponding to each storyboard description text.

[0114] In one implementation, when the executable instruction is executed by the processor 301, the processor 301 executes the following process: Synthesize the storyboard video and the audio segment corresponding to the same text segment to obtain the audio-visual segment corresponding to each text segment; Synthesize multiple audio-visual segments according to the positions of the text segments corresponding to each audio-visual segment in the content text to generate an audio-visual work.

[0115] In one implementation, when the executable instruction is executed by the processor 301, the processor 301 executes the following process: Obtain the target visual style information; The generation of multiple storyboard video generation prompt words based on the multiple storyboard description texts includes: Generate each storyboard video generation prompt word based on each storyboard description text and the target visual style information.

[0116] Among them, the obtaining of the target visual style information includes one of the following: According to the visual style selection operation performed by the user on the visual style selection interface, determine the visual style selected by the user as the target visual style; According to the reference picture upload operation performed by the user on the visual style reference interface, determine the visual style in the reference picture uploaded by the user as the target visual style; Extract the audio features of the audio, input the audio features into a visual style analysis model, and analyze the visual style of the associated video of the audio based on the audio features by the visual style analysis model. Determine the visual style analyzed by the visual style analysis model as the target visual style.

[0117] In one implementation, the video generation large model includes: a video generation large model corresponding to the target platform; when the executable instruction is executed by the processor 301, the processor 301 executes the following process: The generating each storyboard video generation prompt word based on each storyboard description text and the target visual style information includes: Generating each target structure storyboard video generation prompt word based on each storyboard description text and the target visual style information; wherein, the target structure is determined based on the prompt word structure of the video generation large model corresponding to the target platform; The inputting the multiple storyboard video generation prompt words into the video generation large model to enable the video generation large model to generate multiple storyboard videos of the associated video based on the multiple storyboard video generation prompt words includes: Inputting each target structure storyboard video generation prompt word into the video generation large model corresponding to the target platform to enable the video generation large model corresponding to the target platform to generate each storyboard video based on each target structure storyboard video generation prompt word.

[0118] In one implementation, when the executable instruction is executed by the processor 301, the processor 301 executes the following process: Perform audio processing on the audio according to the audio processing operation selected by the user in the audio processing interface, and the audio processing includes one of the following: original sound optimization processing, voice cloning processing, accompaniment mixing processing; The synthesizing the multiple storyboard videos of the associated video with the audio to generate an audio-visual work corresponding to the audio includes: Synthesize the multiple storyboard videos of the associated video with the audio after audio processing to generate an audio-visual work corresponding to the audio.

[0119] Further, based on the same inventive concept, one or more embodiments of the present application also provide a storage medium for storing computer executable instruction information. In a specific embodiment, the storage medium may be a USB flash drive, an optical disc, a hard disk, etc. When the computer executable instruction information stored in the storage medium is executed by the processor, the following process can be realized: Obtain the user's audio and the content text of the audio; Input the content text into a large language model to enable the large language model to generate multiple storyboard description texts for the companion video of the audio based on the content text; Generate multiple storyboard video generation prompts based on the multiple storyboard description texts, and input the multiple storyboard video generation prompts into a video generation large model to enable the video generation large model to generate multiple storyboard videos of the companion video based on the multiple storyboard video generation prompts; Synthesize the multiple storyboard videos of the companion video with the audio to generate an audiovisual work corresponding to the audio.

[0120] In one implementation, when the computer-executable instruction information stored in the storage medium is executed by a processor, the following process can be implemented: Collect the audio input by the user according to the audio input operation performed by the user on the audio input interface; or; Obtain the audio uploaded by the user according to the audio upload operation performed by the user on the audio upload interface.

[0121] Input the audio into an audio content recognition model for audio content recognition to obtain the content text of the audio; or; Obtain the content text of the audio uploaded by the user according to the content text upload operation performed by the user on the text upload interface.

[0122] In one implementation, when the computer-executable instruction information stored in the storage medium is executed by a processor, the following process can be implemented: Input the content text and the storyboard description prompts into a large language model to enable the large language model to generate multiple storyboard description texts for the companion video and the text segments corresponding to each storyboard description text based on the content text and the storyboard description prompts; Among them, the storyboard description prompts are used to guide the large language model to creatively expand and enrich the content text, and generate multiple storyboard description texts for the companion video and the text segments corresponding to each storyboard description text.

[0123] In one implementation, when the computer-executable instruction information stored in the storage medium is executed by a processor, the following process can be implemented: Synthesize the storyboard video and the audio segment corresponding to the same text segment to obtain an audiovisual segment corresponding to each text segment; Synthesize multiple audiovisual segments according to the positions of the text segments corresponding to each audiovisual segment in the content text to generate an audiovisual work.

[0124] In one implementation, when the computer-executable instruction information stored in the storage medium is executed by a processor, the following process can be implemented: Obtain target visual style information; Generating multiple storyboard video generation prompts based on the multiple storyboard description texts includes: Generating each storyboard video generation prompt based on each storyboard description text and the target visual style information.

[0125] Among them, the obtaining of the target visual style information includes one of the following: According to the visual style selection operation performed by the user on the visual style selection interface, determining the visually selected style by the user as the target visual style; According to the reference picture uploading operation performed by the user on the visual style reference interface, determining the visual style in the reference picture uploaded by the user as the target visual style; Extracting the audio features of the audio, inputting the audio features into a visual style analysis model, analyzing the visual style of the accompanying video of the audio by the visual style analysis model based on the audio features, and determining the visual style analyzed by the visual style analysis model as the target visual style.

[0126] In one implementation, when the computer-executable instruction information stored in the storage medium is executed by a processor, the following process can be implemented: Generating each storyboard video generation prompt based on each storyboard description text and the target visual style information includes: Generating each target structure storyboard video generation prompt based on each storyboard description text and the target visual style information; wherein, the target structure is determined based on the prompt structure of the video generation large model corresponding to the target platform; Inputting the multiple storyboard video generation prompts into a video generation large model, so that the video generation large model generates multiple storyboard videos of the accompanying video based on the multiple storyboard video generation prompts includes: Inputting each target structure storyboard video generation prompt into the video generation large model corresponding to the target platform, so that the video generation large model corresponding to the target platform generates each storyboard video based on each target structure storyboard video generation prompt.

[0127] In one implementation, when the computer-executable instruction information stored in the storage medium is executed by a processor, the following process can be implemented: Performing audio processing on the audio according to the audio processing operation selected by the user on the audio processing interface, and the audio processing includes one of the following: original sound optimization processing, voice cloning processing, accompaniment mixing processing; Synthesizing the multiple storyboard videos of the accompanying video with the audio to generate an audiovisual work corresponding to the audio includes: Synthesize the multiple segmented videos of the associated video with the audio after audio processing to generate an audiovisual work corresponding to the audio.

[0128] The specific embodiments of the present application have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0129] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structures of diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is an integrated circuit whose logical function is determined by a user's programming of the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, today, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not just one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing some logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain a hardware circuit that implements the logical method flow.

[0130] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0131] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0132] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing one or more embodiments of the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0133] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0134] Embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable serial-parallel devices for fraud cases to generate a machine, such that the instructions executed by the processor of the computer or other programmable serial-parallel devices for fraud cases generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0135] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable serial-parallel devices for fraud cases to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0136] These computer program instructions can also be loaded onto a computer or other programmable serial-parallel devices for fraud cases, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0137] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0138] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0139] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0140] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0141] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, one or more embodiments of the present application may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, one or more embodiments of the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0142] One or more embodiments of the present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present application may also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0143] Each embodiment in the present application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.

[0144] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A video generation method, characterized in that, The method includes: Obtaining the user's audio and the content text of the audio; Inputting the content text into a large language model, enabling the large language model to generate multiple storyboard description texts for the accompanying video of the audio based on the content text; Generating multiple storyboard video generation prompts based on the multiple storyboard description texts, inputting the multiple storyboard video generation prompts into a video generation large model, enabling the video generation large model to generate multiple storyboard videos of the accompanying video based on the multiple storyboard video generation prompts; Synthesizing the multiple storyboard videos of the accompanying video with the audio to generate an audio-visual work corresponding to the audio.

2. The method according to claim 1, wherein: The obtaining of the user's audio includes one of the following: Collecting the audio input by the user according to the audio input operation performed by the user on the audio input interface; Obtaining the audio uploaded by the user according to the audio upload operation performed by the user on the audio upload interface; The obtaining of the content text of the audio includes one of the following: Inputting the audio into an audio content recognition model for audio content recognition to obtain the content text of the audio; Obtaining the content text of the audio uploaded by the user according to the content text upload operation performed by the user on the text upload interface.

3. The method according to claim 1, wherein The inputting of the content text into a large language model, enabling the large language model to generate multiple storyboard description texts for the accompanying video of the audio based on the content text, includes: Inputting the content text and the storyboard description prompts into a large language model, enabling the large language model to generate multiple storyboard description texts for the accompanying video and the text segments corresponding to each storyboard description text based on the content text and the storyboard description prompts; Wherein, the storyboard description prompts are used to guide the large language model to perform creative expansion and enrichment on the content text, generating multiple storyboard description texts for the accompanying video and the text segments corresponding to each storyboard description text.

4. The method according to claim 3, characterized in that, The synthesizing of the multiple storyboard videos of the accompanying video with the audio to generate an audio-visual work corresponding to the audio, includes: Synthesizing the storyboard videos corresponding to the same text segment with the audio segments to obtain the audio-visual segments corresponding to each text segment; Synthesizing multiple audio-visual segments according to the positions of the text segments corresponding to each audio-visual segment in the content text to generate an audio-visual work.

5. The method according to claim 1, characterized in that The method further includes: Obtaining target visual style information; The generating of multiple storyboard video generation prompts based on the multiple storyboard description texts, includes: Generating each storyboard video generation prompt based on each storyboard description text and the target visual style information; Wherein, the obtaining of the target visual style information includes one of the following: Determining the visually selected style by the user as the target visual style according to the visual style selection operation performed by the user on the visual style selection interface; Determining the visual style in the reference picture uploaded by the user as the target visual style according to the reference picture upload operation performed by the user on the visual style reference interface; Extract the audio features of the audio, input the audio features into a visual style analysis model, and have the visual style analysis model analyze the visual style of the accompanying video of the audio based on the audio features, and determine the visual style analyzed by the visual style analysis model as the target visual style.

6. The method according to claim 5, wherein The video generation large model includes: a video generation large model corresponding to the target platform; The generating each split-screen video generation prompt word based on each split-screen description text and the target visual style information includes: Generating each target structure split-screen video generation prompt word based on each split-screen description text and the target visual style information; wherein, the target structure is determined based on the prompt word structure of the video generation large model corresponding to the target platform; The inputting the multiple split-screen video generation prompt words into the video generation large model to enable the video generation large model to generate multiple split-screen videos of the accompanying video based on the multiple split-screen video generation prompt words includes: Inputting each target structure split-screen video generation prompt word into the video generation large model corresponding to the target platform to enable the video generation large model corresponding to the target platform to generate each split-screen video based on each target structure split-screen video generation prompt word.

7. The method according to claim 1, characterized in that The method further includes: Performing audio processing on the audio according to the audio processing operation selected by the user on the audio processing interface, where the audio processing includes one of the following: original sound optimization processing, voice cloning processing, accompaniment mixing processing; The synthesizing the multiple split-screen videos of the accompanying video with the audio to generate the audio-visual work corresponding to the audio includes: Synthesizing the multiple split-screen videos of the accompanying video with the audio after audio processing to generate the audio-visual work corresponding to the audio.

8. A video generation device, characterized in that, The device includes: An acquisition module, configured to acquire the user's audio and the content text of the audio; A split-screen description generation module, configured to input the content text into a large language model to enable the large language model to generate multiple split-screen description texts of the accompanying video of the audio based on the content text; A split-screen video generation module, configured to generate multiple split-screen video generation prompt words based on the multiple split-screen description texts, input the multiple split-screen video generation prompt words into a video generation large model, and enable the video generation large model to generate multiple split-screen videos of the accompanying video based on the multiple split-screen video generation prompt words; An audio-visual generation module, configured to synthesize the multiple split-screen videos of the accompanying video with the audio to generate the audio-visual work corresponding to the audio.

9. An electronic device, characterized in that, The electronic device includes: A processor; and A memory arranged to store computer-executable instructions, and the executable instructions, when executed, cause the processor to execute the video generation method according to any one of claims 1 to 7.

10. A readable storage medium, characterized in that, The storage medium is used to store computer-executable instructions, and the executable instructions, when executed by the processor, implement the video generation method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Music video generation method and device, equipment and storage medium

    CN121418302A

  • Music video generation method and device, equipment and storage medium

    CN121691838A

  • Digital human audio and video processing method and device, storage medium and program product

    CN122205198A

  • Digital human audio and video processing method and device, storage medium and program product

    CN122205198B