Audio-based video generation method and device, readable medium and electronic equipment
By determining the musical characteristics of the audio data in the video generation method and generating a matching skeleton key point sequence, combined with the action object diagram, the problem of mismatch between the video content and the audio data style in the prior art is solved, the deep matching between the action object and the music feature is achieved, and the expression of the video content is enriched.
Patent Information
- Application Number
- CN202510148311.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-13
AI Technical Summary
In the field of video generation, when the input data provided includes audio data, the prior art usually generates a spectrum dynamic visual effect corresponding to the music data, resulting in limitations in the video content, or directly adding the audio data as background music, resulting in the inability to deeply match the style of the video content and the audio data.
By responding to a video generation request, the music characteristics of the audio data in at least one dimension are determined, and a sequence of bone key points matching the music characteristics is generated, and video data is generated based on the sequence of bone key points and the action object diagram, so that the action changes of the action objects in the video data match the music characteristics of the audio data in the depth.
The action changes of the action objects in the video data are deeply matched with the music characteristics of the audio data, enrich the expression of video content, and improve the flexibility and diversity of video generation.
Smart Images

Figure CN119996725A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of computer technology, and more specifically, to an audio-based video generation method, an audio-based video generation device, a computer-readable storage medium, and an electronic device. Background Art
[0002] This section is intended to provide a background or context to the embodiments of the disclosure that are recited in the claims, and no description herein is admitted to be prior art by inclusion in this section.
[0003] In the field of video generation, when the input data provided includes audio data, the spectral dynamic visual effect corresponding to the music data is usually generated, which limits the content expressed by the video; or it is directly added as background music, so that the style of the video content and the audio data cannot be deeply matched. Summary of the invention
[0004] In this context, embodiments of the present disclosure are intended to provide an audio-based video generation method, an audio-based video generation device, a computer-readable storage medium, and an electronic device.
[0005] According to the first aspect of the embodiment of the present disclosure, there is provided an audio-based video generation method, characterized in that the method comprises: in response to a video generation request, determining the music features of the audio data in at least one dimension; the video generation request comprises the audio data and an action object graph; generating a skeletal key point sequence matching the music features; and generating video data corresponding to the video generation request based on the skeletal key point sequence and the action object graph.
[0006] In an exemplary embodiment, determining the music features of audio data in at least one dimension includes: inputting the audio data into a music feature recognition classifier of the corresponding dimension to obtain output music features.
[0007] In an exemplary embodiment, a skeleton key point sequence matching a music feature is generated, including: inputting the music feature and the noisy skeleton key point sequence into a motion generation model, and obtaining the skeleton key point sequence output by the motion generation model after denoising the noisy skeleton key point sequence based on the music feature.
[0008] In an exemplary embodiment, the training steps of the action generation model are as follows: collect sample videos, and the sample videos correspond to music features of at least one dimension; extract skeleton key points from the sample videos to obtain a sample skeleton key point sequence; construct a sample data set based on the sample skeleton key point sequence and the music features of the corresponding dimension; train the action generation model on the sample data set until the prediction loss of the predicted skeleton key point sequence output by the action generation model meets the convergence condition.
[0009] In an exemplary embodiment, the prediction loss includes a diffusion reconstruction loss, which is used to characterize the consistency between the predicted skeleton key point sequence and the sample skeleton key point sequence.
[0010] In an exemplary embodiment, the prediction loss may also include at least one of a skeletal key point continuity loss and a music and skeletal key point consistency loss: a skeletal key point continuity loss, used to characterize the continuity of the predicted skeletal key point sequence between adjacent frames; a music and skeletal key point consistency loss, used to characterize the consistency between the music amplitude and the change in motion intensity of the predicted skeletal key point sequence in adjacent frames.
[0011] In an exemplary embodiment, generating video data corresponding to a video generation request according to a skeleton key point sequence and an action object graph includes: synthesizing the action object graph and the skeleton key point sequence frame by frame to generate video data corresponding to the video generation request.
[0012] According to the second aspect of the embodiment of the present disclosure, there is provided an audio-based video generation device, characterized in that the device comprises: a music feature extraction module, for determining the music features of audio data in at least one dimension in response to a video generation request; the video generation request comprises audio data and an action object graph; an action sequence generation module, for generating a skeletal key point sequence matching the music features; and a video data generation module, for generating video data corresponding to the video generation request based on the skeletal key point sequence and the action object graph.
[0013] In an exemplary embodiment, the music feature extraction module is specifically used to input audio data into a music feature recognition classifier of a corresponding dimension to obtain output music features.
[0014] In an exemplary embodiment, the action sequence generation module is specifically used to input music features and a noisy skeleton key point sequence into the action generation model, and obtain a skeleton key point sequence output by the action generation model after denoising the noisy skeleton key point sequence based on the music features.
[0015] In an exemplary embodiment, the device also includes a model training module, which is used to collect sample videos, and the sample videos correspond to music features of at least one dimension; extract skeleton key points from the sample videos to obtain a sample skeleton key point sequence; construct a sample data set based on the sample skeleton key point sequence and the music features of the corresponding dimension; and train an action generation model on the sample data set until the prediction loss of the predicted skeleton key point sequence output by the action generation model meets the convergence condition.
[0016] In an exemplary embodiment, the prediction loss includes a diffusion reconstruction loss, which is used to characterize the consistency between the predicted skeleton key point sequence and the sample skeleton key point sequence.
[0017] In an exemplary embodiment, the prediction loss may also include at least one of a skeletal key point continuity loss and a music and skeletal key point consistency loss: a skeletal key point continuity loss, used to characterize the continuity of the predicted skeletal key point sequence between adjacent frames; a music and skeletal key point consistency loss, used to characterize the consistency between the music amplitude and the change in motion intensity of the predicted skeletal key point sequence in adjacent frames.
[0018] In an exemplary embodiment, the video data generation module is specifically used to synthesize the action object graph and the skeleton key point sequence frame by frame to generate video data corresponding to the video generation request.
[0019] According to a third aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, any of the above-mentioned audio-based video generation methods is implemented.
[0020] According to a fourth aspect of an embodiment of the present disclosure, there is provided an electronic device, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any one of the above-mentioned audio-based video generation methods by executing the executable instructions.
[0021] According to the audio-based video generation method, audio-based video generation device, computer-readable storage medium and electronic device provided by the embodiment of the present disclosure. The method determines the music features of audio data in at least one dimension in response to a video generation request, and the video generation request includes audio data and an action object graph; on this basis, a skeleton key point sequence matching the music features can be generated, and then the video data corresponding to the video generation request is generated according to the skeleton key point sequence and the action object graph. In the method, the skeleton key point sequence is generated based on the music features of the audio data, so that the action changes of the action objects in the video data are deeply matched with the music features of the audio data; the audio features can select one or more dimensions, so that the generation of the skeleton key point sequence is more diverse, and the action object graph can be freely provided, so that the matching of the action changes and the music features in the automatically generated video data is more flexible, and the expression of the video content is richer. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, in which:
[0023] Figure 1 A flowchart of a method for generating a video based on audio according to an embodiment of the present disclosure is shown;
[0024] Figure 2 A schematic diagram of skeleton key points of a human object according to an embodiment of the present disclosure is shown;
[0025] Figure 3 A schematic diagram of an action frame generation process according to an embodiment of the present disclosure is shown;
[0026] Figure 4 A flow chart of the training steps of the action generation model according to an embodiment of the present disclosure is shown;
[0027] Figure 5 A schematic diagram of an audio-based video generation device according to an embodiment of the present disclosure is shown;
[0028] Figure 6 A schematic diagram of an implementation architecture of an audio-based video generation device according to an embodiment of the present disclosure is shown;
[0029] Figure 7 A schematic diagram of a storage medium according to an embodiment of the present disclosure is shown;
[0030] Figure 8 A schematic diagram of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0031] The principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present disclosure, and are not intended to limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0032] Those skilled in the art will appreciate that the embodiments of the present disclosure may be implemented as a system, device, apparatus, method or computer program product. Therefore, the present disclosure may be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0033] According to an embodiment of the present disclosure, a method for generating a video based on audio, a device for generating a video based on audio, a computer-readable storage medium, and an electronic device are provided.
[0034] In this document, any number of elements in the drawings is used for illustration rather than limitation, and any naming is used only for distinction and does not have any limiting meaning.
[0035] The principle and spirit of the present disclosure are explained in detail below with reference to several representative embodiments of the present disclosure. SUMMARY OF THE INVENTION
[0037] At present, in the field of video generation, when the input data provided includes audio data, it is usually possible to generate spectral dynamic visual effects corresponding to the music data, such as particle effects, light column effects, etc. Spectral dynamics can visualize the beats and rhythm of music, but the content expressed by the video is still limited; alternatively, audio data can be directly added as video background music, but the style of the video content and the audio data cannot be deeply matched.
[0038] In the implementation of the present disclosure, the video generation request provides audio data and an action object graph, and on this basis, video data can be generated according to the audio data and the action object graph, so that the actions displayed by the objects constructed based on the action object graph in the video data match the music features of the aforementioned audio data. At this time, on the one hand, there is selectivity between the dimensions of the music features and the action object graph, which effectively expands the freedom and flexibility of the video content; on the other hand, matching the action objects in the video with the music features realizes a deep match between the video content and the style of the audio data.
[0039] After introducing the basic principles of the present invention, various non-limiting embodiments of the present invention are described in detail below.
[0040] Exemplary application scenarios
[0041] It should be noted that the following application scenarios are only shown to facilitate understanding of the spirit and principle of the present invention, and the embodiments of the present disclosure are not limited in this respect. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.
[0042] The audio-based video generation method of the disclosed embodiment can be applied to a variety of application scenarios involving audio-based video generation.
[0043] For example, it can involve audio and video production, publishing, sharing, and playback. In this application scenario, the user can trigger video generation in the music production interface, playback interface, details interface, or other related interfaces, and then obtain the corresponding audio data and action object graph to initiate a video generation request to generate video data matching the audio data; or, the user can select the corresponding music data in the video production interface or other related interfaces to trigger video generation. The above scenarios can involve music platforms, video platforms, information platforms, shopping platforms, and social platforms.
[0044] Exemplary Methods
[0045] In combination with the above application scenarios, Figure 1 The audio-based video generation method according to an exemplary embodiment of the present disclosure is described.
[0046] like Figure 1 The step flow of the audio-based video generation method according to an exemplary embodiment of the present disclosure is shown, and the method may include the following steps 101 to 103:
[0047] Step 101: In response to a video generation request, determine the music features of audio data in at least one dimension; the video generation request includes audio data and an action object graph.
[0048] In the disclosed embodiment, the video generation request may refer to a request to generate video data based on audio data and an action object graph, so that the action of the object generated by the action object graph in the video data matches the music features of the music data in the corresponding dimension. The audio data and the action object graph may be selected and provided by the initiator of the video generation request, or a default action object graph or audio data may be preset, and the disclosed embodiment does not specifically limit this.
[0049] The audio data can be selected from existing music, or it can be generated by AI (Artificial Intelligence) based on the provided text description or audio clip, or it can be background music extracted from the video. In response to the video generation request, the audio data can be extracted in one dimension or in different dimensions. The audio features of different dimensions have different emphases on the understanding of audio data, so that it can flexibly adapt to the diverse action choreography needs. For example, the dimensions of music features can include melody, harmony, rhythm, timbre, structure, dynamics, mode, strong and weak beats, paragraph division, singing skills, emotional expression, music intensity, music genre, lyrics content, etc. Furthermore, the dimensions of audio feature extraction can also be specified in the video generation request to meet the more in-depth personalization and customization needs of video generation.
[0050] The action object graph can specify the object of the action choreography in the video data. The object can be a human or animal image, or other images such as a car, an airplane, etc.
[0051] Step 102: Generate a skeleton key point sequence that matches the music features.
[0052] In the disclosed embodiment, the skeleton key point sequence may include a group of skeleton key points arranged in a time series. Each skeleton key point is used to describe the skeleton action at the corresponding time point, and the skeleton key point sequence describes the change of the skeleton action in the time series. The matching relationship between the music feature and the skeleton key point sequence can be obtained based on the existing video data analysis, the audio feature is determined from the background music of the existing video data, and the skeleton key points of the object in the frame image at the time point are extracted from the existing video data to obtain the skeleton key point sequence for matching analysis. Among them, it can be a database for building the matching relationship between the corresponding music feature and the skeleton key point sequence, and the music features of different dimensions can also be classified, stored and managed, so as to query the skeleton key point sequence corresponding to one or more dimensional music features when generating, and the skeleton key point sequence of multiple dimensions can be processed by weighted fusion, voting selection and the like; or, it is also possible to construct a sample data set with the music features extracted from the existing video data and the skeleton key point sequence for deep learning model training, so that the model can fully learn and understand the matching relationship between the music feature and the skeleton key point, and then when the new music feature is input, the matching skeleton key point sequence can be output correspondingly.
[0053] In the existing video data, when the object type is different, the corresponding skeletal key point types can be the same or different. For example, when the object type is human, the skeletal key points can usually include at least the head, torso and limbs, and the limbs include at least the left arm, right arm, left leg and right leg. When the object type is an animal such as a cat, dog, bird, etc., a similar skeletal key point structure can also be used to choreograph its movements; but when the object type is other objects, such as cars, airplanes, etc., other types of skeletal key points can be designed. When supporting multiple types of objects, in the process of generating a skeletal key point sequence that matches the music features, it can also be matched with the object indicated by the action object graph for subsequent fusion processing, such as the skeletal key point sequence can be classified, stored and managed based on the type of the object in the database, or the type of the object can be added as supervisory information in the training of the deep learning model.
[0054] like Figure 2 FIG. 1 is a schematic diagram of the skeleton key points of a human object in an exemplary embodiment of the present disclosure. Figure 2 In the figure, the human object is represented by 24 key points in [0,23], where 201 represents standing action, 202 represents running action, and 203 represents dancing action.
[0055] Taking 24 skeleton key points and generating 5 seconds of 30 frames per second video data as an example, the skeleton key point sequence can be expressed as [5*30, 24].
[0056] Step 103: Generate video data corresponding to the video generation request according to the skeleton key point sequence and the action object graph.
[0057] In the disclosed embodiment, on the basis of obtaining a skeleton key point sequence, the action object graph can be driven based on the skeleton key point sequence to realize the skeleton action at the corresponding time point, and then the frame at the corresponding time point can be generated to generate video data in which the frames are arranged in a time series. The action of the object in the video data is generated based on the skeleton key point sequence, so as to deeply match the music features of one or more dimensions of the audio data. In the disclosed embodiment, the action object graph can be freely provided, and the music features can be selectively and flexibly extracted in multiple dimensions, which effectively enriches the expression content of the video data, is more diverse and flexible, and also improves the matching depth between the action of the object in the video data and the music features to meet the needs of video generation.
[0058] In an exemplary embodiment of the present disclosure, the aforementioned step 101 may include the following step A.
[0059] Step A: input the audio data into the music feature recognition classifier of the corresponding dimension to obtain the output music features.
[0060] In the disclosed embodiment, the music feature recognition classifier can extract the music features of the corresponding dimension based on the input audio data. For the extraction of music features of different dimensions, different encoding methods can be used to train the music feature recognition classifiers of the corresponding dimensions. The audio feature recognition classifier can be implemented using the Transformer structure. The music feature recognition classifier is first constructed with the Transformer structure and the music feature representations of different dimensions are trained. On this basis, in response to the video generation request, the acquired audio data is input into the music feature recognition classifier of the corresponding dimension to obtain the output music features. Transformer is a deep learning model architecture for natural language processing (Natural Language Processing, NLP) and other sequence-to-sequence tasks. Those skilled in the art can apply the Transformer structure to implement the music feature recognition classifier according to actual needs, or make improvements based on the Transformer structure, or use other deep learning model architectures to implement the music feature recognition classifier, and the disclosed embodiment does not make specific restrictions on this.
[0061] In an exemplary embodiment of the present disclosure, the aforementioned step 102 may include the following step B.
[0062] Step B: input music features and a noisy skeleton key point sequence into the action generation model to obtain a skeleton key point sequence output by the action generation model after denoising the noisy skeleton key point sequence based on the music features.
[0063] In the disclosed embodiment, the action generation model can generate a skeleton key point sequence based on the music feature of the input and the noise skeleton key point sequence. The action generation model can denoise the noise skeleton key point sequence based on the music feature based on training learning, so as to output a skeleton key point sequence matching the music feature. The action generation model can be implemented using DDPM (Denoising Diffusion Probabilistic Models, denoising diffusion probability model), and the skeleton key points of the output are controlled on the basis of the noise skeleton key point sequence using the music feature as an input condition, so as to obtain a skeleton key point sequence matching the music feature in dimensions such as rhythm and style. Those skilled in the art can apply DDPM to implement the action generation model according to actual needs, or improve on the basis of DDPM, or implement the action generation model using other deep learning model architectures, and the disclosed embodiment does not specifically limit this.
[0064] In an exemplary embodiment of the present disclosure, the aforementioned step 103 may include the following step C.
[0065] Step C: synthesize the action object graph and the skeleton key point sequence frame by frame to generate video data corresponding to the video generation request.
[0066] In the disclosed embodiment, in the skeleton key point sequence, each skeleton key point can correspond to the action of the object in a frame, so the action object graph and the skeleton key point can be synthesized frame by frame, so that the object indicated by the action object graph generates the corresponding action frame based on the driving of the skeleton key point, thereby forming the video data corresponding to the video generation request in the time series. The frame-by-frame synthesis process can be implemented by using an image synthesis algorithm, and the action object graph is used as an input condition to generate the action frame corresponding to the skeleton key point, such as using Controlnet (control network) to conditionally control the image content generated by the diffusion model, so that the object indicated by the action object graph meets the conditional control of the skeleton key point in the generated action frame. Among them, Controlnet is a neural network that controls the pre-trained image diffusion model, and can receive the input of the conditional image to control the output image of the image diffusion model; the image diffusion model can be a Stable Diffusion (stable diffusion) model, or it can be other image diffusion models. Those skilled in the art can apply Stable Diffusion and Controlnet to perform frame synthesis according to actual needs, or make improvements on this basis, or use other deep learning model architectures to realize frame synthesis, and the disclosed embodiment does not make specific restrictions on this.
[0067] like Figure 3The schematic diagram of the action frame generation process of the exemplary embodiment of the present disclosure is shown, which synthesizes the action object graph 301 with the skeleton key point 302 of the nth frame to generate the corresponding action frame 303. In the action frame 303, the human body object indicated by the action object graph 301 is driven by the skeleton key point 302, thereby changing the human body action.
[0068] like Figure 4 The training process of the action generation model of the exemplary embodiment of the present disclosure is shown, which may include the following steps 401 to 404.
[0069] Step 401: Collect sample videos, where the sample videos correspond to music features of at least one dimension.
[0070] In the disclosed embodiments, the action generation model can be trained on a sample data set, and the sample data set can be constructed based on a sample video. The sample video can be video data of an object performing continuous actions in a time series during audio playback. The sample video can be collected on a related platform, such as collecting MVs (Music Videos, short music videos) that match music on a music platform, or collecting dance performance videos on a video platform; it can also be obtained by live shooting; it can also be amplified based on existing small batches of sample videos, such as changing the type of objects in the sample video through an image synthesis algorithm, etc. The disclosed embodiments do not impose specific restrictions on this.
[0071] Furthermore, the sample video may also correspond to at least one dimension of music features, which are used to characterize the music content of the sample video in this dimension. On this basis, the matching relationship between the continuous actions of the sample video object in the time sequence and the music content is used as the learning basis, and the subsequent training and learning of the matching relationship between the music features and the skeleton key point sequence in this dimension is carried out.
[0072] Step 402: extract skeleton key points from the sample video to obtain a sample skeleton key point sequence.
[0073] In the disclosed embodiment, the sample video includes sample frames that are continuous in time sequence, and the image content of the sample frames includes the action of the object at the time point. On this basis, the skeleton key points can be extracted from the sample frames in the sample video so that the skeleton key points characterize the action of the object in the sample frames. The skeleton key points are extracted on the continuous sample frames to obtain a skeleton key point sequence that is continuous in time sequence. Among them, each sample frame in the sample video can be extracted, or it can be extracted by taking out frames at intervals, and the changes in the image content can also be analyzed to determine the key frames in the sample frames and extract on the key frames, and the disclosed embodiment does not make specific restrictions on this.
[0074] Step 403: construct a sample data set based on the sample skeleton key point sequence and the music features of the corresponding dimensions.
[0075] In the disclosed embodiment, a sample skeleton key point sequence representing continuous actions in a sample video and music features under a dimension may be matched to construct a sample data set. The music features under a dimension in the sample data set are matched with the sample skeleton key points.
[0076] Step 404: train the action generation model on the sample data set until the prediction loss of the predicted skeleton key point sequence output by the action generation model meets the convergence condition.
[0077] In the disclosed embodiment, the action generation model can be trained on a sample data set, and the music features of the corresponding dimension are input into the action generation model to obtain a predicted skeleton key point sequence output by the action generation model, and the action generation model is adjusted by determining the prediction loss of the predicted skeleton key point sequence and the sample skeleton key point sequence. Specifically, the music features and the noisy skeleton key point sequence are input into the action generation model, and the action generation model denoises the noisy skeleton key point sequence based on the music features to output a predicted skeleton key point sequence. When the calculated prediction loss meets the convergence condition, it can be considered that the action generation model converges and can be applied to reasoning tasks. The convergence condition can be determined according to actual business needs, and the disclosed embodiment does not impose specific restrictions on this.
[0078] In an exemplary embodiment of the present disclosure, the prediction loss includes a diffusion reconstruction loss, and the diffusion reconstruction loss is used to characterize the consistency between the predicted skeleton key point sequence and the sample skeleton key point sequence.
[0079] In the disclosed embodiment, the prediction loss may include a basic diffusion reconstruction loss, and consistency calculation may be performed by predicting each predicted skeleton key point in the skeleton key point sequence and the sample skeleton key point at the corresponding time point in the sample skeleton key point sequence. The consistency may be characterized based on the position of the skeleton key point, indicating complete consistency when the predicted skeleton key point coincides with the sample skeleton key point, and indicating inconsistency when the predicted skeleton key point does not coincide with the sample skeleton key point. At this point, the degree of inconsistency may be determined based on the distance between the predicted skeleton key point and the sample skeleton key point. In the process of training the action generation model, parameters are adjusted in the direction of reducing the degree of inconsistency based on the diffusion reconstruction loss for learning.
[0080] In an exemplary embodiment of the present disclosure, the prediction loss may further include at least one of a skeleton key point continuity loss and a music and skeleton key point consistency loss.
[0081] In the disclosed embodiment, the prediction loss can adopt a composite multiple loss function to characterize the performance of the action generation model in the prediction process from multiple angles. For example, it can include the continuity loss of skeleton key points, the consistency loss of music and skeleton key points, etc., to characterize the continuity of the prediction results in time, the degree of matching with the music, etc. Those skilled in the art can choose one or two of them according to actual business needs, or can also choose other loss functions, and the disclosed embodiment does not make specific restrictions on this.
[0082] Skeleton key point continuity loss is used to characterize the continuity of predicted skeleton key point sequences between adjacent frames.
[0083] In the disclosed embodiment, in the composite prediction loss, the continuity of the predicted skeleton key point sequence between adjacent frames can be characterized from the time domain perspective. The continuity calculation can be performed by predicting the predicted skeleton key points corresponding to the time points of any frame in the predicted skeleton key point sequence, and predicting the predicted skeleton key points corresponding to the time points of the corresponding adjacent frames in the predicted skeleton key point sequence. The continuity can be characterized based on the position of the skeleton key points. The action should change between adjacent frames, but the change should be within a certain range so that the action has the continuity of the upper and lower frames. When the predicted skeleton key points of adjacent frames overlap, it means that the action has not changed. When the displacement distance between the predicted skeleton key points of adjacent frames is greater than the continuous threshold, it means that the action continuity is insufficient. When the displacement distance between the predicted skeleton key points of adjacent frames is less than or equal to the continuous threshold, it means that the action has continuity. At this time, the continuity can be determined based on the distance of the predicted skeleton key points of adjacent frames. In the process of training the action generation model, the direction adjustment parameters that characterize the continuity of the predicted skeleton key points of adjacent frames based on the skeleton key point continuity loss are learned.
[0084] The music and skeleton keypoint consistency loss is used to characterize the consistency between the music amplitude and the change of motion intensity of the predicted skeleton keypoint sequence in adjacent frames.
[0085] In the disclosed embodiment, the composite prediction loss may include the loss of consistency between music amplitude and action intensity. Among them, the music amplitude can represent the strength of the music, and the music amplitude and the predicted action intensity changes of the skeleton key points in adjacent frames are compared for consistency, so as to deeply match the music amplitude and action intensity. The action intensity can be characterized based on the distance between the skeleton key points in adjacent frames, so that the size of the music amplitude is associated with the distance between the predicted skeleton key points in adjacent frames. When the action intensity represented by the distance between the predicted skeleton key points of adjacent frames is equal to the music intensity represented by the music amplitude, it indicates complete consistency, and when the action intensity represented by the distance between the predicted skeleton key points of adjacent frames is greater than or equal to the music intensity represented by the sample skeleton key points, it indicates inconsistency. At this time, the degree of inconsistency can be determined based on the predicted skeleton key points and the music amplitude of the adjacent frames. In the process of training the action generation model, the parameters are adjusted in the direction of reducing the degree of inconsistency based on the consistency loss of music and skeleton key points for learning.
[0086] According to the audio-based video generation method of the embodiment of the present disclosure, the music features of the audio data in at least one dimension are determined in response to a video generation request, and the video generation request includes audio data and an action object graph; on this basis, a skeleton key point sequence matching the music features can be generated, and then the video data corresponding to the video generation request is generated according to the skeleton key point sequence and the action object graph. In this method, the skeleton key point sequence is generated based on the music features of the audio data, so that the action changes of the action objects in the video data match the music features of the audio data; the audio features can select one or more dimensions, so that the generation of the skeleton key point sequence is more diverse, and the action object graph can be freely provided, so that the matching of the action changes and the music features in the automatically generated video data is more flexible, and the expression of the video content is richer.
[0087] Exemplary Devices
[0088] After introducing the audio-based video generation method of the exemplary embodiment of the present disclosure, next, refer to Figure 5 An audio-based video generation apparatus according to an exemplary embodiment of the present disclosure will be described.
[0089] It should be noted that other specific details of each functional module of the audio-based video generation device of the embodiment of the present disclosure have been described in detail in the implementation of the above-mentioned audio-based video generation method, and will not be repeated here.
[0090] Figure 5An audio-based video generation device 500 according to an exemplary embodiment of the present disclosure is shown, comprising: a music feature extraction module 501, for determining the music features of audio data in at least one dimension in response to a video generation request; the video generation request comprises audio data and an action object graph; an action sequence generation module 502, for generating a skeleton key point sequence matching the music features; and a video data generation module 503, for generating video data corresponding to the video generation request based on the skeleton key point sequence and the action object graph.
[0091] In an exemplary embodiment, the music feature extraction module 501 is specifically used to input the audio data into a music feature recognition classifier of a corresponding dimension to obtain output music features.
[0092] In an exemplary embodiment, the action sequence generation module 502 is specifically used to input music features and a noisy skeleton key point sequence into the action generation model, and obtain a skeleton key point sequence output by the action generation model after denoising the noisy skeleton key point sequence based on the music features.
[0093] In an exemplary embodiment, the device also includes a model training module, which is used to collect sample videos, and the sample videos correspond to music features of at least one dimension; extract skeleton key points from the sample videos to obtain a sample skeleton key point sequence; construct a sample data set based on the sample skeleton key point sequence and the music features of the corresponding dimension; and train an action generation model on the sample data set until the prediction loss of the predicted skeleton key point sequence output by the action generation model meets the convergence condition.
[0094] In an exemplary embodiment, the prediction loss includes a diffusion reconstruction loss, which is used to characterize the consistency between the predicted skeleton key point sequence and the sample skeleton key point sequence.
[0095] In an exemplary embodiment, the prediction loss may also include at least one of a skeletal key point continuity loss and a music and skeletal key point consistency loss: a skeletal key point continuity loss, used to characterize the continuity of the predicted skeletal key point sequence between adjacent frames; a music and skeletal key point consistency loss, used to characterize the consistency between the music amplitude and the change in motion intensity of the predicted skeletal key point sequence in adjacent frames.
[0096] In an exemplary embodiment, the video data generation module 503 is specifically configured to synthesize the action object graph and the skeleton key point sequence frame by frame to generate video data corresponding to the video generation request.
[0097] According to the audio-based video generation device of the embodiment of the present disclosure, the music features of the audio data in at least one dimension are determined in response to a video generation request, and the video generation request includes audio data and an action object graph; on this basis, a skeleton key point sequence matching the music features can be generated, and then the video data corresponding to the video generation request is generated according to the skeleton key point sequence and the action object graph. In this method, the skeleton key point sequence is generated based on the music features of the audio data, so that the action changes of the action objects in the video data match the music features of the audio data; the audio features can select one or more dimensions, so that the generation of the skeleton key point sequence is more diverse, and the action object graph can be freely provided, so that the matching of the action changes and the music features in the automatically generated video data is more flexible, and the expression of the video content is richer.
[0098] like Figure 6 A schematic diagram of an implementation architecture of an audio-based video generation device according to an exemplary embodiment of the present disclosure is shown. In response to a video generation request, the video generation request includes audio data and an action object graph.
[0099] The music feature extraction module 601 may be composed of music feature recognition classifiers corresponding to multiple dimensions, and the music data is input into the music feature recognition classifier to obtain the music features of the corresponding dimensions.
[0100] The action sequence generation module 602 can be implemented based on the DDPM algorithm, and the music features and the noisy skeleton key point sequence are input. The DDPM algorithm is used to perform denoising and output the skeleton key point sequence with the music features as input conditions.
[0101] The video data generation module 603 can be implemented based on Stable Diffusion and Controlnet architecture, inputting the skeleton key point sequence and the action object graph, and using Controlnet to synthesize the skeleton key point sequence frame by frame with the action object graph as input condition to obtain continuous video data in time series.
[0102] It should be noted that although several modules or units of the audio-based video generation device are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.
[0103] Exemplary Storage Media
[0104] The storage medium according to the exemplary embodiment of the present disclosure is described below.
[0105] In this exemplary embodiment, reference Figure 7 As shown, a program product 700 for implementing the above method according to an exemplary embodiment of the present disclosure is described, such as a portable compact disk read-only memory (CD-ROM) and including program code, and can be run on a device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.
[0106] The program product 700 may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0107] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, in which readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Readable signal media may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0108] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RE, etc., or any suitable combination of the foregoing.
[0109] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user computing device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (FAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).
[0110] Exemplary Electronic Devices
[0111] refer to Figure 8 An electronic device according to an exemplary embodiment of the present disclosure is described.
[0112] Figure 8 The electronic device 800 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0113] like Figure 8 As shown, the electronic device 800 is in the form of a general computing device. The components of the electronic device 800 may include, but are not limited to: at least one processing unit 810, at least one storage unit 820, a bus 830 connecting different system components (including the storage unit 820 and the processing unit 810), and a display unit 840.
[0114] The storage unit stores program codes, which can be executed by the processing unit 810, so that the processing unit 810 performs the steps according to various exemplary embodiments of the present disclosure described in the above “Exemplary Method” section of this specification. For example, the processing unit 810 can perform the following steps: Figure 1 The method steps shown, etc.
[0115] The storage unit 820 may include a volatile storage unit, such as a random access storage unit (RAM) 821 and / or a cache storage unit 822 , and may further include a read-only storage unit (ROM) 823 .
[0116] The storage unit 820 may also include a program / utility 824 having a set (at least one) of program modules 825, such program modules 825 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0117] The bus 830 may include a data bus, an address bus, and a control bus.
[0118] The electronic device 800 can also communicate with one or more external devices 900 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), and this communication can be performed through an input / output (I / O) interface 850. The electronic device 800 also includes a display unit 840, which is connected to the input / output (I / O) interface 850 for display. In addition, the electronic device 800 can also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs) and / or public networks, such as the Internet) through a network adapter 860. As shown in the figure, the network adapter 860 communicates with other modules of the electronic device 800 through a bus 830. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0119] It should be noted that, although several modules or submodules of the device are mentioned in the above detailed description, such division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided into multiple units / modules to be embodied.
[0120] In addition, although the operations of the disclosed method are described in a specific order in the drawings, this does not require or imply that the operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0121] Although the spirit and principle of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the disclosed specific embodiments, and the division of various aspects does not mean that the features in these aspects cannot be combined to benefit, and such division is only for the convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included in the spirit and scope of the attached claims.
Claims
1. A method for generating a video based on audio, characterized in that: The method comprises: In response to a video generation request, determining a music feature of the audio data in at least one dimension; the video generation request includes the audio data and an action object graph; Generating a skeleton key point sequence matching the music feature; The video data corresponding to the video generation request is generated according to the skeleton key point sequence and the action object graph.
2. The method according to claim 1, characterized in that: Determining the music features of the audio data in at least one dimension includes: The audio data is input into a music feature recognition classifier of corresponding dimension to obtain the output music features.
3. The method according to claim 1, characterized in that The generating of a skeleton key point sequence matching the music feature comprises: The music feature and the noisy skeleton key point sequence are input into the action generation model to obtain the skeleton key point sequence output by the action generation model after denoising the noisy skeleton key point sequence based on the music feature.
4. The method according to claim 3, characterized in that The training steps of the action generation model are as follows: Collecting a sample video, wherein the sample video corresponds to a music feature of at least one dimension; Extracting skeleton key points from the sample video to obtain a sample skeleton key point sequence; Constructing a sample data set based on the sample skeleton key point sequence and the music features of the corresponding dimensions; The action generation model is trained on the sample data set until the prediction loss of the predicted skeleton key point sequence output by the action generation model meets the convergence condition.
5. The method according to claim 4, characterized in that The prediction loss includes a diffusion reconstruction loss, and the diffusion reconstruction loss is used to characterize the consistency between the predicted skeleton key point sequence and the sample skeleton key point sequence.
6. The method according to claim 5, characterized in that The prediction loss may also include at least one of a skeleton key point continuity loss and a music and skeleton key point consistency loss: The skeleton key point continuity loss is used to characterize the continuity of the predicted skeleton key point sequence between adjacent frames; The music and skeleton key point consistency loss is used to characterize the consistency between the music amplitude and the change in action intensity of the predicted skeleton key point sequence in adjacent frames.
7. The method according to claim 1, characterized in that The step of generating video data corresponding to the video generation request according to the skeleton key point sequence and the action object graph includes: The action object graph and the skeleton key point sequence are synthesized frame by frame to generate video data corresponding to the video generation request.
8. An audio-based video generation device, characterized in that: The device comprises: A music feature extraction module, configured to determine music features of the audio data in at least one dimension in response to a video generation request; the video generation request comprising the audio data and an action object graph; An action sequence generation module, used to generate a skeleton key point sequence matching the music feature; The video data generation module is used to generate video data corresponding to the video generation request according to the skeleton key point sequence and the action object graph.
9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating a video based on audio according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to execute the audio-based video generation method according to any one of claims 1 to 7 by executing the executable instructions.