Method, storage medium, and electronic device for generating a music video

Audio is classified and separated through network models, harmonics and shock waves are generated, and music videos are automatically generated, solving the problems of high labor costs and low efficiency in the existing technology, and achieving efficient and low-cost music video production.

CN114067840BActive Publication Date: 2025-05-30TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111348161.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-15
Publication Date
2025-05-30
Estimated Expiration
2041-11-15

AI Technical Summary

Technical Problem

When generating music videos in the prior art, the labor cost is high and the efficiency is low, making it difficult to use on a large scale.

Method used

By classifying the audio using the first network model, the second network model separates the audio tracks, generates harmonics and shock waves, and automatically generates music videos based on these features.

Benefits of technology

It reduces labor costs and improves the efficiency of music video production. The generated video matches audio categories and has a high content matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114067840B_ABST
    Figure CN114067840B_ABST
Patent Text Reader

Abstract

The present application discloses a method for generating a music video, including: classifying the target audio by using a first network model to obtain the audio category corresponding to the target audio; performing track separation processing on the target audio by using a second network model to obtain a plurality of separated tracks; generating the harmonics and shock waves of each of the separated tracks, and generating an audio feature vector for each audio frame based on the harmonics and shock waves of each of the separated tracks; generating an increment of the audio feature vector for each audio frame based on the audio feature vector for each audio frame; processing the increment of the audio feature vector for each audio frame by using a third network model corresponding to the audio category to obtain a video frame corresponding to each audio frame; and performing synthesis processing on the video frames corresponding to each audio frame to generate a target dynamic video. The present application also provides a computer-readable storage medium and an electronic device. The solution of the present application can efficiently generate a music video associated with the target audio type, and the generated music video can match the audio features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of multimedia content processing, and specifically, to a method for generating a music video. In addition, this application also relates to related computer-readable storage media and electronic devices. Background Art

[0002] Currently, many songs without official MVs usually use video editing methods to produce music videos, that is, manually sample from movies, TV dramas or animations, and match the sampled videos with the mood or rhythm of the music being played according to the visual effects. This way of generating music videos has a high labor cost and low video production efficiency, and it is difficult to be applied on a large scale. Summary of the Invention

[0003] Therefore, embodiments of the present invention intend to provide a method for generating a music video, as well as a computer-readable storage medium and an electronic device. These solutions can automatically classify the audio style according to the audio data, and automatically generate a music video based on the audio style classification and audio features, thereby reducing the labor cost and improving the efficiency of music video production.

[0004] In an embodiment of the present invention, a method for generating a music video is provided, including the following steps:

[0005] Classify the target audio using a first network model to obtain the audio category corresponding to the target audio;

[0006] Perform audio track separation processing on the target audio using a second network model to obtain multiple separated audio tracks;

[0007] Generate the harmonics and shock waves of each of the separated audio tracks;

[0008] Generate an audio feature vector for each audio frame of the target audio based on the harmonics and shock waves of each of the separated audio tracks;

[0009] Generate an increment of the audio feature vector for each audio frame based on the audio feature vector of each audio frame;

[0010] Process the increment of the audio feature vector of each audio frame using a third network model corresponding to the audio category to obtain a video frame corresponding to each audio frame;

[0011] Perform synthesis processing on the video frames corresponding to each audio frame to generate a target dynamic video.

[0012] In some embodiments of the present invention, the first network model includes an encoding neural network and a projection neural network connected to the output layer of the encoding neural network. The first network model is trained and generated through the following steps:

[0013] Obtain N training audio segments of different music categories, and select two partially overlapping or non-overlapping samples xi and xj from each of the training audio segments respectively;

[0014] Select the samples xi and xj of any one training audio segment for data augmentation processing to obtain the augmented sample xI and the augmented sample xJ, use the augmented sample xI and the augmented sample xJ as positive samples, and use the samples xi and xj of the remaining N-1 training audio segments as negative samples;

[0015] Self-supervise the training of the positive samples and the negative samples by using the contrast loss function to obtain the encoding neural network and the projection neural network.

[0016] In some embodiments of the present invention, the second network model is a waveform-to-waveform model having a semantic segmentation (U-Net) network and a bidirectional long short-term memory (LSTM) network.

[0017] In some embodiments of the present invention, generating the harmonics and shock waves of each of the separated audio tracks includes:

[0018] Convert the time series of each separated audio track into a short-time Fourier transform matrix;

[0019] Process the short-time Fourier transform matrix corresponding to each separated audio track by using a median filter to obtain the initial harmonics and initial shock waves corresponding to each separated audio track;

[0020] Perform an inverse short-time Fourier transform on the initial harmonics and initial shock waves corresponding to each separated audio track, and adjust the time series lengths of the initial harmonics and initial shock waves after the inverse short-time Fourier transform to match the time series length of each separated audio track to generate the harmonics and shock waves of each separated audio track.

[0021] In some embodiments of the present invention, generating the audio feature vector of each audio frame of the target audio based on the harmonics and shock waves of each of the separated audio tracks includes;

[0022] If the separated audio track includes an accompaniment track, generate a pulse feature vector by using the shock wave of the accompaniment track, and generate an action feature vector by using the harmonics of the accompaniment track;

[0023] If the separated audio track includes a vocal track, generate a vocal pitch feature vector by using the harmonics of the vocal track;

[0024] Use the pulse feature vector, the action feature vector, and the vocal pitch feature vector as the audio feature vector of each audio frame.

[0025] In some embodiments of the present invention, generating the pulse feature vector by using the shock wave of the accompaniment track includes:

[0026] Convert the shock wave of the accompaniment track into a spectrogram;

[0027] Multiply the spectrogram by a number of Mel filters to obtain a Mel spectrogram feature matrix;

[0028] Based on the maximum Mel frequency in the Mel spectrogram feature matrix, normalize the Mel spectrogram feature matrix;

[0029] Reduce the dimension of the normalized Mel spectrogram feature matrix to a vector under each audio frame as the pulse feature vector.

[0030] In some embodiments of the present invention, the generating of the action feature vector by using the harmonics of the accompaniment track includes:

[0031] Convert the harmonics of the accompaniment track into a spectrogram;

[0032] Multiply the spectrogram by a number of Mel filters to obtain a harmonic Mel spectrogram feature matrix;

[0033] Perform cepstral analysis on the harmonic Mel spectrogram feature matrix to obtain a Mel frequency cepstral coefficient feature matrix, and calculate the mean of the Mel frequency cepstral coefficient features of each audio frame;

[0034] Use the mean of the Mel frequency cepstral coefficient features of each audio frame to normalize the Mel frequency cepstral coefficient features;

[0035] Reduce the dimension of the normalized Mel frequency cepstral coefficient feature matrix to a vector under each audio frame as the action feature vector.

[0036] In some embodiments of the present invention, the generating of the human voice pitch feature vector by using the harmonics of the human voice track includes:

[0037] Take the absolute value after performing CQT transformation on the harmonics of the human voice track to obtain the absolute value of the CQT transformation at each time point;

[0038] Map the absolute value of the CQT transformation to a chromagram to generate an initial chromagram CQT transformation feature matrix;

[0039] Normalize the initial chromagram CQT transformation feature matrix to generate a chromagram CQT transformation feature matrix;

[0040] Calculate the weighted average chromagram value according to the chromagram values corresponding to each audio frame, where each audio frame corresponds to the chromagram values of T scales;

[0041] Use the weighted average chromagram value corresponding to each audio frame to normalize the chromagram CQT transformation feature matrix;

[0042] The normalized chroma CQT transform feature matrix is reduced to a vector for each audio frame as the human voice pitch feature vector.

[0043] In some embodiments of the present invention, taking the pulse feature vector, the action feature vector, and the human voice pitch feature vector as the audio feature vector of the audio frame includes:

[0044] Using a filter to perform smoothing processing on the pulse feature vector, the action feature vector, and the human voice pitch feature vector along the time axis, and taking the smoothed pulse feature vector, action feature vector, and human voice pitch feature vector as the audio feature vector of the audio frame.

[0045] In some embodiments of the present invention, the increment of the audio feature vector includes one or more of the pulse feature vector increment, the action feature vector increment, the human voice pitch feature vector increment, and the composite audio feature vector increment latent z.

[0046] In some embodiments of the present invention, generating the composite feature vector increment for each audio frame based on the audio feature vector of each audio frame includes:

[0047] Generating a base noise vector for each audio frame;

[0048] Summing up the action feature vector increments of each audio frame between the first audio frame and the current audio frame of the target audio to obtain the cumulative action feature vector increment of the current audio frame;

[0049] Summing up the base noise vector of the current audio frame, the pulse feature vector increment of the current audio frame, the human voice pitch feature vector increment of the current audio frame, and the cumulative action feature vector increment of the current audio frame to generate the composite audio feature vector increment of the current audio frame;

[0050] Loop through the above steps to obtain the composite audio feature vector increment for each audio frame, where the composite audio feature vector increment is used as the audio feature vector increment.

[0051] Further, the pulse feature vector increment, the action feature vector increment, and the human voice pitch feature vector increment of the audio frame are generated in the following manner:

[0052] Construct the basis vectors of the pulse feature vector, the action feature vector, and the human voice pitch feature vector;

[0053] Generate an action random factor at a predetermined time interval;

[0054] Multiplying the basis vector of the pulse feature vector by the pulse feature vector of each audio frame to generate the pulse feature vector increment of each audio frame;

[0055] Multiply the basis vectors of the action feature vectors, the action feature vectors of each audio frame, the action random factors of each audio frame, and the action direction factors of each audio frame to generate the increment of the action feature vector for each audio frame;

[0056] Multiply the basis vector of the pitch feature vector of the human voice and the pitch feature vector of the human voice of each audio frame to generate the increment of the pitch feature vector of the human voice for each audio frame.

[0057] In some embodiments of the present invention, the generation of the basic noise vector for each audio frame includes:

[0058] Generate a normal distribution vector in the order of audio frames based on the standard normal distribution, and truncate the normal distribution vector in the order of audio frames according to the threshold range as the basic noise vector.

[0059] In some embodiments of the present invention, the generation of the basic noise vector for each audio frame includes:

[0060] Generate a truncated normal distribution vector with upper and lower limits of [-2, 2] for [512, number of audio frames] based on the standard normal distribution as the basic noise vector.

[0061] In some embodiments of the present invention, the audio feature vector increment includes a composite audio feature vector increment latent z; wherein, the processing of the audio feature vector increment by using a third network model corresponding to the audio category to obtain a video frame corresponding to each audio frame includes:

[0062] Generate a composite audio feature vector increment matrix latent Z based on the composite audio feature vector increment latent z of each audio frame;

[0063] Select the composite audio feature vector increment corresponding to each audio frame from the composite audio feature vector increment matrix and input it into a third network model corresponding to the audio category to obtain a video frame corresponding to each audio frame.

[0064] In some embodiments of the present invention, the third network model includes a mapping network part and a comprehensive network part; selecting the composite audio feature vector increment corresponding to each audio frame from the audio feature vector increment matrix and inputting it into a third network model corresponding to the audio category to obtain a video frame corresponding to each audio frame includes:

[0065] Input the composite audio feature vector increment of the audio frame into the mapping network part to map and obtain a composite audio feature vector increment mapping vector;

[0066] Input the increment mapping vector of the composite audio feature vector into each layer of the comprehensive network part to generate a video frame corresponding to the audio frame.

[0067] In some embodiments of the present invention, the method further includes:

[0068] Add corresponding synchronization special effects to the video frame according to the intensity of the pulse feature vector corresponding to each audio frame;

[0069] Perform super-resolution optimization on the video frame.

[0070] In some embodiments of the present invention, generating the increment of the audio feature vector for each audio frame based on the audio feature vector of each audio frame includes:

[0071] Construct the basis vectors of the pulse feature vector, the basis vectors of the action feature vector, and the basis vectors of the human voice pitch feature vector;

[0072] Generate an action random factor at a predetermined time interval;

[0073] Multiply the basis vector of the pulse feature vector and the pulse feature vector of each audio frame to generate the increment of the pulse feature vector for each audio frame;

[0074] Multiply the basis vector of the action feature vector, the action feature vector of each audio frame, the action random factor of each audio frame, and the action direction factor of each audio frame to generate the increment of the action feature vector for each audio frame;

[0075] Multiply the basis vector of the human voice pitch feature vector and the human voice pitch feature vector of each audio frame to generate the increment of the human voice pitch feature vector for each audio frame;

[0076] Wherein, the increment of the pulse feature vector, the increment of the action feature vector, and the increment of the human voice pitch feature vector are used as the increment of the audio feature vector.

[0077] In some embodiments of the present invention, multiplying the basis vector of the pulse feature vector and the pulse feature vector of each audio frame to generate the increment of the pulse feature vector for each audio frame includes:

[0078] In the first audio frame, multiply the basis vector of the pulse feature vector by the pulse feature vector of the first audio frame to generate the increment of the pulse feature vector of the first audio frame; in the m-th audio frame, where m is greater than or equal to 2, multiply the basis vector of the pulse feature vector by the pulse feature vector of the m-th audio frame to generate the initial increment of the pulse feature vector of the m-th audio frame;

[0079] In some embodiments of the present invention, generating an action feature vector increment for each audio frame by multiplying the basis vector of the action feature vector, the action feature vector of each audio frame, the action random factor of each audio frame, and the action direction factor of each audio frame includes:

[0080] In the first audio frame, multiplying the basis vector of the action feature vector, the action feature vector of the first audio frame, the action random factor of the first audio frame, and the action direction factor of the first audio frame to generate an action feature vector increment for the first audio frame; in the m-th audio frame, performing a weighted average process on the initial increment of the pulse feature vector of the m-th audio frame and the increment of the pulse feature vector of the (m - 1)-th audio frame to generate an increment of the pulse feature vector of the m-th audio frame; multiplying the basis vector of the action feature vector, the action feature vector of the m-th audio frame, the action random factor of the m-th audio frame, and the action direction factor of the m-th audio frame to generate an initial increment of the action feature vector of the m-th audio frame; performing a weighted average process on the initial increment of the action feature vector of the m-th audio frame and the increment of the action feature vector of the (m - 1)-th audio frame to generate an action feature vector increment for the m-th audio frame;

[0081] In some embodiments of the present invention, generating a pitch feature vector increment for each audio frame by multiplying the basis vector of the pitch feature vector of a person's voice and the pitch feature vector of each audio frame includes:

[0082] In the first audio frame, multiplying the basis vector of the pitch feature vector of a person's voice and the pitch feature vector of the first audio frame to generate a pitch feature vector increment for the first audio frame; in the m-th audio frame, multiplying the basis vector of the pitch feature vector of a person's voice and the pitch feature vector of the m-th audio frame to generate an initial increment of the pitch feature vector of the m-th audio frame; performing a weighted average process on the initial increment of the pitch feature vector of the m-th audio frame and the increment of the pitch feature vector of the (m - 1)-th audio frame to generate a pitch feature vector increment for the m-th audio frame.

[0083] In some embodiments of the present invention, when performing the weighted average process, the weight of the increment of the (m - 1)-th audio frame is 0.75, and the weight of the initial increment of the m-th audio frame is 0.25.

[0084] In some embodiments of the present invention, the method further includes: if the value obtained by adding or subtracting the reaction coefficient of the action feature vector from the absolute value of the audio feature vector increment of the current audio frame is greater than twice the preset truncation value, changing the positive or negative of the action direction factor.

[0085] In some embodiments of the present invention, the action random factor is generated every four seconds, and the random factor takes values between (0.5, 1).

[0086] In some embodiments of the present invention, the third network model corresponding to the audio category is generated through the following steps:

[0087] Obtain video materials corresponding to different audio categories, perform frame extraction on the video materials, scale the frame-extracted video materials to a predetermined size and input them into the adversarial network model for training to generate a third network model corresponding to different audio categories.

[0088] In some embodiments of the present invention, it further includes: adding corresponding synchronization special effects to the video frame according to the intensity of the pulse feature vector corresponding to each audio frame.

[0089] In some embodiments of the present invention, it further includes: a step of performing super-resolution optimization on the video frame.

[0090] In some embodiments of the present invention, the third network model includes a mapping network part and a comprehensive network part; using the third network model corresponding to the audio category to process the audio feature vector increment to obtain a video frame corresponding to each audio frame includes:

[0091] Input the pulse feature vector increment, the action feature vector increment, and the human voice pitch feature vector increment into the mapping network part respectively, and map to obtain a plurality of audio feature vector increment mapping vectors;

[0092] Among the plurality of audio feature vector increment mapping vectors, input the audio feature vector increment mapping vectors corresponding to the action feature vector increment and the human voice pitch feature vector increment into the front network layer of the comprehensive network part, and input the audio feature vector increment mapping vector corresponding to the pulse feature vector increment among the plurality of audio feature vector increment mapping vectors into the rear network layer of the comprehensive network part to generate a video frame corresponding to each audio frame.

[0093] In some embodiments of the present invention, the synthesizing process of the video frame corresponding to each audio frame to generate a target dynamic video includes:

[0094] Use ffmpeg to splice the video frames corresponding to each audio frame to generate the target dynamic video.

[0095] In some embodiments of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, wherein when the program is executed by a processor, it implements the method for generating a music video according to any embodiment of the present invention.

[0096] In some embodiments of the present invention, an electronic device is provided, including: a processor and a memory storing a computer program, and the processor is configured to execute the method for generating a music video according to any embodiment of the present invention when running the computer program.

[0097] An embodiment of the present invention provides a method for generating a music video based on audio feature increments. First, audio data is input into a first network model for audio data classification to determine the type of the audio data. Then, the audio data is processed by splitting tracks. The target track separated is processed to reduce the influence of background noise. Harmonics and shock waves are extracted from the separated track, and an audio feature vector is generated based on the harmonics and shock waves. An audio feature vector increment is generated based on the audio feature vector, and the audio feature vector increment is input into a third network model for calculation to generate video frames. The video frames are spliced together to form a dynamic video. Through the method for generating a music video in the embodiment of the present invention, a video that can be efficiently generated to match the audio data category and reflect the audio content can be obtained. The cost of producing the video is low, the efficiency is high, and the matching degree with the content is good.

[0098] Some other optional features and technical effects of the embodiments of the present invention are described below, and some can be understood by reading this article. BRIEF DESCRIPTION OF THE DRAWINGS

[0099] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. The elements shown are not limited by the proportions shown in the drawings. The same or similar reference numerals in the drawings represent the same or similar elements, where:

[0100] Figure 1 A flowchart showing the method for generating a music video in an embodiment of the present invention is shown;

[0101] Figure 2 A flowchart showing one process of training the first network model in the method for generating a music video in an embodiment of the present invention is shown;

[0102] Figure 3 A flowchart showing another process of training the first network model in the method for generating a music video in an embodiment of the present invention is shown;

[0103] Figure 4 A flowchart showing the process of track separation in the method for generating a music video in an embodiment of the present invention is shown;

[0104] Figure 5a A flowchart showing the process of generating a composite audio feature vector increment in the method for generating a music video in an embodiment of the present invention is shown;

[0105] Figure 5b A flowchart showing the process of generating a composite audio feature vector increment in the method for generating a music video in an embodiment of the present invention is shown;

[0106] Figure 6 A flowchart showing the process of generating video frames in the method for generating a music video in an embodiment of the present invention is shown;

[0107] Figure 7 Shows a schematic structural diagram of the music video generation device according to an embodiment of the present invention;

[0108] Figure 8 Shows a schematic structural diagram of the electronic device according to an embodiment of the present invention. Detailed implementation manners

[0109] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the detailed implementation manners and the accompanying drawings. Herein, the illustrative implementation manners of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.

[0110] In the embodiments of the present invention, "network" has its conventional meaning in the field of machine learning, such as neural network (NN), deep neural network (DNN), convolutional neural network (CNN), recurrent neural network (RNN), other machine learning or deep learning networks, or combinations or modifications thereof.

[0111] In the embodiments of the present invention, "model" has its conventional meaning in the field of machine learning. For example, the model can be a machine learning or deep learning model, such as a machine learning or deep learning model including the above-mentioned network or composed of the above-mentioned network.

[0112] In the embodiments of the present invention, "loss function" and "loss value" have their conventional meanings in the field of machine learning.

[0113] The embodiments of the present invention provide a method, a system, a device, a model, an electronic device and a storage medium for generating a music video. The method, system, device and model can be implemented by means of one or more computers. In some embodiments, the system, device and model can be implemented by software, hardware or a combination of software and hardware. In some embodiments, the electronic device or computer can be implemented by the computer described herein or other electronic devices capable of implementing corresponding functions.

[0114] In the embodiments of the present invention, the music video content includes images, videos and / or audios, which includes parts and / or combinations of images, videos and / or audios. In the embodiments of the present invention, the video content and the audio data are matched. For example, the richness of the video content changes with the richness of the audio rhythm. For another example, when the audio is gentle, the change of the video content is relatively stable.

[0115] As Figure 1 shown, a method for generating a music video according to an embodiment of the present invention includes steps S110-S170.

[0116] S110: Classify the target audio by using a first network model to obtain the audio category corresponding to the target audio.

[0117] In some embodiments of the present invention, a pre-trained first network model is used to process and classify the target audio to obtain the category of the target audio data. The target audio may be music audio data. Specifically, it may be a song, a song segment, or a song combination.

[0118] In some embodiments of the present invention, the audio category can be preset as needed. For example, it may include folk songs, children's songs, pop music, etc.

[0119] In some embodiments of the present invention, for the unlabeled data with a large music library, SimCLR of image contrast learning is applied to the audio field. The unprocessed music waveforms are contrastively learned, and the consistency between different augmented data of the same data is maximized through the contrast loss in the latent space, thereby learning and training the first network model. The first network model includes an encoding neural network (g enc ()) and a projection neural network connected to the output layer of the encoding neural network (g proj ()). As shown in Figure 2 , the first network model is generated through the following steps:

[0120] S111. Obtain N training audio segments of different music categories, and select two partially overlapping or non-overlapping samples xi and xj from each of the training audio segments.

[0121] S112. Select the samples xi and xj of any one training audio segment for data augmentation processing to obtain the augmented samples xI and xJ. The augmented samples xI and xJ are used as positive samples, and the samples xi and xj of the remaining N-1 training audio segments are used as negative samples.

[0122] Among them, a series of data augmentations can be performed according to probabilities, and each augmentation method has an independent probability. As shown in Figure 3 , select a random segment x of size 2N from one complete audio segment, that is, the 2N training audio segments, and randomly select two partially overlapping or non-overlapping samples (such as the x Figure 3 shown, x i,0 , x j,0 or X i,2N , X j,2N , etc.) so that the model can perform local and global inferences. Subsequently, a series of data augmentations are performed according to probabilities, and each augmentation method has an independent probability, x i,0 , x j,0After the above data augmentation, they are used as positive sample pairs. Additionally, 2(N - 1) samples are randomly sampled from the random segment x as negative samples. In the illustrated embodiment, in each iteration of training, two partially overlapping or non - overlapping samples can be continuously selected from the random segment x of size 2N.

[0123] S113. Self - supervised training of the positive samples and the negative samples is performed using a contrast loss function to obtain the encoding neural network and the projection neural network.

[0124] Exemplarily, the encoding neural network (such as Figure 3 the schematically shown genc(﹒)) uses a convolutional neural network (CNN), such as SampleCNN, as the encoder. The audio input is 59049 samples with a sampling rate of 22050Hz. The convolutional neural network (CNN), such as SampleCNN, consists of 9 one - dimensional convolutional blocks. Each convolutional block consists of 3 layers of 1 - D convolutional layers, batch normalization layers, ReLU layers, and 3 layers of max - pooling layers. In this embodiment, the convolutional neural network (CNN), such as SampleCNN, removes the fully - connected layer and the dropout (random abandonment) layer. Thus, for each audio input (such as Figure 3 the shown x i,0 、x j,0 or x i,2N 、x j,2N etc.), a 512 - dimensional feature vector is encoded and generated.

[0125] The encoded 512 - dimensional feature vector can be mapped to a latent space that forms a contrast loss through a projection neural network (such as Figure 3 the shown g proj (﹒)) for iterative update based on the contrast loss to complete the training.

[0126] In a specific implementation, a non - linear layer z i =W (2) ReLU(W (1) h i ) is used as the projection neural network. Additionally, the contrast loss function can be selected as the Normalised Temperature - scaled Cross - entropy Loss, usually referred to as the NT - Xent loss.

[0127] Due to the high cost of manual annotation, a contrastive learning method can be used to perform self-supervised training on music audio to reduce the manual annotation cost. During the training process, different categories of music are downloaded in batches from the music library to construct an audio dataset. For different categories of music, it is necessary to ensure the balance of the sample category distribution to obtain better training results.

[0128] S120: Use the second network model to perform track separation processing on the target audio to obtain multiple separated tracks.

[0129] It can be understood that there are multiple tracks in many music songs. If audio features are calculated from separate vocals, instruments, or accompaniments, the influence of background noise can be reduced, which helps to capture the changing characteristics of the music.

[0130] The second network model is used for track separation, so it can be called a track separation model. In some embodiments, the second network model is a Wave to Wave model with a semantic segmentation (U-Net) network structure and a bidirectional long short-term memory (LSTM) module. As Figure 4 shown in the exemplary embodiment of applying the Wave to Wave model, the process of the track separation model for track separation is as follows: The target audio (such as Figure 4 the target audio waveform shown in the upper left corner) is input into a one-dimensional convolutional layer with 8 layers and a stride of 4 through an encoder Encoder, and after being activated by a Relu layer, it is output to a one-dimensional convolutional layer with both the number of layers and the stride of 1 to multiply the number of channels by 2, and is activated by the gated linear unit GLU of the LSTM module; then it is input into a one-dimensional convolutional layer with 3 layers and a stride of 1 through a decoder (Decoder) and activated by GLU, and then output to a one-dimensional transposed convolutional layer with 8 layers and a stride of 4, and finally activated by Relu to output the separated tracks (such as Figure 4 the 4 separated track waveforms shown in the upper right corner).

[0131] Exemplarily, the types of separated tracks can include accompaniment tracks and vocal tracks.

[0132] S130: Generate the harmonics and shock waves of each of the separated tracks;

[0133] After separating the tracks, it is necessary to separate the harmonics and shock waves of each separated track, and then analyze and process the harmonics or shock waves of each separated track to generate an audio feature vector, as described in step S140 below.

[0134] In some embodiments of the present invention, the tracks can be separated into harmonics and shock waves based on median filtering. The specific steps of generating the harmonics and shock waves of each of the separated tracks include the following:

[0135] Convert the time series of each separated audio track into a short-time Fourier transform (STFT) matrix; process the short-time Fourier transform matrix corresponding to each separated audio track using a median filter to obtain an initial harmonic and an initial shock wave corresponding to each separated audio track; perform an inverse short-time Fourier (iSTFT) transform on the initial harmonic and the initial shock wave corresponding to each separated audio track, and adjust the time series lengths of the initial harmonic and the initial shock wave after the inverse short-time Fourier transform to match the time series length of each separated audio track, generating the harmonic and the shock wave of each separated audio track.

[0136] S140: Generate an audio feature vector for each audio frame of the target audio based on the harmonic and the shock wave of each of the separated audio tracks.

[0137] In an embodiment of the present invention, the target audio can be divided into multiple segments according to a predetermined time length, and each segment is an audio frame. In an embodiment of the present invention, the predetermined time length, that is, the length of the audio frame, can be determined according to the frame rate fps of the video to be generated (such as the reciprocal of the video frame rate), so that the video (video frame) to be generated can correspond exactly to the target audio.

[0138] Separate audio tracks of different types have different audio features, so effective audio features can be extracted according to the type of the separated audio track. For example: for an accompaniment audio track, the Mel spectrum feature of its shock wave can be extracted to reflect the strength of the audio, or the MFCC feature of its harmonic can also be extracted to reflect the change in timbre; for a human voice audio track, the chromatic constant-Q transform of its harmonic can be extracted to reflect the change in pitch.

[0139] In an embodiment of the present invention, the audio feature vector includes: a pulse feature vector, a motion feature vector, and a human voice pitch feature vector.

[0140] In a specific embodiment, generating an audio feature vector for each audio frame based on the harmonic and the shock wave of each separated audio track includes;

[0141] Generate a pulse feature vector using the shock wave of the accompaniment audio track; generate a motion feature vector using the harmonic of the accompaniment audio track; generate a human voice pitch feature vector using the harmonic of the human voice audio track.

[0142] Specifically, generating an audio feature vector for each audio frame of the target audio based on the harmonic and the shock wave of each of the separated audio tracks includes;

[0143] If the separated audio track includes an accompaniment audio track, generate a pulse feature vector using the shock wave of the accompaniment audio track, and generate a motion feature vector using the harmonic of the accompaniment audio track;

[0144] If the separated audio track includes a vocal track, generate a vocal pitch feature vector using the harmonics of the vocal track;

[0145] Use the pulse feature vector, the motion feature vector, and the vocal pitch feature vector as the audio feature vector for each audio frame.

[0146] In some embodiments of the present invention, the step of generating a pulse feature vector using the shock wave of the accompaniment track includes: converting the shock wave of the accompaniment track into a spectrogram; multiplying the spectrogram by a plurality of Mel filters to obtain a Mel spectrogram feature matrix; normalizing the Mel spectrogram feature matrix based on the maximum Mel frequency in the Mel spectrogram feature matrix; reducing the dimensionality of the normalized Mel spectrogram feature matrix to a vector for each audio frame as the pulse feature vector.

[0147] In this embodiment, the conversion relationship between the Mel frequency of the Mel filter and the spectrogram frequency of the spectrogram is:

[0148]

[0149] In some embodiments of the present invention, generating a motion feature vector using the harmonics of the accompaniment track includes: converting the harmonics of the accompaniment track into a spectrogram; multiplying the spectrogram by a plurality of Mel filters to obtain a harmonic Mel spectrogram feature matrix; performing cepstral analysis on the harmonic Mel spectrogram feature matrix to obtain a Mel frequency cepstral coefficient (MFCC) feature matrix, and calculating the mean of the Mel frequency cepstral coefficient features for each audio frame; using the mean of the Mel frequency cepstral coefficient features for each audio frame to normalize the Mel frequency cepstral coefficient features; reducing the dimensionality of the normalized Mel frequency cepstral coefficient feature matrix to a vector for each audio frame as the motion feature vector.

[0150] Specifically, the harmonics in the accompaniment track can be extracted, and then pre-emphasis, framing, windowing, FFT, Mel filter bank, logarithmic operation, and DCT transformation are performed to obtain the MFCC feature matrix of the accompaniment harmonics.

[0151] In some embodiments of the present invention, the method of generating a human voice pitch feature vector using the harmonics of a human voice track includes: performing CQT transformation on the harmonics of the human voice track and taking absolute values ​​to obtain the CQT transformation absolute values ​​at each time point; mapping the CQT transformation absolute values ​​to a chromatogram to generate an initial chromatogram CQT transformation feature matrix; normalizing the initial chromatogram CQT transformation feature matrix to generate a chromatogram CQT transformation feature matrix; calculating a weighted average chromatogram value based on the chromatogram value corresponding to each audio frame, wherein each audio frame corresponds to chromatogram values ​​of T scales; normalizing the chromatogram CQT transformation feature matrix using the weighted average chromatogram value corresponding to each audio frame; and reducing the dimension of the normalized chromatogram CQT transformation feature matrix to a vector under each audio frame as the human voice pitch feature vector.

[0152] In a specific embodiment, when the absolute value of the CQT transformation is mapped to the chromatogram, each time point may correspond to the chromatogram value of N scales, for example, each time point in the chromatogram corresponds to the chromatogram value of 12 scales, and the chromatogram value of each time point may be weighted summed and normalized to obtain the chromatogram constant Q transformation feature matrix. Then, the chromatogram constant Q transformation feature matrix may be reduced to a human voice pitch feature vector, and it can be seen that the human voice pitch feature vector is composed of the chromatogram average values ​​corresponding to each time point.

[0153] In the embodiment of the present invention, the constant Q transform (CQT) refers to a filter group whose center frequencies are distributed according to an exponential law, whose filter bandwidths are different, but whose center frequency to bandwidth ratio is a constant Q.

[0154] In some embodiments of the present invention, in order to avoid signal jitter, the signal can be filtered, for example, the pulse feature vector, the action feature vector and the human voice pitch feature vector are smoothed along the time axis using a filter to update the pulse feature vector, the action feature vector and the human voice pitch feature vector. Typically, signals calculated from audio, such as starting signals or chromatograms, are noisy and unstable, which may cause visual jitter in the final generated music video, or even cause the visual changes to be more dramatic than the corresponding audio. Therefore, the extracted pulse feature vector, action feature vector and human voice pitch feature vector can be smoothed along the time axis using a one-dimensional Gaussian filter to improve the stability of the generated music video.

[0155] S150: Generate an audio feature vector increment of each audio frame based on the audio feature vector of each audio frame.

[0156] The audio feature vector increment includes one or more of: a pulse feature vector increment, a motion feature vector increment, and a human voice pitch feature vector increment.

[0157] In some embodiments of the present invention, a composite audio feature vector increment for each audio frame may be generated based on the audio feature vector increment of each audio frame.

[0158] In some embodiments of the present invention, the composite audio feature vector increment may be denoted as latent z. In some embodiments of the present invention, as Figure 5a shown, step S150 may include:

[0159] S151: Generate a base noise vector for each audio frame;

[0160] S152: Sum up the action feature vector increments of each audio frame between the first audio frame and the current audio frame of the target audio to obtain the cumulative action feature vector increment of the current audio frame;

[0161] S153: Sum up the base noise vector of the current audio frame, the pulse feature vector increment of the current audio frame, the human voice pitch feature vector increment of the current audio frame, and the cumulative action feature vector increment of the current audio frame to generate the composite audio feature vector increment of the current audio frame;

[0162] S154: Loop through the above steps to obtain the composite audio feature vector increment of each audio frame, where the composite audio feature vector increment is used as the audio feature vector increment.

[0163] In one embodiment, step S151 of generating a base noise vector for each audio frame includes: generating a normal distribution vector in the order of audio frames based on the standard normal distribution, and truncating the normal distribution vector in the order of audio frames according to the threshold range as the base noise vector.

[0164] In this embodiment, the pulse feature vector increment, the action feature vector increment, and the human voice pitch feature vector increment of the audio frame are generated in the following manner: constructing the base vectors of the pulse feature vector, the action feature vector, and the human voice pitch feature vector; generating an action random factor at a predetermined interval; multiplying the base vector of the pulse feature vector by the pulse feature vector of each audio frame to generate the pulse feature vector increment of each audio frame; multiplying the base vector of the action feature vector, the action feature vector of each audio frame, the action random factor of each audio frame, and the action direction factor of each audio frame to generate the action feature vector increment of each audio frame; multiplying the base vector of the human voice pitch feature vector by the human voice pitch feature vector of each audio frame to generate the human voice pitch feature vector increment of each audio frame.

[0165] In some embodiments of the present invention, as Figure 5b shown, step S150 may include:

[0166] S151’: Construct the basis vectors of the pulse feature vector, the basis vectors of the action feature vector, and the basis vectors of the human voice pitch feature vector;

[0167] S152’: Generate an action random factor at a predetermined time interval;

[0168] S153’: Multiply the basis vector of the pulse feature vector and the pulse feature vector of each audio frame to generate the pulse feature vector increment of each audio frame;

[0169] S154’: Multiply the basis vector of the action feature vector, the action feature vector of each audio frame, the action random factor of each audio frame, and the action direction factor of each audio frame to generate the action feature vector increment of each audio frame;

[0170] S155’: Multiply the basis vector of the human voice pitch feature vector and the human voice pitch feature vector of each audio frame to generate the human voice pitch feature vector increment of each audio frame.

[0171] In some embodiments, the pulse feature vector increment, the action feature vector increment, and the human voice pitch feature vector increment obtained in steps S151’ to S154’ can be used to determine the composite audio feature vector increment as described in step S153. At this time, this composite audio feature vector increment can be used as the audio feature vector increment as described in step S150.

[0172] In other embodiments, the pulse feature vector increment, the action feature vector increment, and the human voice pitch feature vector increment can be directly used as the audio feature vector increment as described in step S150.

[0173] In some specific embodiments of the present invention, Figure 5a The embodiments or features of Figure 5b can be further combined with the embodiments or features of

[0174] to obtain new embodiments or examples.

[0175] A1. Construct the basic noise and basis vectors. Here, a truncated normal distribution vector with a dimension of [512, number of audio frames] and upper and lower limits of [-2, 2] can be generated based on the standard normal distribution as the basic noise. A sequence of normal distribution vectors is generated according to the order of the audio frame sequence, and the normal distribution vectors are truncated according to the threshold range [-2, 2]. The truncated normal distribution vectors have 512 dimensions, and the sequence of truncated normal distribution vectors is used as the basic noise. Basis vectors of pulse feature vectors, motion feature vectors, and human voice pitch feature vectors with a dimension of 512 are generated from pulse, melody, and human voice response coefficients. In some embodiments, the pulse, melody, and human voice response coefficients can be preset empirical coefficients, for example, they can be determined by statistical analysis of the spectral characteristics of existing audio segments.

[0176] A2. Initialize the random factor sign of the motion feature vector. To achieve the diversity of music video frame changes, a 512-dimensional space is set for the Motion increment, and a random factor (for example, choosing motion_randomness as 0.5) with each dimension size between (1 - motion_randomness, 1) is set, and it is re-initialized every 4 seconds, or every 4 audio frames.

[0177] A3. Generate an increment based on the audio feature vector. In some embodiments, at the first audio frame, the basis vectors of the pulse feature vector, human voice feature vector, etc. are multiplied by the pulse feature vector, motion feature vector, and human voice pitch feature vector to obtain the audio feature vector increment vector add = vector_base × feature_vector, where vector_base is the basis vector and feature_vector is the audio feature vector (which can be a pulse feature vector, motion feature vector, or human voice pitch feature vector). In some embodiments, when the basis vector of the motion feature vector is multiplied by the motion feature vector, it is further multiplied by the random factor and motion direction factor of the motion feature vector to obtain the increment vector corresponding to the current audio frame add = vector_base × feature_vector × sign × rand_factor, where sign is the motion direction factor and rand_factor is the random factor. In the audio frames after the first audio frame, in addition to the above calculations, the current increment is smoothed by taking a weighted average with the previous increment, and the weights can be set as needed. For example, the weight corresponding to the current increment can be set to 0.25, and the weight corresponding to the previous increment can be set to 0.75.

[0178] A4. Synthesize the composite audio feature vector increment latent z of the current audio frame i. Add the base noise to the increment of the impulse feature vector, action feature vector, and human voice pitch feature vector obtained in step A3 to obtain the composite audio feature vector increment. The increment of the impulse feature vector and the increment of the human voice pitch feature vector reflect the visual effects, while the increment of the action feature vector reflects the speed of visual effect deformation, and the increment will be accumulated into the base noise. The specific process can be expressed by the formula: latent z(i) = noise base (i) + motion sum [1:i+1] + pulse add + vocal add , where i represents the i-th audio frame, and noise base (i) represents the noise vector of the i-th audio frame, motion sum [1:i+1] represents the cumulative sum of the increment of the action feature vector from the first audio frame to the (i + 1)-th audio frame, pulse add represents the increment of the impulse feature vector, and vocal add represents the increment of the human voice pitch feature vector. In other words, in this embodiment, the composite audio feature vector increment latent z includes the increment of the impulse feature vector, the increment of the action feature vector, and the increment of the human voice pitch feature vector of the current audio frame, as well as the cumulative sum of the increment of the action feature vector from the first audio frame to the current audio frame.

[0179] In some embodiments of the present invention, the base noise vector, the increment of the impulse feature vector of each audio frame, the increment of the human voice pitch feature vector of each audio frame, and the cumulative sum of the increment of the action feature vector of each audio frame are summed to generate the composite audio feature vector increment of each audio frame.

[0180] Optionally, the action direction factor can be updated according to a predetermined condition. For the composite audio feature vector increment of each audio frame, if the absolute value of the (composite) audio feature vector increment plus or minus the action feature vector response coefficient is greater than twice the truncation value (e.g., the truncation value is 1), then change the positive and negative of the action direction factor. In the embodiments of the present invention, the action feature vector response coefficient can be a preset empirical coefficient, for example, it can be determined according to the statistical spectrum characteristics of existing audio segments.

[0181] In some embodiments of the present invention, the generation of the audio feature vector increment of each audio frame based on the audio feature vector of each audio frame can be combined with the following features:

[0182] Specifically, multiplying the basis vector of the pulse feature vector by the pulse feature vector of each audio frame to generate the increment of the pulse feature vector for each audio frame includes: in the first audio frame, multiplying the basis vector of the pulse feature vector by the pulse feature vector of the first audio frame to generate the increment of the pulse feature vector for the first audio frame. Multiplying the basis vector of the action feature vector, the action feature vector of each audio frame, the action random factor of each audio frame, and the action direction factor of each audio frame to generate the increment of the action feature vector for each audio frame includes: in the first audio frame, multiplying the basis vector of the action feature vector, the action feature vector of the first audio frame, the action random factor of the first audio frame, and the action direction factor of the first audio frame to generate the increment of the action feature vector for the first audio frame. Multiplying the basis vector of the human voice pitch feature vector by the human voice pitch feature vector of each audio frame to generate the increment of the human voice pitch feature vector for each audio frame includes: in the first audio frame, multiplying the basis vector of the human voice pitch feature vector by the human voice pitch feature vector of the first audio frame to generate the increment of the human voice pitch feature vector for the first audio frame;

[0183] Specifically, multiplying the basis vector of the pulse feature vector by the pulse feature vector of each audio frame to generate the increment of the pulse feature vector for each audio frame includes: in the m-th audio frame, where m is greater than or equal to 2, multiplying the basis vector of the pulse feature vector by the pulse feature vector of the m-th audio frame to generate the initial increment of the pulse feature vector for the m-th audio frame; performing weighted average processing on the initial increment of the pulse feature vector for the m-th audio frame and the increment of the pulse feature vector for the (m - 1)-th audio frame to generate the increment of the pulse feature vector for the m-th audio frame. Multiplying the basis vector of the action feature vector, the action feature vector of each audio frame, the action random factor of each audio frame, and the action direction factor of each audio frame to generate the increment of the action feature vector for each audio frame may include: in the m-th audio frame, multiplying the basis vector of the action feature vector, the action feature vector of the m-th audio frame, the action random factor of the m-th audio frame, and the action direction factor of the m-th audio frame to generate the initial increment of the action feature vector for the m-th audio frame; performing weighted average processing on the initial increment of the action feature vector for the m-th audio frame and the increment of the action feature vector for the (m - 1)-th audio frame to generate the increment of the action feature vector for the m-th audio frame. Multiplying the basis vector of the human voice pitch feature vector by the human voice pitch feature vector of each audio frame to generate the increment of the human voice pitch feature vector for each audio frame may include, in the m-th audio frame, multiplying the basis vector of the human voice pitch feature vector by the human voice pitch feature vector of the m-th audio frame to generate the initial increment of the human voice pitch feature vector for the m-th audio frame; performing weighted average processing on the initial increment of the human voice pitch feature vector for the m-th audio frame and the increment of the human voice pitch feature vector for the (m - 1)-th audio frame to generate the increment of the human voice pitch feature vector for the m-th audio frame.

[0184] Exemplarily, when performing weighted average processing, the incremental weight of the (m-1)th audio frame is 0.75, and the weight of the initial increment of the mth audio frame is 0.25.

[0185] S160: Process the incremental audio feature vectors of each audio frame using a third network model corresponding to the audio category to obtain a video frame corresponding to each audio frame. In an embodiment of the present invention, after synthesizing the incremental composite audio feature vectors latent z of each audio frame, an incremental composite audio feature matrix latent Z is obtained.

[0186] Based on custom materials, a third network model corresponding to the audio category is trained. Different music categories correspond to different third network models, and the third network model is used to process the incremental audio feature vectors to obtain video frames. Specifically, the training process of the third network model can be as follows: obtain video materials corresponding to different audio categories, perform frame extraction on the video materials, scale the frame-extracted video materials to a predetermined size and input them into an adversarial network model for training to generate third network models corresponding to different audio categories. Exemplarily, the third network model can select the adversarial network styleGAN2, use ffmpeg to perform frame extraction on the video materials, use openCV to scale the video materials, for example, scale them to a size of 1024*1024, and use styleGAN2-ada for training to generate multiple types of styleGAN2 models.

[0187] Combine Figure 6 The method for generating video frames in some embodiments of the present invention is introduced. For the calculated incremental composite audio feature matrix latent Z, the incremental composite audio feature vector latent z of the current frame is obtained from the incremental composite audio feature matrix latent Z. After inputting the incremental composite audio feature vector into a specific type of styleGAN2 network and mapping it through the mapping network Mapping Network of styleGAN2 to obtain an incremental mapped vector of the composite audio feature vector latent w, the incremental mapped vector of the composite audio feature vector is directly input into each layer of the synthesis network Synthesis Network of styleGAN2, and finally a video frame matching the music features of the current frame is generated.

[0188] In some embodiments of the present invention, video frames can be generated using the composite audio feature vector increment, specifically including: generating a composite audio feature increment matrix based on the composite audio feature vector increment at each moment; selecting the composite audio feature vector increment corresponding to each audio frame from the composite audio feature increment matrix and inputting it into a third network model corresponding to the audio category to obtain the video frame corresponding to each audio frame. In some embodiments of the present invention, corresponding multiple third network models can be provided for a predetermined variety of audio categories, such as the aforementioned folk songs, children's songs, pop music, etc., so that the corresponding third network model can be used according to the audio category determined in step S110.

[0189] In a specific implementation, the third network model includes a Mapping Network part and a Synthesis Network part. Based on this, selecting the composite audio feature vector increment corresponding to each audio frame from the audio feature increment matrix and inputting it into the third network model corresponding to the audio category to obtain the video frame corresponding to each audio frame includes: inputting the composite audio feature vector increment of each audio frame into the Mapping Network part to map to a composite audio feature increment mapped vector; inputting the composite audio feature increment mapped vector into each layer of the Synthesis Network part, and finally generating a video frame corresponding to the audio frame. Among them, the composite audio feature increment mapped vector can be denoted as latent w. In this embodiment, the audio feature vector increment described in step S160 can include the composite audio feature vector increment as described in step S153 and be used as the input of the third network model.

[0190] In some other embodiments of the present invention, video frames can be generated using the pulse feature vector increment, action feature vector increment, and human voice pitch feature vector. Specifically, input the pulse feature vector increment, action feature vector increment, and human voice pitch feature vector increment into the Mapping Network part respectively to map to multiple audio feature increment mapped vectors; input the audio feature increment mapped vectors corresponding to the action feature vector increment and human voice pitch feature vector increment into the front network layer of the Synthesis Network part; input the audio feature increment mapped vector corresponding to the pulse feature vector increment into the rear network layer of the Synthesis Network part, and finally generate a video frame corresponding to the audio frame. Among them, the audio feature increment mapped vector can be denoted as latent w1. In this embodiment, the audio feature vector described in step S160 can include the aforementioned pulse feature vector increment, action feature vector increment, and human voice pitch feature vector increment and be directly used as the input of the third network model.

[0191] In some embodiments of the present invention, after obtaining a video frame, the video frame can be further locally optimized. For example, corresponding synchronization special effects can be added to the video frame according to the intensity of the pulse feature vector corresponding to each audio frame. In addition, the video frame can also be optimized for super resolution.

[0192] Another method is to map the extracted music feature vector increments through the Mapping Network of styleGAN2 respectively to obtain multiple audio feature vector increment mapping vectors. The audio feature vector increment mapping vectors corresponding to the action feature vector increment and the human voice pitch feature vector increment are input into the network layer at the front of the Synthesis Network of styleGAN2 to affect the coarse structure of the generated image. The audio feature vector increment mapping vector corresponding to the pulse feature vector increment is input into the network layer at the rear of the Synthesis Network of styleGAN2 to affect the fine structure of the generated image. Similarly, based on randomly generated random audio feature vector increment mapping vectors, weighted averages are generated by respectively combining the extracted music feature vectors with the random audio feature vector increment mapping vectors and used as inputs to different layers of the corresponding synthesis network to generate video frames.

[0193] To enhance the enthusiasm of local video frames, video special effects can be produced based on Pulse features. For example, image contrast, flash, wave, and swirl special effect functions are defined based on image processing libraries such as PIL, skimage, and openCV. Corresponding synchronization special effects are added to the video frame according to the pulse feature intensity corresponding to each video frame. According to different requirements, the function of the image special effect can be defined by oneself and applied to the video frame.

[0194] To improve the video resolution, the video resolution is optimized based on a super resolution algorithm. For example, the LAPAR image super resolution model is used to optimize the video frame.

[0195] S170: Synthesize the video frames corresponding to each audio frame to generate a target dynamic video. For example, ffmpeg is used to splice the individual video frames to generate the target dynamic video.

[0196] The method for generating a music video in an embodiment of the present invention uses a generative adversarial network to generate a music video, and the visual effect of the generated music video matches the music rhythm. First, a generative adversarial network model for different music categories is constructed using a dataset. After classifying the input audio based on the CLMR model for music contrast learning, the generative adversarial network model corresponding to the music category is selected. The generative adversarial network model can extract the features of the input audio and map them to visual effects, and then output video frames that match the audio features. The visual effect of the video frames matches the music type of the input audio, conforming to the listener's auditory perception of the music and being more able to resonate with the emotions conveyed by the music.

[0197] In some other embodiments of the present invention, refer to Figure 7 , and a music video generation device 100 is provided, including the following modules:

[0198] An audio classification module 110, configured to classify the target audio using a first network model to obtain the audio category corresponding to the target audio;

[0199] An audio track separation module 120, configured to perform audio track separation processing on the target audio using a second network model to obtain multiple separated audio tracks;

[0200] A waveform generation module 130, configured to generate harmonics and shock waves of each of the separated audio tracks;

[0201] An audio feature vector generation module 140, configured to generate an audio feature vector for each audio frame of the target audio based on the harmonics and shock waves of each of the separated audio tracks;

[0202] An audio feature vector increment generation module 150, configured to generate an audio feature vector increment for each audio frame based on the audio feature vector of each audio frame;

[0203] A video frame generation module 160, configured to process the audio feature vector increment of each audio frame using a third network model corresponding to the audio category to obtain a video frame corresponding to each audio frame;

[0204] A video generation module 170, configured to perform synthesis processing on the video frames corresponding to each audio frame to generate a target dynamic video.

[0205] In some embodiments, the music video generation device can incorporate the features of the method for generating a music video in any embodiment, and vice versa, which will not be elaborated here.

[0206] In an embodiment of the present invention, an electronic device is provided, including: a processor and a memory storing a computer program, where the processor is configured to execute the method for generating a music video in any embodiment of the present invention when running the computer program.

[0207] In an embodiment of the present invention, an electronic device is provided, including: a processor and a memory storing a computer program, the processor being configured to execute the method for generating a music video according to any embodiment of the present invention when running the computer program.

[0208] Figure 8 FIG. shows a schematic diagram of an electronic device 800 that can implement the method of the embodiments of the present invention or realize the embodiments of the present invention. In some embodiments, there may be more or fewer electronic devices than shown in the figure. In some embodiments, it can be implemented using a single or multiple electronic devices. In some embodiments, it can be implemented using cloud or distributed electronic devices.

[0209] As Figure 8 shown, the electronic device 800 includes a central processing unit (CPU) 801, which can perform various appropriate operations and processes according to programs and / or data stored in the read-only memory (ROM) 802 and / or programs and / or data loaded from the storage section 808 into the random access memory (RAM) 803. The CPU 801 can be a multi-core processor or can include multiple processors. In some embodiments, the CPU 801 can include a general main processor and one or more special coprocessors, such as a graphics processing unit (GPU), a neural network processing unit (NPU), a digital signal processor (DSP), and so on. In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The CPU 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0210] The above-mentioned processor and memory are jointly used to execute the program stored in the memory, and when the program is executed by a computer, it can implement the steps or functions of the method for generating a music video described in the above embodiments.

[0211] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed, so that the computer program read from it can be installed into the storage section 808 as needed. Figure 8Only some components are schematically shown, which does not mean that the computer system 800 only includes Figure 8 the components shown.

[0212] The systems, devices, modules or units illustrated in the above embodiments can be implemented by a computer or its associated components. The computer can be, for example, a mobile terminal, a smart phone, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a personal digital assistant, a media player, a navigation device, a game console, a tablet computer, a wearable device, a smart TV, an Internet of Things system, a smart home, an industrial computer, a server, or a combination thereof.

[0213] In a preferred embodiment, the training system and method can be implemented or realized at least partially or entirely on a machine learning platform in the cloud or partially or entirely on a self-built machine learning system, such as a GPU array.

[0214] In a preferred embodiment, the evaluation device and method can be implemented or realized in a server, such as in the cloud or a distributed server. In a preferred embodiment, data or content can also be pushed or sent to the interruption based on the evaluation result with the help of the server.

[0215] Although not shown, in an embodiment of the present invention, a storage medium is provided, and the storage medium stores a computer program, and the computer program is configured to execute the method for generating a music video according to any embodiment of the present invention when being run.

[0216] The storage medium in the embodiment of the present invention includes permanent and non-permanent, removable and non-removable articles that can implement information storage by any method or technology. Examples of the storage medium include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technologies, compact disc read only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0217] The methods, programs, systems, devices, etc. in the embodiments of the present invention can be executed or realized in a single or multiple networked computers, and can also be practiced in a distributed computing environment. In the embodiments of this specification, in these distributed computing environments, tasks can be executed by remote processing devices connected through a communication network.

[0218] Those skilled in the art should understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, those skilled in the art can conceive that the implementation of the functional modules / units or controllers and related method steps illustrated in the above embodiments can be achieved in a software, hardware, or a combination of software and hardware manner.

[0219] Unless explicitly stated, the actions or steps of the methods and programs described according to the embodiments of the present invention do not necessarily have to be executed in a specific order and can still achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0220] In this document, multiple embodiments of the present invention have been described. However, for the sake of brevity, the descriptions of each embodiment are not exhaustive, and the same or similar features or parts between the various embodiments may be omitted. In this document, "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean applicable to at least one embodiment or example according to the present invention, rather than all embodiments. The above terms do not necessarily mean referring to the same embodiment or example. Without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0221] The exemplary systems and methods of the present invention have been specifically shown and described with reference to the above embodiments, which are only examples of the best mode for implementing the systems and methods. Those skilled in the art can understand that various changes can be made to the embodiments of the systems and methods described here when implementing the systems and / or methods without departing from the spirit and scope of the present invention defined in the appended claims.

Claims

1. A method for generating a music video, characterized in that, it includes the following steps: Classify the target audio using a first network model to obtain the audio category corresponding to the target audio; Perform track separation processing on the target audio using a second network model to obtain multiple separated tracks; Generate the harmonics and shock waves of each of the separated tracks; Generate the audio feature vector of each audio frame of the target audio based on the harmonics and shock waves of each of the separated tracks; Generate the audio feature vector increment of each audio frame based on the audio feature vector of each audio frame; Process the audio feature vector increment of each audio frame using a third network model corresponding to the audio category to obtain the video frame corresponding to each audio frame; Perform synthesis processing on the video frames corresponding to each audio frame to generate a target dynamic video.

2. The method according to claim 1, characterized in that, the first network model includes an encoding neural network and a projection neural network connected to the output layer of the encoding neural network, and the first network model is generated through the following steps: Obtain N training audio segments, and respectively select two partially overlapping or non-overlapping samples xi and xj from each of the training audio segments; Select the samples xi and xj of any one training audio segment for data augmentation processing to obtain the augmented samples xI and xJ, use the augmented samples xI and xJ as positive samples, and use the samples xi and xj of the remaining N-1 training audio segments as negative samples; Self-supervise the training of the positive samples and the negative samples using a contrast loss function to obtain the encoding neural network and the projection neural network.

3. The method according to claim 1, characterized in that, the second network model is a waveform-to-waveform model with a semantic segmentation network and a bidirectional long short-term memory network.

4. The method according to claim 1, characterized in that, the generation of the harmonics and shock waves of each of the separated tracks includes: Convert the time series of each separated track into a short-time Fourier transform matrix; Process the short-time Fourier transform matrix corresponding to each separated track using a median filter to obtain the initial harmonics and initial shock waves corresponding to each separated track; Perform inverse short-time Fourier transform on the initial harmonics and initial shock waves corresponding to each separated track, and adjust the time series length of the initial harmonics and initial shock waves after each inverse short-time Fourier transform to match the time series length of each separated track to generate the harmonics and shock waves of each separated track.

5. The method according to claim 1, characterized in that, the generation of the audio feature vector of each audio frame of the target audio based on the harmonics and shock waves of each of the separated tracks includes; If the separated track includes an accompaniment track, generate a pulse feature vector using the shock wave of the accompaniment track, and generate an action feature vector using the harmonics of the accompaniment track; If the separated track includes a vocal track, generate a vocal pitch feature vector using the harmonics of the vocal track; Use the pulse feature vector, the motion feature vector, and the human voice pitch feature vector as the audio feature vector for each audio frame.

6. The method according to claim 5, wherein, the generation of the pulse feature vector by using the shock wave of the accompaniment track includes: converting the shock wave of the accompaniment track into a spectrogram; multiplying the spectrogram by a plurality of Mel filters to obtain a Mel spectrogram feature matrix; performing normalization processing on the Mel spectrogram feature matrix based on the maximum Mel frequency in the Mel spectrogram feature matrix; reducing the dimension of the normalized Mel spectrogram feature matrix to a vector for each audio frame as the pulse feature vector.

7. The method according to claim 5, wherein, the generation of the motion feature vector by using the harmonics of the accompaniment track includes: converting the harmonics of the accompaniment track into a spectrogram; multiplying the spectrogram by a plurality of Mel filters to obtain a harmonic Mel spectrogram feature matrix; performing cepstrum analysis on the harmonic Mel spectrogram feature matrix to obtain a Mel frequency cepstral coefficient feature matrix, and obtaining the mean value of the Mel frequency cepstral coefficient features for each audio frame; using the mean value of the Mel frequency cepstral coefficient features for each audio frame to perform normalization processing on the Mel frequency cepstral coefficient features; reducing the dimension of the normalized Mel frequency cepstral coefficient feature matrix to a vector for each audio frame as the motion feature vector.

8. The method according to claim 5, wherein, the generation of the human voice pitch feature vector by using the harmonics of the human voice track includes: performing a CQT transform on the harmonics of the human voice track and taking the absolute value to obtain the absolute value of the CQT transform at each time point; mapping the absolute value of the CQT transform to a chromagram to generate an initial chromagram CQT transform feature matrix; performing normalization processing on the initial chromagram CQT transform feature matrix to generate a chromagram CQT transform feature matrix; calculating a weighted average chromagram value according to the chromagram values corresponding to each audio frame, where each audio frame corresponds to the chromagram values of T scales; using the weighted average chromagram value corresponding to each audio frame to perform normalization processing on the chromagram CQT transform feature matrix; reducing the dimension of the normalized chromagram CQT transform feature matrix to a vector for each audio frame as the human voice pitch feature vector.

9. The method for generating a music video according to claim 5, wherein, using the pulse feature vector, the motion feature vector, and the human voice pitch feature vector as the audio feature vector of the audio frame includes: applying a filter to perform smoothing processing on the pulse feature vector, the motion feature vector, and the human voice pitch feature vector along the time axis, and using the smoothed pulse feature vector, motion feature vector, and human voice pitch feature vector as the audio feature vector of the audio frame.

10. The method according to claim 5, wherein, generating the audio feature vector increment of each audio frame based on the audio feature vector of each audio frame includes: generating a base noise vector for each audio frame; Sum the action feature vector increments of each audio frame between the first audio frame and the current audio frame of the target audio to obtain the cumulative action feature vector increment of the current audio frame; Sum the base noise vector of the current audio frame, the pulse feature vector increment of the current audio frame, the human voice pitch feature vector increment of the current audio frame, and the cumulative action feature vector increment of the current audio frame to generate the composite audio feature vector increment of the current audio frame; Loop through the above steps to obtain the composite audio feature vector increment of each audio frame, where the composite audio feature vector increment serves as the audio feature vector increment.

11. The method according to claim 10, wherein, generating the base noise vector of each audio frame includes: Generating a normal distribution vector in the order of audio frames based on the standard normal distribution, and truncating the normal distribution vector in the order of audio frames according to the threshold range as the base noise vector.

12. The method according to claim 10, wherein, the pulse feature vector increment, action feature vector increment, and human voice pitch feature vector increment of the audio frame are generated by the following method: Construct the basis vectors of the pulse feature vector, the basis vectors of the action feature vector, and the basis vectors of the human voice pitch feature vector; Generate action random factors at predetermined time intervals; Multiply the basis vector of the pulse feature vector by the pulse feature vector of each audio frame to generate the pulse feature vector increment of each audio frame; Multiply the basis vector of the action feature vector, the action feature vector of each audio frame, the action random factor of each audio frame, and the action direction factor of each audio frame to generate the action feature vector increment of each audio frame; Multiply the basis vector of the human voice pitch feature vector by the human voice pitch feature vector of each audio frame to generate the human voice pitch feature vector increment of each audio frame.

13. The method according to claim 10, wherein, processing the audio feature vector increment of each audio frame using the third network model corresponding to the audio category to obtain the video frame corresponding to each audio frame includes: Generating a composite audio feature vector increment matrix based on the composite audio feature vector increment of each audio frame; Select the composite audio feature vector increment corresponding to each audio frame from the composite audio feature vector increment matrix and input it into the third network model corresponding to the audio category to obtain the video frame corresponding to each audio frame.

14. The method according to claim 13, wherein, the third network model includes a mapping network part and a comprehensive network part; selecting the composite audio feature vector increment corresponding to each audio frame from the audio feature vector increment matrix and inputting it into the third network model corresponding to the audio category to obtain the video frame corresponding to each audio frame includes: Inputting the composite audio feature vector increment of the audio frame into the mapping network part to map and obtain a composite audio feature vector increment mapping vector; Inputting the composite audio feature vector increment mapping vector into each layer of the comprehensive network part to generate the video frame corresponding to the audio frame.

15. The method according to claim 1, wherein, it further comprises: adding corresponding synchronization special effects to the video frame according to the intensity of the pulse feature vector corresponding to each audio frame; performing super-resolution optimization on the video frame.

16. The method according to claim 5, wherein, generating an audio feature vector increment for each audio frame based on the audio feature vector of each audio frame, comprising: constructing basis vectors of the pulse feature vector, basis vectors of the motion feature vector, and basis vectors of the human voice pitch feature vector; generating a motion random factor at a predetermined time interval; multiplying the basis vector of the pulse feature vector by the pulse feature vector of each audio frame to generate a pulse feature vector increment for each audio frame; multiplying the basis vector of the motion feature vector, the motion feature vector of each audio frame, the motion random factor of each audio frame, and the motion direction factor of each audio frame to generate a motion feature vector increment for each audio frame; multiplying the basis vector of the human voice pitch feature vector by the human voice pitch feature vector of each audio frame to generate a human voice pitch feature vector increment for each audio frame; wherein, the pulse feature vector increment, the motion feature vector increment, and the human voice pitch feature vector increment are used as the audio feature vector increment.

17. The method according to claim 12 or 16, wherein, the multiplying the basis vector of the pulse feature vector by the pulse feature vector of each audio frame to generate a pulse feature vector increment for each audio frame comprises: at the first audio frame, multiplying the basis vector of the pulse feature vector by the pulse feature vector of the first audio frame to generate a pulse feature vector increment of the first audio frame; at the m-th audio frame, where m is greater than or equal to 2, multiplying the basis vector of the pulse feature vector by the pulse feature vector of the m-th audio frame to generate an initial pulse feature vector increment of the m-th audio frame; the multiplying the basis vector of the motion feature vector, the motion feature vector of each audio frame, the motion random factor of each audio frame, and the motion direction factor of each audio frame to generate a motion feature vector increment for each audio frame comprises: at the first audio frame, multiplying the basis vector of the motion feature vector, the motion feature vector of the first audio frame, the motion random factor of the first audio frame, and the motion direction factor of the first audio frame to generate a motion feature vector increment of the first audio frame; at the m-th audio frame, performing weighted average processing on the initial pulse feature vector increment of the m-th audio frame and the pulse feature vector increment of the (m - 1)-th audio frame to generate a pulse feature vector increment of the m-th audio frame; multiplying the basis vector of the motion feature vector, the motion feature vector of the m-th audio frame, the motion random factor of the m-th audio frame, and the motion direction factor of the m-th audio frame to generate an initial motion feature vector increment of the m-th audio frame; performing weighted average processing on the initial motion feature vector increment of the m-th audio frame and the motion feature vector increment of the (m - 1)-th audio frame to generate a motion feature vector increment of the m-th audio frame; The step of multiplying the basis vectors of the human voice pitch feature vectors by the human voice pitch feature vectors of each audio frame to generate the increment of the human voice pitch feature vectors of each audio frame includes: In the first audio frame, multiplying the basis vectors of the human voice pitch feature vectors by the human voice pitch feature vectors of the first audio frame to generate the increment of the human voice pitch feature vectors of the first audio frame; in the m-th audio frame, multiplying the basis vectors of the human voice pitch feature vectors by the human voice pitch feature vectors of the m-th audio frame to generate the initial increment of the human voice pitch feature vectors of the m-th audio frame; performing weighted average processing on the initial increment of the human voice pitch feature vectors of the m-th audio frame and the increment of the human voice pitch feature vectors of the (m - 1)-th audio frame to generate the increment of the human voice pitch feature vectors of the m-th audio frame.

18. The method according to claim 12 or 16, wherein, it further includes: If the value obtained by adding or subtracting the action feature vector response coefficient to the absolute value of the audio feature vector increment of the current audio frame is greater than twice the preset truncation value, change the positive or negative of the action direction factor.

19. The method according to claim 1, wherein, the third network model corresponding to the audio category is generated through the following steps: Obtain video materials corresponding to different audio categories, perform frame extraction on the video materials, scale the frame-extracted video materials to a predetermined size and input them into the adversarial network model for training to generate the third network model corresponding to different audio categories.

20. The method according to claim 1, wherein, the third network model includes a mapping network part and a comprehensive network part; the step of using the third network model corresponding to the audio category to process the audio feature vector increment to obtain the video frame corresponding to each audio frame includes: Input the pulse feature vector increment, the action feature vector increment, and the human voice pitch feature vector increment into the mapping network part respectively to map and obtain multiple mapped vectors of the audio feature vector increment; Input the mapped vectors of the audio feature vector increment corresponding to the action feature vector increment and the human voice pitch feature vector increment among the multiple mapped vectors of the audio feature vector increment into the front network layer of the comprehensive network part, and input the mapped vector of the audio feature vector increment corresponding to the pulse feature vector increment among the multiple mapped vectors of the audio feature vector increment into the rear network layer of the comprehensive network part to generate the video frame corresponding to each audio frame.

21. A computer-readable storage medium, on which a computer program is stored, wherein, the program, when executed by a processor, implements the method according to any one of claims 1 - 20.

22. An electronic device, wherein, it includes: a processor and a memory storing a computer program, and the processor is configured to execute the method according to any one of claims 1 - 20 when running the computer program.

Citation Information

Patent Citations

  • Audio separation method and device, electronic equipment and storage medium

    CN110503976A

  • Audio and video alignment method and device, equipment and storage medium

    CN113473201A