A method for automatically generating video background music based on the rhythm relationship between audio and video

By establishing the rhythm relationship between video and music, and using the music generation model to generate background music based on the rhythm characteristics of the video, the problem of not being able to automatically generate soundtracks for videos in the prior art is solved, and automated generation, rhythm fit and personalized needs are achieved.

CN113889059BActive Publication Date: 2025-06-17BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111121236.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-24
Publication Date
2025-06-17
Estimated Expiration
2041-09-24

AI Technical Summary

Technical Problem

The prior art cannot automatically generate soundtracks for videos, the search process is cumbersome and cannot meet personalized needs, and there is a risk of copyright infringement.

Method used

By establishing the rhythm relationship between video and music, extracting the rhythm characteristics of video and music, and using the music generation model to generate background music based on the rhythm characteristics of video.

Benefits of technology

It realizes the automatic generation of video background music, perfectly matches the rhythm and video, reduces the difficulty of video production, avoids copyright issues, and meets personalized needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113889059B_ABST
    Figure CN113889059B_ABST
Patent Text Reader

Abstract

The present invention provides a method for automatically generating video background music based on the rhythm relationship between audio and video. The visual rhythm features of the input video are extracted, including visual motion speed features, visual motion saliency features, and the corresponding number of video frames. According to the preset rhythm relationship between the video and the music, the rhythm positions of the visual rhythm features of the input video are automatically replaced with the music rhythm features at the corresponding rhythm positions, including note group density and note group intensity. The converted music rhythm features, together with the music style and instrument type input by the user, are input into a deep learning model to generate video background music. The present invention can quickly generate background music for videos automatically. The generated music can match the video in terms of rhythm, which can facilitate the video production of video editors or ordinary people and obtain personalized video background music.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of music generation and cross-modal, and particularly relates to a method for automatically generating video background music based on the rhythm relationship between audio and video. Background Art

[0002] Video background music generation refers to automatically generating background music according to a video. The existing related technologies cannot automatically generate background music for a video. Instead, they can only retrieve in a music library, which has a large retrieval volume and a complicated process. The retrieval results cannot perfectly fit the video, and cannot meet the personalized needs of users, and there is also a possibility of copyright infringement.

[0003] Therefore, how to provide a method for automatically generating video background music that can automatically generate personalized background music for a video based on the rhythm relationship between audio and video is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0004] In view of this, the present invention provides a method for automatically generating video background music based on the rhythm relationship between audio and video. The present invention establishes three association relationships between the rhythm of the video and the music, and proposes a new music representation form to generate background music according to the rhythm of the video.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] A method for automatically generating video background music based on the rhythm relationship between audio and video, comprising the following steps:

[0007] Obtain the statistical information of the video rhythm features and music rhythm features in the video and music database, and establish the rhythm relationship between the video rhythm features and the music rhythm features. The video rhythm features include the visual motion speed feature, the visual motion saliency feature, and the corresponding number of video frames. The music rhythm features include the note group density, the note group intensity, and the corresponding number of music measures and beats;

[0008] According to the preset rhythm relationship between the video and the music, automatically replace the rhythm position of the visual rhythm features of the input video with the music rhythm features at the corresponding rhythm position, and input them into the music generation model together with the music style and instrument type specified by the user to generate video background music.

[0009] Preferably, the visual motion speed feature includes the average optical flow magnitude of a plurality of video frames; the visual motion saliency feature includes the change amount of the optical flow in different directions between two adjacent video frames.

[0010] Preferably, the note group intensity is the number of notes contained in the note group, and the note group density of the music measure is the number of note groups contained in the music measure.

[0011] Preferably, the rhythm relationship between the video and the music is as follows:

[0012] The number of beats of the music corresponding to the t-th frame of the video, that is, and / or the number of video frames corresponding to the i-th beat of the music, that is, where Tempo is the number of beats per minute, and FPS is the number of video frames contained in the video per second.

[0013] The statistical information of the video rhythm features and the music rhythm features in the video and music database also includes the quantiles of the video rhythm features and the music rhythm features.

[0014] Preferably, automatically replacing the rhythm position where the visual rhythm feature of the input video is located with the music rhythm feature of the corresponding rhythm position specifically includes:

[0015] Establishing the correlation relationship between the visual motion speed feature and the note group density;

[0016] Establishing the correlation relationship between the visual motion saliency feature and the note group intensity;

[0017] Establishing the conversion relationship between the input video frame number and the quantiles of the music bars and beats.

[0018] Preferably, it further includes the following steps:

[0019] Converting the note attributes and rhythm attributes of the music bar into embedding vectors, p k = Embedding k (w k ), k = 1,..., K, where w k is the k-th attribute, Embedding is the embedding vector conversion function, p k is the k-th embedding vector after conversion, the note attributes include duration, pitch, and instrument type; the rhythm attributes include bar start / beat start time, note group density, and note group intensity;

[0020] Combining the embedding vectors and then through linear transformation, the final word vector is obtained, that is, where W in is the linear transformation matrix, is the dimension splicing operation.

[0021] Preferably, the steps for the music generation model to generate the video background music include:

[0022] Training the music generation model: Encoding the notes and the extracted music rhythm features in the music into word vectors, using the first N - 1 word vectors as the input of the deep learning model, predicting and learning the N-th word vector, and repeating the training until the accuracy requirement is met;

[0023] Video background music generation: The video rhythm features extracted from the input video are converted into music rhythm features according to the rhythm relationship between the video rhythm features and the music rhythm features, and then the trained music generation model is used to generate the background music.

[0024] As can be seen from the above technical solutions, compared with the prior art, the beneficial effects of the present invention include:

[0025] The present invention can reduce the difficulty of video production, can automatically generate background music for videos within a few seconds to a few minutes, and the generated music can match the video in terms of rhythm, including the following three aspects:

[0026] 1) The intensity of the music matches the speed of visual movement;

[0027] 2) The musical accents match the salience of visual movement;

[0028] 3) The start and end of the music match the start and end of the video.

[0029] The present invention can facilitate the video production of video editors or ordinary people, avoid the music copyright problem, and can be widely applied to industries such as film and television editing, online live broadcast, and social media. Description of the Drawings

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts;

[0031] Figure 1 It is a flowchart of a method for automatically generating video background music based on the rhythm relationship between audio and video provided by an embodiment of the present invention;

[0032] Figure 2 It is a schematic diagram of a music measure in a method for automatically generating video background music based on the rhythm relationship between audio and video provided by an embodiment of the present invention. Detailed Embodiments

[0033] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0034] See Figure 1 , a method for automatically generating video background music based on the rhythm relationship between audio and video disclosed in this embodiment. Extract rhythm features from the input video, including visual motion speed, visual motion saliency, and corresponding key frame numbers, and input them together with the music attributes specified by the user into the controllable music generation module, then a background music for the input video can be automatically generated.

[0035] The controllable music generation module is a non-volatile computer-readable storage medium storing computer program instructions of the method of this embodiment.

[0036] The music attributes input by the user according to requirements include music style attributes and instrument type attributes.

[0037] The specific execution process of this embodiment is as follows:

[0038] S1. Obtain the statistical information of the video rhythm features and music rhythm features in the video and music database, establish the rhythm relationship between the video rhythm features and the music rhythm features. The video rhythm features include visual motion speed features, visual motion saliency features, and corresponding video frame numbers. The music rhythm features include note group density, note group intensity, and corresponding music measures and beats;

[0039] S2. According to the preset rhythm relationship between the video and the music, automatically replace the rhythm position of the visual rhythm features of the input video with the music rhythm features of the corresponding rhythm position, and input them together with the music style and instrument type specified by the user into the music generation model to generate the video background music.

[0040] It should be noted that rhythm refers to the distribution of events in time. Therefore, first establish the conversion relationship between music and video in time units.

[0041] A video is composed of frames, and the number of frames contained in each second of the video is called FPS (frame per second). Music is usually divided into measures, and the measures are further equally divided into beats (such as four beats in a measure). The number of beats per minute is called Tempo, which controls the rhythm speed of the music.

[0042] In one embodiment, the rhythm relationship between the video and the music is:

[0043] The number of beats of the music corresponding to the t-th frame of the video, that is and / or the number of video frames corresponding to the i-th beat of the music, that is Among them, Tempo is the number of beats per minute, and FPS is the number of video frames contained in each second of the video. According to the above formula, the number of musical measures and beats corresponding to the video frames is obtained. During the music generation process, the rhythm characteristics of a certain video frame are converted into the rhythm characteristics of the corresponding measure and beat to control the music generation process.

[0044] There is a corresponding relationship between music and video in terms of rhythm. When an object moves rapidly, we expect dense notes; when there are significant changes in the picture, such as during a transition, we expect accents to appear. The unity of music and visual rhythm can strengthen the sensory impact and bring a sense of pleasure.

[0045] Based on the above situation, the embodiments of the present invention propose the correlation relationships between the visual motion speed and the density of note groups, and between the visual motion saliency and the intensity of note groups.

[0046] In one embodiment, visual motion can be described by optical flow, which measures the pixel motion between two adjacent frames (frame f and its subsequent frame). The visual motion speed is the average optical flow magnitude of a certain video segment: The visual motion saliency is the comprehensive change of the optical flow between two adjacent frames in different directions.

[0047] In one embodiment, music is composed of notes (denoted by n). Briefly speaking, each note has five attributes: start time, duration, pitch, instrument type, and intensity. As Figure 2 shown, a note group is a set of notes that start sounding simultaneously, that is, N = {n1, n2,...}. The intensity of a note group is the number of notes contained in the note group, that is, S N = |N|. The music is divided into measures. One measure may contain multiple note groups, that is, B = {N1, N2,...}. The note group density of one measure is the number of note groups contained in the measure, that is, D B = |B|.

[0048] In one embodiment, automatically replacing the rhythm position where the visual rhythm characteristics of the input video are located with the rhythm characteristics of the corresponding rhythm position in music specifically includes:

[0049] Establishing the correlation relationship between the visual motion speed characteristics and the note group density;

[0050] Establishing the correlation relationship between the visual motion saliency characteristics and the note group intensity;

[0051] Establishing the quantile conversion relationship between the input video frames and the musical measures and beats.

[0052] In one embodiment, representing music in a way similar to natural language, including two types of word vectors: note attributes and rhythm attributes. It also includes the following steps:

[0053] Convert the note attributes and rhythm attributes of a musical measure into embedding vectors, p k = Embedding k (w k ), k = 1, ..., K, where w k is the k-th attribute, Embedding is the embedding vector conversion function, and p k is the k-th converted embedding vector. The note attributes include duration, pitch, and instrument type; the rhythm attributes include measure start / beat start time, note group density, and note group intensity;

[0054] Combine the embedding vectors through concatenation, and then after a linear transformation, the final word vector is obtained, that is where W in is the linear transformation matrix, is the dimension concatenation operation.

[0055] The concatenation combination form of the embedding vectors can be concatenation. For example, a note has three attributes: duration, pitch, and instrument type. After converting these three attributes into embedding vectors, they are "concatenated" to obtain the word vector of the note.

[0056] In this embodiment, the word vectors are arranged in order and the beat position encoding is added to obtain the final word vector, which is the input of the deep learning model. Among them, the beat position encoding means dividing the whole piece of music into 100 parts, that is, a piece of music has 100 musical measures, divided into 100 parts according to time, and each part is a musical measure. A musical measure is composed of multiple word vectors. Each part of the word vectors within the same musical measure uses the same position encoding, and different musical measures use different beat position encodings. The model learns the association between position and music during training, and realizes the synchronization of the start / end of music and the start / end of video during generation, and better grasps the structure of music.

[0057] Those skilled in the art can understand that the word vectors record the attributes of each note and are converted into MIDI files or audio files using the Muspy software.

[0058] In the first step of this embodiment, the rhythm features of the music library and the video library are extracted, and then the rhythm correspondence between music and video is established according to the statistical information (quantile). The quantile is the sorting position of the current video rhythm feature and the music rhythm feature according to the size of their respective feature values. For example, the visual motion speed 100 should correspond to the note group density 10. It should be noted that the top 10% of the motion speed sorting of each frame of the input video is 100, and the top 10% of the note group density sorting of the music is 10. Then the video frame with the current visual motion speed of 100 corresponds to the note group with the note group density of 10.

[0059] In the second step, train the music generation model. Encode the notes in the music and the extracted music rhythm features into word vectors, and input the first N - 1 word vectors into the deep learning model to enable the model to learn to accurately predict the Nth word vector.

[0060] In the third step, generate the background music for the video. Convert the video rhythm features extracted from the input video into music rhythm features according to the audio - video rhythm relationship obtained in the first step. Then use the music generation model trained in the second step to generate the background music.

[0061] The above has introduced in detail the method for automatically generating the background music for a video based on the audio - video rhythm relationship provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

[0062] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. An automatic video background music generation method based on the rhythm relationship between audio and video, characterized in that, It includes the following steps: Obtain the statistical information of the video rhythm features and the music rhythm features in the video and music database, and establish the rhythm relationship between the video rhythm features and the music rhythm features. The video rhythm features include the visual motion speed feature, the visual motion saliency feature, and the corresponding number of video frames. The music rhythm features include the note group density, the note group intensity, and the corresponding number of music bars and beats. The visual motion speed feature includes the average optical flow magnitude of a number of video frames. The visual motion saliency feature includes the change amount of the optical flow in different directions between two adjacent video frames. According to the preset rhythm relationship between the video and the music, automatically replace the rhythm position of the visual rhythm feature of the input video with the music rhythm feature of the corresponding rhythm position, and input it into the music generation model together with the music style and instrument type specified by the user to generate the video background music.

2. The automatic video background music generation method based on the rhythm relationship between audio and video according to claim 1, characterized in that, The note group intensity is the number of notes contained in the note group, and the note group density of the music bar is the number of note groups contained in the music bar.

3. The automatic video background music generation method based on the rhythm relationship between audio and video according to claim 1, characterized in that, The rhythm relationship between the video and the music is as follows: The number of beats of the music corresponding to the t-th frame of the video, i.e., and / or the number of video frames corresponding to the i-th beat of the music, i.e., where Tempo is the number of beats per minute and FPS is the number of video frames contained in the video per second.

4. The automatic video background music generation method based on the rhythm relationship between audio and video according to claim 1, characterized in that, The statistical information of the video rhythm features and the music rhythm features in the video and music database includes the quantiles of the video rhythm features and the music rhythm features. Automatically replacing the rhythm position of the visual rhythm feature of the input video with the music rhythm feature of the corresponding rhythm position specifically includes: Establish the association relationship between the visual motion speed feature and the note group density; Establish the association relationship between the visual motion saliency feature and the note group intensity; Establish the conversion relationship between the number of input video frames and the music bars and beats.

5. The automatic video background music generation method based on the rhythm relationship between audio and video according to claim 1, characterized in that, It also includes the following steps: Convert the note attributes and rhythm attributes of a musical measure into an embedded vector, p k = Embedding k (w k ), k = 1, ..., K, where w k is the k-th attribute, Embedding is the embedded vector conversion function, and p k is the k-th converted embedded vector. The note attributes include duration, pitch, and instrument type; the rhythm attributes include measure start / beat start time, note group density, and note group intensity; After combining the embedded vectors through concatenation and then performing a linear transformation, the final word vector is obtained, i.e., where W in is the linear transformation matrix, is the dimension concatenation operation.

6. The automatic video background music generation method based on the rhythm relationship between audio and video according to claim 1, characterized in that, The steps for the music generation model to generate the video background music include: Train the music generation model: Encode the notes in the music and the extracted music rhythm features into word vectors, use the first N - 1 word vectors as the input of the deep learning model, predict and learn the Nth word vector, and perform repeated training until the accuracy requirement is met; Video background music generation: Convert the video rhythm features extracted from the input video into music rhythm features according to the rhythm relationship between the video rhythm features and the music rhythm features, and then use the trained music generation model to generate the background music.