Video background music generation system

By designing a video background music generation system, using video description generation and sentiment analysis modules to generate background music text descriptions, and generating music through Transformer decoder, the problems of low matching degree of music and video content and insufficient music quality in the prior art are solved, and high-quality and harmonious video background music generation are achieved.

CN120048231APending Publication Date: 2025-05-27HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510119054.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing video background music generation method fails to effectively correlate video emotions, resulting in low matching between music and video content, and poor music quality, low sound quality, repetition of music clips, and limited length.

Method used

A video background music generation system is designed, including a video description generation module, a video sentiment analysis module, a text fusion module and a music generator. The system generates background music text descriptions by extracting and analyzing video content and emotions, and generates high-quality background music using Transformer decoder.

Benefits of technology

The high-quality matching of video content and emotions with music is achieved, ensuring that the generated music is harmonious with the video content and emotions, and the music quality is improved, avoiding the problems of repetition of music clips and limited length.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048231A_ABST
    Figure CN120048231A_ABST
Patent Text Reader

Abstract

The invention discloses a video background music generation system, and belongs to the technical field of cross-modal music generation. The invention aims to solve the problem that the matching degree of music and video content is low because the generation of existing video background music is not associated with video emotion expression. Comprising a video description generation module used for carrying out video content feature extraction on an input video to obtain video content text description; the video sentiment analysis module is used for extracting sentiment features of the input video to obtain video sentiment category text description; the text fusion module is used for fusing the video content text description, the video emotion category text description and the music type text description input by the user to obtain background music text description; and the music generator is used for generating target background music according to the background music text description. The method and the device are applied to the background music generation of the short video in the user production content mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a video background music generation system, and belongs to the technical field of cross-modal music generation. Background Art

[0002] Music plays a vital role in video creation, it can improve the overall quality of the video and enhance the audience's sense of immersion.

[0003] With the rapid development of social platforms, the need to find suitable music for videos has extended from professional to amateur use. However, the difficulty of this task and the possible copyright issues make automatic generation of background music for videos of great value to the majority of creators, but it is challenging because the correspondence between music and video is not a deterministic one-to-one mapping, but a more complex mapping related to the content and emotion, and the music that needs to be created is not only coherent and melodic, but also harmonious with the given video in terms of content and emotion.

[0004] Existing video-music generation methods can be divided into several categories. The first category is severely limited in video types, such as solving the problem of motion-to-music generation through motion features in videos, generating music for music performance videos, extracting the coordinates of key points of the human skeleton from videos, and using the Transformer's multi-head attention module to fuse the information of the two modalities of video and audio. The second category is not limited to video types, but the generated music has limitations, such as only exploring the rhythmic relationship between the video and background music, lacking content connection or only focusing on the emotional matching between music and video, and having problems in music quality such as poor richness, low sound quality, repetitive music clips, and limited length. Summary of the invention

[0005] Aiming at the problem that the generation of existing video background music is not associated with the emotional expression of the video, resulting in a low matching degree between the music and the video content, the present invention provides a video background music generation system.

[0006] A video background music generation system of the present invention comprises:

[0007] A video description generation module is used to extract video content features from the input video and obtain a text description of the video content;

[0008] The video sentiment analysis module is used to extract sentiment features from the input video and obtain a text description of the video sentiment category;

[0009] A text fusion module, used to fuse the video content text description, the video emotion category text description and the music type text description input by the user to obtain a background music text description;

[0010] A music generator for generating a target background music according to a background music text description.

[0011] According to the video background music generation system of the present invention, the video description generation module includes:

[0012] An action recognition network TSP for extracting video features of an input video to obtain a video frame-level feature map;

[0013] A multi-scale feature encoder for extracting different-level multi-scale feature sequences from the video frame-level feature map;

[0014] A deformable Transformer encoder for encoding different-level multi-scale feature sequences and the input position encoding to obtain multi-scale frame-level encoded features;

[0015] An ECA_ASPP module for enhancing the multi-scale frame-level encoded features to obtain enhanced multi-scale frame-level encoded features;

[0016] A deformable Transformer decoder for querying and decoding the enhanced multi-scale frame-level encoded features according to an event query to obtain event-level query decoded features and reference points;

[0017] A description generation head for obtaining text descriptions of multiple event queries by using a deformable soft attention mechanism according to the event-level query decoded features and reference points;

[0018] A localization head for obtaining a bounding box prediction and binary classification of each event query according to the event-level query decoded features and reference points to obtain a tuple of the currently detected event query; wherein the bounding box prediction obtains a two-dimensional relative offset between the current event and the reference event, and the binary classification obtains the foreground confidence of the current event;

[0019] An event counter for globally aggregating the event-level query decoded features to obtain a global feature vector representing the entire content of the input video, and predicting the number of events in the video according to the global feature vector;

[0020] A video content prediction module for sorting according to the text descriptions of multiple event queries, the tuple of the currently detected event query, and the number of events in the video based on the foreground confidence to obtain a video content text description.

[0021] According to the video background music generation system of the present invention, the video emotion analysis module includes:

[0022] A face detection module for performing face detection on all frames of the input video to obtain a face image containing only faces;

[0023] The facial expression classification module is used to extract features from facial images and obtain text descriptions of video emotion categories through a classification head.

[0024] According to the video background music generation system of the present invention, the music generator uses an Encodec encoder to obtain audio tokens based on the background music text description, and uses a T5 text encoder to obtain text tokens based on the background music text description; the audio tokens and text tokens are used by an autoregressive Transformer decoder to generate target audio tokens, and then an Encodec decoder is used to generate the target background music.

[0025] According to the video background music generation system of the present invention, the multi-scale feature encoder is implemented based on the SCDown_PSA module;

[0026] The video frame-level feature map uses four temporal convolutional layers with a stride of 2 and a kernel size of 3 to obtain four different levels of feature sequences; among them, the first-level feature sequence is added element-wise to the second-level feature sequence after passing through the scDown module to obtain a new second-level feature sequence; the new second-level feature sequence is added element-wise to the third-level feature sequence after passing through the scDown module to obtain a new third-level feature sequence; the new third-level feature sequence is added element-wise to the fourth-level feature sequence after passing through the scDown module, and the fused feature obtained is then passed through the PSA module to obtain a new fourth-level feature sequence; the first-level feature sequence, the new second-level feature sequence, the new third-level feature sequence, and the new fourth-level feature sequence form different-level multi-scale feature sequences.

[0027] According to the video background music generation system of the present invention, the process by which the PSA module obtains the new fourth-level feature sequence includes:

[0028] The PSA module uses the SPC module to divide the channels, performs multi-scale feature extraction on the spatial information of the input fused feature on each channel, then uses a squeeze-and-excitation module to extract the channel attention of different-scale features to obtain channel attention vectors at different scales; then uses Softmax to calibrate the features of the channel attention vectors at different scales to obtain the attention weights after multi-scale channel interaction; then performs an element-wise dot product operation on the multi-scale features extracted on each channel based on the attention weights after multi-scale channel interaction to obtain the new fourth-level feature sequence after attention weighting of the multi-scale feature information;

[0029] The scDown module downsamples the input feature sequence in the spatial dimension through a 3x3 convolutional layer with a stride of 2 and reduces the number of channels of the input feature sequence through a 1x1 convolution.

[0030] In the video background music generation system according to the present invention, in the ECA_ASPP module, the ASPP module uses dilated convolutions with different dilation rates to obtain the global information of multi-scale frame-level encoded features. The global information of the multi-scale frame-level encoded features is passed through the adaptive average pooling layer, 1x1 convolutional layer, and Sigmoid activation function of the ECA module, and then multiplied element-wise with the global information of the multi-scale frame-level encoded features to obtain enhanced multi-scale frame-level encoded features;

[0031] The dilation rates of the dilated convolutions with different dilation rates are 1x1; 3x3, rate = 6; 3x3, rate = 12; 3x3, rate = 18 in sequence.

[0032] In the video background music generation system according to the present invention, the event counter includes a max pooling layer and a fully connected layer with softmax activation.

[0033] In the video background music generation system according to the present invention, the network structure of the face detection module includes three parts: backbone, neck, and head:

[0034] The 0th and 1st layers of the backbone part are 3x2 convolutional layers, the 2nd layer is a C2f layer, repeated three times, the 3rd layer is a 3x2 convolutional layer, the 4th layer is a C2f layer, repeated six times, the 5th layer is a 3x2 GSConv layer, the 6th layer is a C2f layer, repeated 3 times, the 7th layer is a 3x2 convolutional layer, the 8th layer is a C2f layer, repeated 3 times, the 9th layer is an SPPF layer, and the 10th layer is an MSCAM layer;

[0035] The 11th layer of the neck part is an upsampling layer, the 12th layer is to concatenate the outputs of the 6th and 11th layers, the 13th layer is a VoVGSCSP layer, the 14th layer is an upsampling layer, the 15th layer is to concatenate the outputs of the 4th and 14th layers, the 16th layer is a VoVGSCSP layer, and the 17th layer is a DCNv4 layer;

[0036] The 18th layer of the head part is a 3x2 convolutional layer, the 19th layer is to concatenate the outputs of the 13th and 18th layers, the 20th layer is a VoVGSCSP layer, the 21st layer is a DCNv4 layer, the 22nd layer is a 3x2 convolutional layer, the 23rd layer is to concatenate the outputs of the 10th and 22nd layers, the 24th layer is a VoVGSCSP layer, and the 25th layer is a DCNv4 layer;

[0037] The DCNv4 layer includes two branches, each branch includes a 3x1 convolutional layer and a corresponding 1x1 convolutional layer connected thereto; one branch is used to calculate the bounding box loss, and the other branch is used to calculate the classification loss;

[0038] The C2f layer performs a convolution on the input data, doubling the number of channels to twice the number of input channels, and then uses the Bottleneck module to gradually extract features. The Bottleneck module contains multiple convolutional layers and uses residual connections. After the output data of multiple Bottleneck modules is concatenated with the input data, the concatenated feature map is compressed through a convolution operation, and a feature map with the target number of channels is output.

[0039] The GSConv layer downsamples the input through ordinary convolution, then uses depthwise convolution (DWConv), concatenates the results of the two convolutions, and then performs a shuffle operation.

[0040] The SPPF module contains an initial convolutional layer for feature compression, then performs multiple serial max-pooling operations, and finally fuses features of different scales through a convolutional layer.

[0041] The MSCAM layer is composed of CAB, SAB, and MSCB.

[0042] The VoVGSCSP layer contains a convolutional layer, two GSConv layers. The output results of the two GSConv layers are concatenated with the output result of the convolutional layer, and then an output is obtained through a convolutional layer.

[0043] According to the video background music generation system of the present invention, the facial expression classification module sequentially has three feature extraction layers for hierarchical feature extraction. Each feature extraction layer sequentially includes a deformable convolutional layer (DCNv4), a two-dimensional convolutional layer, a batch normalization layer, and a ReLU activation function. A max-pooling layer is configured after the ReLU activation function.

[0044] The max-pooling layer of the third feature extraction layer is connected to the squeeze-and-excitation module, then sequentially connected to the MSCA module, three residual modules, and an adaptive average pooling layer, and finally the possible facial expression probabilities are output through a linear fully connected layer.

[0045] Among them, the MSCA module extracts feature information of different scales through multiple convolutional operations with different kernel sizes and paddings, then integrates the feature information of different scales through channel mixing by a convolutional layer, and performs element-wise multiplication on the feature after channel mixing and the input feature of the MSCA module to implement the convolutional attention mechanism.

[0046] The beneficial effects of the present invention: The system of the present invention combines the understanding of video content at the two levels of video description generation and emotional analysis of the video with user needs through a language model as text conditions, and introduces them into music generation. The music generator generates background music oriented to video content, ensuring the quality of the generated music, the matching degree of the music with the given video, and guaranteeing the richness of the music. The present invention is mainly applied to generating background music for short videos in the user-generated content mode. Brief Description of the Drawings

[0047] Figure 1 is the overall structural block diagram of the video background music generation system described in the present invention;

[0048] Figure 2 is the network architecture diagram of the video description generation module;

[0049] Figure 3 is the structural schematic diagram of the ECA_ASPP module;

[0050] Figure 4 is the structural schematic diagram of the SCDown_PSA module;

[0051] Figure 5 is the structural schematic diagram of the PSA module;

[0052] Figure 6 is the structural schematic diagram of the face detection module;

[0053] Figure 7 is Figure 6 the structural schematic diagram of the GSConv layer in;

[0054] Figure 8 is the structural schematic diagram of the MSCAM layer;

[0055] Figure 9 is Figure 8 the structural schematic diagram of the MSCB in;

[0056] Figure 10 is Figure 8 the structural schematic diagram of the CAB in;

[0057] Figure 11 is Figure 8 the structural schematic diagram of the SAB in;

[0058] Figure 12 is the structural schematic diagram of the VoVGSCSP layer;

[0059] Figure 13 is the structural schematic diagram of the face expression classification module;

[0060] Figure 14 is the structural schematic diagram of the MSCA in the face expression classification module. Detailed Implementation Manner

[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0062] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0063] The present invention will be further described below in conjunction with the accompanying drawings, but it is not a limitation of the present invention.

[0064] Specific embodiments I. In combination with Figure 1 As shown, the present invention provides a video background music generation system, including

[0065] A video description generation module, configured to extract video content features from the input video to obtain a video content text description;

[0066] A video emotion analysis module, configured to extract emotion features from the input video to obtain a video emotion category text description;

[0067] A text fusion module, configured to fuse the video content text description, the video emotion category text description, and the music type text description input by the user to obtain a background music text description;

[0068] A music generator, configured to generate a target background music according to the background music text description.

[0069] This embodiment constructs a deep learning network model. By inputting a video, it can achieve that the generated music not only matches the video content and emotion, but also has high quality.

[0070] Further, in combination with Figure 2 As shown, the video description generation module includes:

[0071] An action recognition network TSP, configured to extract video features of the input video to obtain a video frame-level feature map;

[0072] A multi-scale feature encoder, configured to extract different-level multi-scale feature sequences from the video frame-level feature map;

[0073] A deformable Transformer encoder, configured to encode different-level multi-scale feature sequences and the input position encoding to obtain multi-scale frame-level encoded features;

[0074] The ECA_ASPP module is used to enhance the multi-scale frame-level encoded features to obtain enhanced multi-scale frame-level encoded features;

[0075] The deformable Transformer decoder is used to query and decode the enhanced multi-scale frame-level encoded features according to the event queries to obtain event-level query decoded features and reference points;

[0076] The description generation head is used to adopt a deformable soft attention mechanism to obtain text descriptions of multiple event queries according to the event-level query decoded features and reference points;

[0077] The localization head is used to obtain the bounding box prediction and binary classification of each event query according to the event-level query decoded features and reference points to obtain a tuple of currently detected event queries; wherein the bounding box prediction obtains the two-dimensional relative offset between the current event and the reference event, and the binary classification obtains the foreground confidence of the current event;

[0078] The event counter is used to globally aggregate the event-level query decoded features to obtain a global feature vector representing the entire input video content, and predict the number of events in the video according to the global feature vector;

[0079] The video content prediction module is used to sort based on the foreground confidence according to the text descriptions of multiple event queries, the tuples of currently detected event queries, and the number of events in the video to obtain a video content text description.

[0080] The video description generation module is mainly composed of a two-layer deformable Transformer. The deformable Transformer replaces the self-attention module in the Transformer encoder and the cross-attention module in the Transformer decoder with a deformable attention (multi-scale deformable attention, MSDAtt) module, which can enable the model to have a faster convergence speed and better representation ability in object detection. MSDAtt alleviates the problem of slow self-attention convergence of the Transformer when processing image feature maps by focusing on a set of sparse sampling points around the reference point. Given a multi-scale feature map where query element q j , the normalized reference point p j ∈[0, 1] 2 , MSDAtt outputs a context vector on the feature map of L scales through the weighted sum of K×L sampling points according to Formula 1 and Formula 2.

[0081]

[0082]

[0083] where and A jkl are the attention weight and position of the k-th sampling key for the j-th query element at the l-th scale respectively. W is the mapping matrix of the key elements. φ t maps the normalized reference point to the feature map at the l-th scale. Δp jkl is the sampling offset with respect to φ l (p j ). A jkl and Δp jkl are obtained by linear mapping of the query elements. The ECA_ASPP module is added after the deformable encoder as shown in Figure 3 Figure.

[0084] Combined as shown in Figure 2 Figure, the video description generation module first encodes the features of the video, uses the pre-trained action recognition network TSP to extract frame-level features, and changes the time dimension of the feature map to K = 512 through the function interpolation method for convenient batch processing.

[0085] The decoding network contains a deformable Transformer decoder and three parallel heads, one description head for generating descriptions, one localization head for predicting the confidence scores of event boundaries, and one event counter for predicting the appropriate event numbers. The purpose of the decoder is to directly query the event-level features from the frame features conditioned on N = 10 learnable embeddings as well as their corresponding scale reference points p j . p j is predicted from q j through linear mapping with a sigmoid activation function. The event queries and reference points serve as initial guesses for the event features and positions (center points), which will be continuously refined at each decoding layer. The output query features and reference points are respectively denoted as

[0086] The encoded features are input into the deformable Transformer decoder, and two parallel prediction heads predict the boundaries and descriptions for each event query, and the event counter predicts the number of events N set from a global perspective. Select the N set events with the highest confidence to ensure that a complete and relevant story can be obtained finally. The role of the localization head is to perform bounding box prediction and binary classification for each event query. The purpose of bounding box prediction is to predict the two-dimensional relative offset between the true annotation segment and the reference point, and the purpose of binary classification is to generate the foreground confidence of each event query. Both bounding box prediction and binary classification are implemented by a multi-layer perceptron, obtaining a set of tuples representing the detected events represent the start time and end time, where is the positioning confidence of the event query The description generation head uses deformable soft attention (DSA) to concentrate the soft attention weights in a small area around the reference point. When generating the t-th word w t , first generate K sampling points from each f l , with the language query h jt and the event query as conditions, following Formula 1 and Formula 2, where h jt represents the hidden state in the LSTM. Then, regard the K×L sample points as keys / values, and regard as the query in the soft attention. Since the sampling points are distributed around the reference point , the output feature z jt of DSA is restricted in a relatively small area. The LSTM combines the context feature z jt , the event query feature and the previous word w j,t-1 as inputs. The probability of the next word w jt is obtained by h jt through a fully connected layer with softmax activation. Finally, an LSTM obtains a sentence where M j is the sentence length.

[0087] Furthermore, the video emotion analysis module includes:

[0088] A face detection module, used to perform face detection on all frames of the input video, remove most of the background, and obtain a face image containing only the face;

[0089] A face expression classification module, used to extract features from the face image and obtain a text description of the video emotion category through the classification head.

[0090] In this embodiment, the music generator uses an Encodec encoder to obtain audio tokens according to the background music text description, and uses a T5 text encoder to obtain text tokens according to the background music text description; the audio tokens and text tokens are used by an autoregressive Transformer decoder to generate target audio tokens, and then the Encodec decoder is used to generate the target background music.

[0091] Combined with Figure 4 as shown, the multi-scale feature encoder is implemented based on the SCDown_PSA module;

[0092] The video frame-level feature maps are obtained by using four temporal convolutional layers with a stride of 2 and a kernel size of 3 to obtain four different-level feature sequences; among them, the first-level feature sequence is added element-wise to the second-level feature sequence after passing through the scDown module to obtain a new second-level feature sequence; the new second-level feature sequence is added element-wise to the third-level feature sequence after passing through the scDown module to obtain a new third-level feature sequence; the new third-level feature sequence is added element-wise to the fourth-level feature sequence after passing through the scDown module, and the fused feature obtained is then passed through the PSA module to obtain a new fourth-level feature sequence; the first-level feature sequence, the new second-level feature sequence, the new third-level feature sequence, and the new fourth-level feature sequence form different-level multi-scale feature sequences.

[0093] To better utilize multi-scale features to predict multi-scale events, use the SCDown_PSA module as shown in Figure 4 . First, use L = 4 temporal convolutional layers on the feature maps to obtain 4 different-level feature sequences; then, use the PSA module as shown in Figure 5 as the new fourth-level feature sequence, and input it together with the position embedding into the deformable Transformer encoder to extract the relationship between frames at multiple scales and output the frame-level features.

[0094] Combined with Figure 5 as shown, the process of the PSA module obtaining the new fourth-level feature sequence includes:

[0095] The PSA module uses the SPC module to divide the channels, performs multi-scale feature extraction on the spatial information of the input fused feature on each channel, then uses the Squeeze and Excitation (SE) module to extract the channel attention of different-scale features to obtain the channel attention vectors at different scales; then uses Softmax to calibrate the features of the channel attention vectors at different scales to obtain the attention weights after multi-scale channel interaction; then performs an element-wise dot product operation on the multi-scale features extracted on each channel based on the attention weights after multi-scale channel interaction to obtain the new fourth-level feature sequence after attention weighting of the multi-scale feature information.

[0096] The scDown module performs spatial downsampling on the input feature sequence through a 3x3 convolutional layer with a stride of 2, reduces the number of channels of the input feature sequence through a 1x1 convolution, reduces the width and height of the input feature map, reduces the number of channels of the input feature map through a 1x1 convolution, and at the same time preserves the spatial information of the feature map after downsampling. The spatial downsampling and channel downsampling are performed separately to avoid information loss in traditional downsampling.

[0097] Combined with Figure 3As shown, in the ECA_ASPP module, the ASPP module uses dilated convolutions with different dilation rates to obtain the global information of multi-scale frame-level encoded features. The global information of the multi-scale frame-level encoded features is passed through the adaptive average pooling layer, 1x1 convolutional layer, and Sigmoid activation function of the ECA module, and then multiplied element-wise with the global information of the multi-scale frame-level encoded features to obtain enhanced multi-scale frame-level encoded features;

[0098] The dilation rates of the dilated convolutions with different dilation rates are 1x1; 3x3, rate = 6; 3x3, rate = 12; 3x3, rate = 18 in sequence.

[0099] The ASPP module parallelly uses dilated convolutions with different dilation rates, and then through an adaptive average pooling operation, compresses the spatial dimension of the feature map to 1x1 to obtain global information, and fuses the outputs of the dilated convolutions with different dilation rates and the output of the adaptive average pooling. The output of the ASPP module is then passed through the adaptive average pooling layer, 1x1 convolutional layer, and Sigmoid activation function and multiplied element-wise with the output of the ASPP module.

[0100] In this embodiment, the event counter includes a max pooling layer and a fully connected layer with softmax activation. The event counter contains a max pooling layer and a fully connected layer with softmax activation. First, the most prominent information of the event query is compressed into a global feature vector, and then a fixed-size vector r is predicted len , where each value represents the probability of a specific number. In the inference stage, the predicted number of events is given by N set = argmax(r len ). The final output is obtained by selecting the top N set events with accurate boundaries and good captions from the N event queries. The confidence of each event query is calculated by formula 3:

[0101]

[0102] where is the probability of generating a word, the adjustment factor γ is used to correct the influence of the caption length, and μ is a balance coefficient.

[0103] The ActivityNet Captions dataset is used for training. During the training process, the model generates a set of N events with locations and captions. To match the predicted events with the annotated data in the global scheme, the Hungarian algorithm is used to find the best bipartite matching result. The matching cost is defined as C = α giou L giou + α cls L cls, where L giou represents the generalized IoU between the predicted time period and the ground truth segment, and L cls represents the focal loss between the predicted classification score and the true label. Matching pairs are selected to calculate the set prediction loss, which is the weighted sum of the gIoU loss, classification loss, focal loss, and description generation loss, as shown in Equation 4:

[0104] L = β giou L giou + β cls L cls + β ec L ec + β cap L cap , Equation 4

[0105] where L cap measures the cross-entropy between the predicted word probability and the true label, and is normalized by the description generation length. L ec is also the cross-entropy loss between the predicted count distribution and the true label. Prediction heads are added to each layer of the Transformer decoder, and the final loss is the sum of the set prediction losses of all decoder layers.

[0106] Furthermore, as shown in combination with Figures 6 to 12 , the network structure of the face detection module includes three parts: backbone, neck, and head:

[0107] For the backbone part, the 0th and 1st layers are 3x2 convolutional layers, the 2nd layer is a C2f layer, repeated three times, the 3rd layer is a 3x2 convolutional layer, the 4th layer is a C2f layer, repeated six times, the 5th layer is a 3x2 GSConv layer, the 6th layer is a C2f layer, repeated three times, the 7th layer is a 3x2 convolutional layer, the 8th layer is a C2f layer, repeated three times, the 9th layer is an SPPF layer, and the 10th layer is an MSCAM layer;

[0108] For the neck part, the 11th layer is an upsampling layer, the 12th layer is the concatenation of the outputs of the 6th and 11th layers, the 13th layer is a VoVGSCSP layer, the 14th layer is an upsampling layer, the 15th layer is the concatenation of the outputs of the 4th and 14th layers, the 16th layer is a VoVGSCSP layer, and the 17th layer is a DCNv4 layer;

[0109] For the head part, the 18th layer is a 3x2 convolutional layer, the 19th layer is the concatenation of the outputs of the 13th and 18th layers, the 20th layer is a VoVGSCSP layer, the 21st layer is a DCNv4 layer, the 22nd layer is a 3x2 convolutional layer, the 23rd layer is the concatenation of the outputs of the 10th and 22nd layers, the 24th layer is a VoVGSCSP layer, and the 25th layer is a DCNv4 layer;

[0110] The DCNv4 layer consists of two branches, each branch including a 3x1 convolutional layer and a corresponding 1x1 convolutional layer connected thereto; one branch is used for calculating the bounding box loss, and the other branch is used for calculating the classification loss;

[0111] The C2f layer performs a convolution on the input data, doubles the number of channels to twice the number of input channels, and then uses the Bottleneck module to gradually extract features; the Bottleneck module contains multiple convolutional layers and uses residual connections. After the output data of multiple Bottleneck modules is concatenated with the input data, the concatenated feature map is compressed through a convolution operation, and a feature map with the target number of channels is output;

[0112] Such as Figure 7 As shown, the GSConv layer downsamples the input through ordinary convolution, then uses DWConv depth convolution, concatenates the results of the two convolutions, and then performs a shuffle operation;

[0113] The SPPF module contains an initial convolutional layer for feature compression, then performs multiple serial max-pooling operations, and finally fuses features of different scales through a convolutional layer;

[0114] The MSCAM layer consists of CAB, SAB, and MSCB; MSCAM is a module that combines multi-scale convolution and attention mechanism. As Figure 8 shown, in which CAB processes the input feature map using global average pooling and global max pooling, extracts channel information, generates weight coefficients for the channels through two fully connected layers, and finally applies the weights to the original feature map using the Sigmoid function to achieve weighting for each channel. SAB first performs average pooling and max pooling on the input feature map to generate two-channel feature maps. Then, these two feature maps are input into a convolutional layer through concatenation to generate a spatial attention map. Finally, the spatial attention map is applied to the original feature map using the Sigmoid function. MSCB contains multiple depth convolutional layers, and each convolutional layer performs parallel convolution operations using different-sized convolutional kernels. The convolution results of these different scales are fused together to provide rich multi-scale feature information.

[0115] Combined with Figure 9 As shown, the VoVGSCSP layer contains a convolutional layer, two GSConv layers. The output results of the two GSConv layers are concatenated with the output result of the convolutional layer, and then an output is obtained through a convolutional layer.

[0116] DCNv4 removes the softmax normalization in the spatial aggregation of deformable convolution to enhance its dynamicity and expressive ability; optimizes memory access to minimize redundant operations for acceleration.

[0117] Training is carried out on the Wider Face dataset. The loss calculation is divided into two parts: classification and regression. The Sigmoid function is used to calculate the probability of each category, and the BCE Loss is used to calculate the global category loss. The regression loss calculation is divided into two parts: CIou_Loss and Distribution Focal Loss.

[0118] Combined Figure 13 As shown, the facial expression classification module sequentially has three feature extraction layers for hierarchical feature extraction; each feature extraction layer sequentially includes a deformable convolutional layer DCNv4, a two-dimensional convolutional layer, a batch normalization layer, and a ReLU activation function. A max pooling layer is configured after the ReLU activation function; the batch normalization layer can stabilize learning and improve training efficiency; the max pooling layer can reduce the spatial dimension.

[0119] The max pooling layer of the third feature extraction layer is connected to a squeeze-and-excitation module to help the network assign higher weights to effective features during the learning process, and then sequentially connected to an MSCA module, three residual modules, and an adaptive average pooling layer. Finally, the possible facial expression probabilities are output through a linear fully connected layer;

[0120] Among them, the MSCA module extracts feature information of different scales through convolutional operations with multiple different kernel sizes and paddings, and integrates information of different scales by channel mixing of these features. Specifically, channel mixing is performed through the last convolutional layer to integrate feature information of different scales, and the features after channel mixing are multiplied element-wise with the input features of the MSCA module to implement the convolutional attention mechanism.

[0121] The feature extraction process of the convolutional network is expressed as Equation 5:

[0122] X FE = CNet(X), Equation 5

[0123] where x is the original image sample, and X FB is the feature map output by the convolutional network.

[0124] The core of the SE network is the SE block, which performs two main functions: squeeze, using global average pooling to compress the spatial data from each channel into a global descriptor; excitation, adopting an S-shaped activation gating mechanism to capture channel dependencies.

[0125] X FE The squeeze of is as Equation 6:

[0126]

[0127] where H and W are the width and height of the feature map, X FE(c, i, j) represents the activation at position (i, j) in channel c.

[0128] The output z passes through the MSCA module as Figure 14 shown. First, it passes through convolutional operations with multiple kernel sizes and paddings to extract feature information at different scales. By performing channel mixing on these features, different-scale information is integrated, and the channel mixing operation is completed by the last convolutional layer. Finally, by performing element-wise multiplication of the channel-mixed features with the input features, a convolutional attention mechanism is achieved. The module selectively emphasizes or suppresses the input features by assigning different weights to the features of different channels.

[0129] The compressed output z is processed through two fully connected layers: a dimensionality reduction layer, followed by a dimensionality increase layer, with a ReLU activation in between. The excitation operation can be expressed as Formulas 7, 8, and 9.

[0130] X s1 = ReLU(W 1 z), Formula 7

[0131] X s2 = W 2 X s1 , Formula 8

[0132] s = sigmoid(X s2 ), Formula 9

[0133] where W 1 is the weight matrix for dimensionality reduction, X s1 is the output of the first dimensionality reduction layer, W 2 is the weight matrix that expands it back to the original number of channels, i.e., c, and X s2 is the output of the first dimensionality expansion layer. The s obtained from the excitation stage adjusts the weights per channel to adjust the original input features Figure X FE . This scaling operation is performed separately for each element and is represented by Y as Formula 10:

[0134] Y = s · X FE , Formula 10

[0135] The function in the residual block can be defined as Formula 11:

[0136] F(x) = H(x) + x, Formula 11

[0137] In the residual block, x is the input, H(x) is the output of the stacked layers, and F(x) is the final output.

[0138] The Adaptive Average Pooling (AAP) layer pools XRB As the output feature map of the residual block, is the output of the AAP operation, which can be expressed as Equation 12:

[0139]

[0140] Finally, the probability distribution of possible facial expressions is output through a linear layer, and this module is trained on the FER-2013 dataset.

[0141] The text of the user's demand for the generated music, the video description text generated by the video description generation module, and the video emotion category text output by the video emotion analysis module are called ChatGPT for integration and used as the input of the music generator.

[0142] The EnCodec encoder in the music generator is a convolutional autoencoder, which consists of a 1D convolutional layer and 4 convolutional blocks. Each convolutional block contains a residual unit and a downsampling layer with a strided convolution, and the kernel size K is twice the stride S. Whenever downsampling occurs, the number of channels doubles. After the convolutional blocks, there is a two-layer LSTM for sequence modeling, and finally a one-dimensional convolutional layer with a kernel size of 7, and the sizes of K are 4, 4, 5, and 8 respectively. The EnCodec encoder processes 32kHz mono audio, and the output frame rate is 50Hz. The embedding is quantized using RVQ with four quantizers, and the codebook size of each quantizer is 2048.

[0143] The latent space in the EnCodec model is quantized using Residual Vector Quantization (RVQ) and adversarial reconstruction loss. Given a reference audio random variable (d is the audio duration, f s is the sampling rate), EnCodec encodes it into a continuous tensor with a frame rate of f r <<f s Then this representation is quantized into where K is the number of codebooks used in RVQ and M is the size of the codebook. After quantization, K parallel discrete token sequences will be obtained, each with a length of T = d·f r representing the audio samples. In RVQ, each quantizer encodes the quantization error left by the previous quantizer, so the quantization values of different codebooks are generally not independent. The first codebook is the most important codebook, and the main problem with the representation Q obtained from the EnCodec model is that there are K codebooks at each time step. Here, it is changed to an arbitrary codebook interleaving pattern, Ω = {(t, k): {1,..., d·f r}, k ∈ {1, ..., K}) is the set of all time step and codebook index pairs. The codebook pattern is a sequence P = (P 0 , P 1 , P 2 , ... P s ), where for all 0 < s ≤ S, the Q model is established by parallelly predicting all positions in P s , conditioned on all positions in P 0 , P 1 , ..., P s-1 such that each codebook index appears at most once in any P s .

[0144] The formula is:

[0145] P s = {(s, k): k ∈ {1, ..., K}} Formula 13

[0146] Introduce a "delay" between codebooks:

[0147] P s = {(s - k + 1, k): k ∈ {1, ..., K}, s - k ≥ 0}, Formula 14

[0148] In the text encoding part, given a text description matching the input audio X, a conditional tensor is calculated using a T5 encoder. The discrete tokens obtained by RVQ are added to the positional encoding and input into the decoder of an autoregressive Transformer with L layers, each layer containing a causal attention block, and then a cross-attention block that takes the conditional tensor as input. The layer ends with a fully connected block composed of a linear layer, ReLU, and a linear layer. The attention and fully connected blocks are connected by residual connections, and layer normalization is applied to each block before adding them to the residual connections.

[0149] Although the present invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed, as long as they do not depart from the spirit and scope of the present invention as defined by the appended claims. It should be understood that different dependent claims and the features described herein can be combined in a manner different from that described in the original claims. It should also be understood that the features described in connection with individual embodiments can be used in other described embodiments.

Claims

1. A video background music generation system, characterized in that: include, A video description generation module is used to extract video content features from the input video and obtain a text description of the video content; The video sentiment analysis module is used to extract sentiment features from the input video and obtain a text description of the video sentiment category; A text fusion module, used to fuse the video content text description, the video emotion category text description and the music type text description input by the user to obtain a background music text description; The music generator is used to generate target background music according to the background music text description.

2. The video background music generation system according to claim 1, characterized in that: The video description generation module includes: The action recognition network TSP is used to extract video features from the input video and obtain a video frame-level feature map; A multi-scale feature encoder is used to extract multi-scale feature sequences of different levels from the video frame-level feature map; The deformable Transformer encoder is used to encode the multi-scale feature sequences at different levels and the positional encoding of the input to obtain multi-scale frame-level encoding features; The ECA_ASPP module is used to enhance the multi-scale frame-level coding features to obtain enhanced multi-scale frame-level coding features; A deformable Transformer decoder is used to query and decode the enhanced multi-scale frame-level encoded features according to the event query to obtain the event-level query decoded features and reference points; Description generation head, which uses a deformable soft attention mechanism to decode features and reference points from event-level queries to obtain text descriptions for multiple event queries; The positioning head is used to obtain the bounding box prediction and binary classification of each event query based on the event-level query decoding features and reference points, and obtain the tuple of the currently detected event query; the bounding box prediction obtains the two-dimensional relative offset between the current event and the reference event, and the binary classification obtains the foreground confidence of the current event; An event counter, used to globally aggregate the event-level query decoding features to obtain a global feature vector representing the entire input video content, and predict the number of events in the video based on the global feature vector; The video content prediction module is used to obtain a text description of the video content by sorting the text descriptions of multiple event queries, the tuple of the currently detected event query and the number of events in the video based on the foreground confidence.

3. The video background music generation system according to claim 2, characterized in that: The video sentiment analysis module includes: A face detection module is used to perform face detection on all frames of the input video to obtain a face image containing only faces; The facial expression classification module is used to extract features from facial images and obtain a text description of the video emotion category through the classification head.

4. The video background music generation system according to claim 3, characterized in that: The music generator uses the Encodec encoder to obtain audio tags according to the background music text description, and uses the T5 text encoder to obtain text tags according to the background music text description; the audio tags and text tags are generated by the autoregressive Transformer decoder to generate target audio tags, and then the Encodec decoder generates target background music.

5. The video background music generation system according to claim 2, characterized in that: The multi-scale feature encoder is implemented based on the SCDown_PSA module; The video frame-level feature map uses four temporal convolutional layers with a stride of 2 and a convolution kernel size of 3 to obtain feature sequences at four different levels; The first-level feature sequence passes through the scDown module and is then element-wise added to the second-level feature sequence to obtain a new second-level feature sequence; the new second-level feature sequence passes through the scDown module and is then element-wise added to the third-level feature sequence to obtain a new third-level feature sequence; the new third-level feature sequence passes through the scDown module and is element-wise added to the fourth-level feature sequence, and the obtained fusion features are then passed through the PSA module to obtain a new fourth-level feature sequence; the first-level feature sequence, the new second-level feature sequence, the new third-level feature sequence and the new fourth-level feature sequence form multi-scale feature sequences at different levels.

6. The video background music generation system according to claim 5, characterized in that: The process of the PSA module obtaining a new fourth-level feature sequence includes: The PSA module uses the SPC module to divide the channels, extracts multi-scale features from the spatial information of the input fusion features on each channel, and then uses the compression and excitation module to extract the channel attention of features of different scales to obtain channel attention vectors at different scales; then uses Softmax to perform feature calibration on the channel attention vectors at different scales to obtain the attention weights after multi-scale channel interaction; then, based on the attention weights after multi-scale channel interaction, the multi-scale features extracted on each channel are element-wise multiplied to obtain a new fourth-level feature sequence after multi-scale feature information attention weighting; The scDown module downsamples the input feature sequence in spatial dimension through a 3x3 convolution layer with a step size of 2, and reduces the number of channels of the input feature sequence through a 1x1 convolution.

7. The video background music generation system according to claim 2, characterized in that: In the ECA_ASPP module, the ASPP module uses dilated convolutions with different dilation rates to obtain the global information of multi-scale frame-level coding features. The global information of the multi-scale frame-level coding features is point-multiplied with the global information of the multi-scale frame-level coding features after the adaptive average pooling layer, 1x1 convolution layer and Sigmoid activation function of the ECA module to obtain the enhanced multi-scale frame-level coding features. The dilation rates of the dilated convolutions with different dilation rates are 1x1; 3x3, rate=6; 3x3, rate=12; 3x3, rate=18 respectively.

8. The video background music generation system according to claim 2, characterized in that: The event counter consists of a max pooling layer and a fully connected layer with softmax activation.

9. The video background music generation system according to claim 3, characterized in that: The network structure of the face detection module includes three parts: backbone, neck and head: The backbone part has 3x2 convolutional layers in layers 0 and 1, C2f layer in layer 2, repeated three times, 3x2 convolutional layer in layer 3, C2f layer in layer 4, repeated six times, 3x2 GSConv layer in layer 5, C2f layer in layer 6, 3x2 convolutional layer in layer 7, C2f layer in layer 8, repeated three times, SPPF layer in layer 9, and MSCAM layer in layer 10. The 11th layer of the neck part is upsampling, the 12th layer is the concatenation of the outputs of the 6th and 11th layers, the 13th layer is the VoVGSCSP layer, the 14th layer is upsampling, the 15th layer is the concatenation of the outputs of the 4th and 14th layers, the 16th layer is the VoVGSCSP layer, and the 17th layer is the DCNv4 layer; The 18th layer of the head part is a 3x2 convolution layer, the 19th layer is the concatenation of the outputs of the 13th and 18th layers, the 20th layer is a VoVGSCSP layer, the 21st layer is a DCNv4 layer, the 22nd layer is a 3x2 convolution layer, the 23rd layer is the concatenation of the outputs of the 10th and 22nd layers, the 24th layer is a VoVGSCSP layer, and the 25th layer is a DCNv4 layer; The DCNv4 layer includes two branches, each of which includes a 3x1 convolutional layer and a corresponding 1x1 convolutional layer; one branch is used to calculate the bounding box loss, and the other branch is used to calculate the classification loss; The C2f layer performs a convolution on the input data, making the number of channels twice the number of input channels, and then uses the Bottleneck module to gradually extract features; the Bottleneck module contains multiple convolution layers and uses residual connections. After the output data of multiple Bottleneck modules are concatenated with the input data, the concatenated feature map is compressed again through the convolution operation to output the feature map of the target number of channels; The GSConv layer downsamples the input using a normal convolution, then uses a DWConv deep convolution, concatenates the results of the two convolutions, and then performs a shuffle operation; The SPPF module contains an initial convolutional layer for feature compression, followed by multiple serial maximum pooling operations, and finally convolutional layers for feature fusion at different scales. The MSCAM layer consists of CAB, SAB, and MSCB; The VoVGSCSP layer contains a convolutional layer and two GSConv layers. The output results of the two GSConv layers are concatenated with the output results of the convolutional layer, and then the output is obtained through another convolutional layer.

10. The video background music generation system according to claim 3, characterized in that: The facial expression classification module has three feature extraction layers in sequence for hierarchical feature extraction; each feature extraction layer includes a deformable convolution layer DCNv4, a two-dimensional convolution layer, a batch normalization layer and a ReLU activation function in sequence, and a maximum pooling layer is configured after the ReLU activation function; The maximum pooling layer of the third feature extraction layer is connected to the compression and excitation modules, and then sequentially connected to the MSCA module, three residual modules and the adaptive average pooling layer, and finally outputs the possible facial expression probability through the linear fully connected layer; The MSCA module extracts feature information of different scales through multiple convolution operations with different kernel sizes and padding, and then integrates feature information of different scales through channel mixing through the convolution layer. The channel mixed features are element-wise multiplied with the input features of the MSCA module to realize the convolutional attention mechanism.